Keyword mining method and system based on multiple factors and bilingual channels

By using a keyword mining method based on multi-factor and dual semantic channels, the problems of chaotic data structure and keyword duplication on e-commerce platforms were solved, achieving accurate and differentiated keyword mining and improving the traffic structure and sales conversion of apparel companies.

CN121961701AActive Publication Date: 2026-05-01ZHEJIANG ZIBUYU ELECTRONIC COMMERCE CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG ZIBUYU ELECTRONIC COMMERCE CO LTD
Filing Date
2026-03-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, the search term report data structure of e-commerce platforms is chaotic and lacks structured classification, resulting in a mixture of high-traffic terms, long-tail terms, and brand terms. It is difficult to assess the breadth and depth of coverage, manual screening is inefficient, and static term libraries cannot actively discover potential new terms. The keyword construction results for the same type of products have high repetition.

Method used

We employ a keyword mining method based on multi-factor and dual semantic channels. Through a data reconstruction module, we divide products into different combinations to generate a long table structure. We utilize multi-dimensional matching scores and vectorized encoding, combined with an intelligent recommendation module for deduplication and re-sorting, dynamically select the reference range for keyword mining, and apply the MPS multi-factor scoring model and differentiated damping attenuation technology to accurately mine high-conversion long-tail keywords.

Benefits of technology

It solves the problems of internal competition and homogenized recommendations for apparel companies when optimizing keywords for a massive number of products, ensuring that each product discovers keywords with precise traffic and differentiated positioning, achieving the optimal overall traffic structure and maximizing sales conversion for the store, and improving the system's recall rate and business effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121961701A_ABST
    Figure CN121961701A_ABST
Patent Text Reader

Abstract

The invention provides a keyword mining method and system based on multiple factors and bilingual channels, and belongs to the technical field of data processing, and the method specifically comprises the steps that a file obtaining module is responsible for intelligently screening out a target worksheet set containing effective data, a slicing module is responsible for performing partition processing on worksheet content to obtain multiple data slices, a header of each data slice and sample data are input into a large language model workflow, dimension types and structural features of the data are intelligently identified, an execution module is responsible for generating management strategies and cue words, calling a large language model to dynamically generate targeted data processing function codes, executing the generated function codes through a function execution engine, and executing the data processing function codes through a data processing engine. And the data slices are subjected to standardized conversion, and all successfully processed data are merged and exported into the standardized spreadsheet, so that the standardized processing efficiency of the spreadsheet is improved.
Need to check novelty before this filing date? Find Prior Art

Description

A Keyword Mining Method and System Based on Multi-Factor and Dual Semantic Channels Technical Field

[0001] This invention belongs to the field of data processing technology, and in particular relates to a keyword mining method and system based on multi-factor and dual semantic channels. Background Technology

[0002] With the development of cross-border e-commerce, accurate keyword targeting is crucial for merchants to acquire traffic. Currently, when merchants utilize search term reports (such as ABA reports) provided by e-commerce platforms, they generally encounter the following problems: chaotic data structure and lack of classification: raw data is usually in the form of "wide tables," lacking structured classification, resulting in a mix of high-traffic terms, long-tail terms, and brand terms, making it difficult to assess the breadth and depth of coverage, and leading to low efficiency in manual screening.

[0003] The thesaurus is static and lacks semantic expansion: existing report analysis is mostly based on static review of historical data (existing thesaurus), which cannot actively discover potential new words that are "semantically related" to high-performing words but do not appear in the report (incremental word expansion).

[0004] To address the aforementioned technical issues, the invention patent application CN202411505725.6, "Method, System, and Computing Device for Optimizing Product Titles," utilizes a large language model based on product information to select optimized traffic keywords from third-party candidate traffic keywords and adds these optimized traffic keywords to the original title, thereby generating a new title. This approach better helps merchants filter traffic-driving keywords, more conveniently optimize product titles, and enhance product competitiveness. However, the above technical solution has the following technical problems: For the same type of product, if the same reference product's search data and conversion data are used to construct keywords, it may lead to a high degree of duplication in the keyword construction results. Therefore, how to determine the keyword mining strategy based on the reference product data of the combined products and the composition data of the same type of product, thereby reducing the risk of duplication from the root, has become an urgent technical problem to be solved.

[0005] Therefore, there is an urgent need for a keyword mining method and system based on multi-factor and dual semantic channels. Summary of the Invention

[0006] To achieve the objectives of this invention, the following technical solution is adopted: Specifically, this application provides a keyword mining system based on multi-factor and dual semantic channels, specifically including: a data reconstruction module, which divides products into different combinations based on product type, determines keyword mining strategies for products in the combination using similar product data and product data, obtains reference products for the products using the mining strategies, converts the reference product data from the e-commerce platform into a long table structure with "search term-product" granularity, and associates it with product brand and type metadata, and calculates a comprehensive match for each "search term-product" pair based on multiple dimensions. The system comprises a score and an index building module, which vectorizes the search terms and product titles in the long table to generate semantic vectors for the search terms and product titles, and stores them together with the matching scores in the retrieval engine. An intelligent recommendation module, responding to user-input queries, performs parallel matching retrieval based on text relevance and similarity retrieval based on dual-channel semantic vectors through the retrieval engine to obtain recall results. After deduplication of the recall results, they are reordered according to the matching scores to output the final keyword recommendation list. The module also uses the keyword mining and processing method for the products in the combination to determine the method for updating the keyword mining and processing results for the products.

[0007] The beneficial effects of this invention are as follows: by using similar product data and product data of the products in the combination, the keyword mining strategy for the products in the combination is determined. This solves the problems of "internal competition" and "homogenized recommendations" faced by apparel companies when optimizing massive amounts of product keywords. By intelligently analyzing the density of similar products and the richness of internal product lines, the most suitable keyword mining reference range is dynamically selected, thereby ensuring that keywords that can capture precise traffic and form a differentiated positioning are mined for each apparel product. This avoids the technical problem of overly similar keyword mining results caused by using the same similar products, and ultimately achieves the optimization of the overall traffic structure of the store and maximizes sales conversion.

[0008] Through its original MPS multi-factor scoring model, the characteristics of four orthogonal dimensions—conversion efficiency, position confidence, ranking long-tail effect, and market size—are non-linearly coupled, overcoming the one-sidedness of single-indicator evaluation and constructing a quantitative system that more closely approximates real business value.

[0009] By constructing independent vector spaces for both search terms and product titles, and establishing a hybrid retrieval architecture, a leap from "literal matching" to "semantic and intent matching" has been achieved. This enables the discovery of potential keywords that differ in form but are highly semantically related, effectively solving the challenges of cold start and incremental keyword expansion.

[0010] The scoring model innovatively applies differentiated dynamic damping attenuation technology, using strong attenuation for product position and very weak attenuation for search term ranking. This effectively suppresses "traffic hegemony" while ensuring the value of top keywords, accurately discovering high-conversion long-tail keywords that are ignored by traditional methods, and improving the overall recall rate and commercial utility of the system.

[0011] By utilizing keyword mining and processing methods for products in a combination, a method for updating the keyword mining and processing results of products is determined. This ensures the timeliness of the update processing for a small number of "updated products" with weak reference product data foundations and requiring high-frequency monitoring when reference keywords are updated. Furthermore, by combining the overlap between "non-updated products" with a large amount of reference product data and updated products, the reliability of keyword update processing when updated products respond to changes in reference product keywords is determined. The timeliness of the update processing is used to determine the keyword update processing strategy for non-updated products, thereby achieving absolute tilt in resource allocation and maintaining the basic health of the keyword database for the same type of product at extremely low cost.

[0012] Furthermore, the dimensions include product performance factors, product position factors, search term ranking factors, and search term volume factors.

[0013] Secondly, this application provides a keyword mining method based on multi-factor and dual-semantic channel, applied to the aforementioned keyword mining system based on multi-factor and dual-semantic channel. Specifically, it includes: S1 dividing products into different combinations based on product type; determining a keyword mining strategy for products in the combination using similar product data and product data; when the mining strategy is an optimized mining strategy, proceeding to the next step; S2 determining reference products based on the analysis results of the similarity between the product image and similar products; determining products in the combination where all reference products are used for keyword mining based on the reference product data and the similarity between the reference product and the reference products of the product; and using these as target products; determining a keyword mining method for the target product based on the composition data of the target product in the combination and the similarity between the target product and the reference products of the target product; S3 using the keyword mining method to mine keywords for the product and obtain keyword mining results; and determining a method for updating the keyword mining results of the product using the keyword mining method for the products in the combination.

[0014] Furthermore, the products can be divided into different groups, specifically including grouping products of the same type into the same group.

[0015] Furthermore, similar products are those of the same type as the product in question.

[0016] Furthermore, the method for determining the keyword mining strategy for the products in the combination is as follows: based on the similar product data of the product, determine the number of similar products of the product; based on the product data in the combination, determine the number of products in the combination; based on the number of similar products of the product in the combination and the number of products in the combination, determine the keyword mining strategy for the products in the combination.

[0017] Furthermore, the method for determining the keyword mining and processing method for the product is as follows: based on the composition data of the target product in the combination, determine the proportion of the target product in the combination; based on the similarity between the product and the reference product of the target product, determine the number of overlaps between the product and the reference product of the target product; based on the proportion of the target product in the combination and the number of overlaps between the product and the reference product of the target product, determine the keyword mining and processing method for the product.

[0018] Other features and advantages will be set forth in the following description, and the objects and other advantages of the invention are realized and obtained through the structures particularly pointed out in the description and the drawings.

[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0020] The above and other features and advantages of the present invention will become more apparent from a detailed description of exemplary embodiments thereof with reference to the accompanying drawings.

[0021] Figure 1 is a framework diagram of a keyword mining system based on multi-factor and dual semantic channels; Figure 2 is a flowchart of a keyword mining method based on multi-factor and dual semantic channels; Figure 3 is a flowchart of a method for determining the keyword mining strategy for products in a combination; Figure 4 is a flowchart of a method for determining the target product. Detailed Implementation

[0022] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0023] Example 1, as shown in Figure 1, provides a keyword mining system based on multi-factor and dual semantic channels, specifically including: a data reconstruction module, which divides products into different combinations based on product type, determines keyword mining strategies for products in the combination using similar product data and product data, obtains reference products for the products using the mining strategies, converts the reference product data from the e-commerce platform into a long table structure with "search term-product" granularity, and associates it with product brand and type metadata, and calculates a comprehensive matching score for each "search term-product" pair based on multiple dimensions; an index building module; and an index construction module. The first module performs vectorized encoding on the search terms and product titles in the long table, generating semantic vectors for the search terms and product titles, and stores them in the retrieval engine along with the matching scores. The second module, in response to the user's input query terms, performs parallel matching retrieval based on text relevance and similarity retrieval based on dual-channel semantic vectors through the retrieval engine to obtain recall results. After deduplication of the recall results, they are reordered according to the matching scores to output the final keyword recommendation list. The third module uses the keyword mining and processing method of the products in the combination to determine the method for updating the keyword mining and processing results of the products.

[0024] The beneficial effects of this invention are as follows: by using similar product data and product data of the products in the combination, the keyword mining strategy for the products in the combination is determined. This solves the problems of "internal competition" and "homogenized recommendations" faced by apparel companies when optimizing massive amounts of product keywords. By intelligently analyzing the density of similar products and the richness of internal product lines, the most suitable keyword mining reference range is dynamically selected, thereby ensuring that keywords that can capture precise traffic and form a differentiated positioning are mined for each apparel product. This avoids the technical problem of overly similar keyword mining results caused by using the same similar products, and ultimately achieves the optimization of the overall traffic structure of the store and maximizes sales conversion.

[0025] Through its original MPS multi-factor scoring model, the characteristics of four orthogonal dimensions—conversion efficiency, position confidence, ranking long-tail effect, and market size—are non-linearly coupled, overcoming the one-sidedness of single-indicator evaluation and constructing a quantitative system that more closely approximates real business value.

[0026] By constructing independent vector spaces for both search terms and product titles, and establishing a hybrid retrieval architecture, a leap from "literal matching" to "semantic and intent matching" has been achieved. This enables the discovery of potential keywords that differ in form but are highly semantically related, effectively solving the challenges of cold start and incremental keyword expansion.

[0027] The scoring model innovatively applies differentiated dynamic damping attenuation technology, using strong attenuation for product position and very weak attenuation for search term ranking. This effectively suppresses "traffic hegemony" while ensuring the value of top keywords, accurately discovering high-conversion long-tail keywords that are ignored by traditional methods, and improving the overall recall rate and commercial utility of the system.

[0028] By utilizing keyword mining and processing methods for products in a combination, a method for updating the keyword mining and processing results of products is determined. This ensures the timeliness of the update processing for a small number of "updated products" with weak reference product data foundations and requiring high-frequency monitoring when reference keywords are updated. Furthermore, by combining the overlap between "non-updated products" with a large amount of reference product data and updated products, the reliability of keyword update processing when updated products respond to changes in reference product keywords is determined. The timeliness of the update processing is used to determine the keyword update processing strategy for non-updated products, thereby achieving absolute tilt in resource allocation and maintaining the basic health of the keyword database for the same type of product at extremely low cost.

[0029] Furthermore, the dimensions include product performance factors, product position factors, search term ranking factors, and search term volume factors.

[0030] In one possible specific implementation, an original wide table containing the search term, the top N ranked products (top_product_asin_1 to N), and their corresponding click share and conversion share are obtained from an e-commerce platform. Through a traversal transformation algorithm, each row of the wide table data is broken down into N independent records, forming a long table with (search term, product ASIN) as the primary key. During this process, each record is clearly associated with product brand (top_brand) and category (top_category) information, resolving the ambiguity in the original data.

[0031] Click share refers to the percentage of clicks received by a specific product (usually your product or its competitors) in the search results page after a consumer searches for a specific search term within a specific statistical period and e-commerce platform (such as the Amazon Brand Analytics reporting period), out of the total number of clicks received by all products under that search term. It directly measures the product's exposure conversion attractiveness or click competitiveness under a specific search term.

[0032] Conversion share refers to the percentage of actual purchases (or order conversions) achieved by a specific product through that search term within the same search term and statistical period, out of the total number of purchases achieved by all products under that search term. It measures the product's ultimate sales monetization ability or conversion efficiency under a specific search term.

[0033] This metric is designed for deeper quantification from the perspective of "commercial value realization." Clicks are important, but traffic that converts clicks into actual purchases is high-value traffic. Conversion share reveals a product's ability to meet users' genuine purchase intent in a specific search context and successfully complete a transaction in a competitive environment. It is directly related to sales revenue and is a key indicator for measuring the purity of keyword commercial value. Similar to click share, using the "share" format also ensures the comparability of product conversion efficiency across search terms with different traffic volumes.

[0034] MPS score calculation: A matching score (MPS) is calculated for each record in the long table, which is defined as the product of four factors: MPS = PPF × PDF × RDF × VF.

[0035] Product Performance Factor (PPF): PPF = (click_share × w_click) + (conversion_share × w_conv), where w_conv (e.g., set to 3) is greater than w_click (e.g., set to 1) to give higher weight to conversion behavior and filter high-net-worth intent keywords.

[0036] Product Position Factor (PDF): The exponential decay model PDF = exp(-λ_pos × (product_position - 1)) is adopted, where the decay coefficient λ_pos can be 0.3, simulating the non-linear decay of user attention as the ranking position decreases.

[0037] Search Rank Factor (RDF): The RDF model is adopted with a small decay, RDF = exp(-λ_rank × (search_rank- 1)), where the decay coefficient λ_rank takes a minimum value such as 0.001, so as to smoothly retain long-tail ranking terms and avoid precipitous cut-off of their value.

[0038] Search term volume factor (VF): Incorporating external monthly search volume data, VF = log(1 + search_volume). Utilizing the compression properties of the logarithmic function, it balances the numerical span between top-ranking keywords and long-tail keywords, ensuring the model's numerical stability and comparability across different traffic volumes.

[0039] In another embodiment, an index is created whose mapping includes the following core fields: the search_term and product_title fields for traditional text retrieval; the search_term_vector and product_title_vector fields for vector retrieval; and metadata fields such as MPS score, brand, and category for filtering and display.

[0040] Hybrid retrieval and ranking: Multi-path recall: When the user enters the query term Q, the system executes two-path retrieval in parallel.

[0041] Path A (Text Recall): Use the BM25 algorithm to perform full-text matching recall on the search_term field.

[0042] Path B (Semantic Recall): After vectorizing Q, perform approximate nearest neighbor search on the search_term_vector and product_title_vector fields respectively to recall semantically similar search terms and related products.

[0043] Deduplication of results: Merge the two recall results, utilize the field folding function of the search engine, use search_term as the folding key, and retain only the record with the highest MPS score in each group to achieve keyword-level deduplication.

[0044] Fine-ranking output: The top K candidate results after deduplication are subjected to a second fine-ranking process. The final ranking score is dominated by the offline calculated MPS score, with a slight combination of the relevance score from real-time retrieval, for example: Final_Score = MPS × (1 + ε × Relevance_Score). The recommended keyword list is output in descending order of the final score.

[0045] In one possible embodiment, the messy original data is organized—the "wide table" is transformed into a "long table": the original "wide table" is constructed based on reference products, with the search term "study light", the first ranked product: product A (60% click-through rate, 20% conversion rate), the second ranked product: product B (30% click-through rate, 10% conversion rate), and the third ranked product: product C (10% click-through rate, 5% conversion rate).

[0046] The converted "long table" looks like a clear Excel spreadsheet: Table 1: Converted Table

[0047] Step 2: Scientific Scoring – “MPS Multifactor Scoring Model”: Patented technology: Calculates a comprehensive score for each row (a combination) in the table. MPS = PPF × PDF × RDF × VF.

[0048] Using the data from "study light - product A" to calculate: Product performance factor: mainly looking at conversion value.

[0049] PPF = Click Share * 1 + Conversion Share * 3 (assuming conversion weight is 3 times that of click); Calculation: 60% * 1 + 20% * 3 = 0.6 + 0.6 = 1.2; Product Position Factor: Considers attention decay. The probability of a product ranked 1st and 3rd is different from that of being noticed by users.

[0050] PDF = exp(-0.3 * (ranking - 1)), calculation (ranking 1st): exp(-0.3 * 0) = 1.0 (ranking 1st, full weight); Search term ranking factor: consider long-tail value. The term "study light" may rank 5000th on the platform's overall ranking, but as long as it is effective, it should not be discarded.

[0051] RDF = exp(-0.001 * (rank - 1)), calculated (assuming the total ranking of this search term is 5000): exp(-0.001*4999) ≈ exp(-5) ≈ 0.0067 (the decay is very small, preserving value); Search term size factor: considering market size. The monthly search volume for "study light" may be very large, requiring logarithmic smoothing.

[0052] VF = log(1 + monthly search volume), calculated (assuming monthly search volume is 100,000): log(1+100,000) ≈ log(100,001) ≈ 11.5.

[0053] Final MPS score: MPS = 1.2 × 1.0 × 0.0067 × 11.5 ≈ 0.092.

[0054] In layman's terms: This score combines four factors: "high conversion rate (high PPF)," "ranked first (high PDF)," "useful despite being a niche keyword (RDF not zero)," and "significant market size (high VF)." The system calculates this for all keywords in this way to fairly compare top-performing keywords and potential long-tail keywords.

[0055] Step 3: Intelligent Association – “Dual-channel Semantic Vector Retrieval”: This is the key to achieving intelligent word expansion, solving the problem of “how to associate ‘study light’ with ‘readinglamp’”.

[0056] Traditional method (literal matching): If you search for "study light", the system can only find words that completely contain these words, such as "good study light".

[0057] Patented method (semantic matching): Channel 1: Word-to-word association: The system converts "study light" into a mathematical vector (which can be understood as "semantic coordinates"), and then searches for the word closest to it in the vector space, such as "reading lamp".

[0058] Channel Two: Keyword and Product Association: The system will also match the vector for "study light" with the vectors for all product titles. If a product titled "LED Desk Lamp for Students" has a very similar vector, and this product is associated with the search term "desk light," then "desk light" will also be recalled.

[0059] Effect: In this way, the system can "fill in" and mine keywords with different shapes but highly related meanings from massive amounts of data, greatly expanding the range of words to be selected.

[0060] Step 4: Deduplication and Sorting – “Hybrid Search and Final Recommendation”: When you enter the seed keyword “desk lamp” into the system for a query, the system will simultaneously launch two engines: Text search engine: quickly find all words that literally contain “desk lamp”, such as “led desklamp”.

[0061] Semantic search engine: Through the above two channels, find words with similar meanings, such as "study light", "reading lamp", and "workspace light".

[0062] Next comes "refinement": deduplication and merging: the term "study light" may be recalled multiple times from multiple paths, and the system will merge them into one.

[0063] Sort by MPS: This is the decisive step. The system no longer simply sorts by relevance or search volume, but instead sorts by the MPS score, which we calculated earlier and which best represents business value.

[0064] Final output: You will get a recommendation list similar to the table below, which directly tells you which related keywords have the highest value and are most worth prioritizing: Table 2 Keyword Recommendation Results

[0065] Example 2: Specifically, as shown in Figure 2, this application provides a keyword mining method based on multi-factor and dual-semantic channels, applied to the aforementioned keyword mining system based on multi-factor and dual-semantic channels. Specifically, it includes: S1: Based on the type of the product, dividing the products into different combinations; using similar product data and product data within the combinations, determining the keyword mining strategy for the products in the combinations; when the mining strategy is an optimized mining strategy, proceeding to the next step; S2: Determining the reference product for the product based on the analysis results of the similarity between the product image and similar products; and based on the reference product data and similar product data... S3. Based on the similarity between the product and the reference products in the same combination, identify the product for keyword mining using all reference products and designate it as the target product. Then, using the composition data of the target product in the combination and the similarity between the product and the reference products of the target product, determine the keyword mining method for that product. S4. Use the keyword mining method to mine the keywords of the product to obtain the keyword mining results. Finally, use the mining results and the keyword mining method to determine the method for updating the keyword mining results of the product.

[0066] Furthermore, the products can be divided into different groups, specifically including grouping products of the same type into the same group.

[0067] Furthermore, similar products are those of the same type as the product in question.

[0068] Furthermore, as shown in Figure 3, the method for determining the keyword mining strategy for the products in the combination is as follows: The core decision-making objective of this embodiment is to solve the problems of "internal competition" and "homogenized recommendations" faced by apparel companies when optimizing massive amounts of product keywords. Its core logic lies in dynamically selecting the most suitable keyword mining reference range by intelligently analyzing the density of external market competitors and the richness of internal product lines. This ensures that keywords that can capture precise traffic and form a differentiated positioning are mined for each apparel product, ultimately achieving the optimization of the overall store traffic structure and maximizing sales conversion.

[0069] This method follows a closed-loop decision-making process of "market insight → self-assessment → strategy selection → execution mining". First, assess the level of competition the target product faces on the platform; second, assess the SKU depth of similar products within the company; finally, based on the two-dimensional analysis of "external competition and internal abundance", intelligent decision-making determines whether to adopt a "broad-spectrum reference" strategy for rapid coverage or to initiate a "refined grouping" strategy for in-depth mining to drive the precise allocation of marketing resources.

[0070] S11 Based on the similar product data of the product, determine the number of similar products of the product; S12 Based on the product data in the combination, determine the number of products in the combination; S13 Based on the number of similar products of the product in the combination and the number of products in the combination, determine the keyword mining strategy for the products in the combination.

[0071] It is understandable that if the number of similar products of the product is less than the preset threshold for the number of similar products, then the number of similar products of the product is small. If the reference products for keyword mining of different products are classified in a secondary manner, the reference value will be low due to the small number. Therefore, the keyword mining strategy for the products in the combination is to use all similar products for keyword mining.

[0072] Additionally, it should be noted that if the number of similar products is not less than a preset threshold for the number of similar products, then the number of products in the combination is obtained, and it is determined whether the number of products in the combination is greater than the preset threshold for the number of products. If yes, then the keyword mining strategy for the products in the combination is determined to be an optimized mining strategy; otherwise, then the keyword mining strategy for the products in the combination is determined to be keyword mining processing using all similar products.

[0073] In one possible embodiment, a cross-border e-commerce apparel company that sells "yoga pants" is used as an example to illustrate the entire decision-making process.

[0074] Product bundle definition: All "yoga pants" category products are grouped into the same bundle.

[0075] Similar products definition: On e-commerce platforms, all other brands' products that belong to the same "yoga pants" category as your own products.

[0076] Step S1: Assess the external market competition environment (number of similar products). The system automatically calls the platform API or category data to count the total number of active products for sale under the "Yoga Pants" category (excluding its own products).

[0077] Example result: A search revealed approximately 50,000 yoga pants available for sale on the platform. This enormous number indicates extremely fierce market competition.

[0078] Step S2: Inventory the scale of the internal product line (number of products in the bundle) and count the number of all "yoga pants" SKUs in the company's own product library.

[0079] Example result: The company has a wide product line with 120 different yoga pants (covering high-waisted, cropped, printed, solid color, and different materials).

[0080] Step S3: Strategy Decision Based on Two-Dimensional Analysis: The system compares the above two data points with preset thresholds and automatically generates a data mining strategy. Assume the preset thresholds are: similar product quantity threshold = 10,000, product quantity threshold = 50.

[0081] Scenario A: Blue Ocean Competition or Single Product Line → Use "Broad-Spectrum Reference Strategy" Judgment Criteria: Number of Similar Products < 10,000 or Number of Products in the Combination < 50.

[0082] Strategic logic: When external competition is not intense, you can boldly refer to all competitor keywords; or when you have few products, you don't need to worry about internal competition, and the goal is to quickly acquire mainstream market traffic.

[0083] Example: If a company only sells 5 basic black yoga pants, it can directly refer to 50,000 competitors across the entire network, dig out high-traffic generic keywords such as "yoga pants" and "workout leggings" to quickly establish brand awareness.

[0084] Scenario B: Intense competition and a rich product line → Activate the "Optimization Mining Strategy": Judgment criteria: Number of similar products ≥ 10,000 and number of products in the combination ≥ 50.

[0085] Strategic Logic: This is precisely the situation faced by the example company (50,000 competing products, 120 self-operated products). If all 120 products compete for the broad term "yoga pants," it will lead to severe internal traffic cannibalization and wasted advertising budget. Refined segmentation must be implemented.

[0086] Decision: Since both conditions are met, the system automatically determines to "optimize the mining strategy".

[0087] Furthermore, when the mining strategy is an optimized mining strategy, different reference products for keyword mining processing are determined, and keyword mining processing is carried out in a targeted manner.

[0088] Furthermore, the reference product is a currently listed product whose similarity to the product image meets the requirements.

[0089] Specifically, as shown in Figure 4, the method for determining the target product is as follows: This embodiment aims to solve the problem of selecting the starting point for product keyword mining: In a product portfolio, which products should be prioritized for in-depth and global keyword mining with reference to all their visual competitors in order to most efficiently build a unique and comprehensive keyword base library? Its core decision-making objective is to quantitatively evaluate the "scarcity" of each product's external visual competitors and its "network independence" from other internal products, selecting products with a simple competitive environment and low risk of internal keyword homogenization as "target products," thereby ensuring that the initial batch of keyword assets generated has high value and low redundancy.

[0090] This method follows a screening process of "assessing competitor scarcity → analyzing internal network overlap → hierarchical filtering to determine targets". First, products are sorted according to the number of visual competitors to identify unique styles in the market. Second, the overlap between each product's competitors and those of other products in the portfolio is analyzed to assess its independence. Finally, a four-level decision funnel is used to progressively filter products based on four dimensions: "absolute scarcity", "segment independence", "global share control", and "network influence", ultimately determining the most suitable target product list as the benchmark for global keyword mining.

[0091] S21 uses the reference product data of the product to determine the quantity of reference products for the product, and uses the quantity of reference products to determine the ranking result of the product in the combination; quantifies and ranks the scarcity of competitors: "the quantity of reference products for the product" refers to the total number of other brand competitors that are visually highly similar to the target product in appearance, style, and design, and are currently on sale, as identified by image search or image similarity algorithms on the e-commerce platform. This quantity directly reflects the intensity of direct and homogeneous competition faced by the product in the open market.

[0092] Visual similarity is a key factor influencing consumer perception and search comparison. The fewer visual competitors a product has, the less reference data is available for keyword mining. Therefore, when mining keywords, it's crucial to comprehensively consider the needs of the reference products. Targeting such products for in-depth analysis helps to gain a competitive edge in a niche market, and the keywords discovered are more likely to include blue ocean terms, avoiding competition with a vast number of generic terms repeated by competitors.

[0093] This step establishes preliminary screening criteria, transforming "design uniqueness" into a quantifiable "competitive scarcity" indicator, and ranking all products within the portfolio to provide an objective and comparable data foundation for subsequent screening.

[0094] Specific example: The system scans through the image retrieval interface to obtain the number of visual competitors for each product: Product P (Pioneer Model): 65 reference products.

[0095] Product Q (Niche Style): 120 reference products.

[0096] Product R (Basic): 2200 reference products.

[0097] The products are sorted from least to most in quantity as follows: Product P (65) -> Product Q (120) -> Product R (2200) -> ...

[0098] S22 determines the number of reference products that overlap between the product and the reference products of different products in the same combination, based on the similarity between the product's reference product and the reference products of the products in the same combination; it analyzes the overlap of the internal competitor network: "the number of reference products that overlap between a product and the reference products of different products in the combination" refers to the size of the intersection between the product's own visual competitor set and the visual competitor set of another product in the combination. It accurately measures the degree of overlap of the visual competitor customer groups covered and competed for by two products in the external market.

[0099] Simply having a small number of competitors is insufficient to determine whether a product is suitable as a "target" for comprehensive keyword mining. If a product's few competitors also highly overlap with several of its own products, then keywords mined based on that product may severely intersect with the organic keyword flow of other products, leading to internal traffic competition. The purpose of analyzing overlap is to assess the "potential intrusiveness" or "unique contribution" of the keyword assets generated by this product if it were selected as a target to other members of the portfolio.

[0100] This step shifts from external competition analysis to internal ecosystem synergy diagnosis. It constructs a "competitive overlap network" among products to identify those products that are not only unique externally but also relatively independent internally. This is a crucial step in avoiding internal homogenization of keyword assets.

[0101] Specific example (continued from S21): The system further analyzes the overlap between the competitor set of product P and the competitor set of other products in the combination: Among the 65 competitors of product P, the number of competitors that overlap with product Q is 5.

[0102] Of the 65 competing products of product P, 10 overlap with the number of competing products of product R.

[0103] S23 determines whether the product is the target product based on the sorting result of the product in the combination and the number of times the product overlaps with the reference product of different products in the combination.

[0104] It should be noted that the target product is the product for which keyword mining was performed using all reference products.

[0105] It is understood that the sorting result of the goods in the combination is obtained by sorting from the fewest to the most based on the number of reference goods in the combination to which the goods belong.

[0106] Specifically, determining whether a product is a target product includes: S231, based on the product's ranking result in the combination, determining whether the product's ranking result in the combination is before a preset position; if yes, then the product is determined to be a target product; otherwise, proceeding to step S232; Initial selection—capturing absolutely scarce products: "Product ranking result in the combination" refers to the order of the reference products in ascending order based on the number of reference products obtained in step S21. "Before a preset position" is an absolute order threshold, such as the top 3. Products with extremely scarce reference products have a small amount of reference data when extracting keywords. Therefore, if secondary segmentation is performed, the keyword extraction results will be less comprehensive and accurate.

[0107] The preset positions are the top 2. Product P (ranked 1st) is directly selected as the target product because it has very few competitors (65). Product Q (ranked 2nd, 120 products) is also selected.

[0108] S232 determines reference products that do not overlap with any other product in the combination based on the number of reference products that overlap with the product in the combination, and treats these as independent products. It then determines whether the proportion of independent products among the product's reference products is greater than a preset independent product proportion threshold. If yes, the product is determined as the target product; otherwise, proceed to step S233. Re-selection—Identifying Independent Track Products: "Independent products" refer to those competitors in the current product's reference product set that do not overlap with the reference product sets of any other product in the combination. "Independent product proportion" is the percentage of independent competitors out of the total number of competitors.

[0109] Some products may not have the fewest total number of competitors, but their competitor pool has extremely low overlap with other products within the same category. This means they are actually carving out a completely independent competitive track. Setting these products as targets ensures that the resulting keyword pool has minimal overlap with the keyword pools of other existing target products, thereby maximizing the overall coverage of the keyword library.

[0110] This rule supplements S231, expanding from "scarcity" to "independent relationship," enabling the discovery of "track pioneer" products that, despite being in a competitive environment, have highly differentiated positioning.

[0111] For product R not selected in S231, calculate its independent product ratio. Of its 2200 competitors, 2000 overlap with other internal products such as product P or product Q, leaving only 200 independent competitors, a ratio of 9.1%. Assuming the preset independent ratio threshold is 60%, then 9.1% < 60%, which does not meet the condition, so proceed to S233.

[0112] S233 Obtain the percentage of the number of target products in the combination, and determine whether the percentage of the number of target products in the combination is greater than the preset target product percentage threshold. If yes, determine that the product is not a target product. If no, proceed to step S234; Regulation - control the overall scale of target products: "Percentage of the number of target products in the combination" refers to the percentage of the number of products that have been identified as targets to the total number of products in the combination.

[0113] A company's operational resources (such as the cost of in-depth data analysis and the focus of advertising budgets) are limited. When there are too many target products, it is already possible to ensure that a large number of products in the portfolio can undergo keyword mining using relatively comprehensive reference products. The mining results are already reliable enough. Therefore, if the reliability is sufficient, products with few independent products can be directly determined not to be target products.

[0114] Specific decision: Currently, products P and Q have been identified as targets (2 products). Assuming the total number of products in the combination is 10, the target percentage is 20%. Set the preset target product percentage threshold to 25%. 20% < 25%, the upper limit is not triggered, therefore product R is still eligible to enter the final evaluation (S234).

[0115] S234 determines the influence coefficient of the product on the products in the combination based on the number of overlaps with the reference products in the combination and the proportion of the number of reference products in the product. Based on the influence coefficient of the product on the products in the combination, it determines whether the product is a target product.

[0116] It is understood that when the average value of the influence coefficient of the product on the products in the combination is less than the preset influence coefficient threshold, the product is determined to be the target product.

[0117] Final Review – Assessing Internal Network Influence: The "Influence Coefficient of a Product on Products in the Combination" is a comprehensive metric used to quantify the "dilution risk" or "homogenization influence" that a product may cause to the keyword uniqueness of other products due to the high overlap between its competitors and those of other internal products. Its calculation is typically based on the proportion of the number of reference products the product overlaps with to the total number of reference products for those other products, then the average value is taken. The lower the average value, the "purer" the product's competitor pool, and the smaller the threat to the keyword uniqueness of other products.

[0118] For products that have passed the first three rounds of screening, this is the final and most meticulous balancing act. Even if a product has many competitors and a low independent ratio, if its competitor pool has little overlap with its peers (i.e., a low influence coefficient), it means the additional information it brings is still relatively unique. Conversely, if the influence coefficient is high, it means its competitors severely overlap with others. Setting it as a target for global keyword mining will result in keywords that largely overlap with existing target products, limiting its value. Therefore, the final screening is conducted based on the core objective of maintaining the uniqueness of the overall keyword asset. This ensures that each newly added target product brings significant, non-repetitive incremental value to the keyword library, minimizing internal redundancy.

[0119] Specific decision: Calculate the influence coefficient of product R. Its significant overlap with products P and Q results in an average influence coefficient as high as 0.55. Assuming a preset influence coefficient threshold of 0.2, 0.55 > 0.2. Therefore, product R is ultimately determined not to be a target product.

[0120] This embodiment constructs a progressive intelligent selection system for target products, moving from the surface to the core, and from static indicators to dynamic network relationships. It innovatively transforms the visual competitive landscape of products into calculable scarcity indicators and network relationship data. Through a rigorous four-tiered filtering logic—"absolute scarcity priority → independent track supplementation → global scale control → final review of network influence"—it accurately identifies key products from a massive pool of goods that represent specific market segments without polluting the internal keyword ecosystem, serving as the starting point for in-depth analysis.

[0121] This application lays a solid foundation for unique keyword assets: by selecting target products with scarce competitors and low internal overlap for the first round of mining, it can ensure that the generated core keyword library is inherently highly unique and has low internal conflict, laying a high-quality foundation for the entire product's keyword system and effectively preventing internal traffic competition during operation: from the source, it avoids the problem of homogenization of subsequent keyword recommendations caused by selecting highly overlapping products as targets, thereby preventing potential internal consumption of advertising bidding between products and mutual influence of traffic.

[0122] Specifically, as shown in Figure 4, the method for determining the keyword mining and processing method for the product is as follows: This embodiment aims to solve the core contradiction of keyword strategy in multi-product operation: how to ensure that the keyword sets mined from different products have high uniqueness and minimize the overlap between them through systematic rules. Its core objective is not to actively seek differentiated markets, but to forcibly divert data from the source through a "defensive" or "isolated" reference product screening mechanism, so that the keyword mining process of different products is based on as different as possible competitor samples, thereby automatically generating keyword assets with inherent differences in the results, and completely avoiding the keyword homogenization problem caused by overlapping reference sources.

[0123] This method follows a control process of "global monitoring → source isolation diagnosis → representative allocation within the cluster". First, it monitors the density of "target products" that already have priority in mining within the monitoring portfolio. Second, for each product to be processed, it analyzes the overlap between its external competitor pool and the target product, and assesses the size of its independent competitor pool. Finally, for clusters with highly overlapping competitors among ordinary products, a "single representative" mining system is adopted. The core operation of the entire process is the intelligent allocation and isolation of the "reference product" data source, rather than post-processing deduplication of the mining results.

[0124] S31 determines the proportion of target products in the combination based on the composition data of the target products in the combination; Keyword explanation: "Proportion of target products in the combination" refers to the percentage of products that have been designated by the system for global, deep keyword mining out of the total number of products in the combination. Keyword mining of these target products will be unrestricted, using all their reference products.

[0125] This is the master switch for activating different control levels. If the target product's proportion is too high (e.g., exceeding 50%), it means that most products within the portfolio already have independent, prioritized keyword mining channels. At this point, to prevent the remaining few ordinary products from competing with these "protagonists" for reference resources, the system will enforce the strictest "reference source isolation" rule: ordinary products must not only completely avoid the target product's reference pool, but also, when highly similar to each other, can only be mined by one of them, eliminating any possible overlap and duplication from the source. Establishing isolation awareness from the top level of the portfolio prevents unnecessary internal data source competition in a resource-biased ecosystem, ensuring the depth and purity of keyword mining for the target products.

[0126] The combination contains 5 products, with 2 being target products A and B. The target product ratio = 2 / 5 = 40%. Assuming the preset target product ratio threshold is 50%, 40% is not greater than the threshold. Therefore, the highest level of mandatory isolation mode is not triggered, and the personalized diagnostic process for product C begins.

[0127] S32 determines the overlap between the product and the reference products of the target product based on the similarity between the product and the reference products of the target product; quantifies the conflict with the reference source of the "target product": the "overlap between the product and the reference products of the target product" refers to the size of the intersection between the external competitor set of the current ordinary product and the external competitor set of a certain target product. It directly measures how much of the potential data source of keyword mining the two have in common.

[0128] This is the first crucial step in achieving "reference source isolation." If a regular product shares a large number of identical competitor references with a target product, the keywords mined from these common competitors will inevitably be highly similar, thus undermining uniqueness. Therefore, it is essential to identify this overlap and provide data support for subsequent "avoidance" operations.

[0129] The abstract goal of "avoiding keyword overlap" is transformed into the quantifiable and actionable task of "reducing the intersection of reference product sets".

[0130] Specific example (continued from S31): Product C has 2000 reference products (competitors). Analyzing the overlap with the target product: the number of reference products overlapping with target product A is 120, and the number of reference products overlapping with target product B is 30. These overlapping reference products are the risk source of keyword homogenization.

[0131] S33 determines the keyword mining and processing method for the product based on the proportion of the target product in the combination and the number of overlaps between the product and the reference product of the target product.

[0132] It should be noted that if the proportion of the target product in the combination is greater than the preset target product proportion threshold, then the keyword mining processing method for the product is to no longer consider the reference products of the target product, and only consider the reference products that belong to multiple products excluding the target product when performing keyword mining processing on the product with the fewest reference products.

[0133] Furthermore, if the proportion of the target product in the combination is not greater than a preset target product proportion threshold, the following steps are also included: S331 Obtain the number of reference products that overlap between the product and the target product. Based on the proportion of the number of reference products that overlap between the target product and the product in the total number of reference products of the product, determine the influence coefficient of the target product on the product. Determine whether there is a target product whose influence coefficient is greater than a preset influence coefficient threshold. If so, all reference products except the reference products of the target product should be considered when performing keyword mining processing on the product. If not, proceed to step S332; Diagnosis—whether it is "covered" by a certain target product: "Influence coefficient of the target product on the product" refers to the proportion of the number of reference products that overlap between a single target product and the current ordinary product to the total number of reference products of that ordinary product. The higher the coefficient, the more the competitor pool of the target product "covers" most of the ordinary product's field of vision.

[0134] If a target product has an extremely high influence coefficient (e.g., exceeding 30%), it means that most of the keywords for ordinary products likely originate from the same competitor pool as the target product. To ensure uniqueness, the most direct method is to instruct the ordinary product to completely avoid all reference products of the target product during keyword mining, using all remaining reference products as data sources. This is the most lenient isolation measure, providing protective isolation for ordinary products overshadowed by a strong target product. It does not require consideration of overlap with reference products of products other than the target product, making it the most effective means of ensuring keyword uniqueness.

[0135] Calculate the influence coefficient of target product A on product C: 120 / 2000 = 0.06 (6%). Assuming the preset influence coefficient threshold is 0.3 (30%), 0.06 < 0.3. This isolation strategy is not triggered; proceed to the next step.

[0136] S332 removes reference products that overlap with the target product as available reference products, and determines whether the proportion of available reference products in the reference products of the product is less than a preset available proportion threshold. If so, all reference products except those of the target product need to be considered when performing keyword mining processing on the product. If not, proceed to step S333; Diagnosis – Sufficient Independent Data Source: “Available reference products” refers to the set of completely independent competing products remaining after removing all parts that overlap with any target product from the total reference products of ordinary products. “Usable reference product proportion” is the size proportion of this independent set.

[0137] Even if a product isn't covered by a single target product, if the independent competitor pool of a typical product is too small (e.g., <50%), it indicates a limited number of "clean" data sources available for differentiated keyword mining. Forcing it to use only this portion could lead to insufficient keyword mining. In this case, as a trade-off, allowing it to consider all reference products that overlap with the target product during mining, thereby expanding the scope of reference products, without needing to consider overlap with reference products other than the target product, is the most comprehensive way to ensure keyword uniqueness. This strikes a balance between ensuring uniqueness and guaranteeing the feasibility of mining, avoiding data source depletion due to excessive isolation.

[0138] Specific decision: The number of available reference products for product C = 2000 - 120 - 30 = 1850, accounting for 1850 / 2000 = 92.5%. Assuming the preset available percentage threshold is 50%, 92.5% > 50%, indicating that product C has extremely rich independent data sources. This step is not triggered; proceed to the final diagnosis.

[0139] S333: Based on the number of reference products that overlap with different products, determine the similarity coefficient between the product and the reference products of different products. Determine whether there are any products with a similarity coefficient greater than a preset similarity coefficient threshold. If yes, proceed to step S334. If no, all reference products except the target product's reference products need to be considered when performing keyword mining processing on the product. S334: Group the products with similarity coefficients greater than the preset similarity coefficient threshold into the same group. Determine whether the number of products in the group is greater than a preset product number threshold. If no, all reference products except the target product's reference products need to be considered only when performing keyword mining processing on the product with the fewest preset number of reference products. If yes, all reference products except the target product's reference products need to be considered only when performing keyword mining processing on the product with the fewest reference products.

[0140] Diagnosis and Decision Making – Handling Internal Overlap Between Ordinary Goods: The "similarity coefficient" measures the similarity of reference product sets between two ordinary goods. Specifically, it is determined by the ratio of the number of overlapping reference products in the combination of the two goods to the number of reference products in the goods with the fewest reference products among the two goods. Goods with a "similarity coefficient greater than a preset threshold" constitute a highly cohesive cluster, meaning that their external competitive views highly overlap.

[0141] This is crucial for preventing keyword homogenization among ordinary products. If multiple ordinary products reference almost the same competitor pool, even if they all avoid the target product, the keywords they generate will still be largely duplicated. To address this issue, the system groups these products into a "cluster" and implements a "representative system" within the cluster: only one product in the cluster (usually the one with the fewest reference products, as its perspective may be the most unique) is designated to use all available reference products for keyword mining, while other products in the cluster directly share its mining results or only use a subset assigned to them. In this way, the entire cluster generates only one set of keywords based on a common reference source, fundamentally avoiding internal duplication.

[0142] The principle of "reference source isolation" is extended from "target product - ordinary product" to "ordinary product - ordinary product". Through clustered management and representative mining, reference source merging and isolation are also achieved at the ordinary product level, ensuring that the keyword mining process of any two products will not be based on highly overlapping data sources.

[0143] In one possible embodiment, S333 (cluster identification): Suppose the system finds that the similarity coefficients between products C, D, and E are all 0.8 (threshold 0.7), and the three are classified into the same cluster.

[0144] S334 (representing specification): This combination has 3 members (C, D, E). Assuming the preset combination size threshold is 3, where 3 is not greater than 3, the overlap risk is low. The system specifies that for each reference product, only the two products in the combination with the fewest number of reference products containing the reference product need to be considered when mining keywords.

[0145] For example, for a certain reference product F, if the reference products of products C, D, and E all contain F, then since the number of reference products of C and D is the smallest, reference product F only needs to be considered when keyword mining is performed on products C and D, that is, it is included as part of the keyword mining process.

[0146] This embodiment constructs a keyword uniqueness assurance system centered on "intelligent isolation and allocation of reference product resource pools." Through rigorous three-level diagnostic logic (coverage diagnosis, sufficiency diagnosis, and cluster diagnosis), it dynamically allocates suitable, and as non-overlapping as possible, reference product data sources to each product. Its core idea is to solve the problem from the "input end" rather than the "output end," naturally guaranteeing the uniqueness of the "finished product" (keywords) by controlling the exclusivity of the mined "raw materials" (reference products).

[0147] Specifically, the method for determining the update method of the keyword mining and processing results of the product is as follows: This embodiment targets the business scenario of the apparel industry, which features numerous styles, rapid iteration, and obvious long-tail characteristics. It aims to establish a keyword mining and update mechanism that is "core-driven and radiating updates." Its core objective is to prioritize a small number of "updated products" with weak reference product data foundations and requiring high-frequency monitoring, ensuring the extreme timeliness of their keywords. Updates of other "non-updated products" with solid data foundations are designed as incidental behaviors triggered by the update events of updated products. That is, updates are only performed when the reliability of the updated product's response to keyword updates for reference products is poor. This achieves absolute resource allocation bias, maintaining the basic health of the overall product keyword database at extremely low cost.

[0148] This method follows a unidirectional, "strong core, weak radiation" process: monitoring and forcibly updating core products → assessing the impact of core updates on radiation products → updating radiation products only when the impact is significant and the evidence is conclusive. System resources are continuously and proactively invested in monitoring and updating "updated products"; while for "non-updated products," the system adopts a passive and cautious approach, only initiating updates for the latter when the updates from "updating merchants" cannot accurately and comprehensively reflect market changes. The entire process ensures that updates begin and end at the core, with radiation updates merely a cautious extension of core updates.

[0149] S41, based on the keyword mining method, determines the proportion of reference products considered during keyword mining for the product, and uses this proportion as a reference ratio. The "reference ratio" refers to the percentage of reference products actually adopted and used for analysis when executing a predetermined keyword mining strategy, out of the theoretically obtainable total number of reference products for a given product. For example, a dress may have 200 visual or category competitors, but due to strategy exclusion, only 150 are mined, resulting in a reference ratio of 75%.

[0150] The reference ratio is a core quantitative indicator for measuring the "data completeness" and "decision confidence" of keyword mining results. A lower ratio means that the sample data upon which the conclusions are based is insufficient, and the keyword library may not be comprehensive enough. Therefore, products with low reference ratios have the highest risk of keyword incompleteness and require the highest priority and most frequent system maintenance.

[0151] This step provides objective and quantifiable screening criteria for the entire update system. It transforms the strategic decision of "which products need to be prioritized for maintenance" into an automated judgment based on "data completeness" scores, ensuring the scientific and consistent allocation of resources and laying the foundation for building a "core-radial" structure.

[0152] The system calculates the reference ratio for each dress in the store. Product F (a collaboration tie-dye dress) has a unique design, but only 60 reference products are available, some of which are used for data mining, resulting in a reference ratio of 60%, less than 70%. Due to weak data foundation, the system determines it should be marked as an "updated product" (core). Product X (the classic little black dress), on the other hand, has 1800 reference products and a reference ratio greater than 70%, thus it is classified as a "non-updated product" (outlier).

[0153] S42 determines, based on the reference ratio, which products require updating of the keyword mining results and designates them as updated products. In this method, "updated products" specifically refers to those products with a reference ratio below a preset threshold (e.g., 70%). The system assigns them the highest update priority; updating them is mandatory and a top priority given the inaccurate and incomplete keyword generation results.

[0154] Given resource constraints, it is neither possible nor necessary to perform deep updates on all products at the same frequency. Defining products with a low reference ratio as "updated products" acknowledges the fragility of their data foundation and the comprehensiveness bias of their keywords. Through continuous, high-frequency monitoring and updates, the system can reliably update the keywords for these products.

[0155] Product F (the tie-dye collaboration) has been officially marked as a core "updated product" by the system. The system's backend monitoring thread will continuously scan for keyword changes in its 36 reference products and be ready to trigger the update process at any time. Meanwhile, Products X and Y are marked as "non-updated products," and the system will not actively perform periodic scans on them.

[0156] S43 determines the update method for keyword mining and processing of the product based on the reference ratio and the updated product data in the combination.

[0157] Furthermore, the updated product is a product whose reference ratio is less than a preset reference ratio threshold.

[0158] Furthermore, based on the reference ratio and the updated product data in the combination, the update method for mining and processing the keywords of the product is determined, specifically including: Case 1: If the product is an updated product, keyword mining and processing will be performed whenever the keywords of the reference products under consideration change; once any keyword change of the core "updated product" (product F) is detected, the system will immediately and unconditionally trigger a complete keyword re-mining process for product F.

[0159] This directly reflects the "strong core" principle. Since the core product data foundation is weak, the comprehensiveness of the keyword database is improved through dynamic updates. Immediate response is the bottom line requirement for maintaining the usability and competitiveness of its keyword database.

[0160] Of the 36 reference products for Product F, 15 updated their keywords (for example, adding trending terms such as "summer oil painting feel" and "new Chinese tie-dye"). The system detected this change and automatically recalculated all keywords for Product F within one hour, generating a new keyword library containing the latest trending terms.

[0161] Scenario 2: If the product is not an updated product, determine whether the proportion of updated products in the combination is greater than a preset updated product proportion threshold. If so, when the proportion of reference products whose keywords have changed among the reference products considered for the product is greater than a preset change proportion threshold, keyword mining processing is performed. If not, proceed to the next step. "Proportion of updated products in the combination" is an evaluation indicator used to quantify whether the keywords of the products in the combination can be reliably and comprehensively updated when the keywords of the reference products change. It represents the reliability and timeliness of the keyword update processing of the products in the combination.

[0162] Specific example (continuing from S43 case 1): The updated product in the combination is F, and there are a total of 3 products. The proportion of the updated products in the combination is 0.33, which is less than 0.5. When the keywords of the reference product change, the timeliness of the update is poor, so it is necessary to proceed to the next step.

[0163] The reference products that both the product and the updated product need to consider when performing keyword mining are used as the product's associated reference products. It is determined whether the proportion of associated reference products among the reference products that need to be considered when performing keyword mining is greater than a preset proportion threshold. If so, keyword mining is performed when the proportion of reference products whose keywords have changed among the reference products considered for the product is greater than a preset change proportion threshold. If not, keyword mining is performed when the proportion of reference products whose keywords have changed among the reference products considered for the product is greater than a second preset change proportion threshold.

[0164] It should be noted that the preset change percentage threshold is greater than the second preset change percentage threshold.

[0165] Additional update decision based on representativeness: The system presets a low "preset change percentage threshold" (e.g., 2.0%). The calculated radiation impact is then compared with this threshold.

[0166] If the proportion of related reference products considered during keyword mining for a product exceeds a threshold (20%), it is considered that the core update event can comprehensively represent the market changes that the product needs to address. In this case, the update of the updated product can be considered to reflect the keyword update requirements when reference products change. For the non-updated product, due to its large reference proportion, its keyword extraction results are sufficiently reliable and comprehensive. Only when the proportion of reference products with changed keywords among the considered reference products exceeds a preset change proportion threshold is it necessary to fully utilize the reference products considered during keyword mining for the product to perform keyword mining.

[0167] If the proportion of related reference products considered during keyword mining for a product is less than the threshold (20%), it is considered that the core update event cannot fully and accurately reflect the keyword update needs when reference products change, i.e., it cannot effectively and comprehensively reflect changes in the market. The update needs when reference products change are not met. In this case, the system needs to further check the total change ratio of the reference product's own pool of reference products. As long as its own change ratio exceeds another lower "second preset change ratio threshold" (e.g., 1%), a separate, complete update (accompanying update) will be triggered.

[0168] This rule strictly adheres to the "weak radiation" principle. Only when the impact of changes to the reference product is sufficiently broad (high impact), and the updated product's keyword mining fails to fully reflect the impact of these changes, will the system initiate keyword mining and updating for that product under fixed conditions. This minimizes intervention in non-updated products, ensuring the long-term stability of the keyword strategy for non-updated products, while simultaneously not overlooking truly significant market changes that have a global impact.

[0169] Specific decision (continuing from S43, scenario 2): For product Y (when performing keyword mining, the proportion of related reference products among the reference products to be considered should not exceed 20%): The representativeness of the updated product is insufficient. The system then checks the total change proportion of product Y's 500 reference products. If it finds that more than 1% (5) of its own reference products have changed, reaching the "passive update threshold," an independent update of product Y is triggered. Otherwise, it is completely ignored.

[0170] This embodiment successfully constructed a keyword update intelligent scheduling system characterized by "core-driven, cautiously radiating" approach. It scientifically defines core products through a "reference ratio" and builds a complete evaluation and decision-making chain around update events for these core products. The system prioritizes resources to ensure the timely update of products with fewer reference products during keyword mining. Updates of non-updating products are only executed as a cautious extension of the core update process when the updated product fails to update its keywords according to the changed reference products—that is, when the updated product's keywords do not accurately reflect the changes in the reference product's keywords.

[0171] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0172] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0173] The above description is merely one or more embodiments of this specification and is not intended to limit this specification. Various modifications and variations can be made to the one or more embodiments of this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of one or more embodiments of this specification should be included within the scope of the claims of this specification.

Claims

1. A keyword mining system based on multi-factor and dual semantic channels, characterized in that, Specifically, it includes: The data reconstruction module categorizes products into different combinations based on their type. Using similar product data and other product data within each combination, it determines a keyword mining strategy for that combination. This strategy yields reference products, which are then converted into a long table structure with "search term-product" granularity. Product brand and type metadata are associated with this long table, and a comprehensive matching score is calculated for each "search term-product" pair based on multiple dimensions. The index building module vectorizes the search terms and product titles in the long table, generating semantic vectors for the search terms and product titles. These vectors, along with the matching scores, are stored in the retrieval engine. The intelligent recommendation module responds to user-input queries by performing parallel matching retrieval based on text relevance and similarity retrieval based on dual-channel semantic vectors. After deduplication, the retrieved results are reordered based on the matching scores, resulting in a final keyword recommendation list. The module also uses the keyword mining processing method for the products within the combination to determine the method for updating the keyword mining results.

2. The keyword mining system based on multi-factor and dual semantic channels as described in claim 1, characterized in that, The dimensions include product performance factors, product location factors, search term ranking factors, and search term volume factors.

3. A keyword mining method based on multi-factor and dual-semantic channel, applied to the keyword mining system based on multi-factor and dual-semantic channel as described in any one of claims 1-2, specifically comprising: Based on product type, products are divided into different combinations. Using similar product data and product data within each combination, a keyword mining strategy is determined. If the mining strategy is an optimized mining strategy, the process proceeds to the next step. Reference products are determined based on the similarity analysis between the product image and similar products. Based on the reference product data and the similarity between the reference products and the reference products in the combination, products for which keyword mining is performed using all reference products are identified and designated as target products. A keyword mining method is determined based on the composition data of the target product in the combination and the similarity between the target product and its reference products. The keyword mining method is then used to mine keywords for the products, yielding keyword mining results. Finally, a method for updating the keyword mining results is determined using the keyword mining method for the products in the combination.

4. The keyword mining method based on multi-factor and dual semantic channels as described in claim 3, characterized in that, Dividing products into different groups, specifically including grouping products of the same type into the same group.

5. The keyword mining method based on multi-factor and dual semantic channels as described in claim 3, characterized in that, Similar products to the product are products of the same type as the product.

6. The keyword mining method based on multi-factor and dual semantic channels as described in claim 3, characterized in that, The method for determining the keyword mining strategy for the products in the combination is as follows: based on the similar product data of the product, determine the number of similar products of the product; based on the product data in the combination, determine the number of products in the combination; based on the number of similar products of the product in the combination and the number of products in the combination, determine the keyword mining strategy for the products in the combination.

7. The keyword mining method based on multi-factor and dual semantic channels as described in claim 6, characterized in that, If the number of similar products is less than a preset threshold for the number of similar products, the keyword mining strategy for the products in the combination is determined to be to use all similar products for keyword mining.

8. The keyword mining method based on multi-factor and dual semantic channels as described in claim 3, characterized in that, The method for determining the keyword mining and processing method for the product is as follows: based on the composition data of the target product in the combination, determine the proportion of the target product in the combination; based on the similarity between the product and the reference product of the target product, determine the number of overlaps between the product and the reference product of the target product; based on the proportion of the target product in the combination and the number of overlaps between the product and the reference product of the target product, determine the keyword mining and processing method for the product.

9. The keyword mining method based on multi-factor and dual semantic channels as described in claim 3, characterized in that, The method for determining the update method of the keyword mining and processing results of the product is as follows: based on the keyword mining and processing method, determine the proportion of the reference products considered when the product is mining and processing the keyword, and use the proportion of the reference products considered when the product is mining and processing the keyword as the reference proportion. Based on the reference ratio, determine the products whose keyword mining results need to be updated and use them as the updated products; based on the reference ratio and the updated product data in the combination, determine the update method for the keyword mining process of the products.

10. The keyword mining method based on multi-factor and dual semantic channels as described in claim 9, characterized in that, The updated products are those whose reference ratio is less than a preset reference ratio threshold.

Citation Information

Patent Citations

  • Commodity keyword determination method and device based on big data, medium and equipment

    CN115203395A

  • Commodity picture label generation method and device, equipment, medium and product

    CN116796027A

  • Method and system for optimizing commodity titles and computing equipment

    CN119669444A

  • SEO automatic setting system and method in station building CMS

    CN120780938A

  • Intelligent SEO keyword detection system constructed based on knowledge graph

    CN121212159A