Synonym expansion search method and device, equipment and medium
By extracting keywords in store grouping on the e-commerce platform and combining industry feature tags for synonyms, the problem of inaccurate synonyms expansion in the existing technology is solved, and a more accurate and suitable synonyms for industry characteristics is achieved, which significantly improves the accuracy and user experience of search results.
Patent Information
- Application Number
- CN202510289412.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
The product search method of existing e-commerce platforms is based on a global synonym database, resulting in inaccurate synonyms expansion, affecting the accuracy of search results and user experience.
By extracting the product information in the store group and the keywords in the user history query statement, keywords with high correlation with product information are selected to form a keyword collection for each store group. Then synonyms are expanded based on the industry feature labels of each store group, a synonym word set related to the industry characteristics is generated, and synonyms whose confidence scores are lower than the preset threshold are filtered out through the confidence scoring model.
The generated synonyms word set is more accurate and industry-friendly, significantly improving the accuracy and recall of search results, and providing users with a more search experience that meets their needs.
Smart Images

Figure CN120218055A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of e-commerce technology, and in particular, to a synonym expansion search method, its device, equipment, and medium. Background Art
[0002] Product search, as one of the important functions of an e-commerce platform, is an important means for consumers to shop. However, when a user searches for products in a store, there may be a description difference between the input query statement (query) and the retrieval fields such as the product title and description in the store, resulting in a low recall rate. To improve the accuracy of search results and the user experience, it is necessary to perform synonym expansion on the user query statement to eliminate the semantic difference between the query statement and the product description.
[0003] However, the product search in the prior art is based on a global synonym library, that is, all stores share the same synonym library. Since the products sold in each store of an e-commerce independent site vary greatly, and there are significant differences in the same expansion requirements for the same keyword in different industries, if all stores use the same synonym library, it may lead to inaccurate synonym expansion, thereby affecting the search results.
[0004] Therefore, there is an urgent need for a synonym expansion search method that can combine industry characteristics to improve the accuracy of search results and the user experience in an e-commerce platform. Summary of the Invention
[0005] The primary objective of this application is to solve at least one of the above problems and provide a synonym expansion search method, its device, equipment, and medium.
[0006] To meet the various objectives of this application, the following technical solutions are adopted in this application:
[0007] A synonym expansion search method provided to meet one of the objectives of this application includes the following steps:
[0008] Extract the product information in the store cluster and the keywords in the user's historical query statements corresponding to the queries executed in the store cluster, and screen out the keywords with relatively high relevance to the product information to form a keyword set for each store cluster;
[0009] Based on the industry feature tags of each store cluster, perform synonym expansion on the corresponding keyword set to obtain a synonym set containing multiple synonym pairs, where the synonym pair includes a keyword and its corresponding synonym;
[0010] Input the synonym sets of each store cluster into a confidence score model to obtain the confidence scores of the multiple synonym pairs, and filter out the synonym pairs with confidence scores lower than a preset confidence threshold from the synonym set;
[0011] Based on the target query statement input by the user and the synonym set corresponding to the store clustering, recall some products and feedback them to the user.
[0012] On the other hand, a synonym expansion search device provided to meet one of the purposes of the present application includes:
[0013] A keyword set construction module, configured to extract keywords in the product information in the store clustering and the user's historical query statements corresponding to the query executed in the store clustering, screen out the keywords with relatively high relevance to the product information, and construct a keyword set for each store clustering;
[0014] A synonym set construction module, configured to perform synonym expansion on the corresponding keyword set based on the industry feature tags of each store clustering to obtain a synonym set including multiple synonym pairs, where the synonym pairs include keywords and their corresponding synonyms;
[0015] A synonym pair filtering module, configured to input the synonym sets of each store clustering into a confidence score model to obtain the confidence scores of the multiple synonym pairs, and filter out the synonym pairs with confidence scores lower than a preset confidence threshold from the synonym sets;
[0016] A product recall module, configured to recall some products and feedback them to the user based on the target query statement input by the user and the synonym set corresponding to the store clustering.
[0017] On the other hand, a computer device provided to meet one of the purposes of the present application includes a central processing unit and a memory, and the central processing unit is configured to call and run a computer program stored in the memory to execute the steps of the synonym expansion search method described in the present application.
[0018] On the other hand, a computer-readable storage medium provided to meet another purpose of the present application stores a computer program implemented according to the synonym expansion search method in the form of computer-readable instructions. When the computer program is called and run by the computer, it executes the steps included in the corresponding method.
[0019] The technical solution of the present application has many advantages, including but not limited to the following aspects:
[0020] By combining industry feature tags for store clustering, this application can generate a synonym set highly relevant to industry features, ensuring the accuracy and industry adaptability of synonym expansion. Specifically, this application extracts keywords from product information in store clustering and users' historical query statements, filters out the keywords with relatively high relevance to product information, and forms a keyword set for each store cluster. On this basis, it uses industry feature tags to perform targeted synonym expansion on the keyword set to generate synonym pairs highly relevant to industry characteristics. This technical feature not only considers the semantics of the keywords themselves but also combines the industry features of different stores, enabling the generated synonyms to better cover users' search intentions in specific industry scenarios, thus significantly improving the accuracy and recall rate of search results and providing users with a search experience more in line with their needs.
[0021] In addition, this application screens the generated synonym pairs through a confidence score model, filtering out the synonym pairs with a confidence score lower than the preset threshold, thereby ensuring the quality and reliability of the final synonym set used for search. This screening mechanism effectively avoids the noise problem caused by synonym expansion, further optimizing the accuracy of search results and improving the user search experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and / or additional aspects and advantages of this application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:
[0023] Figure 1 is a schematic flowchart of a typical embodiment of the synonym expansion search method of this application;
[0024] Figure 2 is a schematic flowchart of the process of dividing stores into corresponding store clusters in the embodiment of this application;
[0025] Figure 3 is a schematic flowchart of the process of forming a keyword set in the embodiment of this application;
[0026] Figure 4 is a schematic flowchart of the process of calling a large language model to form a synonym set in the embodiment of this application;
[0027] Figure 5 is a schematic flowchart of the iterative training process of the confidence score model in the embodiment of this application;
[0028] Figure 6 is a schematic flowchart of the process of product recall in the embodiment of this application;
[0029] Figure 7 is a schematic flowchart of the process of synonym set expansion in the embodiment of this application;
[0030] Figure 8 The principle block diagram of the synonym expansion search device of the present application;
[0031] Figure 9 The structural schematic diagram of a computer device adopted by the present application. Detailed implementation manners
[0032] A synonym expansion search method of the present application can be programmed as a computer program product and deployed to run in a server. For example, in the exemplary application scenario of the present application, it can be deployed in the server of an e-commerce platform. Among them, the e-commerce platform can be an e-commerce platform that provides open independent station services. An independent station refers to a new official website (website) built on a SaaS technology platform, with an independent domain name, private content, data, and rights, having independent operation sovereignty and operation entity responsibility, supported by social cloud computing capabilities, and capable of independently and freely docking third-party software tools, publicity and promotion media and channels.
[0033] In the business scenario of an e-commerce platform, accurate product search is a key link to improve user experience and operation efficiency. Since the types of products sold in each store on the e-commerce platform are different, and there are significant differences in the same expansion requirements for the same keyword in different industries. For example, the synonym of "apple" in the 3C industry may be "mobile phone", while the synonym of "apple" in the fresh food industry may be "fruit". At this time, when searching for "apple" in a 3C industry store, the synonym expansion based on the global synonym library may wrongly expand "apple" to "fruit", resulting in no recalled products, thereby reducing the accuracy of search results and user experience. Based on this, it is necessary to construct a synonym library related to industry characteristics for each industry's corresponding store to match search results.
[0034] In a typical embodiment of the present application, by extracting the product information in the store cluster and the keywords in the user's historical query statements corresponding to the store cluster execution query, the keywords with relatively high relevance to the product information are screened out and used as the basic words for synonym expansion of the corresponding store cluster. The store cluster is a group divided according to different industry characteristics to distinguish different stores, that is, each store cluster includes multiple stores with the same industry characteristics. Industry characteristics such as electronic products, pet supplies, clothing and accessories, etc. In one embodiment, the products in each store are classified into industry categories, the number of products in each industry category of all the listed products in each store is counted, the industry category corresponding to the highest number of products is determined, and the store is divided into the store cluster corresponding to this industry category.
[0035] In a typical embodiment of the present application, after screening out keywords to form a keyword set for each store cluster, based on the industry feature tags of each store cluster, the keyword set is expanded with synonyms. The expanded synonyms have a strong correlation with the industry feature tags. For example, if the industry feature tag corresponding to the store cluster is "electronic products" and the keyword set contains "Apple", the expanded synonyms may include "iPhone", "iPad", "MacBook", "Apple mobile phone", "Apple tablet", "Apple computer", etc. These expanded synonyms can more comprehensively cover the electronic product-related vocabulary involved in the store cluster, thereby improving the accuracy and recall rate of the search.
[0036] A synonym expansion search method of the present application can be programmed as a computer program product and deployed to run on a server. For example, in an exemplary application scenario of the present application, it can be deployed and implemented in the server of an e-commerce customer service platform.
[0037] Please refer to Figure 1 , in a typical embodiment, the synonym expansion search method of the present application includes the following steps:
[0038] Step S5100: Extract the product information in the store cluster and the keywords in the user's historical query statement corresponding to the query executed in the store cluster, and screen out the keywords with relatively high relevance to the product information to form a keyword set for each store cluster;
[0039] In an e-commerce platform or an independent website, each store contains a large amount of product information, which includes but is not limited to text data such as product titles and product descriptions. When a user enters a store and triggers the corresponding search control, the user inputs a query statement to search for the required products in the store.
[0040] Before the implementation of the typical embodiments of this application, it is necessary to classify each store in the e-commerce platform or independent website according to industry characteristics and divide each store into the corresponding store clusters. For example, stores mainly selling different categories of products such as electronic products, food, and books are respectively divided into the corresponding store clusters. For the specific division method, please refer to the following specific implementation manners and will not be elaborated here. In this step, the product information in the store cluster and the keywords in the user's historical query statements corresponding to the store cluster query are extracted. Specifically, the product information of all the stores listed in the store cluster is segmented to obtain the keywords corresponding to the product information. When merchants list products, these product information are input through a visualization component and then stored in the corresponding database of the e-commerce platform. At the same time, the user's historical query statements corresponding to the store cluster query are segmented. The user's historical query statements can be obtained from the user behavior logs of the e-commerce platform. In one embodiment, by setting a preset time period, the query statements within this time period are retrieved from the user's historical behavior logs as the user's historical query statements. The extraction order of the product information and the keywords of the user's historical query statements does not affect the embodiment of the innovative spirit of this application.
[0041] For the extracted keywords, the keywords with relatively high relevance to the product information are screened out. In one embodiment, the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm is used to calculate the relevance scores of each keyword, the keywords are sorted according to the relevance scores, and a preset relevance score threshold is used to screen the keywords to screen out multiple keywords with relatively high relevance scores to form a keyword set. In another embodiment, by setting a preset number, the preset number of keywords with the highest relevance scores are obtained to form a keyword set.
[0042] Step S5200: Based on the industry characteristic tags of each store cluster, perform synonym expansion on the corresponding keyword set to obtain a synonym set containing multiple synonym pairs, where the synonym pair includes a keyword and its corresponding synonym;
[0043] Before the implementation of the typical embodiments of this application, multiple industry characteristic tags are preset according to the characteristics of different stores in the e-commerce platform. The industry characteristic tag is the core attribute of the store cluster and is used to reflect the main business field of the store. In one embodiment, the industry characteristic tag associated with the store cluster is the industry characteristic tag corresponding to the stores in the store cluster, where the industry characteristic tag is determined by the first-level category of the products in the store. For the convenience of understanding, the industry characteristic tag here is set corresponding to the first-level category. For example, if the number of products belonging to the first-level category "electronic products" in a store is the largest, the store is divided into the store cluster with the industry characteristic tag of "electronic products".
[0044] In this step, based on the industry feature tags of each store cluster, the keyword set of the store cluster is expanded with synonyms, and the expanded synonyms have a high correlation with the industry feature tags. In one embodiment, a large language model is called to generate synonyms related to the industry feature tags based on the industry feature tags of each store cluster, and a preset number of keywords are determined as the result of this synonym expansion. For the specific expansion steps, please refer to the following detailed implementation manners and will not be elaborated here.
[0045] In another embodiment, the open-source synonym thesaurus can be directly crawled through web crawler technology to obtain the synonyms corresponding to the keyword set. The open-source synonym thesaurus contains a large number of synonym pairs that have been manually annotated and verified. After obtaining the synonyms, a set of synonyms with relatively high relevance is selected according to the preset number to form a synonym set.
[0046] In another embodiment, by obtaining the historical search terms of the user in the word query session, based on the semantic similarity of the historical search terms, a preset number of search term pairs with relatively high similarity are selected to form a synonym set. For the specific steps, please refer to the following detailed implementation manners and will not be elaborated here.
[0047] In another embodiment, the synonym sets obtained by any one or more of the above three expansion methods can be merged as the final synonym set. During the merging process, operations such as deduplication of synonyms need to be performed to ensure the quality of the final synonym set.
[0048] Through the above various synonym expansion methods, this application can generate a high-quality and highly relevant synonym set according to the industry characteristics of different store clusters.
[0049] Step S5300: Input the synonym sets of each store cluster into the confidence score model to obtain the confidence scores of the multiple synonym pairs, and filter out the synonym pairs with confidence scores lower than the preset confidence threshold from the synonym sets;
[0050] Due to the complexity and diversity of language, the automatically generated synonyms may have noise, that is, some of the expanded synonyms may have a weak semantic association with the original keywords or even be completely irrelevant. Therefore, in order to ensure the quality of synonym expansion, it is necessary to screen the generated synonym pairs and retain the high-quality synonyms.
[0051] The confidence score model is a model built based on machine learning or deep learning techniques. Its main function is to evaluate the input synonym pairs and output a confidence score. This confidence score reflects the model's confidence in the accuracy of the synonym pairs, and it is a value between 0 and 1, where 1 represents complete confidence and 0 represents complete disbelief. The core task of this confidence score model is to evaluate the quality of the synonym pairs, that is, to judge whether the extended synonyms are highly semantically related to the original keywords.
[0052] The confidence score model can be implemented in various ways. A common method is based on machine learning algorithms such as logistic regression, support vector machines, or deep learning models (such as neural networks). These models are trained with a large amount of labeled data to learn how to distinguish high-quality synonym pairs from low-quality ones. Before training the confidence score model, a labeled dataset needs to be prepared. This dataset contains a large number of synonym pairs, and each pair is labeled as "is a synonym" or "is not a synonym". For example, for the keyword "mobile phone", "smartphone" is labeled as a synonym, while "fruit" is labeled as a non-synonym. These labeled data can be obtained through manual annotation or generated semi-automatically. The quality of the labeled data directly affects the performance of the model, so it is necessary to ensure the accuracy and diversity of the labeled data.
[0053] After the confidence score model is trained, the synonym sets generated by clustering each store are input into the confidence score model. The confidence score model will calculate a confidence score for each synonym pair according to the learned features and rules. This score is a value between 0 and 1, indicating the probability that the synonym pair is a high-quality synonym. To screen out high-quality synonym pairs, a preset confidence threshold is set. The confidence threshold can be adjusted according to actual needs. For example, if you want to screen out very high-quality synonym pairs, the threshold can be set higher (such as 0.8 or 0.9); if you want to retain more synonym pairs to improve the recall rate, the threshold can be set lower (such as 0.5 or 0.6). Multiple experiments need to be conducted based on the test data and business requirements to determine the optimal confidence threshold. Through the above process, the high-quality screening of the synonym sets is achieved, ensuring the semantic relevance and accuracy of the extended synonyms.
[0054] Step S5400: Based on the target query statement input by the user and the synonym set corresponding to the store clustering, recall some products and feedback them to the user.
[0055] In one embodiment, when a user enters a query statement in a certain store on an e-commerce platform, the query statement is first segmented to extract the keywords therein. Based on the extracted keywords, synonyms related to these keywords are matched from the synonym set of the corresponding store cluster of the store. Then, a preset number of highly relevant synonyms are obtained, and they form an extended keyword collection with the keywords of the target query statement. After determining the extended keyword collection, an interaction is performed with the product database through the database query interface to execute a product recall operation for multi-keyword matching. Specifically, first, a composite query condition is constructed based on the extended keyword set. For example, for the extended keywords "laptop", "gaming laptop", and "high performance", the query condition is "the title or description contains 'laptop' or 'gaming laptop' or 'high performance'". In one embodiment, the inverted index technology is adopted to pre-establish an inverted index table for product titles, descriptions, and attribute tags, map the keywords to the list of product IDs containing the keywords, and quickly obtain the candidate product set. After recalling the product set, the products are sorted according to the preset relevance sorting rules to filter out the high-quality product recommendation list that best meets the user's needs. For the specific recall steps, please refer to the following specific embodiments, which will not be elaborated here.
[0056] In another embodiment, the extended keyword set is semantically matched with the product information in the product database, and the semantic similarity algorithm (such as cosine similarity) is used to calculate the semantic relevance between the product title, description, etc. and the extended keyword set, and the products with higher relevance are recalled. During the recall process, the historical behavior data of the user (such as click and purchase records) can be combined to perform personalized sorting on the products, and the products with a higher matching degree with the user's preferences are recommended first. Finally, the sorted product list is fed back to the user to ensure that the user can quickly find the product that best meets their needs.
[0057] Based on the keywords of the target query statement, the corresponding synonyms are matched in the synonym set of the corresponding store cluster, and the obtained synonyms are related to the industry characteristics of the store cluster. Using this extended keyword collection to recall products can significantly improve the accuracy and recall rate of search results.
[0058] According to the typical embodiments of the present application, it can be known that the technical solution of the present application has many advantages, including but not limited to the following aspects:
[0059] By combining industry feature tags for store clustering, this application can generate a synonym set highly relevant to industry features, ensuring the accuracy and industry adaptability of synonym expansion. Specifically, this application extracts keywords from product information in store clustering and user historical query statements, filters out keywords with high relevance to product information, and constructs a keyword set for each store cluster. On this basis, targeted synonym expansion is performed on the keyword set using industry feature tags to generate synonym pairs highly relevant to industry characteristics. This technical feature not only considers the semantics of the keywords themselves but also combines the industry characteristics of different stores, enabling the generated synonyms to better cover users' search intentions in specific industry scenarios, thus significantly improving the accuracy and recall rate of search results and providing users with a search experience more in line with their needs.
[0060] In addition, this application screens the generated synonym pairs through a confidence score model, filtering out synonym pairs with a confidence score lower than a preset threshold, thereby ensuring the quality and reliability of the final synonym set used for search. This screening mechanism effectively avoids the noise problem caused by synonym expansion, further optimizing the accuracy of search results and improving the user search experience.
[0061] In a further embodiment, please refer to Figure 2 , before extracting keywords from product information in store clustering and user historical query statements corresponding to queries executed in the store cluster, filtering out keywords with high relevance to the product information, and constructing a keyword set for each store cluster, the following steps are included:
[0062] Step S6100: Based on a preset industry category, call a product classification model to classify each product in the store into an industry category;
[0063] Since the product types on the e-commerce platform involve multiple industries, multiple industry categories are preset according to the characteristics of different stores. Among them, the characteristics of the store are determined by the types of products sold. For example, if the type of product with the largest quantity sold in the store is "electronic products", then "electronic products" is used as the industry category of the store. In one embodiment, a product classification model is trained to classify products in the store. This product classification model is based on a deep learning algorithm. The product information is input into the product classification model, and the model can automatically learn information such as the text description and title of the product, thereby realizing the automatic classification of the product. Specifically, through a large amount of labeled data, including the text description of the product, the product title, and the corresponding industry category, the model learns the mapping relationship between the product text description, the product title, and the industry category. Through the trained product classification model, all products in the store can be accurately classified into the preset industry categories.
[0064] In another embodiment, some products may belong to multiple industry categories simultaneously, or the text descriptions or product titles of some products are not clear enough, making it difficult for the product classification model to accurately classify them. Based on this problem, a confidence threshold is set. When the classification confidence of the product classification model for a certain product is lower than this confidence threshold, the product is marked as "pending review" and further reviewed and classified manually.
[0065] Step S6200: Count the number of products in each store under each industry category, and calculate the proportion of the number of products in each industry category in that store.
[0066] After completing the product classification, count the number of products in each store under each industry category, and calculate the proportion of the number of products in each industry category in that store. First, traverse all the products in each store, and classify the products into the corresponding industry categories according to the results of the product classification model. For example, for store A, count the number of products in each industry category such as "electronic products", "household items", and "clothing and shoes". Then calculate the proportion of the number of products in each industry category in that store. Specifically, for each industry category, divide the number of its products by the total number of products in that store to obtain the proportion of the number of products in that industry category. For example, if store A has a total of 1000 products and 600 products are in the "electronic products" category, then the proportion of the number of products in the "electronic products" category is 60%. In this way, the product distribution of each store under each industry category can be clearly obtained, providing data support for subsequent store clustering.
[0067] Step S6300: For each store, determine the industry category with the highest number of products and a corresponding product proportion exceeding the preset proportion threshold, and divide the store into the store cluster corresponding to that industry category.
[0068] After counting and calculating the proportion of the number of products in each industry category of each store in the previous step, the stores are grouped according to a preset proportion threshold. First, for each store, find the industry category with the highest number of products, and check whether the proportion of the number of products in this industry category exceeds the preset proportion threshold. For example, if the preset proportion threshold is 50%, and the proportion of the number of products in the "electronic products" category of store A is 60%, then the condition is met, and the store is further classified into the store group corresponding to this industry category. For example, store A is classified into the "electronic products" store group. If the proportion of the industry category with the highest number of products in a certain store does not exceed the preset threshold, it can be judged that the business distribution of this store is relatively wide and it is not suitable to be classified into a single industry category group. At this time, the store can be classified into a "comprehensive category" store group, or processed according to other business rules. In this way, stores can be efficiently classified into the corresponding industry category groups, providing a clear industry classification basis for subsequent keyword extraction and synonym expansion.
[0069] In one embodiment, the setting of the preset proportion threshold takes multiple factors into comprehensive consideration. For example, for a vertical e-commerce platform focusing on a specific field, the threshold can be set relatively high, such as 70% or 80%, to ensure the high accuracy of store grouping. For a comprehensive e-commerce platform, since the business scope of the store may be relatively wide, the threshold can be set relatively low, such as 40%, to improve the flexibility and coverage of store grouping. In addition, the platform can also dynamically adjust the preset proportion threshold according to information such as the scale of the store to adapt to the business characteristics of different stores.
[0070] In this embodiment, by calling a product classification model to classify each product in the store into an industry category, and then determining the industry category with the highest number of products in each store and a proportion exceeding the preset proportion threshold, the store can be accurately classified into the corresponding industry category group, ensuring the rationality and effectiveness of store grouping. This method of store grouping based on industry categories provides a clear industry classification basis for subsequent synonym expansion.
[0071] In a further embodiment, please refer to Figure 3 Extract the product information in the store group and the keywords in the user's historical query statements corresponding to the queries executed in the store group, and screen out the keywords with relatively high relevance to the product information to form a keyword set for each store group, including the following steps:
[0072] Step S5110: Obtain text data, where the text data includes the product titles, product description texts of the products in the store group, and the user's historical query statements corresponding to the queries executed in this store.
[0073] The product title and product description text are stored in the database of the e-commerce platform, and relevant fields are obtained through the database query interface. For the user's historical query statements, the search records of the user in a specific store group are extracted from the user behavior log. The search record includes the query statement input by the user, and the user's historical query statements reflect the user's search intent and demand for products. To ensure the integrity and accuracy of the data, preprocessing is performed on the obtained text data, such as removing duplicate data, filtering invalid characters (such as special symbols, spaces, etc.), and handling missing values. For example, if the description text of a product is empty, the terminal device will mark it as "missing" and attempt to supplement relevant information from other data sources (such as the OCR recognition result of the product picture). In addition, duplicate removal and screening are also performed on the user's historical query statements to remove duplicate query statements and invalid queries (such as blank queries or irrelevant queries). Implementing this step can efficiently obtain and organize high-quality text data, providing reliable data support for subsequent keyword extraction and screening.
[0074] Step S5120: Perform word segmentation on the text data to obtain multiple target word segments;
[0075] Perform word segmentation on the text data obtained in the previous step to split it into multiple target word segments. Word segmentation can split a continuous text sequence into meaningful lexical units. Use word segmentation tools (such as Jieba, HanLP, etc.) to perform word segmentation on the product title, product description text, and the user's historical query statements. For example, for the product title "New High-Performance Smart Phone", the word segmentation result is "New", "High-Performance", "Smart", "Phone". For the user query statement "Laptop suitable for playing games", the word segmentation result is "Suitable", "For playing games", "Of", "Laptop". During the word segmentation process, the text is segmented according to the preset dictionary and rules to ensure the accuracy of word segmentation. For example, for proper nouns or industry terms (such as "Laptop"), they are segmented as a whole instead of being split into "Notes" and "Computer". In addition, the terminal device will also perform post-processing on the word segmentation result, such as removing stop words (such as meaningless words like "Of", "Is", etc.). In this way, the text data can be efficiently split into multiple target word segments, providing basic data support for subsequent keyword screening and synonym expansion.
[0076] Step S5130: Use an information retrieval algorithm to calculate the relevance scores of the target word segments with the product titles and product description texts of the products in the store, and screen out a preset number of target word segments with higher relevance scores to form the keyword sets for each store group.
[0077] Calculate the relevance scores of each target token with the product titles and product description texts of the products in the store to filter out keywords that are highly relevant to the product information. Use an information retrieval algorithm to calculate the relevance scores. Common information retrieval algorithms include TF-IDF (Term Frequency-Inverse Document Frequency), BM25, cosine similarity, etc. For example, for the target token "smartphone", calculate its relevance scores with all product titles and description texts. The higher the score, the stronger the relevance of the token to the product information. Taking the TF-IDF algorithm as an example, the TF-IDF algorithm is a statistics-based method used to evaluate the importance of a word for a document in a document collection. The TF-IDF value increases proportionally with the frequency of the word in the document, but at the same time decreases inversely with the frequency of the word in the corpus. That is, TF-IDF tends to filter out common words and retain important words. In the scenario of the embodiments of the present application, assume that a target token "iPhone" appears frequently in a certain product title but has a low frequency in the entire product database. Then the TF-IDF value of the word "iPhone" will be high.
[0078] When calculating the relevance scores, traverse each target token and match it with the titles and description texts of all products in the store. To filter out target tokens with higher relevance, set a preset quantity or relevance threshold. For example, it can be set to only retain the top 10 target tokens with the highest relevance scores, or set a relevance threshold, and tokens exceeding this threshold are retained.
[0079] In addition to the TF-IDF algorithm, the BM25 algorithm can also be used. BM25 is a ranking function based on a probabilistic model. This algorithm is based on multiple factors such as the frequency of words, the length of the document, and the distribution of words, and can more accurately evaluate the relevance of words to the document. The information retrieval algorithm used does not affect the manifestation of the creative spirit of the present application. In this way, keywords that are highly relevant to the product information can be efficiently extracted, providing basic data support for subsequent synonym expansion and search.
[0080] In this embodiment, keywords that are highly relevant to the product are filtered out to ensure the accuracy and practicality of the keyword set, providing high-quality basic data support for subsequent synonym expansion and search, thereby significantly improving the accuracy of search results and the user experience.
[0081] In a further embodiment, please refer to Figure 4 , based on the industry feature tags of each store cluster, perform synonym expansion on the corresponding keyword set to obtain a synonym set containing multiple synonym pairs. The synonym pairs include keywords and their corresponding synonyms, and the steps are as follows:
[0082] Step S5210: Input the keywords in the keyword set and the industry feature tags of the corresponding store clusters into the industry synonym generation template to obtain the corresponding industry synonym generation instructions;
[0083] After the keyword set is screened in the above steps, the keywords are expanded with synonyms according to the industry feature tags of the store clusters. First, input the keywords in the keyword set and the industry feature tags of the corresponding store clusters into the industry synonym generation template. The industry synonym generation template is a predefined structured text framework that contains a fixed instruction format and dynamic insertion points for generating industry synonym generation instructions. For example, for the "electronic products" store cluster, the industry feature tags may include "intelligent devices", "high performance", "portability", etc. Input the keyword "smartphone" and the industry feature tag "intelligent devices" into the template to generate the corresponding industry synonym generation instruction, such as "Generate industry synonyms related to'smartphone' applicable to the 'intelligent devices' field". The industry synonym generation template is pre-designed by industry experts based on industry knowledge and semantic analysis, which can ensure that the generated synonyms are highly relevant to the industry characteristics.
[0084] When generating the industry synonym generation instructions, the industry synonym generation template considers multiple factors. In one embodiment, a suitable synonym expansion strategy is selected according to the industry feature tags. For example, in the "electronic products" field, synonyms related to technology are mainly expanded, such as "smartphone", "mobile phone", "5G phone", etc.; while in the "clothing" field, synonyms related to styles or functions are emphasized, such as "casual clothing", "sportswear", "formal wear", etc.
[0085] The implementation of this step can combine the keyword set with the industry feature tags of the store clusters to generate accurate synonym expansion instructions. This not only considers the semantic features of the keywords but also combines the industry background, and the generated synonyms are highly relevant to the industry characteristics.
[0086] Step S5220: Input the industry synonym generation instructions into the large language model to control the large language model to generate industry synonyms corresponding to the keywords, forming a synonym set.
[0087] Input the industry synonym generation instruction into the large language model, and utilize the natural language generation ability of the large language model to generate industry synonyms corresponding to the keywords. Based on deep learning technology, the large language model can understand the context information in the instruction and generate synonyms related to the keywords according to the industry feature tags. For example, for the industry synonym generation instruction "Generate industry synonyms related to'smartphone' and applicable to the'smart device' field", the large language model may generate synonyms such as "mobile phone", "5G phone", "smart terminal", etc. The generated synonyms are not only semantically similar to the keywords but also meet the requirements of the industry feature tags. Pair the generated synonyms with the original keywords to form synonym pairs, such as "smartphone - mobile phone", "smartphone - 5G phone", etc. Efficiently construct a synonym set highly relevant to the industry characteristics.
[0088] In one embodiment, the context awareness ability of the large language model can also be utilized to provide richer background information for generating the industry synonym generation instruction. For example, add fragments of product description text or user query statements to the industry synonym generation instruction to help the large language model better understand the semantic background of the keywords. For example, the industry synonym generation instruction can be extended to: "According to the product description 'Professional sports shoes, suitable for running and fitness use', generate synonyms for the keyword'sports shoes' in the'sports equipment' field". Such an instruction with context can guide the large language model to generate more accurate synonyms.
[0089] In another embodiment, in order to further optimize the synonym generation process, an artificial review mechanism can also be introduced. For example, have professionals review the synonyms generated by the large language model, remove irrelevant or low-quality synonyms, and at the same time supplement some synonyms that are important in practice but not generated by the large language model. This combination of human and machine can effectively improve the quality and practicality of the synonym set.
[0090] This step utilizes the powerful semantic generation ability of the large language model to generate high-quality industry synonyms for the keywords. These synonyms are not only semantically related to the original keywords but also conform to the industry characteristics of the store clustering, can effectively expand the search intent of users, and enhance the relevance of search results and user experience.
[0091] In this embodiment, by combining the keyword set with the industry feature tags of the store clustering, generating accurate synonym expansion instructions can ensure that the synonym expansion is highly relevant to the industry characteristics. Utilizing the natural language generation ability of the large language model, synonyms that are semantically similar to the keywords and conform to the industry background can be generated according to the industry feature tags, significantly improving the accuracy and practicality of the synonyms.
[0092] In a further embodiment, please refer to Figure 5, before filtering out the synonym pairs with confidence scores lower than the preset confidence threshold from the synonym set, the following steps are included when inputting the synonym sets of each store group into the confidence score model to obtain the confidence scores of the multiple synonym pairs:
[0093] Step S7100: Train a confidence score model based on a training data set, where the training data set includes initial labeled synonym pairs, and the confidence score model is used to predict the confidence scores of synonym pairs;
[0094] Prepare a training data set, which includes initial labeled synonym pairs. These synonym pairs are generated by industry experts or through automated tools and are manually reviewed and labeled to ensure their accuracy and reliability. For example, in the field of electronic products, "i Phone" and "Apple mobile phone" are high-quality synonyms, while in the food field, "Apple" and "Fruit" are more appropriate synonyms. The annotators need to make accurate annotations according to the specific industry and semantic background. Use the training data set to train a confidence score model, which is constructed based on machine learning algorithms (such as logistic regression, random forest, or deep learning models) and can learn the semantic relationships and industry characteristics between synonym pairs, so as to predict the confidence scores of new synonym pairs. During the training process, the training data set is divided into a training set and a validation set, and the performance of the model is optimized through cross-validation and hyperparameter tuning. For example, use metrics such as accuracy, recall, and F1 score to evaluate the prediction effect of the model, and adjust the model parameters to improve its prediction accuracy.
[0095] Taking a neural network as an example, the training process includes the following stages. First, input the training data set into the neural network, and the network will perform forward propagation according to the annotation results to calculate the confidence scores of each pair of words. Then, through the backpropagation algorithm, according to the difference between the annotation results and the model predictions, adjust the weights and biases of the network to optimize the performance of the model. This process will be repeated multiple times until the performance of the model reaches a predetermined threshold or converges. During the training process, it is also necessary to evaluate the performance of the model through the validation set to avoid overfitting or underfitting.
[0096] In this way, a high-precision confidence score model can be trained to provide reliable support for the subsequent screening of synonym pairs.
[0097] Step S7200: Input the synonym set into the confidence score model to predict the confidence scores of the synonym pairs in the synonym set;
[0098] Input the synonym pairs in the synonym set into the trained confidence score model to predict the confidence score for each pair. The confidence score reflects the semantic relevance and industry suitability of the synonym pair. The higher the score, the higher the confidence in the pair. For example, for the synonym pair "smartphone - mobile phone", the confidence score model may predict a relatively high confidence score (e.g., 0.95), while for "smartphone - tablet", the confidence score model may predict a relatively low confidence score (e.g., 0.65). Based on the output of the confidence score model, generate a confidence score for each synonym pair and store it in association with the pair.
[0099] Step S7300: Based on the entropy method, screen out the synonym pairs with the entropy value of the confidence score distribution higher than the preset entropy threshold to form an uncertain sample set.
[0100] After obtaining the confidence scores of the synonym pairs in the previous step, use the entropy method to analyze the confidence score distribution to screen out the synonym pairs with relatively high uncertainty. The entropy method is a method for measuring the uncertainty of data distribution. The higher the entropy value, the more dispersed the data distribution and the greater the uncertainty. First, calculate the entropy value of the confidence score distribution for each synonym pair. For example, for a certain pair, if its confidence score distribution is relatively dispersed under different models or different data subsets, its entropy value is higher. Then, according to the preset entropy threshold, screen out the synonym pairs with entropy values higher than the threshold to form an uncertain sample set.
[0101] This application introduces the entropy method. By quantifying the distribution uncertainty of the confidence scores, screen out the synonym pairs that need further review to form an uncertain sample set. The formation of the uncertain sample set is for further processing in subsequent embodiments, such as manual review or model iterative training. Please refer to the specific subsequent embodiments.
[0102] Step S7400: Screen out the synonym pairs that are manually marked as synonyms from the uncertain sample set, and merge the screened synonym pairs into the training data set to iteratively train the confidence score model until the prediction accuracy of the confidence score model exceeds the preset accuracy threshold.
[0103] After obtaining the set of uncertain samples in the previous step, these samples are manually reviewed to screen out the word pairs that are actually synonyms. In one embodiment, industry experts or annotation teams review each word pair in the set of uncertain samples to determine whether it is a valid synonym. The synonym word pairs that pass the review are merged into the training dataset to expand the scale and diversity of the training data. The updated training dataset is used to retrain the confidence score model, and the prediction accuracy of the model is verified through cross-validation and performance evaluation (such as accuracy, recall, etc.). If the prediction accuracy of the confidence score model does not reach the preset accuracy threshold (such as 90%), repeat the above steps, continue to screen the uncertain samples, conduct manual review, and iteratively train the model until the prediction accuracy of the model exceeds the preset threshold. Gradually optimize the performance of the confidence score model to ensure that it can accurately predict the confidence scores of synonym word pairs and provide reliable support for subsequent synonym screening and application.
[0104] Specifically, after the manual review is completed, the annotators will mark the review results as "is a synonym" or "is not a synonym". The word pairs marked manually will be screened out to form a set of high-quality training samples. The merged training dataset will contain the initial marked synonym word pairs and the new manually marked word pairs. These data will be used to retrain the confidence score model to improve its performance. During the iterative training process, the confidence score model will learn new features and patterns, so as to better handle the uncertain synonym word pairs. If the prediction accuracy of the confidence score model does not reach the preset threshold, iterative training needs to continue. The steps include further expanding the scale of the set of uncertain samples, increasing the number of samples for manual review, or adjusting the structure and parameters of the confidence score model. For example, if it is found that the confidence score model still has errors when processing certain specific types of word pairs, the word pairs of this type can be specifically added to the set of uncertain samples for more detailed manual review and marking.
[0105] In this embodiment, the set of uncertain samples with a higher entropy value of the confidence score distribution is screened out by using the entropy method, which can identify the synonym word pairs that need further review and avoid the influence of low-quality or incorrect synonym word pairs on the search results. Then, the word pairs in the set of uncertain samples are manually reviewed and merged into the training dataset to continuously optimize the confidence score model and improve its prediction accuracy. This iterative training mechanism ensures the continuous improvement of the confidence score model, enabling it to more accurately evaluate the semantic relevance and industry suitability of synonym word pairs. Finally, by filtering out the synonym word pairs with a confidence score lower than the preset threshold, the high quality of the synonym word set can be ensured.
[0106] In a further embodiment, please refer to Figure 6, recall some products based on the target query statement input by the user and the synonym set of the corresponding store group, including the following steps:
[0107] Step S5410: Based on the keywords of the target query statement, match a preset number of synonyms from the synonym set of the corresponding store group to form an extended keyword set;
[0108] After the user inputs the target query statement, first perform word segmentation on the target query statement to extract its keywords. For example, for the query statement "Notebook computers suitable for playing games", the word segmentation results are "suitable", "playing games", and "notebook computer". Then match the synonyms related to these keywords from the synonym set of the corresponding store group. For example, for the keyword "notebook computer", the synonym set may contain synonyms such as "portable computer", "thin and light laptop", and "gaming laptop". Select the most relevant synonyms according to the preset number (such as the first 5 synonyms) to form an extended keyword set. The preset number can be adjusted according to actual needs. For example, if you want the search results to be more comprehensive, you can set a larger number; if you want the search results to be more accurate, you can set a smaller number. The synonym set of the store group is generated through synonym expansion and confidence evaluation in the previous steps, and contains high-quality synonyms related to the store's products. These synonyms can help better understand the potential intention of the user's query and cover more relevant products.
[0109] In this step, based on the target query statement input by the user and the synonym set of the store group, generate a high-quality extended keyword set. This not only expands the user's search intention but also ensures the relevance and accuracy of the search results through the semantic relevance of synonyms, efficiently expanding the keywords of the user's query statement and providing richer search conditions for subsequent product recall.
[0110] Step S5420: Recall a set of products that match the extended keyword set from the product database;
[0111] Interact with the product database through the database query interface to perform a product recall operation for multi-keyword matching. The product database stores the metadata of all products in the store group, including product titles, detailed descriptions, attribute tags (such as brand, model, price, functional features), and user behavior data (such as click-through rate, purchase rate).
[0112] Match each keyword in the extended keyword set with the product information in the product database. The matching process is based on multiple fields such as product title, description, tags, etc. In one embodiment, the inverted index technology is adopted to pre - establish an inverted index table for product titles, descriptions, and attribute tags, mapping keywords to a list of product IDs that contain the keyword. For example, in the inverted index, "laptop" corresponds to the product ID list [1001, 1002, 1005], and "gaming laptop" corresponds to [1002, 1005, 1008]. By merging these lists and removing duplicates, a candidate product set can be quickly obtained.
[0113] In another embodiment, since the product database may be very large, the recall process needs to be optimized to ensure fast response. In addition to the inverted index, other technologies such as caching mechanisms can also be used. For example, for frequently queried keywords, the matching results can be cached for quick return next time.
[0114] Step S5430: Sort the product set according to the preset relevance sorting rules, and screen out a preset number of products to form a product recommendation list and feedback it to the user.
[0115] Sort the products according to the preset relevance sorting rules to screen out a high - quality product recommendation list that best meets the user's needs. The relevance sorting rules comprehensively consider multiple factors, including the matching degree of the product with the query keyword, the fit degree of the product attributes with the user's needs, user behavior data (such as click - through rate, purchase rate, rating), and the real - time status of the product (such as inventory, promotion information). For example, for the query "laptop suitable for playing games", calculate the matching degree of each product with the keywords "laptop", "gaming laptop", "high performance", and combine the product attributes (such as graphics card model, screen refresh rate, processor performance) for weighted scoring. At the same time, give priority to recommending products with high click - through rate, high purchase rate, or high rating to improve user satisfaction with the recommendation results. In addition, adjust according to the real - time status of the product. For example, give priority to recommending products with inventory, on promotion, or time - limited offers. Finally, sort the products in descending order according to the comprehensive score, screen out a preset number (such as the top 10) of products to form a recommendation list, and display it to the user through a graphical user interface (GUI).
[0116] In one embodiment, in order to further optimize the quality of the product recommendation list, some dynamic adjustment mechanisms can also be introduced. For example, if a certain product is not clicked by the user for a long time in the recommendation list, its sorting weight can be reduced, and it can be replaced by other products that are more likely to attract the user's attention, so that the recalled product set is transformed into a high - quality and highly relevant product recommendation list.
[0117] In this embodiment, an extended keyword set is generated based on the target query statement input by the user and the synonym set corresponding to the store cluster. The synonym set has a high correlation with the corresponding store cluster. Synonyms are matched from the synonym set to ensure that the extended synonyms are highly relevant to the industry characteristics. Then, the extended keyword set obtained in this step is used to recall products, which can significantly improve the accuracy and recall rate of search results.
[0118] In a further embodiment, please refer to Figure 7 , based on the industry characteristic tags of each store cluster, the corresponding keyword set is extended with synonyms to obtain a synonym set containing multiple synonym pairs. After the synonym pairs include the keyword and its corresponding synonym, it includes:
[0119] Step S8100: Obtain the historical search terms of the user in a single query session. The historical search terms include the search terms corresponding to multiple query statements input by the user within the query period of the query session.
[0120] In the search service of an e-commerce platform, the user's search behavior is often not isolated but has a certain continuity and relevance. The user may input multiple query statements in a single query session. When the user conducts a search, all the search query statements of the user in a single query session are recorded. The query session starts when the user logs in or starts searching and ends when the user logs out or has no operation for a long time, that is, this is taken as a query period.
[0121] In one embodiment, the historical search terms of the user in a single query session are obtained through the user behavior log system of the e-commerce platform. The user behavior log system will record information such as the query statement input by the user each time, the search timestamp, and the click behavior. All the query statements within the query session are extracted from the user behavior log system and split into multiple search terms. To ensure the integrity and accuracy of the data, the search terms are de-duplicated and filtered, removing duplicate search terms and invalid search terms (such as blank queries or irrelevant queries). In this way, the historical search terms of the user in a single query session are obtained, providing basic data support for subsequent semantic similarity calculation and synonym expansion.
[0122] Step S8200: Calculate the semantic similarity of the historical search terms and filter out the search term pairs with a semantic similarity higher than a preset similarity threshold.
[0123] Calculate the semantic similarity of the search terms obtained in the previous step to filter out pairs of words with similar semantics. Semantic similarity refers to the degree of proximity in semantics between two text segments, which can be calculated through various natural language processing techniques. Common methods include those based on word embeddings (such as Word2Vec, GloVe), those based on pre-trained language models (such as BERT, RoBERTa), or those based on traditional text similarity metrics (such as cosine similarity, Jaccard similarity), etc. For example, for the search terms "laptop" and "portable computer", calculate the cosine similarity between their word vectors or sentence vectors to obtain a similarity score (such as 0.92). If the similarity score is higher than a preset similarity threshold (such as 0.8), then mark this pair of words as a pair of words with similar semantics. This step identifies pairs of search terms with similar semantics used by the user in the query session, providing data support for subsequent synonym expansion.
[0124] Step S8300: Merge the pair of search terms into the synonym word set.
[0125] Deduplicate and verify the pair of search terms to ensure that they do not duplicate the pairs of words in the existing synonym word set. Merge the new pair of search terms with the existing synonym word set to form an updated synonym word set. In one embodiment, during the merging process, generate a confidence score for each new pair of words and record its source (such as the user query session). To dynamically expand the synonym word set, making it more rich and accurate, providing more comprehensive semantic support for subsequent search expansion and recommendation.
[0126] In this embodiment, by obtaining the historical search terms of the user in a single query session, it is possible to capture the continuity and relevance in the user's search process, comprehensively reflecting the user's search intent and needs. Calculating the semantic similarity of the historical search terms and filtering out pairs of search terms with similar semantics can effectively identify synonymous or near-synonymous expressions used by the user during the search process, further enriching the content of the synonym word set. The implementation of this embodiment not only improves the practicality and coverage of the synonym word set, but also provides richer data support for subsequent searches.
[0127] Please refer to Figure 8, A synonym expansion search device provided to meet one of the purposes of the present application is a functional embodiment of the synonym expansion search method of the present application. On the other hand, a synonym expansion search device provided to meet one of the purposes of the present application includes a keyword set construction module 5100, a synonym set construction module 5200, a synonym pair filtering module 5300, and a product recall module 5400. Among them, the keyword set construction module 5100 is used to extract keywords in the product information in the store cluster and the user's historical query statements corresponding to the queries executed in the store cluster, screen out the keywords with relatively high relevance to the product information, and construct a keyword set for each store cluster; the synonym set construction module 5200 is used to perform synonym expansion on the corresponding keyword set based on the industry characteristic tags of each store cluster to obtain a synonym set containing multiple synonym pairs, and the synonym pair includes a keyword and its corresponding synonym; the synonym pair filtering module 5300 is used to input the synonym sets of each store cluster into a confidence score model to obtain the confidence scores of the multiple synonym pairs, and filter out the synonym pairs with confidence scores lower than a preset confidence threshold from the synonym set; the product recall module 5400 is used to recall some products based on the target query statement input by the user and the synonym set of the corresponding store cluster and feedback them to the user.
[0128] In a further embodiment, before the keyword set construction module 5100, it includes: an industry category classification sub-module, which is used to classify each product in the store into industry categories by calling a product classification model based on a preset industry category; a product quantity statistics sub-module, which is used to count the number of products in each store under each industry category and calculate the proportion of the number of products in each industry category in the store; a store cluster division sub-module, which is used for each store to determine the industry category with the highest number of products and the corresponding product quantity proportion exceeding a preset proportion threshold, and divide the store into the store cluster corresponding to the industry category.
[0129] In a further embodiment, the keyword set construction module 5100 includes: a text data acquisition sub-module, which is used to acquire text data, and the text data includes the product title, product description text of the products in the store cluster, and the user's historical query statements corresponding to the queries executed in the store; a word segmentation sub-module, which is used to perform word segmentation processing on the text data to obtain a plurality of target word segments; a target word segment screening sub-module, which is used to calculate the relevance scores of the target word segments with the product titles and product description texts of the products in the store by using an information retrieval algorithm, and screen out a preset number of target word segments with relatively high relevance scores to construct a keyword set for each store cluster.
[0130] In a further embodiment, the synonym set construction module 5200 includes: an instruction generation sub-module, configured to input the keywords in the keyword set and the industry feature tags of the corresponding store clusters into an industry synonym generation template to obtain corresponding industry synonym generation instructions; an industry synonym generation sub-module, configured to input the industry synonym generation instructions into a large language model to control the large language model to generate industry synonyms corresponding to the keywords, thereby constructing a synonym set.
[0131] In a further embodiment, the synonym pair filtering module 5300 includes: a model training sub-module, configured to train a confidence score model based on a training data set, where the training data set includes initial labeled synonym pairs, and the confidence score model is used to predict the confidence scores of the synonym pairs; a confidence score prediction sub-module, configured to input the synonym set into the confidence score model to predict the confidence scores of the synonym pairs in the synonym set; a sample set construction sub-module, configured to filter out synonym pairs with an entropy value of the confidence score distribution higher than a preset entropy value threshold based on the entropy value method to construct an uncertain sample set; a model iterative training sub-module, configured to filter out synonym pairs manually labeled as synonyms from the uncertain sample set, and merge the filtered synonym pairs into the training data set to iteratively train the confidence score model until the prediction accuracy of the confidence score model exceeds a preset accuracy threshold.
[0132] In a further embodiment, the product recall module 5400 includes: an extended keyword set construction sub-module, configured to match a preset number of synonyms from the synonym set of the corresponding store cluster based on the keywords of the target query statement to construct an extended keyword set; a product set recall sub-module, configured to recall a product set matching the extended keyword set from a product database; a product recommendation list construction sub-module, configured to sort the product set according to a preset relevance ranking rule, and filter out a preset number of products to form a product recommendation list and feedback it to the user.
[0133] In a further embodiment, after the synonym set construction module 5200, it includes: a historical search term acquisition sub-module, configured to acquire the historical search terms of the user in a single query session, where the historical search terms include the search terms corresponding to multiple query statements input by the user during the query period of the query session; a search term pair filtering sub-module, configured to calculate the semantic similarity of the historical search terms and filter out search term pairs with a semantic similarity higher than a preset similarity threshold; a search term pair merging sub-module, configured to merge the search term pairs into the synonym set.
[0134] To solve the above technical problems, an embodiment of the present application also provides a computer device. As Figure 9As shown, it is a schematic internal structure diagram of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected through a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database can store a control information sequence. When the computer-readable instructions are executed by the processor, the processor can implement a synonym expansion search method. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device can store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the synonym expansion search method of the present application. The network interface of the computer device is used to connect and communicate with a terminal. Those skilled in the art can understand that Figure 9 the structure shown is only a block diagram of some structures related to the solution of the present application and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0135] In this embodiment, the processor is used to execute Figure 8 the specific functions of each module and its sub-modules in. The memory stores the program codes and various types of data required to execute the above-mentioned modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. The memory in this embodiment stores the program codes and data required to execute all modules / sub-modules in the synonym expansion search device of the present application. The server can call the program codes and data of the server to execute the functions of all sub-modules.
[0136] The present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the synonym expansion search method according to any embodiment of the present application.
[0137] Those of ordinary skill in the art can understand that to implement all or part of the processes in the methods of the above embodiments of the present application, it can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0138] Those skilled in the art can understand that the various operations, methods, steps, measures, and solutions in the processes discussed in this application can be alternated, changed, combined, or deleted. Further, other steps, measures, and solutions in the various operations, methods, and processes discussed in this application can also be alternated, changed, rearranged, decomposed, combined, or deleted. Further, the steps, measures, and solutions in the prior art that are the same as those in the various operations, methods, and processes open-sourced in this application can also be alternated, changed, rearranged, decomposed, combined, or deleted.
[0139] The above are only some implementation manners of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A synonym expansion search method, characterized in that: The steps include: Extracting product information in the store group and keywords in the user's historical query statements corresponding to the query executed in the store group, screening out keywords with high relevance to the product information, and forming a keyword set for each store group; Based on the industry characteristic labels of each store cluster, synonym expansion is performed on the corresponding keyword set to obtain a synonym word set containing multiple synonym word pairs, wherein the synonym word pairs include keywords and their corresponding synonyms; Inputting the synonym word set of each store group into the confidence scoring model to obtain the confidence scores of the multiple synonym word pairs, and filtering out the synonym word pairs whose confidence scores are lower than a preset confidence threshold from the synonym word set; Based on the target query statement entered by the user and the synonym word set corresponding to the store cluster, some products are recalled and fed back to the user.
2. The synonym expansion search method according to claim 1, characterized in that: Extracting product information in the store group and keywords in the user's historical query statements corresponding to the store group execution query, screening out keywords with high relevance to the product information, and forming a keyword set for each store group, including: Based on the preset industry categories, the product classification model is called to classify the industry categories of each product in the store; Count the number of products in each industry category in each store, and calculate the proportion of products in each industry category in the store; For each store, determine the industry category with the highest number of products and a corresponding product quantity ratio exceeding a preset ratio threshold, and assign the store to the store group corresponding to the industry category.
3. The synonym expansion search method according to claim 1, characterized in that: Extract the product information in the store group and the keywords in the user's historical query statements corresponding to the query executed in the store group, filter out the keywords with high relevance to the product information, and form a keyword set for each store group, including: Acquire text data, where the text data includes product titles and product descriptions of products in the store group and user historical query statements corresponding to queries executed in the store; Performing word segmentation processing on the text data to obtain multiple target word segments; An information retrieval algorithm is used to calculate the relevance scores between the target segment words and the product titles and product description texts of the products in the store, and a preset number of target segment words with higher relevance scores are screened out to form a keyword set for each store cluster.
4. The synonym expansion search method according to claim 1, characterized in that: Based on the industry characteristic labels of each store cluster, the corresponding keyword set is expanded with synonyms to obtain a synonym word set containing multiple synonym word pairs, wherein the synonym word pairs include keywords and their corresponding synonyms, including: Input the keywords in the keyword set and the industry characteristic tags of the corresponding store grouping into the industry synonym generation template to obtain the corresponding industry synonym generation instructions; The industry synonym generation instruction is input into the large language model to control the large language model to generate industry synonyms corresponding to the keywords to form a synonym word set.
5. The synonym expansion search method according to claim 1, characterized in that: Inputting the synonym word set of each store group into the confidence scoring model to obtain the confidence scores of the multiple synonym word pairs, and filtering out the synonym word pairs whose confidence scores are lower than a preset confidence threshold from the synonym word set, including: Training a confidence scoring model based on a training data set, wherein the training data set includes initial labeled synonym word pairs, and the confidence scoring model is used to predict confidence scores of the synonym word pairs; Inputting the synonym word set into a confidence scoring model to predict confidence scores of synonym word pairs in the synonym word set; Based on the entropy method, synonym word pairs with confidence score distribution entropy values higher than the preset entropy threshold are selected to form an uncertain sample set; Synonymous word pairs that are manually marked as synonyms are screened out from the uncertain sample set, and the screened out synonymous word pairs are merged into the training data set to iteratively train the confidence scoring model until the prediction accuracy of the confidence scoring model exceeds a preset accuracy threshold.
6. The synonym expansion search method according to claim 1, characterized in that: Based on the target query sentence entered by the user and the synonym word set corresponding to the store group, some products are recalled and fed back to the user, including: Based on the keywords of the target query sentence, a preset number of synonyms are matched from the synonym word set corresponding to the store cluster to form an extended keyword set; Recalling a set of commodities matching the expanded keyword set from a commodity database; The product set is sorted according to a preset relevance sorting rule, and a preset number of products are screened out to form a product recommendation list to be fed back to the user.
7. The synonym expansion search method according to any one of claims 1 to 6, characterized in that: Based on the industry characteristic labels of each store cluster, the corresponding keyword set is expanded with synonyms to obtain a synonym word set containing multiple synonym word pairs, wherein the synonym word pair includes a keyword and its corresponding synonym, and includes: Acquire historical search terms of a user in a single query session, wherein the historical search terms include search terms corresponding to multiple query statements input by the user within a query period of the query session; Calculating semantic similarity of the historical search terms, and screening out search term pairs whose semantic similarity is higher than a preset similarity threshold; The search word pairs are merged into the synonym word set.
8. A synonym expansion search device, characterized in that: include: A keyword set formation module is used to extract the product information in the store group and the keywords in the user's historical query statements corresponding to the query executed in the store group, filter out the keywords with high relevance to the product information, and form a keyword set for each store group; A synonym word set construction module is used to perform synonym expansion on the corresponding keyword set based on the industry characteristic labels of each store group, and obtain a synonym word set containing multiple synonym word pairs, wherein the synonym word pairs include keywords and their corresponding synonyms; A synonym word pair filtering module is used to input the synonym word set of each store group into the confidence scoring model to obtain the confidence scores of the multiple synonym word pairs, and filter out the synonym word pairs whose confidence scores are lower than a preset confidence threshold from the synonym word set; The product recall module is used to recall some products and give them back to the user based on the target query statement entered by the user and the synonym word set corresponding to the store grouping.
9. A computer device comprising a central processing unit and a memory, characterized in that: The central processing unit is used to call and run the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: It stores a computer program implemented according to the method described in any one of claims 1 to 7 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.
Citation Information
Cited By
Serialization recommendation method based on large language model
CN120873293A