Commodity word normalization method, device, equipment and storage medium
Patent Information
- Application Number
- CN202610895612.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]但是,基于静态阈值的哈希方法虽然计算效率高,但是保留的词对中两个商品词的语义关联度可能较低,因而导致最终得到归一化商品词的准确率较低;基于语义模型的全量比对方法虽然归一化商品词的确定准确率较高,但是在商品词较多时,该种方法确定归一化商品词的效率较差,远不及基于静态阈值的哈希方法
[0017]上述商品词归一化方法、装置、设备及存储介质,通过筛选出初始词对集中海明距离小于等于海明距离阈值的词对,得到中间词对集;筛选出所述中间词对集中语义相似性大于预设的相似性阈值的词对,得到目标词对集;基于所述目标词对集中各商品词的词频,确定标准代表词。上述实施中,通过海明距离阈值筛除初始词对集中海明距离较大的词对,得到相对于初始词对集词对数量大量减少的中间词对集,然后,通过预设的相似性阈值进一步筛除中间词对集中语义相似性较小的词对,从而得到相对于中间词对集词对数量进一步减少的目标词对集,目标词对集相较于初始词对集其词对数量大幅减少,且词对中两个商品词的文本相似性与语义相似性得到了大幅提高,数量大幅减少可有效提升归一化商品词(标准代表词)的计算效率,而文本相似性与语义相似性提高可有效提升归一化商品词(标准代表词)的计算准确率。
Smart Images

Figure CN122819237A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of language processing technology, and in particular to a method, apparatus, device and storage medium for normalizing commodity terms. Background Technology
[0002] When users search for products on shopping websites or software, they first enter custom product terms. For the same product, different users may use different custom product terms. To improve the accuracy and recall of product search results, different custom product terms can be converted into a unified product description, which is to achieve product term normalization.
[0003] Currently, product term normalization methods include hashing methods based on static thresholds and full comparison methods based on semantic models. The hashing method based on static thresholds works by retaining the word pair if the Hamming distance between any two product terms is less than a preset threshold, and then determining the normalized product term based on the retained word pairs. The full comparison method based on semantic models works by retaining the word pair if the similarity between any two product terms is less than a preset threshold, and then determining the normalized product term based on the retained word pairs.
[0004] However, while hashing methods based on static thresholds are computationally efficient, the semantic correlation between the two product words in the retained word pairs may be low, resulting in a lower accuracy of the final normalized product words. Semantic model-based full-scale comparison methods, while achieving higher accuracy in determining normalized product words, become less efficient when dealing with a large number of product words, far inferior to hashing methods based on static thresholds. In summary, it is currently difficult to balance computational efficiency and accuracy in determining normalized product words. Summary of the Invention
[0005] To facilitate a balance between computational efficiency and accuracy in determining normalized product terms, this application provides a product term normalization method, apparatus, device, and storage medium.
[0006] Firstly, this application provides a product term normalization method, including:
[0007] Filter out word pairs in the initial word pair set whose Hamming distance is less than or equal to the Hamming distance threshold to obtain the intermediate word pair set;
[0008] The target word pair set is obtained by filtering out word pairs in the intermediate word pair set whose semantic similarity is greater than a preset similarity threshold;
[0009] Based on the word frequency of each product word in the target word pair set, a standard representative word is determined.
[0010] Secondly, this application provides a product term normalization device, comprising:
[0011] The distance filtering module is used to filter out word pairs in the initial word pair set whose Hamming distance is less than or equal to the Hamming distance threshold, thus obtaining the intermediate word pair set.
[0012] The similarity filtering module is used to filter out word pairs in the intermediate word pair set whose semantic similarity is greater than a preset similarity threshold, thereby obtaining the target word pair set;
[0013] The normalization module is used to determine the standard representative word based on the word frequency of each product word in the target word set.
[0014] Thirdly, this application provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the method described above.
[0015] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described method.
[0016] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0017] The above-mentioned product term normalization method, apparatus, equipment, and storage medium obtain an intermediate word pair set by filtering out word pairs in the initial word pair set whose Hamming distance is less than or equal to the Hamming distance threshold; obtain a target word pair set by filtering out word pairs in the intermediate word pair set whose semantic similarity is greater than a preset similarity threshold; and determine a standard representative word based on the word frequency of each product term in the target word pair set. In the above implementation, word pairs with large Hamming distances in the initial word pair set are filtered out using the Hamming distance threshold, resulting in an intermediate word pair set with a significantly reduced number of word pairs compared to the initial word pair set. Then, word pairs with low semantic similarity in the intermediate word pair set are further filtered out using a preset similarity threshold, resulting in a target word pair set with a further reduced number of word pairs compared to the intermediate word pair set. The target word pair set has a significantly reduced number of word pairs compared to the initial word pair set, and the textual and semantic similarity between the two product words in the word pair is significantly improved. The significant reduction in the number of word pairs can effectively improve the computational efficiency of normalized product words (standard representative words), while the improvement in textual and semantic similarity can effectively improve the computational accuracy of normalized product words (standard representative words).
[0018] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart of a product term normalization method provided in the embodiments of this application;
[0021] Figure 2 This is a schematic diagram of the structure of a product term normalization device provided in the embodiments of this application;
[0022] Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application;
[0023] Figure 4 This is an internal structural diagram of a computer-readable storage medium provided in an embodiment of this application. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this disclosure.
[0025] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings herein are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, apparatus, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0026] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0027] Example 1
[0028] Figure 1 This is a flowchart of a product term normalization method provided in Embodiment 1 of this application, for reference. Figure 1 The method can be executed by a device that performs the method, which can be implemented in software and / or hardware, and the method includes:
[0029] S110. Select word pairs in the initial word pair set whose Hamming distance is less than or equal to the Hamming distance threshold to obtain the intermediate word pair set.
[0030] It should be noted that when users select products in shopping apps, websites, and mini-programs, they often enter product terms in the search box to search for corresponding products. The shopping system can generate corresponding transaction data by collecting the product terms entered by users in the past. In the transaction data, product terms of different categories are divided into different category groups. For example, in the transaction data, "skirt", "dress", "dress", and "long dress" are classified under the category group "skirt products". Multiple category groups can be classified through transaction data. Each time a product term entered by a user is obtained, a category group corresponding to that product term can be further associated.
[0031] Taking one of the product category terms as an example, by combining any two product terms in the product category term group, multiple word pairs are obtained. Each word pair contains two product terms, and the multiple word pairs formed by combining them are recorded as the initial word pair set.
[0032] Taking a word pair in the initial word pair set as an example, in this embodiment, the hash fingerprints of the two product words in the word pair are calculated respectively by a preset hash algorithm. For example, the hash algorithm can be the SimHash algorithm, and the hash fingerprint can be a 32-bit or 64-bit binary number, which is not limited. Through the hash fingerprints corresponding to the two product words respectively, the Hamming distance between the two product words can be further determined. The Hamming distance is used to characterize the number of different bits of the binary number at the same position in the two hash fingerprints. The smaller the Hamming distance, the higher the text similarity between the corresponding two product words.
[0033] In this implementation, the Hamming distance for each word pair in the initial word pair set can be calculated, and the Hamming distance for each word pair may be high or low. Since a semantic similarity calculation model is needed to further calculate the semantic similarity of word pairs in the initial word pair set, and semantic similarity calculation is inefficient, this embodiment aims to first select word pairs with high textual similarity in the initial word pair set before performing semantic similarity calculation. If the textual similarity of word pairs is high, the probability of high semantic similarity is also high. To select word pairs with high textual similarity in the initial word pair set, this embodiment presets a Hamming distance threshold based on historical experience data. This Hamming distance threshold is used to compare with the Hamming distances corresponding to each word pair in the initial word pair set, in order to select word pairs in the initial word pair set whose Hamming distance is less than or equal to the Hamming distance threshold, that is, to select word pairs with high textual similarity, and the selected word pairs are recorded as the intermediate word pair set.
[0034] Through the above implementation, the intermediate word pair set obtained by filtering out only word pairs with high text similarity from the initial word pair set significantly reduces the number of word pairs compared to the initial word pair set. This effectively reduces the number of word pairs that need to be calculated for semantic similarity in the subsequent steps, thereby improving the calculation efficiency of the normalized product words finally calculated by this scheme. Furthermore, since the text similarity between the two product words in the filtered word pairs is relatively poor, while the text similarity between the two product words in the retained word pairs is relatively high, the calculation accuracy of the normalized product words finally calculated can also be improved.
[0035] S120. Select word pairs in the intermediate word pair set whose semantic similarity is greater than a preset similarity threshold to obtain the target word pair set.
[0036] It should be noted that, based on the intermediate word pair set, this embodiment intends to further filter out only word pairs with high semantic similarity from the intermediate word pair set, and calculate normalized product words based on word pairs with high semantic similarity. To this end, it is necessary to first calculate the semantic similarity between the two product words in each word pair in the intermediate word pair set. In order to calculate the semantic similarity between the two product words, this embodiment has a preset semantic similarity calculation model. For example, the semantic similarity calculation model is the BERTScore semantic model. Taking a word pair in the intermediate word pair set as an example, by inputting the word pair into the BERTScore semantic model for processing, the BERTScore semantic model can calculate the semantic similarity between the two product words in the word pair. The semantic similarity is a normalized score used to measure the degree of semantic similarity between the two product words.
[0037] In this embodiment, after calculating the semantic similarity of each word pair in the intermediate word pair set, word pairs with high semantic similarity need to be selected. To determine whether the semantic similarity is high, this embodiment presets a similarity threshold based on historical experience data. This similarity threshold is used to compare with the semantic similarity of each word pair in the intermediate word pair set to measure whether the semantic similarity is high.
[0038] Specifically, word pairs with semantic similarity greater than a preset similarity threshold are selected from the intermediate word pair set and used as the target word pair set.
[0039] It should be noted that the target word pair set selected from the intermediate word pair set through the preset similarity threshold has a significantly reduced number of word pairs compared to the initial word pair set and the intermediate word pair set, and the semantic similarity between the two product words in the word pair is significantly improved. The significant reduction in the number of pairs can improve the computational efficiency of the subsequent calculation of normalized product words, while the improved semantic similarity can improve the computational accuracy of the normalized product words, because more accurate normalized product words need to be determined from word pairs with high semantic similarity.
[0040] S130. Based on the word frequency of each product word in the target word pair set, determine the standard representative word.
[0041] The target word pair set includes multiple word pairs, each containing two product words. That is, the target word pair set includes various different product words. If one of the product words appears most frequently in the transaction data (highest word frequency), it indicates that the product word has the highest probability of being the product that the user actually wants to search for. In other words, the product word with the highest word frequency in the target word pair set has the highest probability of being the normalized product word that the user is searching for. The product word with the highest word frequency in the target word pair set, that is, the normalized product word, is denoted as the standard representative word.
[0042] By implementing the above methods, the product words with the highest frequency in the target word pair are used as the normalized product words for the user's search, which can effectively ensure the accuracy of the normalized product words for the user's search.
[0043] It should be noted that this embodiment obtains an intermediate word pair set by filtering out word pairs in the initial word pair set whose Hamming distance is less than or equal to the Hamming distance threshold; it then obtains a target word pair set by filtering out word pairs in the intermediate word pair set whose semantic similarity is greater than a preset similarity threshold; and finally, it determines a standard representative word based on the word frequency of each product word in the target word pair set. In the above implementation, word pairs with large Hamming distances in the initial word pair set are filtered out using the Hamming distance threshold, resulting in an intermediate word pair set with a significantly reduced number of word pairs compared to the initial word pair set. Then, word pairs with small semantic similarity in the intermediate word pair set are further filtered out using a preset similarity threshold, resulting in a target word pair set with a further reduced number of word pairs compared to the intermediate word pair set. The target word pair set has a significantly reduced number of word pairs compared to the initial word pair set, and the textual and semantic similarity of the two product words in the word pairs is significantly improved. The significant reduction in the number of word pairs can effectively improve the calculation efficiency of normalized product words (standard representative words), while the improvement in textual and semantic similarity can effectively improve the calculation accuracy of normalized product words (standard representative words).
[0044] Example 2
[0045] This application provides a product term normalization method in Embodiment 2, which optimizes the step in Embodiment 1 of "screening out word pairs in the intermediate word pair set whose semantic similarity is greater than a preset similarity threshold to obtain a target word pair set". It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:
[0046] S210. Select word pairs in the initial word pair set whose Hamming distance is less than or equal to the Hamming distance threshold to obtain the intermediate word pair set.
[0047] S221. In response to the existence of short word pairs in the intermediate word pair set, determine the short word pair threshold corresponding to the short word pair; wherein, a short word pair is a word pair in which the number of characters of at least one product word is less than or equal to the number of characters threshold.
[0048] It should be noted that the intermediate word pair set includes multiple word pairs, and the word pair types specifically include short word pairs and long word pairs; among them, short word pairs are word pairs in which the number of characters of at least one product word is less than or equal to the character count threshold, while long word pairs are the remaining word pairs in the intermediate word pair set besides short word pairs; for example, the character count threshold is 3, but in other embodiments, the specifics are not limited.
[0049] Long word pairs contain more characters in their two product terms, while short word pairs contain fewer characters in their two product terms. For example, a long word pair might be "[short dress, floor-length dress]", and a short word pair might be "[skirt, dress]". It should be noted that when using the SimHash algorithm to hash the two product terms in a word pair, the fewer the number of identical characters between the two product terms, the greater the Hamming distance between the hash fingerprints of the two product terms.
[0050] Comparatively, because longer word pairs contain more characters in their product terms, the proportion of identical characters between the two product terms is relatively small, which in turn tends to result in a larger Hamming distance between the hash fingerprints corresponding to the two product terms. For example, in the long word pair [short dress, floor-length dress], the proportion of identical characters is only 1 / 4, so the Hamming distance between "short dress" and "floor-length dress" is relatively large. Conversely, because shorter word pairs contain fewer characters in their product terms, the proportion of identical characters between the two product terms is relatively large, which in turn tends to result in a smaller Hamming distance between the hash fingerprints corresponding to the two product terms. For example, in the short word pair [skirt, dress], the proportion of identical characters is 1 / 2, so the Hamming distance between "skirt" and "dress" is relatively small.
[0051] In step S210, the initial word pair set is similar to the intermediate word pair set, including both short and long word pairs. However, the same Hamming distance threshold is used for filtering this initial word pair set; for example, this Hamming distance threshold is 4. This Hamming distance threshold is suitable for filtering long word pairs in the initial word pair set. However, as the above analysis shows, relatively speaking, the Hamming distance of long word pairs is larger, while the Hamming distance of short word pairs is smaller. In step S210, because the same Hamming distance threshold is used to filter the initial word pair set, some short word pairs, even if their Hamming distance is large, will be filtered out and retained in the intermediate word pair set. The text similarity between the two product words in the short word pair with a large Hamming distance is relatively low, while this application aims to retain only short word pairs with high text similarity. Therefore, this embodiment needs to further screen the short word pairs in the intermediate word pair set to select the short word pairs with smaller Hamming distances. In order to further select short word pairs with smaller Hamming distances from the intermediate word pair set, it is necessary to determine a threshold smaller than the Hamming distance threshold shown in step S210, which is denoted as the short word pair threshold. For example, the Hamming distance threshold shown in step S210 is 4. In this embodiment, based on historical experience data, the short word pair threshold is set to 3.
[0052] S222. Screen out short word pairs in the intermediate word pair set whose Hamming distance is greater than the short word pair threshold to obtain a candidate word pair set.
[0053] It should be noted that, based on the number of words in the product terms in the intermediate word pair set, the word pairs in the intermediate word pair set can be divided into long word pairs and short word pairs. For the short word pairs, short word pairs with a Hamming distance greater than the short word pair threshold are filtered out. Then, the remaining word pairs in the intermediate word pair set are used as the candidate word pair set. Compared with the intermediate word pair set, the candidate word pair set has fewer short word pairs, and the Hamming distance of the short word pairs is smaller.
[0054] It should be noted that, based on the intermediate word pair set where the number of word pairs has already decreased compared to the initial word pair set, further filtering out short word pairs in the intermediate word pair set whose Hamming distance is greater than the short word pair threshold can further reduce the number of word pairs that need to be calculated for semantic similarity in the subsequent process. Since semantic similarity calculation is relatively time-consuming, reducing the number of word pairs that need to be calculated for semantic similarity in the subsequent process can effectively improve the efficiency of the subsequent semantic similarity calculation, thereby facilitating the determination of normalized product words.
[0055] S223. Select word pairs from the candidate word pair set whose semantic similarity is greater than a preset similarity threshold to obtain the target word pair set.
[0056] In a similar manner to step S120 in Embodiment 1, the semantic similarity between the two corresponding product words in each word pair of the candidate word pair set can also be calculated using the BERTScore semantic model. Then, the candidate word pair set is further filtered using a preset similarity threshold to select word pairs with higher semantic similarity as the target word pair set. For example, semantic similarity is a normalized score with a value between 0 and 1, and the preset similarity threshold is 0.88.
[0057] It should be noted that further filtering of the candidate word pair set through similarity thresholds can not only further reduce the number of word pairs used to calculate normalized product words, thereby improving the calculation efficiency of normalized product words, but also improve the semantic similarity of the word pairs used to normalize product words, thereby improving the calculation accuracy of normalized product words.
[0058] S230. Based on the word frequency of each product word in the target word pair set, determine the standard representative word.
[0059] Example 3
[0060] This application provides a product term normalization method in Embodiment 3, which supplements the step after "screening out word pairs in the candidate word pair set whose semantic similarity is greater than a preset similarity threshold to obtain the target word pair set" in Embodiment 2. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:
[0061] S310. Select word pairs in the initial word pair set whose Hamming distance is less than or equal to the Hamming distance threshold to obtain the intermediate word pair set.
[0062] S321. In response to the existence of short word pairs in the intermediate word pair set, determine the short word pair threshold corresponding to the short word pair; wherein, a short word pair is a word pair in which the number of characters of at least one product word is less than or equal to the number of characters threshold.
[0063] S322. Screen out short word pairs in the intermediate word pair set whose Hamming distance is greater than the short word pair threshold to obtain a candidate word pair set.
[0064] S323. Select word pairs from the candidate word pair set whose semantic similarity is greater than a preset similarity threshold to obtain the target word pair set.
[0065] S330. Based on the word frequency of each product word in the target word pair set, determine the standard representative word.
[0066] S340. Based on the target word pair set and the candidate word pair set, determine the word pair pass rate.
[0067] It should be noted that the target word pair set consists of word pairs selected from the candidate word pair set. In order to measure the degree to which word pairs in the candidate word pair set pass the screening and become word pairs in the target word pair set, this embodiment intends to calculate the word pair pass rate based on the above-mentioned target word pair set and candidate word pair set. For example, the word pair pass rate is the ratio of the number of word pairs in the target word pair set (first number) to the number of word pairs in the candidate word pair set (second number).
[0068] S350. Based on the pass rate of the word pair and the preset benchmark value, determine the pass rate fluctuation range.
[0069] It should be noted that if the calculated pass rate fluctuates significantly compared to the preset benchmark value, it indicates that the preset Hamming distance threshold in step S310 is inadequate, resulting in a large number of word pairs entering the target word pair set with low semantic similarity, or a small number of word pairs entering the target word pair set with high semantic similarity.
[0070] In order to measure the fluctuation range of the pass rate of word pairs, this embodiment presets a benchmark value, which is set based on historical experience data. For example, the benchmark value is 85%. The fluctuation range of the pass rate is the difference between the currently calculated pass rate and the benchmark value, and the difference is recorded as the fluctuation range of the pass rate.
[0071] S360. Based on the throughput fluctuation amplitude and the preset amplitude threshold, dynamically adjust the Hamming distance threshold.
[0072] In order to measure whether the above-mentioned pass rate fluctuation is too large, this embodiment presets an amplitude threshold. If the pass rate fluctuation is greater than the amplitude threshold, it means that the current pass rate fluctuation is large, which means that the Hamming distance threshold preset in step S310 of this solution is not good. At this time, the Hamming distance threshold can be adjusted to a suitable value according to the preset adjustment strategy.
[0073] In this way, after each calculation of the candidate word pair set and the target word pair set, the Hamming distance threshold between the candidate word pair set and the target word pair set can be dynamically adjusted to ensure that there are not too many word pairs in the candidate word pair set with low semantic similarity entering the target word pair set, or that there are not too few word pairs in the candidate word pair set with high semantic similarity entering the target word pair set. This makes it easier to improve the accuracy of the normalized product words calculated subsequently.
[0074] Example 4
[0075] This application provides a product term normalization method in Embodiment 4, which optimizes the "determining the pass rate fluctuation range based on the pass rate of the word pair and a preset benchmark value" in Embodiment 3. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:
[0076] S410. Select word pairs in the initial word pair set whose Hamming distance is less than or equal to the Hamming distance threshold to obtain the intermediate word pair set.
[0077] S421. In response to the existence of short word pairs in the intermediate word pair set, determine the short word pair threshold corresponding to the short word pair; wherein, a short word pair is a word pair in which the number of characters of at least one product word is less than or equal to the number of characters threshold.
[0078] S422. Screen out short word pairs in the intermediate word pair set whose Hamming distance is greater than the short word pair threshold to obtain a candidate word pair set.
[0079] S423. Select word pairs from the candidate word pair set whose semantic similarity is greater than a preset similarity threshold to obtain the target word pair set.
[0080] S430. Based on the target word pair set and the candidate word pair set, determine the word pair pass rate.
[0081] S440. Calculate the mean of at least one historical pass rate to obtain a baseline value.
[0082] It should be noted that in response to receiving a product term input by a user, the normalization step corresponding to that product term can be executed once using this method. If this is the first execution of the product term normalization method, only step S430 needs to be executed to calculate the pass rate of the corresponding word pair. If this is not the first execution of the product term normalization method, the time when the product term normalization method is started is taken as the time base point. The historical pass rates obtained from each execution of the product term normalization method before this time base point are obtained and abbreviated as the historical pass rate. Then, the average of each historical pass rate is calculated as the benchmark value used for subsequent difference calculation with the word pair pass rate calculated this time.
[0083] S450. Based on the pass rate of the word pair and the preset benchmark value, determine the pass rate fluctuation range.
[0084] S460. Based on the throughput fluctuation amplitude and the preset amplitude threshold, dynamically adjust the Hamming distance threshold.
[0085] S470. Based on the word frequency of each product word in the target word pair set, determine the standard representative word.
[0086] Example 5
[0087] This application provides a product term normalization method in Embodiment 5, which optimizes the "dynamic adjustment of the Hamming distance threshold based on the pass rate fluctuation amplitude and the preset amplitude threshold" in Embodiment 3. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:
[0088] S510. Select word pairs in the initial word pair set whose Hamming distance is less than or equal to the Hamming distance threshold to obtain the intermediate word pair set.
[0089] S521. In response to the existence of short word pairs in the intermediate word pair set, determine the short word pair threshold corresponding to the short word pair; wherein, a short word pair is a word pair in which the number of characters of at least one product word is less than or equal to the number of characters threshold.
[0090] S522. Screen out short word pairs in the intermediate word pair set whose Hamming distance is greater than the short word pair threshold to obtain a candidate word pair set.
[0091] S523. Select word pairs from the candidate word pair set whose semantic similarity is greater than a preset similarity threshold to obtain the target word pair set.
[0092] S530. Based on the word frequency of each product word in the target word pair set, determine the standard representative word.
[0093] S540. Based on the target word pair set and the candidate word pair set, determine the word pair pass rate.
[0094] S550. Based on the pass rate of the word pair and the preset benchmark value, determine the pass rate fluctuation range.
[0095] S561. Calculate the absolute value of the fluctuation range of the pass rate to obtain the absolute value of the fluctuation range.
[0096] The pass rate fluctuation range is the difference between the pass rate fluctuation range calculated in step S550 and the preset benchmark value, and the absolute value of the difference is recorded as the absolute value of the fluctuation range.
[0097] S562. In response to the absolute value of the fluctuation amplitude being greater than a preset amplitude threshold and the throughput fluctuation amplitude being negative, the Hamming distance threshold is lowered according to a preset adjustment strategy.
[0098] The pass rate fluctuation may be greater than or less than the preset benchmark value. If the pass rate fluctuation is less than the preset benchmark value, the pass rate fluctuation is negative. This indicates that the pass rate calculated by this method has decreased compared to the benchmark value. If the absolute value of the fluctuation is greater than the preset amplitude threshold, it indicates that the pass rate calculated by this method has not only decreased compared to the benchmark value, but the decrease is too large. The general reason for this situation is that the Hamming distance threshold preset in S510 is too large (the choice of the Hamming distance threshold is too lenient), which causes a large number of word pairs with low text similarity to enter the intermediate word pair set, and then causes a large number of word pairs with low semantic similarity to enter the candidate word pair set. At this time, it is necessary to tighten the Hamming distance threshold, that is, to lower the preset Hamming distance threshold. In order to lower the Hamming distance threshold, this embodiment has a preset lowering strategy. For example, the lowering strategy is to reduce the preset Hamming distance threshold by 1, thereby obtaining the lowered new Hamming distance threshold.
[0099] It should be noted that when the absolute value of the fluctuation amplitude is greater than the preset amplitude threshold and the pass rate fluctuation amplitude is negative, the problem of the preset Hamming distance threshold being too lenient can be solved by reducing the Hamming distance threshold. This prevents a large number of word pairs with low text similarity from entering the intermediate word pair set and a large number of word pairs with low semantic similarity from entering the candidate word pair set, thereby improving the efficiency and accuracy of the finally calculated normalized product words.
[0100] S563. In response to the absolute value of the fluctuation amplitude being greater than a preset amplitude threshold and the throughput fluctuation amplitude being positive, the Hamming distance threshold is increased according to a preset adjustment strategy.
[0101] If the fluctuation range of the pass rate is greater than the preset benchmark value, then the fluctuation range of the pass rate is positive. This indicates that the pass rate calculated by this method in this execution has increased compared to the benchmark value. If the absolute value of the fluctuation range is also greater than the preset range threshold, it indicates that the pass rate calculated by this method in this execution has not only increased compared to the benchmark value, but the increase is too large. The general reason for this situation is that the Hamming distance threshold preset in S510 is too small (the selection of the Hamming distance threshold is too strict), resulting in only a small number of word pairs with low text similarity entering the intermediate word pair set, and thus only a small number of word pairs with low semantic similarity entering the candidate word pair set. At this time, it is necessary to relax the Hamming distance threshold, that is, to increase the preset Hamming distance threshold. In order to increase the Hamming distance threshold, this embodiment has a preset increase strategy. For example, the increase strategy is to add 1 to the preset Hamming distance threshold to obtain the new increased Hamming distance threshold.
[0102] It should be noted that when the absolute value of the fluctuation amplitude is greater than the preset amplitude threshold and the fluctuation amplitude of the pass rate is positive, increasing the Hamming distance threshold can solve the problem that the preset Hamming distance threshold is too strict. This prevents only a small number of word pairs with low text similarity from entering the intermediate word pair set and only a small number of word pairs with low semantic similarity from entering the candidate word pair set, thereby improving the accuracy of the finally calculated normalized product words.
[0103] Example 6
[0104] This application provides a product term normalization method in Embodiment Six, which optimizes the "determining a standard representative term based on the word frequency of each product term in the target term set" in Embodiment One. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:
[0105] S610. Select word pairs in the initial word pair set whose Hamming distance is less than or equal to the Hamming distance threshold to obtain the intermediate word pair set.
[0106] S620. Select word pairs in the intermediate word pair set whose semantic similarity is greater than a preset similarity threshold to obtain the target word pair set.
[0107] S631. Semantically associate each target word pair in the target word pair set to obtain connected word groups.
[0108] It should be noted that the textual and semantic similarity of each word pair in the target word pair set is high, and the same product words often appear in different word pairs. For example, taking multiple word pairs in the target word pair set as an example, these multiple word pairs include: [dress, skirt], [skirt, long dress]. Subsequently, we want to use the product word with the highest word frequency in the target word pair set as the normalized product word. However, since the product words in the target word pair set are scattered in different word pairs, it is difficult to directly count different kinds of product words from the target word pair set. Therefore, in this embodiment, we first perform semantic association on each target word pair in the target word pair set, thereby counting multiple different kinds of product words.
[0109] For example, by statistically identifying the different types of product words appearing in the target word pair set, and combining the statistically identified product words into a word group, semantic association of each target word pair in the target word pair set can be achieved; for example, if the word pairs in the target word pair set include [dress, skirt] and [skirt, long dress], then the word group generated after semantic association is {"dress", "skirt", "long dress"}, and the word group generated after semantic association is denoted as a connected word group.
[0110] S632. Based on the word frequency of each product word in the connected word group, determine the standard representative word.
[0111] In this embodiment, the frequency of each product word in the connected word group in the transaction data is taken as the word frequency of the corresponding product word. In order to ensure that the normalized product word obtained in the end has the highest probability of being used to represent the product word entered by the user when searching for products, the product word with the highest word frequency in the connected word group is taken as the normalized product word, and the normalized product word is recorded as the standard representative word.
[0112] Example 7
[0113] This application provides a product term normalization method in Embodiment 7, which supplements the steps preceding "screening out word pairs in the initial word pair set whose Hamming distance is less than or equal to the Hamming distance threshold to obtain an intermediate word pair set" in Embodiment 1. It should be noted that for parts not detailed in this embodiment, please refer to the descriptions in other embodiments. The method includes:
[0114] S710. Clean the initial product words to obtain cleaned product words, classify the cleaned product words, and generate single-category word groups.
[0115] Specifically, the product terms entered by users during their historical product searches are recorded as initial product terms. These initial product terms need to be stored in the product categories associated with them. It should be noted that, due to their personalized and non-standardized nature, the initial product terms entered by users are not suitable for direct storage in the product categories associated with them. Therefore, the initial product terms need to be cleaned to achieve a standardized expression, making them eligible for direct storage in the product categories associated with them.
[0116] For example, an initial product term is "dress!!". After washing the initial product term, the corresponding washed product term can be obtained, which is "dress".
[0117] Furthermore, each washed product term can be semantically associated with its corresponding product category. For example, "dress" can be categorized under the product category of skirts, coded as 630790. Each washed product term contained under a product category is recorded as a single-category term group.
[0118] S720. Combine the washed product words in the single-category word groups associated with the user's search terms in pairs to obtain an initial word pair set.
[0119] Specifically, the product terms entered by the user are recorded as the user's search terms. Taking a single product category term as an example, the user's search terms can be associated with the corresponding single product category term through semantic association and other methods. Assuming that the single product category term has a total of 100 washed product terms, by combining each washed product term in the single product category term in pairs, 100×100 word pairs can be obtained, and each word pair includes 2 washed product terms. All word pairs generated by combining them in pairs are collectively referred to as the initial word pair set.
[0120] It should be noted that by cleaning the initial product terms, the standardization of the product terms in each word pair in the generated initial word pair set can be improved, thereby facilitating the improvement of the standardization and accuracy of the normalized product terms generated subsequently based on the initial word pair set.
[0121] S730. Select word pairs in the initial word pair set whose Hamming distance is less than or equal to the Hamming distance threshold to obtain the intermediate word pair set.
[0122] S740. Select word pairs in the intermediate word pair set whose semantic similarity is greater than a preset similarity threshold to obtain the target word pair set.
[0123] S750. Based on the word frequency of each product word in the target word pair set, determine the standard representative word.
[0124] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0125] Example 8
[0126] Based on the same inventive concept, this embodiment also provides a product term normalization device for implementing the product term normalization method described above. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more product term normalization device embodiments provided below can be found in the limitations of the product term normalization method above, and will not be repeated here.
[0127] In this embodiment, as Figure 2 As shown, a product term normalization device is provided, comprising:
[0128] The distance filtering module is used to filter out word pairs in the initial word pair set whose Hamming distance is less than or equal to the Hamming distance threshold, thus obtaining the intermediate word pair set.
[0129] The similarity filtering module is used to filter out word pairs in the intermediate word pair set whose semantic similarity is greater than a preset similarity threshold, thereby obtaining the target word pair set;
[0130] The normalization module is used to determine the standard representative word based on the word frequency of each product word in the target word set.
[0131] Each module in the aforementioned product term normalization device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0132] It should be noted that this embodiment obtains an intermediate word pair set by filtering out word pairs in the initial word pair set whose Hamming distance is less than or equal to the Hamming distance threshold; it then obtains a target word pair set by filtering out word pairs in the intermediate word pair set whose semantic similarity is greater than a preset similarity threshold; and finally, it determines a standard representative word based on the word frequency of each product word in the target word pair set. In the above implementation, word pairs with large Hamming distances in the initial word pair set are filtered out using the Hamming distance threshold, resulting in an intermediate word pair set with a significantly reduced number of word pairs compared to the initial word pair set. Then, word pairs with small semantic similarity in the intermediate word pair set are further filtered out using a preset similarity threshold, resulting in a target word pair set with a further reduced number of word pairs compared to the intermediate word pair set. The target word pair set has a significantly reduced number of word pairs compared to the initial word pair set, and the textual and semantic similarity of the two product words in the word pairs is significantly improved. The significant reduction in the number of word pairs can effectively improve the calculation efficiency of normalized product words (standard representative words), while the improvement in textual and semantic similarity can effectively improve the calculation accuracy of normalized product words (standard representative words).
[0133] In an optional embodiment, in order to filter out word pairs in the intermediate word pair set whose semantic similarity is greater than a preset similarity threshold and obtain a target word pair set, the similarity filtering module is specifically used to: in response to the existence of short word pairs in the intermediate word pair set, determine the short word pair threshold corresponding to the short word pair; wherein, a short word pair is a word pair in which the number of characters of at least one product word is less than or equal to the number of characters threshold;
[0134] Short word pairs whose Hamming distance is greater than the short word pair threshold are filtered out from the intermediate word pair set to obtain a candidate word pair set;
[0135] The target word pair set is obtained by filtering out word pairs in the candidate word pair set whose semantic similarity is greater than a preset similarity threshold.
[0136] In an optional embodiment, after filtering out word pairs in the candidate word pair set whose semantic similarity is greater than a preset similarity threshold to obtain the target word pair set, the product word normalization device further includes:
[0137] The pass rate determination module is specifically used to determine the pass rate of word pairs based on the target word pair set and the candidate word pair set;
[0138] The amplitude calculation module is specifically used to determine the amplitude of the pass rate fluctuation based on the pass rate of the word pair and a preset benchmark value;
[0139] The threshold adjustment module is specifically used to dynamically adjust the Hamming distance threshold based on the throughput fluctuation amplitude and a preset amplitude threshold.
[0140] In an optional embodiment, before determining the pass rate fluctuation range based on the pass rate of the word pair and a preset benchmark value, the product word normalization device further includes:
[0141] The baseline value calculation module is specifically used to calculate the average of at least one historical pass rate to obtain the baseline value.
[0142] In an optional embodiment, regarding the dynamic adjustment of the Hamming distance threshold based on the throughput fluctuation amplitude and a preset amplitude threshold, the threshold adjustment module is specifically used for:
[0143] Calculate the absolute value of the fluctuation range of the pass rate to obtain the absolute value of the fluctuation range;
[0144] In response to the absolute value of the fluctuation amplitude being greater than a preset amplitude threshold and the throughput fluctuation amplitude being negative, the Hamming distance threshold is lowered according to a preset reduction strategy.
[0145] In response to the absolute value of the fluctuation amplitude being greater than a preset amplitude threshold, and the throughput fluctuation amplitude being positive, the Hamming distance threshold is increased according to a preset adjustment strategy.
[0146] In an optional embodiment, in determining the standard representative word based on the word frequency of each product word in the target word pair set, the normalization module is specifically used for:
[0147] Semantic association is performed on each target word pair in the target word pair set to obtain connected word groups;
[0148] Based on the word frequency of each product word in the connected word group, a standard representative word is determined.
[0149] In an optional embodiment, before filtering out word pairs in the initial word pair set whose Hamming distance is less than or equal to the Hamming distance threshold to obtain the intermediate word pair set, the product word normalization device further includes:
[0150] The phrase generation module is specifically used to clean the initial product words to obtain cleaned product words, classify the cleaned product words, and generate single-category phrases;
[0151] The word pair determination module is specifically used to combine the washed product words in the single-category word groups associated with the user's search terms in pairs to obtain an initial word pair set.
[0152] Example 9
[0153] In this embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows. Figure 3As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores data. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a product term normalization method.
[0154] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the computer device to which the present disclosure is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0155] Example 10
[0156] In this embodiment, a computer-readable storage medium is provided, such as... Figure 4 As shown, a computer program is stored thereon, and when the computer program is executed by the processor, it implements the steps in the above-described method embodiments.
[0157] Example 11
[0158] In this embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0159] It should be noted that the information collected is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions, necessary confidentiality measures have been taken, and it does not violate public order and good morals. Corresponding operation portals are provided for users to choose to authorize or refuse.
[0160] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this disclosure can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this disclosure may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this disclosure may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0161] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0162] The embodiments described above are merely illustrative of several implementations of this disclosure, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent disclosure. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this disclosure, and these all fall within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the appended claims.
Claims
1. A method for normalizing product terms, characterized in that, include: Filter out word pairs in the initial word pair set whose Hamming distance is less than or equal to the Hamming distance threshold to obtain the intermediate word pair set; The target word pair set is obtained by filtering out word pairs in the intermediate word pair set whose semantic similarity is greater than a preset similarity threshold; Based on the word frequency of each product word in the target word pair set, a standard representative word is determined.
2. The method according to claim 1, characterized in that, The process of filtering out word pairs in the intermediate word pair set whose semantic similarity is greater than a preset similarity threshold yields a target word pair set, including: In response to the existence of short word pairs in the intermediate word pair set, a short word pair threshold corresponding to the short word pair is determined; wherein, a short word pair is a word pair in which the number of characters of at least one product word is less than or equal to the number of characters threshold; Short word pairs whose Hamming distance is greater than the short word pair threshold are filtered out from the intermediate word pair set to obtain a candidate word pair set; The target word pair set is obtained by filtering out word pairs in the candidate word pair set whose semantic similarity is greater than a preset similarity threshold.
3. The method according to claim 2, characterized in that, After selecting word pairs from the candidate word pair set whose semantic similarity is greater than a preset similarity threshold to obtain the target word pair set, the process further includes: Based on the target word pair set and the candidate word pair set, the word pair pass rate is determined; Based on the pass rate of the word pair and a preset benchmark value, the pass rate fluctuation range is determined; The Hamming distance threshold is dynamically adjusted based on the throughput fluctuation amplitude and the preset amplitude threshold.
4. The method according to claim 3, characterized in that, Before determining the pass rate fluctuation range based on the pass rate of the word pair and a preset benchmark value, the method further includes: Calculate the mean of at least one historical pass rate to obtain a baseline value.
5. The method according to claim 3, characterized in that, The step of dynamically adjusting the Hamming distance threshold based on the throughput fluctuation amplitude and a preset amplitude threshold includes: Calculate the absolute value of the fluctuation range of the pass rate to obtain the absolute value of the fluctuation range; In response to the absolute value of the fluctuation amplitude being greater than a preset amplitude threshold and the throughput fluctuation amplitude being negative, the Hamming distance threshold is lowered according to a preset reduction strategy. In response to the absolute value of the fluctuation amplitude being greater than a preset amplitude threshold, and the throughput fluctuation amplitude being positive, the Hamming distance threshold is increased according to a preset adjustment strategy.
6. The method according to claim 1, characterized in that, The process of determining standard representative words based on the word frequency of each product word in the target word pair set includes: Semantic association is performed on each target word pair in the target word pair set to obtain connected word groups; Based on the word frequency of each product word in the connected word group, a standard representative word is determined.
7. The method according to claim 1, characterized in that, Before obtaining the intermediate word pair set by filtering out word pairs in the initial word pair set whose Hamming distance is less than or equal to the Hamming distance threshold, the process also includes: The initial product terms are cleaned to obtain cleaned product terms, which are then categorized to generate single-category word groups. The washed product terms in the single-category terms associated with the user's search terms are combined in pairs to obtain an initial set of term pairs.
8. A product term normalization device, characterized in that, The device includes: The distance filtering module is used to filter out word pairs in the initial word pair set whose Hamming distance is less than or equal to the Hamming distance threshold, thus obtaining the intermediate word pair set. The similarity filtering module is used to filter out word pairs in the intermediate word pair set whose semantic similarity is greater than a preset similarity threshold, thereby obtaining the target word pair set; The normalization module is used to determine the standard representative word based on the word frequency of each product word in the target word set.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.