Commodity transaction platform data exchange method

Through the bidirectional maximum matching algorithm and cosine similarity calculation, a cross-language entry mapping library is established, which solves the problem of describing similar products in different languages ​​in cross-border e-commerce platforms, and improves the accuracy of search results and transaction success rate.

CN120218082APending Publication Date: 2025-06-27厦门工学院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510535217.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In cross-border e-commerce platforms, users in different countries use different languages ​​to describe similar products, resulting in incomplete search results, reducing transaction success rate.

Method used

A two-way maximum matching algorithm is used to segment the product description, and group it according to languages, a cross-language entry mapping library is established, the semantic similarity of entries between different languages ​​is calculated through cosine similarity, a cross-language correspondence relationship is established, and the cross-language product matching results are found based on the keywords in the user's search request.

Benefits of technology

It improves the accuracy and coverage of cross-language product matching, enhances transaction success rate, and solves the limitations of the existing technology when processing multilingual data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218082A_ABST
    Figure CN120218082A_ABST
Patent Text Reader

Abstract

The invention provides a commodity transaction platform data exchange method, which relates to the technical field of e-commerce, and comprises the following steps of: performing word segmentation on commodity description through a bidirectional maximum matching algorithm, grouping according to languages, establishing a cross-language corresponding relation for entries with semantic similarity, and establishing a cross-language corresponding relation for the entries with semantic similarity; and all cross-language commodity matching results are searched according to keywords in the search request of the user, so that transaction is promoted. Thus, the technical problem that in a cross-border e-commerce platform, due to the fact that users in different countries describe similar commodities in different languages, search results are incomplete, and the transaction success rate is reduced can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of e-commerce, and more specifically, to a method for data exchange on a commodity trading platform. Background Art

[0002] With the rapid development of e-commerce, commodity trading platforms have become an important channel for global consumers and merchants to conduct commodity transactions and information interactions. In recent years, the demand for cross-language commodity transactions has been increasing. Especially in the international e-commerce environment, the semantic understanding of commodity descriptions and user search requests in different languages has become an important challenge. To address this issue, many commodity trading platforms have gradually introduced multi-language data processing technologies, including natural language processing (NLP), machine translation, and semantic analysis, etc., hoping to achieve efficient cross-language data exchange and commodity information matching through technical means. However, the existing technologies still have certain limitations in aspects such as accurate word segmentation, semantic similarity calculation, and cross-language entry mapping. For example, rule-based word segmentation technologies are difficult to handle complex commodity descriptions and multi-language characteristics, and traditional entry matching methods cannot accurately capture the semantic relationships between different languages, resulting in limitations in the accuracy and coverage of user search results. In addition, current cross-language commodity matching technologies usually rely on static dictionaries or translation systems, and this approach shows obvious deficiencies when dealing with real-time dynamic transaction data.

[0003] The deficiencies of the existing technologies in multi-language data processing on commodity trading platforms are mainly reflected in the following aspects: First, the word segmentation technology cannot effectively handle professional terms, abbreviations, and language mixing phenomena in commodity descriptions, resulting in insufficient integrity and accuracy of the word segmentation results, thus affecting subsequent semantic analysis and data matching; Second, the existing semantic similarity calculation methods lack pertinence in grouping and comparing multi-language data, and most rely on a single translation model or shallow semantic matching algorithms, unable to fully explore the deep semantic relationships between different languages; Third, there is a lack of real-time performance and adaptability for dynamic transaction data, and the update of traditional entry mapping libraries is slow, unable to quickly respond to cross-language search requests. In addition, in the process of keyword expansion of user search requests and commodity matching, the existing technologies usually only rely on simple keyword matching or translation results, and fail to fully utilize the semantic information in commodity transaction data, resulting in low accuracy and diversity of search results. Summary of the Invention

[0004] To solve the above technical problems, the present invention is proposed. The present invention provides a method for data exchange on a commodity trading platform, which can, to a certain extent, solve the technical problem that in cross-border e-commerce platforms, due to users in different countries using different languages to describe the same kind of commodities, the search results are incomplete, reducing the transaction success rate.

[0005] According to one aspect of the present invention, there is provided a method for data exchange in a commodity trading platform, which includes: Obtain multilingual transaction data in the commodity trading platform; Based on the multilingual transaction data, use the bidirectional maximum matching algorithm to segment the commodity description, and screen the segmentation results of successful transactions according to the transaction identifier to generate a term mapping library; Group the segmentation results in the term mapping library according to the language tags, calculate the semantic similarity of terms between different languages using cosine similarity, and establish a cross-language correspondence for terms with the semantic similarity when the semantic similarity is greater than a preset threshold; Receive a search request from a user, search for semantically similar terms in the cross-language correspondence according to the keywords in the search request, expand the search scope based on the terms, and output cross-language commodity matching results.

[0006] Further, the bidirectional maximum matching algorithm includes two sub-processes: forward maximum matching and backward maximum matching; When the number of segmentation results of the forward maximum matching and the backward maximum matching is inconsistent, select the result with fewer segmentation numbers as the final segmentation result.

[0007] Further, match the segmented results with a preset commodity domain dictionary, and when there is a corresponding entry for the segmentation result in the commodity domain dictionary, mark the segmentation result as a valid term; Screen the valid terms in the transaction records with the transaction identifier value of 1, form a term mapping unit with the corresponding language tag and the search keyword of the valid term, and store the term mapping unit in the term mapping library.

[0008] Further, the forward maximum matching and the backward maximum matching respectively compare the current string to be matched with the terms in the commodity domain dictionary from the left or right side of the text; if the match is successful, it is segmented into independent terms and shifted to the right, otherwise, the backtracking step size is dynamically adjusted according to the dictionary word length frequency to gradually shorten the match until the match is successful or the text length is reduced to a single character, and finally a complete segmentation sequence is output.

[0009] Further, the dynamic adjustment of the backtracking step size includes two control parameters: a continuous matching counter and an adjustment coefficient; Confirm a new backtracking step size according to the continuous matching counter and the adjustment coefficient. If no matching term is found yet, enter the single-character segmentation mode.

[0010] Further, the single-character segmentation pattern divides characters into numeric, alphabetic, punctuation, and other categories by character type judgment, combines and segments them according to adjacent character types, combines consecutive characters belonging to the same language into one entry, and applies post-processing rules to optimize the segmentation result, finally outputting a segmentation result that better meets the requirements of the field.

[0011] Further, the calculation of semantic similarity of entries includes two stages: construction of word vectors and calculation of similarity; In the stage of constructing word vectors, the entries in the entry mapping library are classified by language grouping, and entry feature vectors are constructed based on the term frequency-inverse document frequency matrix; the matrix is dynamically updated when the number of entries changes; The similarity calculation extracts the feature vectors of the entries to be compared. If the vector dimensions are inconsistent, they are aligned by zero padding, and the cosine similarity formula is used to calculate the similarity value between the two entries.

[0012] Further, the establishment of cross-language correspondence relationships includes three stages: similarity screening, transfer verification, and relationship storage; The transfer verification ensures that entry a in language A and entry b in language B are the most similar entries to each other through two-way verification, and after passing the verification, they are marked as valid word pairs.

[0013] Further, the two-way verification includes forward verification, reverse verification, and consistency confirmation; The forward verification starts from the source entry a and checks whether the target entry b ranks first or second in its sequence of similar entries and the similarity difference from the first place is less than a preset value; The reverse verification takes the target entry b as the source entry and checks whether it meets similar conditions in the sequence of similar entries in the source language; The consistency confirmation finally confirms the candidate word pairs that pass the forward and reverse verifications. Those that pass both ways are marked as strong correspondence relationships, and those that pass one way are marked as weak correspondence relationships; if the source entry forms weak correspondence relationships with multiple target entries, the word pair with the highest similarity is retained, and the rest are marked as relationships to be observed; weak correspondence relationships can be upgraded to strong correspondence relationships when the similarity is stable after three consecutive updates of transaction data.

[0014] Further, finding semantically similar entries in the cross-language correspondence relationships includes: Extracting the keywords in the search request and determining their languages; if the language is not marked, the language is inferred through character features. If the inference fails, it is marked as an unknown language and skipped for processing; For keywords with a determined language, search for corresponding nodes in the cross - language graph structure. If there is no matching node, determine whether it is an isolated node. For isolated nodes, return an empty result; for non - isolated nodes, extract the edges and sort them by similarity; filter out semantically similar entries according to the target language. If the target language is not specified, return all target node entries that meet the threshold; When the out - edge is empty or the similarity is lower than the threshold, mark it as having no similar entries. If multiple target nodes meet the conditions, select the best entry according to the priorities of similarity and out - degree value; If there is no match between the keyword and the target language, try to establish a mapping relationship using the cross - language combination mode. If the mapping is successful, update the graph structure and add it to the result set; if it fails, return an empty result; Group and output the search results by entry and language. If there is more than one target language, execute the above logic separately and merge the results to ensure coverage of all language requirements.

[0015] Compared with the prior art, a data exchange method for a commodity trading platform provided by the present invention tokenizes the commodity description through a bidirectional maximum matching algorithm, groups it according to the language, establishes a cross - language correspondence relationship for entries with the semantic similarity, and then searches for all commodity matching results in the cross - language based on the keywords in the user search request to promote transactions. In this way, it can solve the technical problem in cross - border e - commerce platforms that due to users in different countries using different languages to describe the same kind of commodities, the search results are incomplete, reducing the transaction success rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings: Figure 1 FIG. is a flowchart of the data exchange method for a commodity trading platform according to an embodiment of the present invention.

[0017] Figure 2 FIG. is a flowchart of the dynamic adjustment of the backtracking step length in the data exchange method for a commodity trading platform according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] Next, exemplary embodiments according to the present invention will be described in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments of the present invention. It should be understood that the present invention is not limited by the exemplary embodiments described herein.

[0019] Figure 1Flowchart of the method for data exchange on a commodity trading platform according to an embodiment of the present invention. As Figure 1 shown, in the method for data exchange on a commodity trading platform, it includes: S1: Obtain multilingual transaction data in the commodity trading platform, where the multilingual transaction data includes the user's search keywords, commodity descriptions, transaction completion indicators, and language tags of the languages to which they belong; Obtain multilingual transaction data in the commodity trading platform, where the multilingual transaction data includes the user's search keywords, commodity descriptions, transaction completion indicators, and language tags of the languages to which they belong; wherein, the search keywords record the actual search terms input by the user on the commodity trading platform, the commodity descriptions include commodity titles, specification parameters, and seller-defined description texts, the transaction completion indicators mark whether a transaction is completed in the form of a Boolean value, where the value 1 indicates that the transaction is successful and the value 0 indicates that the transaction is not successful, the language tags use two-letter codes of the ISO 639-1 standard to identify the languages to which the data belongs, including but not limited to en for English, zh for Chinese, ja for Japanese, and de for German; the acquisition of the multilingual transaction data is completed by calling the data interface of the commodity trading platform, and the acquisition time range is set to the transaction records within the most recent 90 days to ensure data timeliness; perform data cleaning on the obtained multilingual transaction data, remove data entries containing illegal characters, duplicate records, and incomplete information, and store the cleaned valid data entries in the cache area of the data processing unit to prepare for subsequent word segmentation processing.

[0020] S2: Based on the multilingual transaction data, use the bidirectional maximum matching algorithm to segment the commodity descriptions, and screen the word segmentation results of successful transactions according to the transaction completion indicators to generate a term mapping library; The bidirectional maximum matching algorithm includes two sub - processes: forward maximum matching and reverse maximum matching. The forward maximum matching starts scanning from the left side of the product description for word segmentation, and the reverse maximum matching starts scanning from the right side of the product description for word segmentation. When the number of word segmentation results of the forward maximum matching and the reverse maximum matching is inconsistent, the result with fewer word segments is selected as the final word segmentation result. When performing word segmentation on the product description, the maximum word length is set to 10 characters, and the rationality of the segmentation position is judged sequentially through a sliding window method. The size of the sliding window gradually decreases from the maximum word length to 1 character. The segmented result is matched with a preset product domain dictionary. When there is a corresponding entry for the word segmentation result in the product domain dictionary, the word segmentation result is marked as a valid entry. The valid entries in the transaction records with the transaction identification value of 1 are screened, and the valid entries, their corresponding language tags, and the search keywords are combined to form an entry mapping unit, and the entry mapping unit is stored in the entry mapping library. The entry mapping library is stored in the form of key - value pairs, where the valid entry is used as the primary key, and the language tag and the search keyword are used as the corresponding attribute values.

[0021] More specifically, the forward maximum matching starts from the first character on the left side of the product description, sets the initial segmentation length as the maximum word length of 10 characters, and takes 10 consecutive characters from the left side of the product description as the current string to be matched. The forward maximum matching first compares the current string to be matched with the entries in the product domain dictionary. When a completely matching entry is found for the current string to be matched in the product domain dictionary, the current string to be matched is segmented into an independent entry, and the segmentation position is shifted to the end position of the current string to be matched. When no matching entry is found for the current string to be matched in the product domain dictionary, the backtracking step size is determined according to the frequency distribution of the entry lengths in the product domain dictionary. The calculation of the backtracking step size is as Figure 2As shown, the method is as follows: Count the occurrence frequencies of the lengths of all entries in the commodity domain dictionary, and use the entry length with the highest occurrence frequency as the preferred backtracking step value. When the preferred backtracking step value is greater than 3, set the backtracking step to 3. When the preferred backtracking step value is less than or equal to 3, set the backtracking step to the preferred backtracking step value. Based on the backtracking step, intercept the corresponding number of characters from the right side of the current string to be matched, and re-match the remaining string for entries. When no matching entry is found after three consecutive matches using the backtracking step, reduce the backtracking step to 1 and continue character-by-character matching. When the length of the current string to be matched is reduced to 1 character and no matching entry is found in the commodity domain dictionary, regard this single character as an independent entry and move the segmentation position to the next character. After each successful segmentation of an independent entry, the forward maximum matching will restart the above matching process from the next character after the segmentation position and reset the segmentation length to the maximum word length of 10 characters. When the segmentation position reaches the end of the commodity description, the forward maximum matching process ends and the complete word segmentation sequence is output.

[0022] The reverse maximum matching starts from the last character on the right side of the product description. The initial segmentation length is set to the maximum word length of 10 characters, and 10 consecutive characters from the right side of the product description are used as the current string to be matched. The reverse maximum matching first compares the current string to be matched with the entries in the product domain dictionary. When a completely matching entry is found in the product domain dictionary for the current string to be matched, the current string to be matched is segmented into an independent entry, and the segmentation position is shifted to the starting position of the current string to be matched. When no matching entry is found in the product domain dictionary for the current string to be matched, the backtracking step size is determined according to the frequency distribution of the entry lengths in the product domain dictionary. The calculation method of the backtracking step size is as follows: count the occurrence frequencies of all entry lengths in the product domain dictionary, and use the entry length with the highest occurrence frequency as the preferred backtracking step size value. When the preferred backtracking step size value is greater than 3, the backtracking step size is set to 3. When the preferred backtracking step size value is less than or equal to 3, the backtracking step size is set to the preferred backtracking step size value. Based on the backtracking step size, the corresponding number of characters is intercepted from the left side of the current string to be matched, and the remaining string is re-matched for entries. When no matching entry is found after three consecutive matches using the backtracking step size, the backtracking step size is reduced to 1, and character-by-character matching continues. When the length of the current string to be matched is reduced to 1 character and no matching entry is found in the product domain dictionary, the single character is used as an independent entry, and the segmentation position is shifted to the previous character. After each successful segmentation of an independent entry, the reverse maximum matching re-sets a segmentation window with a maximum word length of 10 characters starting from the segmentation position, and repeats the above matching process for the new current string to be matched. When the segmentation position reaches the first character of the product description, the reverse maximum matching process ends, and the complete word segmentation sequence is output. The word segmentation sequence output by the reverse maximum matching needs to reverse the order of the entries to keep the order consistent with the original text.

[0023] For example, taking the product description "Professional digital single-lens reflex camera" as an example, the relevant entries and their length statistics included in the product domain dictionary are as follows: digital camera (4 characters) appears 85 times, single-lens reflex camera (4 characters) appears 76 times, camera (2 characters) appears 156 times, digital (2 characters) appears 92 times, single-lens reflex (2 characters) appears 89 times, professional level (3 characters) appears 45 times. Based on the frequency distribution of the above entry lengths, it can be seen that the total occurrence frequency of entries with a length of 2 is the highest, and the preferred backtracking step size value is set to 2. Since the preferred backtracking step size value is less than or equal to 3, the actual backtracking step size is set to 2. When the forward maximum matching starts processing from the left side, the process is as follows: First, take the maximum word length of 10 characters, "Professional-level digital single", for matching, and no corresponding entry is found; Adopt a backtracking step size of 2, truncate 2 characters from the right, and use "Professional-level digital" for matching, still no corresponding entry is found; Adopt a backtracking step size of 2 again, and use "Professional-level series" for matching, still no corresponding entry is found; After the third time of adopting a backtracking step size of 2 and still not succeeding in matching, reduce the backtracking step size to 1; Backtrack character by character until "Professional-level" is matched, and split it into independent entries; Re-adopt the maximum word length matching for the remaining text "Digital single-lens reflex camera", and "Digital single-lens reflex" is matched; Finally, match the remaining text "Camera", and obtain the complete word segmentation sequence: "Professional-level / Digital single-lens reflex / Camera"; When the reverse maximum matching starts to process from the right, the process is as follows: First, take the maximum word length of 10 characters, "Digital single-lens reflex camera", for matching, and no corresponding entry is found; Adopt a backtracking step size of 2, truncate 2 characters from the left, and use "Single-lens reflex camera" for matching, and the matching is successful; Re-adopt the maximum word length matching for the remaining text "Professional-level series", and no corresponding entry is found; Adopt a backtracking step size of 2 for matching, and finally match "Professional-level" and "Digital"; Reverse the order of the word segmentation sequence to obtain: "Professional-level / Digital / Single-lens reflex camera".

[0024] Based on the above example, the two-way maximum matching algorithm with the backtracking step size optimization has the following technical effects: Improve word segmentation efficiency: The backtracking step size is dynamically determined based on the dictionary frequency distribution, highly adapted to the dictionary features, reducing the number of invalid matching times and improving the execution efficiency of the word segmentation process; when the word length distribution in the product description is similar to the word length distribution in the dictionary, the backtracking step size optimization can reduce the number of matching times by more than 40%; Maintain word segmentation accuracy: The backtracking step size sets a maximum limit and a decay mechanism, and automatically degrades to character-by-character matching after three consecutive matching failures, ensuring that the accuracy of the word segmentation result is not affected while improving the efficiency; after testing, the word segmentation accuracy rate after adding the backtracking step size optimization is consistent with the traditional algorithm, both reaching more than 95%; Reduce system resource consumption: The backtracking step size optimization reduces the number of string truncation and comparison operations, reducing the memory occupancy and CPU usage rate during the processing; when processing long texts, the resource occupancy can be reduced by more than 30%; Improve system scalability: The calculation method of the backtracking step size can be adjusted according to the dictionary features of different languages, with good scalability. For character-based languages such as English, the backtracking step size can be associated with the average word length in the dictionary. For ideographic languages such as Chinese, the backtracking step size can be associated with the common word length in the dictionary.

[0025] On the other hand, the commodity domain dictionary is stored and organized using a two-way index structure, including a forward index table and a reverse index table. The forward index table creates a first-level index according to the first character of the entry, classifies entries with the same first character into the same index entry, creates a second-level index according to the entry length under each first-level index entry, and the entries in the second-level index are arranged in descending order of length. The reverse index table creates a first-level index according to the last character of the entry, classifies entries with the same last character into the same index entry, creates a second-level index according to the entry length under each first-level index entry, and the entries in the second-level index are arranged in descending order of length. The commodity domain dictionary also includes a statistical table of entry length frequencies, which records the distribution characteristics of entry lengths and is used to calculate the backtracking step size. When performing the forward maximum matching, the forward index table is used to search for entries. First, the corresponding first-level index entry is located according to the first character of the current string to be matched, and then the match is performed in the second-level index according to the length of the string to be matched. When performing the reverse maximum matching, the reverse index table is used to search for entries. First, the corresponding first-level index entry is located according to the last character of the current string to be matched, and then the match is performed in the second-level index according to the length of the string to be matched. The commodity domain dictionary also sets a language identifier for different languages. When performing entry matching, the search is preferentially conducted within the scope of entries in the corresponding language. When no matching entry is found within the preferred scope, the search is extended to the entire language scope.

[0026] The entry length frequency statistics table includes four statistical dimensions: length value range, number of entries, frequency proportion, and cumulative frequency. Among them, the length value range records the possible lengths of the entries, with a value range from 1 to the maximum word length of 10. The number of entries records the number of entries corresponding to each length value range. The frequency proportion is the number of entries in this length value range divided by the total number of entries. The cumulative frequency is the sum of the frequency proportions of this length value range and all length value ranges less than this length value range. The entry length frequency statistics table dynamically maintains the entry length distribution data of the most recent 1 million successful matches. When calculating the backtracking step length based on the entry length frequency statistics table, first obtain the length of the current string to be matched, and use this length as the initial search position. When the frequency proportion corresponding to the initial search position is greater than 30%, use this length value range as the preferred backtracking step length. When the frequency proportion corresponding to the initial search position is less than or equal to 30%, search downward for the nearest length value range with a frequency proportion greater than 30%, and use this length value range as the preferred backtracking step length. When the difference between the cumulative frequency of the preferred backtracking step length and the cumulative frequency of its previous length value range is greater than 50%, increase the preferred backtracking step length by 1. When no length value range with a frequency proportion greater than 30% is found, set the backtracking step length to 3. When the preferred backtracking step length is greater than 3, set the final backtracking step length to 3. When the preferred backtracking step length is less than or equal to 3, set the final backtracking step length to the value of the preferred backtracking step length. After the entry length frequency statistics table processes 100,000 matches, it performs an update calculation on the frequency distribution to ensure the dynamic adaptability of the backtracking step length.

[0027] When the current string to be matched is not found in the commodity domain dictionary, enter the backtracking step length dynamic adjustment process. The backtracking step length dynamic adjustment process includes two control parameters: a continuous match counter and an adjustment coefficient. The continuous match counter is used to record the number of consecutive match failures at the current backtracking step length, with an initial value of 0. The adjustment coefficient is used to calculate the next backtracking step length, with an initial value of 1.0. After each match failure, increment the continuous match counter by 1. When the continuous match counter reaches 3, calculate the new backtracking step length based on the adjustment coefficient: if the current backtracking step length is greater than 1, multiply the adjustment coefficient by 0.7 and set the new backtracking step length to the integer value obtained by multiplying the current backtracking step length by the adjustment coefficient. If the calculated new backtracking step length is less than 1, set the new backtracking step length to 1. When using the new backtracking step length for matching, reset the continuous match counter to 0. When the backtracking step length is reduced to 1 and no matching entry is still found, enter the single-character segmentation mode. If a match is successful for the first time after using the new backtracking step length, multiply the adjustment coefficient by 1.2 for subsequent calculation of the backtracking step length, but ensure that the adjusted backtracking step length does not exceed the initial backtracking step length. The value range of the adjustment coefficient is restricted between 0.3 and 1.2, and take the boundary value when exceeding the range.

[0028] The single-character segmentation mode is activated when no matching entry is found after the backtracking step size is reduced to 1, and it includes two processing stages: character type judgment and combined segmentation. In the character type judgment stage, the current character to be processed is first classified into four types: numeric, alphabetic, punctuation, and other. When a numeric character is detected, it is judged whether its adjacent character is also numeric. If the adjacent character is numeric, the consecutive numeric characters are combined into a numeric entry. When an alphabetic character is detected, it is judged whether its adjacent character is also alphabetic. If the adjacent character is alphabetic, the consecutive alphabetic characters are combined into an alphabetic entry. When a punctuation character is detected, it is separately segmented into a punctuation entry. When an other character is detected, it enters the combined segmentation stage. In the combined segmentation stage, while performing single-character segmentation, the language attribute of the segmented characters is recorded. When the continuously segmented characters belong to the same language, these characters are combined into an entry. The single-character segmentation mode also includes post-processing rules: when there is an isolated entry with a length of 1 in the segmentation result, the language attributes of its adjacent entries on the left and right are judged. If it is the same as the language attribute of any adjacent entry, it is merged with the corresponding adjacent entry. When there are consecutive single-character entries in the segmentation result, it is judged whether these entries conform to the common combination patterns in the commodity domain dictionary. If they conform, they are combined into an entry.

[0029] Preferably, by constructing a commodity domain dictionary with a two-way index structure, fast positioning of entry lookup is achieved, and the time complexity of entry matching is reduced from O(n) to O(log n), where n is the total number of entries in the dictionary. Secondly, based on the entry length frequency statistics table, the backtracking step size is dynamically calculated, enabling the word segmentation process to adaptively adjust the segmentation strategy. Compared with the traditional algorithm with a fixed step size, the average number of matching times can be reduced by 40%. Thirdly, through the dynamic adjustment mechanism of the backtracking step size, when consecutive matching fails, it quickly converges to a suitable segmentation position, avoiding the frequent character-by-character backtracking in the traditional algorithm and improving the word segmentation efficiency. Fourthly, by introducing character type judgment and combined segmentation strategies in the single-character segmentation mode, the word segmentation result is more in line with the semantic characteristics of commodity descriptions, improving the accuracy of word segmentation.

[0030] In addition, the two-way index structure of the commodity domain dictionary supports parallel processing, and forward matching and reverse matching can be performed simultaneously, improving the system throughput. The entry length frequency statistics table can automatically adapt to the entry distribution characteristics of different languages and different commodity categories through a dynamic update mechanism, and has strong domain migration ability.

[0031] S3: Group the word segmentation results in the word mapping library according to the language tags, calculate the semantic similarity of the words between different languages using cosine similarity. When the semantic similarity is greater than a preset threshold, establish a cross-language correspondence relationship for the words with the semantic similarity; The calculation of the semantic similarity of the words includes two stages: word vector construction and similarity calculation. In the word vector construction stage, first group the words in the word mapping library by language, and classify the words with the same language tag into the same language group. For each language group, construct a term frequency-inverse document frequency matrix based on the search keywords of this language. The term frequency counts the number of times the word appears in the corresponding search keywords, and the inverse document frequency counts the logarithm of the reciprocal of the ratio of the number of search keywords containing the word to the total number of keywords. When the number of words in the word mapping library changes, recalculate the term frequency-inverse document frequency matrix. Use the term frequency-inverse document frequency vector of each word as the feature vector of this word. In the similarity calculation stage, for two words in different languages whose similarity needs to be calculated, extract their respective feature vectors. When the dimensions of the two feature vectors are inconsistent, expand the feature vector with the lower dimension to the same dimension as the higher dimension vector by zero padding. Calculate the dot product of the two feature vectors, and divide the dot product result by the product of the norms of the two vectors to obtain the cosine similarity value. When the cosine similarity value is greater than the preset threshold of 0.8, mark these two words as a pair of semantically similar words. When a word finds multiple similar words in the target language, select the word with the largest cosine similarity as its corresponding semantically similar word. If the frequency of the word to be calculated appears too low in the search keywords of the target language, set the values of each dimension of the feature vector of this word to the mean value of the values of each dimension of the feature vectors of the words in this language to ensure the reliability of the calculation.

[0032] The construction process of the term frequency-inverse document frequency matrix includes three stages: term frequency statistics, document frequency calculation, and matrix generation. In the term frequency statistics stage, all search keywords in the same language are grouped according to the entry mapping relationship, and each search keyword is regarded as an independent document. When performing term frequency statistics on a certain entry, all search keywords in this language are traversed, and the number of times this entry appears in each search keyword is counted. The obtained number is the term frequency value of this entry corresponding to this search keyword. When the same entry appears multiple times in the same search keyword, these occurrence times are accumulated as the term frequency value of this entry. In the document frequency calculation stage, the distribution of each entry in all search keywords is counted, the number of search keywords containing this entry is calculated, and this number is divided by the total number of search keywords in this language to obtain the document frequency. After calculating the document frequency, take its reciprocal and perform natural logarithm operation to obtain the inverse document frequency. In the matrix generation stage, the entries are used as row indices, and the search keywords are used as column indices to construct a term frequency matrix. After generating the term frequency matrix, multiply the term frequency value of each entry by its corresponding inverse document frequency to obtain the term frequency-inverse document frequency value. If a certain entry does not appear in a specific search keyword, the term frequency-inverse document frequency value at this position is set to 0. After the term frequency-inverse document frequency matrix is generated, normalize each non-zero element in the matrix by dividing each element by the square root of the sum of the squares of the elements in this row. If all elements in a certain row are 0, then all elements in this row are set to the average term frequency-inverse document frequency value of the entries in this language.

[0033] The cross - language correspondence establishment process includes three stages: similarity screening, transfer verification, and relationship storage. In the similarity screening stage, the cosine similarity results between calculated entries are first obtained. When the cosine similarity value between two entries in different languages is greater than the preset threshold of 0.8, these two entries form a candidate pair. In the transfer verification stage, two - way verification is performed on each candidate pair. When it is verified that entry a in language A and entry b in language B form a candidate pair, it is checked whether the most similar entry of entry b in language A is entry a. When the two - way verification passes, the candidate pair is marked as a valid pair. In the relationship storage stage, a graph structure is used to store the correspondence between entries. Entries in different languages are used as nodes of the graph, and the correspondence between valid pairs is used as the edges of the graph. When adding a new valid pair to the graph, it is first checked whether these two entries already exist in the graph. If they already exist, only the corresponding relationship edge is added. If they do not exist, the entry nodes are added first and then the corresponding relationship edge is added. When a certain entry node in the graph is connected to multiple entry nodes in other languages, the similarity values corresponding to these connection edges are recorded. If, when adding a new corresponding relationship edge, it is found that the entry node already has a connection edge in the same language, the similarity values of the new and old edges are compared, and the edge with the higher similarity value is retained. When the construction of the graph is completed, the out - edges of each entry node are sorted and stored in descending order of similarity value. If there are isolated entry nodes in the graph, these nodes are marked as nodes to be supplemented, and the corresponding relationships of these nodes are preferentially processed in the subsequent entry mapping process.

[0034] The two-way verification process includes three stages: forward verification, reverse verification, and consistency confirmation. In the forward verification stage, candidate word pairs composed of the source entry a in language A and the target entry b in language B are first obtained. All similar entries of entry a in language B are extracted and sorted in descending order of similarity to form a sequence of similar entries. When entry b ranks first in the sequence of similar entries, the candidate word pair is marked as passing the forward verification. When entry b ranks second in the sequence of similar entries and the similarity difference from the entry ranking first is less than 0.05, the candidate word pair is also marked as passing the forward verification. In the reverse verification stage, entry b is used as the source entry, and all its similar entries in language A are extracted and sorted in descending order of similarity to form a sequence of similar entries. When entry a ranks first in the sequence of similar entries, the candidate word pair is marked as passing the reverse verification. When entry a ranks second in the sequence of similar entries and the similarity difference from the entry ranking first is less than 0.05, the candidate word pair is also marked as passing the reverse verification. In the consistency confirmation stage, the candidate word pairs that pass both the forward verification and the reverse verification are finally confirmed. When a candidate word pair passes both the forward verification and the reverse verification, the word pair is marked as a strong correspondence relationship. When a candidate word pair passes only one-way verification, the word pair is marked as a weak correspondence relationship. If the source entry forms weak correspondence relationships with multiple target entries, the word pair with the highest similarity is selected and retained, and the remaining word pairs are marked as relationships to be observed. When the similarity of a candidate word pair remains stable after three consecutive updates of transaction data, the weak correspondence relationship is upgraded to a strong correspondence relationship.

[0035] S4: Receive the search request of the user, search for semantically similar entries in the cross-language correspondence relationship according to the keywords in the search request, expand the search scope based on the entries, and output cross-language product matching results.

[0036] First, extract the keyword set in the search request and process each keyword one by one. For each keyword, determine its belonging language. If the language of the keyword is not marked, infer its possible language based on the character features of the keyword (such as letters, numbers, or special characters). When the language inference of the keyword fails, mark the keyword as an unknown language and skip the subsequent processing of the keyword.

[0037] When the language of the keyword is determined, find the node corresponding to the keyword in the graph structure of the cross-language correspondence relationship. If there is no node in the graph that exactly matches the keyword, determine whether the keyword is an isolated node. If it is an isolated node, return an empty result and mark that the keyword has no semantically similar entries. If it is not an isolated node, proceed to the next step.

[0038] When a node corresponding to a keyword is found in the graph, all outgoing edges of the node are extracted and sorted from high to low according to the similarity value; the target node is searched from the outgoing edges in turn. If the language of the target node is consistent with the target language of the current search request, the entry corresponding to the target node is marked as a semantically similar entry; if the target language is not clear, the entries corresponding to all target nodes are returned as a set of semantically similar entries.

[0039] When the out-edge of a keyword is empty or the similarity values ​​of all target nodes are lower than a preset threshold (for example, 0.8), the keyword is marked as having no semantically similar terms in the target language; if there are multiple similarity values ​​of the target nodes that exceed the preset threshold, the term corresponding to the target node with the highest similarity is selected; if the similarity values ​​are the same, the target terms are sorted in descending order according to their out-degree values ​​(number of connected edges) in the graph, and the target terms with higher out-degree values ​​are selected as semantically similar terms.

[0040] If the keyword does not match the language of multiple target nodes, it is marked as a cross-language mapping failure and enters subsequent processing; in the subsequent processing stage, check whether the keyword belongs to the node to be supplemented in the graph. If so, try to establish a mapping relationship using common cross-language combination patterns (such as common translation pairs in dictionaries); if the mapping relationship is successfully established, add new edges and update the graph structure, and add similar entries of the keyword to the result set; if the mapping fails, return an empty result and mark the keyword as having no match.

[0041] Finally, for each keyword search result, the successfully matched terms and their languages ​​are grouped and a set of semantically similar terms is output. If the request contains target requirements in multiple languages, the above search logic is executed for each target language separately, and the result sets are merged to ensure that each target language has corresponding semantically similar terms.

[0042] For example, taking the user search keyword "mobile phone" as an example, the relevant entries and their similarity statistics contained in the cross-language correspondence diagram are as follows: the English entry "mobile phone" corresponding to mobile phone (Chinese) has a similarity of 0.95; the corresponding German entry "Handy" has a similarity of 0.92; the corresponding French entry "téléphone portable" has a similarity of 0.85; based on the above similarity distribution, it can be seen that the similarities of the corresponding entries in English and German are both higher than the preset threshold of 0.8, and the similarity of the French entry is lower than the threshold, so only the entries in English and German are selected as the extended keyword set.

[0043] When the extended keyword set is ["手机" (Chinese), "mobile phone" (English), "Handy" (German)], the process is as follows: First, search for the original keyword "mobile phone" in the Chinese product database, and the following products are matched: Product A: Huawei P50 mobile phone; Product B: Xiaomi 11 mobile phone; Product C: iPhone 14 mobile phone; Then, search for the extended keyword "mobile phone" in the English product database, and the following products are matched: Product D: Samsung Galaxy S21 mobile phone; Product E: iPhone 14 mobile phone; Next, search for the extended keyword "Handy" in the German product database, and the following products are matched: Product F: Samsung Galaxy S21 Handy; Product G: iPhone 14 Handy; Deduplicate all the matching results: Product C, Product E, and Product G all represent the same product iPhone 14, but from different language databases; retain the matching result of Product E according to the keyword "mobile phone" with the highest similarity (similarity 0.95), and deduplicate the rest.

[0044] Product D and Product F both represent the same product Samsung Galaxy S21, but from different language databases; retain the matching result of Product F according to the keyword "Handy" with the highest similarity (similarity 0.92), and deduplicate the rest.

[0045] Finally, the merged and deduplicated cross-language product matching results are: Product A: Huawei P50 mobile phone (Chinese); Product B: Xiaomi 11 mobile phone (Chinese); Product E: iPhone 14 mobile phone (English, from the keyword "mobile phone" with the highest similarity); Product F: Samsung Galaxy S21 Handy (German, from the keyword "Handy" with the highest similarity).

[0046] The same applies to the above search keywords in other languages. Similarly, based on the cross-language correspondence graph, the search scope is extended. First, search for matching results in product databases of different languages, and then process the search results according to the deduplication and multi-language merging logic based on similarity. Finally, a cross-language product matching list is generated.

[0047] In summary, the data exchange method for a commodity trading platform according to an embodiment of the present invention is elucidated. It performs word segmentation on commodity descriptions through a bidirectional maximum matching algorithm, groups them according to languages, establishes cross-language correspondence relationships for the entries with the semantic similarity, and then searches for all cross-language commodity matching results based on the keywords in the user search request to promote transactions. In this way, it can solve the technical problem in a cross-border e-commerce platform that due to users in different countries using different languages to describe the same type of commodity, the search results are incomplete, reducing the transaction success rate.

[0048] Here, those skilled in the art can understand that the specific operations of each step in the above data exchange method for a commodity trading platform have been introduced in detail in the description of the Figures 1 to 2 data exchange method for a commodity trading platform above, and therefore, the repeated description thereof will be omitted.

[0049] In summary, the data exchange method for a commodity trading platform according to an embodiment of the present invention is elucidated. It performs word segmentation on commodity descriptions through a bidirectional maximum matching algorithm, groups them according to languages, establishes cross-language correspondence relationships for the entries with the semantic similarity, and then searches for all cross-language commodity matching results based on the keywords in the user search request to promote transactions. In this way, it can solve the technical problem in a cross-border e-commerce platform that due to users in different countries using different languages to describe the same type of commodity, the search results are incomplete, reducing the transaction success rate.

Claims

1. A commodity trading platform data exchange method, characterized in that: include: Obtain multilingual transaction data from commodity trading platforms; Based on the multilingual transaction data, segment the product description using a bidirectional maximum matching algorithm, and filter the segmentation results of successful transactions according to the transaction identifier to generate a term mapping library; The word segmentation results in the term mapping library are grouped according to the language labels, and the semantic similarity of terms between different languages ​​is calculated using cosine similarity, and when the semantic similarity is greater than a preset threshold, a cross-language correspondence is established for the terms with the semantic similarity; A search request from a user is received, semantically similar terms are searched in the cross-language correspondence according to keywords in the search request, a search scope is expanded based on the terms, and cross-language commodity matching results are output.

2. The commodity trading platform data exchange method according to claim 1, characterized in that: The bidirectional maximum matching algorithm includes two sub-processes: forward maximum matching and reverse maximum matching; When the number of word segmentation results of the forward maximum match and the reverse maximum match is inconsistent, the result with a smaller number of word segmentations is selected as the final word segmentation result.

3. The commodity trading platform data exchange method according to claim 2, characterized in that: Match the segmented result with a preset commodity domain dictionary, and when the segmented result has a corresponding entry in the commodity domain dictionary, mark the segmented result as a valid entry; Valid entries in the transaction record whose transaction identifier is a value of 1 are screened, the valid entries, the corresponding language tags, and the search keywords are combined into an entry mapping unit, and the entry mapping unit is stored in the entry mapping library.

4. The commodity trading platform data exchange method according to claim 2, characterized in that: The forward maximum match and the reverse maximum match compare the current string to be matched with the entries in the commodity field dictionary from the left or right side of the text respectively; if the match is successful, it is split into independent entries and shifted right, otherwise the backtracking step is dynamically adjusted according to the dictionary word length frequency to gradually shorten the match until the match is successful or the text length is reduced to a single character, and finally a complete word segmentation sequence is output.

5. The commodity trading platform data exchange method according to claim 4, characterized in that: The dynamic adjustment of the backtracking step length includes two control parameters: a continuous matching counter and an adjustment coefficient; A new backtracking step is determined according to the continuous matching counter and the adjustment coefficient. If a matching entry is still not found, the single character segmentation mode is entered.

6. The commodity trading platform data exchange method according to claim 5, characterized in that: The single-character segmentation mode divides characters into numbers, letters, punctuation and other categories by character type judgment, and combines and segments them according to adjacent character types, and combines consecutive characters belonging to the same language into one entry. At the same time, post-processing rules are applied to optimize the segmentation results, and finally output word segmentation results that better meet field requirements.

7. The commodity trading platform data exchange method according to claim 1, characterized in that: The semantic similarity calculation of the term includes two stages: word vector construction and similarity calculation; The word vector construction stage classifies the terms in the term mapping library by language grouping, and constructs the term feature vector based on the term frequency-inverse document frequency matrix; dynamically updates the matrix when the number of terms changes; The similarity calculation extracts the feature vectors of the terms to be compared, and if the vector dimensions are inconsistent, they are aligned by zero padding, and the cosine similarity formula is used to calculate the similarity value of the two terms.

8. The commodity trading platform data exchange method according to claim 1, characterized in that: The establishment of cross-language correspondence includes three stages: similarity screening, transfer verification and relationship storage; The transfer verification ensures that the term a of language A and the term b of language B are the most similar terms to each other through two-way verification, and are marked as valid word pairs after passing the verification.

9. The commodity trading platform data exchange method according to claim 9, characterized in that: The two-way verification includes forward verification, reverse verification and consistency confirmation; The forward verification starts from the source term a and checks whether the target term b ranks first or second in its similar term sequence and whether the similarity difference with the first term is less than a preset value; The reverse verification uses the target term b as the source term to check whether it satisfies similar conditions in a similar term sequence in the source language; The consistency confirmation performs final confirmation on the candidate word pairs that have passed the forward and reverse verification, and marks the two-way pass as a strong correspondence, and the one-way pass as a weak correspondence; if the source term forms a weak correspondence with multiple target terms, the word pair with the highest similarity is retained, and the rest are marked as relationships to be observed; the weak correspondence can be upgraded to a strong correspondence when the similarity is stable after three consecutive transaction data updates.

10. The commodity trading platform data exchange method according to claim 1, characterized in that: Searching for semantically similar terms in the cross-language correspondence includes: Extract keywords from the search request and determine their language. If the language is not marked, infer the language through character features. If the inference fails, mark it as an unknown language and skip processing. For the keywords determined by the language, the corresponding nodes are searched from the cross-language graph structure. If there is no matching node, it is determined whether it is an isolated node. An empty result is returned for an isolated node, and edges are extracted from non-isolated nodes and sorted by similarity. Semantically similar terms are filtered out according to the target language. If the target language is not clear, all target node terms that meet the threshold are returned. When the outgoing edge is empty or the similarity is lower than the threshold, it is marked as no similar terms; if multiple target nodes meet the conditions, the best term is selected according to the priority of similarity and outgoing value; If there is no match between the keyword and the target language, try to use the cross-language combination mode to establish a mapping relationship; if the mapping is successful, update the graph structure and add it to the result set; if it fails, return an empty result; The search results are grouped and output by terms and languages. If there is more than one target language, the above logic is executed separately and the results are merged to ensure that all language requirements are covered.