Commodity data matching method, electronic equipment and storage medium
By acquiring text information input by users, determining its reasonableness and question tendencies, and generating first and second type document sets, the problem of low keyword search accuracy in e-commerce platforms is solved, and more accurate product data matching is achieved.
Patent Information
- Application Number
- CN202410567808.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-09
- Publication Date
- 2025-11-14
AI Technical Summary
Existing e-commerce platforms rely solely on keyword searches for products, resulting in search results that do not match users' actual needs and have low accuracy.
By acquiring text information input by users, determining its reasonableness and question tendency, generating a first type of document set using preset matching thresholds and document category identifiers, and then matching a second type of document set in product data documents by combining keywords in the text information, finally merging to generate product data that matches user needs.
It enables comprehensive matching of product data based on user needs, improves the accuracy of retrieval, and ensures that search results better match users' actual needs.
Smart Images

Figure CN120952898A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of product search technology, and in particular to a product data matching method, electronic device, and storage medium. Background Technology
[0002] With the rapid development of the internet and the ever-growing demand for online shopping, e-commerce has experienced rapid growth. When purchasing goods on an e-commerce platform, users need to search for product information to generate a purchase desire.
[0003] However, current e-commerce platforms all use keyword search for product searches. Users can only submit keywords, and the system matches these keywords with its database using a search engine, then returns related products to the user. Because it only supports keyword searches, the search results may not match the user's actual needs. Therefore, how to accurately search for related product information according to user needs has become a pressing problem to be solved. Summary of the Invention
[0004] This invention provides a product data matching method, electronic device, and storage medium to solve the problem that in the prior art, searching for products based solely on keywords results in a product matching structure that does not match the user's actual needs and has low accuracy.
[0005] According to one aspect of the present invention, a method for matching commodity data is provided, wherein the method includes:
[0006] Obtain text information input by the user, and determine the reasonableness of the text information and the questioning tendency;
[0007] The document category identifier to be matched is determined based on the reasonableness and the questioning tendency.
[0008] Based on the document category identifier to be matched and the preset matching threshold, the product data document associated with the text information is determined as the first type of document, and a first product data document set is generated;
[0009] Based on the keywords contained in the text information, a second set of product data documents is generated by matching the second type of documents associated with the text information in all product data documents;
[0010] The first type of document in the first product data document set and the second type of document in the second product data document set are merged to obtain product data that matches the text information.
[0011] According to another aspect of the present invention, a product data matching device is provided, wherein the device comprises:
[0012] The text acquisition module is used to acquire text information input by the user and determine the reasonableness of the text information and the questioning tendency.
[0013] The identifier determination module is used to determine the document category identifier to be matched based on the reasonableness and the questioning tendency.
[0014] The first set determination module is used to determine the product data document associated with the text information as the first type document based on the document category identifier to be matched and the preset matching threshold, and generate the first product data document set;
[0015] The second set determination module is used to match second type documents associated with the text information in all product data documents based on the keywords contained in the text information to generate a second product data document set;
[0016] The product data determination module is used to merge the first type of document in the first product data document set and the second type of document in the second product data document set to obtain product data that matches the text information.
[0017] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0018] At least one processor; and
[0019] A memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the product data matching method according to any embodiment of the present invention.
[0021] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the commodity data matching method according to any embodiment of the present invention.
[0022] The technical solution of this invention involves acquiring user-input text information, determining the reasonableness and question tendency of the text information, determining the document category identifier to be matched based on the reasonableness and question tendency, identifying product data documents associated with the text information as first-type documents based on the document category identifier and a preset matching threshold, generating a first product data document set, matching second-type documents associated with the text information in all product data documents based on keywords contained in the text information, generating a second product data document set, and merging the first-type documents in the first product data document set and the second-type documents in the second product data document set to obtain product data matching the text information. This achieves comprehensive detection and matching of product data based on text information and keywords, accurately searching for product-related information according to user needs, and improving the accuracy of product data retrieval.
[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a flowchart of a product data matching method provided in Embodiment 1 of the present invention;
[0026] Figure 2 This is a flowchart of another product data matching method provided in Embodiment 2 of the present invention;
[0027] Figure 3 This is a schematic diagram of the structure of a product data matching device according to Embodiment 4 of the present invention;
[0028] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the product data matching method of this invention. Detailed Implementation
[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0031] Example 1
[0032] Figure 1 This is a flowchart of a product data matching method according to Embodiment 1 of the present invention. This embodiment is applicable to situations where product data is matched according to user needs on an e-commerce platform. The method can be executed by a product data matching device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0033] S110. Obtain the text information input by the user and determine the reasonableness of the text information and the questioning tendency.
[0034] Here, textual information refers to information content presented in written form, which may include, but is not limited to, natural language statements and keywords. Reasonableness can be understood as the degree of logical and legal compliance of the textual information, serving as an objective evaluation standard. For example, if the textual information contains illegal words, its reasonableness can be considered unreasonable; or, if the word segmentation phrases in the textual information do not fall within the scope of the product's business, its reasonableness can be considered unreasonable; or, if the textual information contains syntactic errors or semantic contradictions, its reasonableness can be considered unreasonable; or, if the textual information contains prohibited words, its reasonableness can be considered unreasonable. Question tendency can be understood as the category of the user's question. For example, the categories of question tendency include at least: product-oriented inquiries, product recommendation inquiries, and product application method inquiries.
[0035] In this embodiment, user-input text information can be extracted, and semantic analysis can be performed to determine the questioning tendency of the text information. In actual operation, a pre-trained Natural Language Processing (NLP) model can be used to perform semantic analysis on the text information to determine the questioning tendency; alternatively, a neural network model can be used for semantic understanding to determine the questioning tendency; or, the text information can be segmented to determine the context or grammatical structure of the segmented phrases and thus the questioning tendency. When identifying phrases in the text information, it is determined whether each phrase conforms to grammatical rules and whether any phrases fall within the scope of the goods being sold. If each phrase conforms to grammatical rules and a phrase falls within the scope of the goods being sold, the text information is considered reasonable; otherwise, it is considered unreasonable. In a specific implementation, the categories of goods within the scope of the goods being sold can be pre-defined. For example, pens and rulers belong to stationery. When the text information contains "stationery," the phrase can be considered to fall within the scope of the goods being sold. Simultaneously, the semantics of the phrase can be determined. If a phrase with the same semantic meaning exists within the scope of business operations, the phrase can be considered to fall within the scope of business operations. In one embodiment, the phrase can be matched against a preset prohibited word database. If it is determined that the phrase does not belong to the preset prohibited word database, the legality of the text information is determined. If the text information is determined to be legal, and in another embodiment, the segmented phrases can be matched against a preset prohibited word database, if it is determined that the segmented phrases do not belong to the preset prohibited word database, the legality of the text information is determined to be legal. In one embodiment, when the text information is a natural language statement, the natural language statement can be segmented, and the reasonableness of the segmented phrases can be determined; when the text information is a keyword, the reasonableness can be directly determined through the keyword.
[0036] S120. Determine the document category identifier to be matched based on the reasonableness and question tendency.
[0037] The document category identifier to be matched can be understood as an identifier that identifies the document category. Documents stored on e-commerce platforms may belong to different categories, such as product introduction documents, application scenario-related documents, and product description documents. For each category of documents, the same document category identifier can be assigned to documents of the same category. Furthermore, the category of the question can be associated with the document category identifier. For example, the document category identifier to be matched may include, but is not limited to, product introduction document identifiers, application scenario-related document identifiers, and product description document identifiers.
[0038] In this embodiment, the existence of a document category identifier to be matched can be determined first based on its reasonableness, and then the document category identifier to be matched can be determined again based on the question's tendency. In actual operation, if the reasonableness is deemed unreasonable, it can be assumed that the text information does not fall within the scope of the product's business, meaning that the corresponding product information cannot be found on the e-commerce platform. In this case, it can be determined that there is no document category identifier matching the text information. The identifier information associated with the question's tendency category can be pre-configured and stored. When the reasonableness is determined to be reasonable, the identifier information associated with the question's tendency category can be extracted as the document category identifier to be matched.
[0039] S130. Based on the document category identifier to be matched and the preset matching threshold, determine the product data document associated with the text information as the first type of document, and generate the first product data document set.
[0040] The preset matching threshold can be understood as a critical value for determining whether a product data document is associated with text information. When the matching value between the text information and the product data document is greater than or equal to the preset matching threshold, the product data document can be considered as associated with the text information; when the matching value between the text information and the product data document is less than the preset matching threshold, the product data document can be considered as associated with the text information. The first type of document refers to the product data document associated with text information.
[0041] In this embodiment, product data documents with document category identifiers to be matched can be extracted as product data documents to be matched. The matching values between the text information and the product data documents to be matched are determined, and product data documents with matching values greater than or equal to a preset matching threshold are designated as first-type documents. In actual operation, the text information can be vectorized to generate text information vectors, and the product data documents to be matched can be vectorized to generate document vectors. The cosine similarity between the text information vectors and each document vector is determined, and product data documents with cosine similarity greater than or equal to a preset matching threshold are designated as first-type documents. These first-type documents are then stored as a first set of product data documents.
[0042] S140. Based on the keywords contained in the text information, match the second type of documents associated with the text information in all product data documents to generate a second product data document set.
[0043] In this embodiment, when the text information is a natural language statement, keywords contained in the text information can be extracted first. For example, the text information can be segmented to obtain segmented word groups, the semantic information of the segmented word groups can be determined, and segmented word groups existing in a preset thesaurus, as well as word groups existing in the preset thesaurus with the same semantic information as the segmented word groups, can be combined as keywords of the text information. Then, keywords are sequentially searched in all product data documents, and product data documents containing keywords are treated as second-type documents associated with the text information, generating a second set of product data documents. In one embodiment, keywords may include product names.
[0044] S150. Merge the first type of documents in the first product data document set and the second type of documents in the second product data document set to obtain product data that matches the text information.
[0045] In this embodiment, first-type documents from a first set of product data documents and second-type documents from a second set of product data documents can be extracted. These first-type and second-type documents are then merged to form product data that matches the text information. In practice, duplicate product data documents may exist in the first and second-type documents. Duplicate first-type documents can be deleted, or only one document of the same type can be retained. Simultaneously, the first and second-type documents can be sorted. Since documents found through keyword matching may have higher matching values, all second-type documents can be sorted before all first-type documents. Furthermore, since there may be a large number of second-type documents matched through keyword matching, they can be sorted according to the frequency of keyword occurrences, retaining only the top-ranking documents. Additionally, the first-type documents can be sorted and output according to their matching values, allowing users to more easily obtain the product data that matches the text information they require.
[0046] In one embodiment, the weights of keywords in the first type of document and the second product data document can be determined, and the first and second type of documents can be arranged according to the keyword weights. Meanwhile, the product data matching the text information can be a single product data document or multiple product data documents. In actual operation, paragraphs containing keywords are extracted from the first type of document and the second product data document, and the paragraphs are arranged in descending order of weight. These paragraphs are then concatenated to form the product data matching the text information. Alternatively, the first and second type of documents can be arranged in descending order of weight as the product data matching the text information.
[0047] In this embodiment of the invention, by acquiring text information input by the user, the reasonableness and question tendency of the text information are determined. Then, based on the reasonableness and question tendency, a document category identifier to be matched is determined. Based on the document category identifier to be matched and a preset matching threshold, product data documents associated with the text information are determined as first type documents, generating a first product data document set. Based on the keywords contained in the text information, second type documents associated with the text information are matched among all product data documents to generate a second product data document set. The first type documents in the first product data document set and the second type documents in the second product data document set are merged to obtain product data that matches the text information. This achieves comprehensive detection and matching of product data based on text information and keywords, accurately searches for product-related information according to user needs, and improves the accuracy of product data retrieval.
[0048] In one embodiment, before obtaining the text information input by the user and determining the reasonableness of the text information and the question's tendency, the method further includes:
[0049] Build a knowledge base for product data documents and store all product data documents in the knowledge base;
[0050] All product data documents are preprocessed with text, and the product data documents are divided into at least one document category according to the pre-stored question tendency. A document category identifier is then created for each document category.
[0051] In this embodiment, a knowledge base for product data documents can be constructed, storing all product data documents within it. Simultaneously, operations such as word segmentation, stop word removal, stemming, or lemmatization are performed on the product data documents to categorize them into at least one document category based on question tendency, and a document category identifier is established for each category. In one embodiment, product data documents can be further divided into product introduction documents, product recommendation documents, product application method documents, and technical tutorial documents, etc. For example, a massive database of documents such as "User Manuals for Specific XX Products," "Smart Home Popular Science Documents," and "Smart Home Renovation Phase Popular Science Documents" can be constructed to facilitate subsequent matching of product documents. Product application method documents refer to the methods of using the product, such as "The Correct Way to Use a Toothbrush" or "The Usage and Precautions of a Magnetic Stirrer"; technical tutorial documents refer to the specified methods of solutions, such as "How to Properly Perform Oral Care" or "How to Plan a Smart Home Renovation Solution in Advance During the Renovation Phase."
[0052] Example 2
[0053] Figure 2 This is a flowchart of another product data matching method provided by Embodiment 2 of the present invention. This embodiment is a further optimization and extension based on the above embodiments, and can be combined with various optional technical solutions in the above embodiments. Figure 2 As shown, the method includes:
[0054] S201. Perform semantic analysis on the text information to determine the questioning tendency of the text information.
[0055] The categories of questions should include at least: product-specific inquiries, product recommendation inquiries, product application method inquiries, and technical tutorial inquiries.
[0056] In this embodiment, the semantics of the text information can be determined through a pre-trained natural language processing model or by analyzing the grammatical structure of the sentence, and then the questioning tendency of the text information can be determined according to the semantics. Categories of questioning tendency can include product-oriented inquiries, product recommendation inquiries, and product application method inquiries, etc.
[0057] S202. The text information is segmented into word groups, and the word groups are matched with a preset prohibited word library. If it is determined that the word groups do not belong to the preset prohibited word library, the legality of the text information is determined to be legal.
[0058] The preset prohibited word database can be understood as a pre-set database that prohibits searching for certain product-related information.
[0059] In this embodiment, boundary markers can be added between words in the text information to perform word segmentation and obtain word groups. In actual operation, the grammatical structure of the sentence can be analyzed, and the text information can be segmented according to the grammatical structure; alternatively, the forward maximum matching method can be used to determine the word segmentation method for the text information. Each word group can be queried sequentially in a preset prohibited word database. If it is determined that a word group does not belong to the preset prohibited word database, the legality of the text information is determined to be legal; if it is determined that a word group contains a word group from the preset prohibited word database, the legality of the text information is considered to be illegal.
[0060] S203. Determine the semantic information of the segmented word groups, and combine the segmented word groups existing in the preset thesaurus, as well as the word groups existing in the preset thesaurus that have the same semantic information as the segmented word groups, as keywords of the text information.
[0061] The preset vocabulary refers to a collection of pre-set word data.
[0062] In this embodiment, the semantic information of segmented word groups can be determined, and word groups with the same semantic information as the segmented word groups can be determined according to the sentence information. Segmented word groups and word groups with the same semantic information as segmented word groups are searched in a preset thesaurus. All segmented word groups existing in the preset thesaurus, as well as word groups with the same semantic information as segmented word groups existing in the preset thesaurus, are used as keywords in the text information.
[0063] S204. Generate keyword vectors by vectorizing the keywords into text, and determine the word vectors of the phrases in the preset product thesaurus.
[0064] The pre-set product terminology database can be understood as a collection of terms related to the products offered on the platform. The product business scope refers to the range of products offered by the platform.
[0065] In this embodiment, phrases can be extracted from a preset product thesaurus, and the phrases and keywords can be vectorized into text to obtain keyword vectors and phrase vectors. In actual operation, one-hot encoding, Word2Vec encoding, or GloVe encoding can be used to vectorize the phrases and words into text to generate keyword vectors and phrase vectors, respectively.
[0066] S205. Determine the first cosine similarity between the keyword vector and each word group vector. If the first cosine similarity is greater than or equal to the preset first matching threshold, determine that the text information conforms to the scope of business operations.
[0067] The preset first matching threshold refers to a pre-set threshold used to determine whether text information conforms to the scope of business operations. In this embodiment, the first cosine similarity between the keyword vector and each word group vector can be calculated, and the relationship between the first cosine similarity and the preset first matching threshold can be compared. If the first cosine similarity is greater than or equal to the preset first matching threshold, the text information is determined to conform to the scope of business operations.
[0068] S206. When the legality of the keywords is legal and the keywords are within the scope of business operations, the reasonableness of the text information is determined to be reasonable.
[0069] In this embodiment, if the keywords are deemed legal and also fall within the scope of the business, the text information can be considered reasonable.
[0070] S207. When the legality of keywords is illegal or the keywords do not conform to the business scope of the goods, the reasonableness of the text information is determined to be unreasonable.
[0071] In the embodiment, when the reasonableness level is determined to be unreasonable or the keywords do not conform to the business scope of the goods, the reasonableness level of the text information can be considered unreasonable.
[0072] S208. If the reasonableness level is deemed unreasonable, determine that there is no document category identifier that matches the text information.
[0073] In one embodiment, if the reasonableness level is determined to be unreasonable, it can be assumed that there is no product data document matching the text information, i.e., there is no document category identifier matching the text information. In one embodiment, when the reasonableness level is unreasonable, a pre-set text can be displayed to the user, informing the user that a product data document matching the text information cannot be found.
[0074] S209. If the reasonableness level is reasonable, extract the identification information associated with the category of question tendency from the preset configuration information as the category identifier of the document to be matched.
[0075] In this embodiment, the preset configuration information can store the correspondence between question tendency categories and identification information. When the reasonableness level is determined to be reasonable, the preset configuration information can be obtained, and the identification information associated with the question tendency category can be extracted from the preset configuration information. This identification information is then used as the document category identifier to be matched.
[0076] S210. Use the product data document with the document category identifier to be matched as the document to be matched.
[0077] Here, the document to be matched can be understood as a product data document within the text information matching range. For example, when the category identifier of the document to be matched is a product introduction document, the document to be matched can be a product introduction document; when the category identifier of the document to be matched is an application scenario related document, the document to be matched can be an application scenario related document; and when the category identifier of the application scenario related document is a product description document, the document to be matched can be an application scenario related document.
[0078] In this embodiment, the document category identifiers of all product data documents can be extracted, and product data documents with the document category identifier to be matched can be identified as the documents to be matched.
[0079] S211. Vectorize the text information to generate a text information vector, and vectorize the document to be matched to generate a document vector.
[0080] In this embodiment, the text information and the document to be matched can be vectorized using One-hot encoding, Word2Vec encoding, or GloVe encoding to generate text information vectors and document vectors, respectively.
[0081] S212. Determine the second cosine similarity between the text information vector and each document vector.
[0082] In this embodiment, the dot product between the text information vector and each document vector can be calculated separately. The dot product is the sum of the products of corresponding elements of the two vectors, mainly used to measure the similarity between the two vectors in the same direction. Then, the vector length of the text information vector and each document vector is calculated separately, that is, the distance of the vector from the origin to its endpoint. This is mainly used to normalize the vectors and eliminate the influence of vector length on the similarity calculation. Finally, the cosine value of the angle between the text information vector and each document vector is calculated separately, and the similarity value corresponding to the cosine value of the angle is determined as the second cosine similarity.
[0083] S213. The product data documents with a second cosine similarity greater than or equal to a preset second matching threshold are taken as first type documents, and the first type documents are stored as a first product data document set.
[0084] In this embodiment, the relationship between the second cosine similarity and the preset matching threshold can be determined. Product data documents with a second cosine similarity greater than or equal to the preset second matching threshold are designated as first type documents, and the first type documents are merged into a first product data document set.
[0085] S214. Extract keywords from the text information, match the keywords with all product data documents, and obtain the target product data document.
[0086] In this embodiment, keywords in the text information can be identified, keywords can be searched in all product data documents, and product data documents containing keywords can be selected as target product data documents.
[0087] S215. Treat the target product data document as a second type of document, and store the second type of document as a second product data document collection.
[0088] S216. Identify identical first-type and second-type documents, and delete duplicate first-type or second-type documents.
[0089] In this embodiment, the first type of document and the second type of document may contain the same document. The first type of document and the second type of document can be compared, and the same first type of document and the second type of document can be extracted and one of them can be retained to prevent the duplicate output of product data documents.
[0090] S217. Extract the keyword paragraphs to which the keywords belong from the first type of document and the second product data, respectively.
[0091] In this embodiment, all first-type documents in the first product data document set and all second-type documents in the second product data document set can be extracted. The keywords contained in the first-type documents and the second product data documents can be determined, and the paragraphs to which the keywords belong in the first-type documents and the second product data documents can be determined. The paragraphs containing the keywords are designated as keyword paragraphs.
[0092] S218. Determine the preset weights of each keyword in the first type of document and the second product data document, and determine the frequency of each keyword in the keyword paragraph.
[0093] In the embodiments, the frequency of occurrence of keywords in each keyword paragraph can be determined, and the preset weight of each keyword can be extracted.
[0094] S219. Determine the sum of the products of the frequency of each keyword in each keyword paragraph and the corresponding preset weight as the weight sum.
[0095] In this embodiment, for each keyword paragraph, the preset weight of each keyword can be multiplied by the number of times it appears, and the products obtained in each keyword paragraph can be summed as the weight sum.
[0096] S220. Arrange each keyword paragraph according to its weight and in descending order, and then concatenate the keyword paragraphs to form product data that matches the text information.
[0097] In this embodiment, keyword paragraphs can be ordered from highest to lowest weight, and then concatenated to generate a product data document, which serves as the product data matching the text information. In actual operation, the sorting and concatenation of paragraphs can be accomplished by a pre-set NLP model.
[0098] In the embodiments of the invention, by determining the reasonableness of text information based on the legality of keywords and whether it conforms to the scope of business of the goods, it is possible to quickly determine whether the text information can be searched for corresponding product data documents; by vectorizing the text information to generate text information vectors and vectorizing the documents to be matched to generate document vectors, the first type of document is determined according to the cosine similarity between the determined text information vector and each document vector, and the second type of document is searched according to the keywords in the text information, thereby improving the comprehensiveness of product data matching; by extracting the corresponding paragraphs from the first type of document and the second type of document and sorting and outputting them, the reasonableness of product data output is improved, thereby improving the user experience.
[0099] In one embodiment, after concatenating the keyword paragraphs to obtain product data that matches the text information, the method further includes:
[0100] Output questions that are related to product data.
[0101] In this embodiment, after identifying product data that matches the text information, questions associated with the product data can be predicted and generated for output. In actual operation, a preset NLP model can be used to associate related questions based on the content of the product data, allowing users to expand upon the information.
[0102] Example 3
[0103] This embodiment uses an NLP pre-trained model to determine the reasonableness of text information and questioning tendency, and applies it to an e-commerce platform as an example to further explain the matching method for product data.
[0104] Build a knowledge base related to the main products of the e-commerce platform. When building this knowledge base, product data documents can be divided into two categories: product description documents and application scenario-related documents. For example, if the e-commerce platform mainly sells smart home products, it can build a massive amount of product data documents such as "User Manuals for Specific XX Products," "Smart Home Popular Science Documents," and "Popular Science Documents for the Renovation Stage of Building a Smart Home."
[0105] The product data documents undergo text preprocessing, including word segmentation, stop word removal, stemming, and lemmatization. Then, the Doc2Vec model and other methods are used to vectorize the product data documents, generating document vectors. The product data documents are then categorized according to predefined question preferences, and each type of document is labeled with a document category identifier. The main purpose is to create a unique identifier with document feature vectors for each document, facilitating subsequent citation and indexing.
[0106] The process begins by acquiring user-input text information and applying a pre-trained NLP model to perform semantic analysis. This analysis determines the text's reasonableness, input type, and question tendency. For example, it considers whether the search term aligns with the mall's business scope, whether it's a keyword or natural language input, and whether the question tendency is product recommendation or application method. Based on reasonableness and question tendency, document category identifiers are determined, and the corresponding documents are extracted to identify their document vectors. The text information is then vectorized to generate text information vectors. The dot product between each text information vector and each document vector is calculated. The dot product is the sum of the products of corresponding elements of two vectors and primarily measures the similarity between two vectors in the same direction. Next, the vector lengths of the text information vector and each document vector are calculated (distance from the origin to the destination). This normalizes the vectors and eliminates the influence of vector length on similarity calculations. Finally, the cosine of the angle between the text information vector and each document vector is calculated, and the cosine similarity is determined as the cosine similarity. Product data documents with a cosine similarity greater than or equal to a preset matching threshold are classified as first-type documents. Keywords are extracted from the text information, and second-type documents associated with the text information are matched across all product data documents based on these keywords.
[0107] Determine keyword weights, prioritize second-type documents with higher keyword weights, and then prioritize all second-type documents at the top of all first-type documents. Sort all second-type product data documents in the second product data document set according to the frequency of keyword occurrences from highest to lowest, retaining a predetermined number of second-type documents with the highest keyword occurrences. Determine the frequency of keyword occurrences in the text information of first-type documents and sort them according to the frequency of occurrences from highest to lowest. Generate an index from multiple documents or document fragments and concatenate the document content. The refined index and new documents need to be linked to products in the online store for output. At this point, when a user searches for "how to decorate a smart home," we will return an introduction to smart homes, preparations needed during decoration, necessary equipment purchases, recommended equipment brands, and a list of relevant products.
[0108] Example 4
[0109] Figure 3 This is a schematic diagram of the structure of a product data matching device provided in Embodiment 4 of the present invention. Figure 3 As shown, the device includes: a text acquisition module 31, an identifier determination module 32, a first set determination module 33, a second set determination module 34, and a product data determination module 35.
[0110] The text acquisition module 31 is used to acquire text information input by the user and determine the reasonableness of the text information and the questioning tendency.
[0111] The identifier determination module 32 is used to determine the document category identifier to be matched based on the reasonableness and question tendency.
[0112] The first set determination module 33 is used to determine the product data documents associated with the text information as the first type of documents based on the document category identifier to be matched and the preset matching threshold, and generate the first product data document set.
[0113] The second set determination module 34 is used to match second type documents associated with the text information in all product data documents based on the keywords contained in the text information to generate a second product data document set.
[0114] The product data determination module 35 is used to merge the first type of documents in the first product data document set and the second type of documents in the second product data document set to obtain product data that matches the text information.
[0115] In this embodiment of the invention, a text acquisition module acquires text information input by the user, determines the reasonableness of the text information and the question tendency, an identifier determination module determines the document category identifier to be matched based on the reasonableness of the text information and the question tendency, a first set determination module determines the product data documents associated with the text information as first type documents based on the document category identifier to be matched and a preset matching threshold, and generates a first product data document set, a second set determination module matches second type documents associated with the text information in all product data documents based on the keywords contained in the text information, and generates a second product data document set, a product data determination module merges the first type documents in the first product data document set and the second type documents in the second product data document set to obtain product data that matches the text information, thereby realizing comprehensive detection and matching of product data according to text information and keywords, accurately searching for product-related information according to user needs, and improving the accuracy of product data retrieval.
[0116] In one embodiment, the text acquisition module 31 includes:
[0117] The extraction tendency determination unit is used to perform semantic analysis on text information and determine the question tendency of the text information. The categories of question tendency include at least: product-oriented questions, product recommendation questions, product application method questions, and technical tutorial directions.
[0118] The legality determination unit is used to perform word segmentation on the text information to obtain word groups, match the word groups with a preset prohibited word library, and determine the legality of the text information as legal when it is determined that the word groups do not belong to the preset prohibited word library;
[0119] The keyword extraction unit is used to determine the semantic information of word segments and combine word segments existing in the preset thesaurus, as well as word segments existing in the preset thesaurus with the same semantic information as the word segments, as keywords of the text information.
[0120] The vector determination unit is used to vectorize keywords into keyword vectors and determine the word vectors of phrases in the preset product thesaurus.
[0121] The scope determination unit is used to determine the first cosine similarity between the keyword vector and each word group vector. When the first cosine similarity is greater than or equal to the preset first matching threshold, the text information is determined to be consistent with the business scope of the goods.
[0122] The first reasonableness determination unit is used to determine the reasonableness of text information when the legality of the keyword is legal and the keyword is in line with the business scope of the product.
[0123] The second reasonableness determination unit is used to determine the reasonableness of text information as unreasonable when the legality of the keywords is illegal or the keywords do not conform to the business scope of the goods.
[0124] In one embodiment, the identifier determination module 32 includes:
[0125] The first identifier determination unit is used to determine, when the reasonableness level is unreasonable, that there is no document category identifier to be matched that matches the text information;
[0126] The second identifier determination unit is used to extract identifier information associated with the category of the question tendency from the preset configuration information as the identifier of the document category to be matched, provided that the reasonableness level is reasonable.
[0127] In one embodiment, the first set determination module 33 includes:
[0128] The matching document determination unit is used to identify product data documents with a document category identifier to be matched as documents to be matched.
[0129] The document vector generation unit is used to vectorize text information into text information vectors and to vectorize the document to be matched into document vectors.
[0130] The similarity determination unit is used to determine the second cosine similarity between the text information vector and each document vector;
[0131] The first set determination unit is used to identify product data documents with a second cosine similarity greater than or equal to a preset second matching threshold as first type documents and store the first type documents as a first product data document set.
[0132] In one embodiment, the second set determination module 34 includes:
[0133] The target document determination unit is used to extract keywords from the text information, match the keywords with all product data documents, and obtain the target product data document.
[0134] The second set determination unit is used to treat the target product data document as a second type of document and store the second type of document as a second product data document set.
[0135] In one embodiment, the product data determination module 35 includes:
[0136] The document deletion unit is used to identify identical first-type and second-type documents and delete duplicate first-type or second-type documents.
[0137] The paragraph extraction unit is used to extract the keyword paragraphs to which the keywords belong in the first type of document and the second product data, respectively.
[0138] The frequency determination unit is used to determine the preset weight of each keyword in the first type of document and the second product data document, and to determine the frequency of each keyword in the keyword paragraph;
[0139] The weighting and determination unit is used to determine the sum of the products of the frequency of each keyword in each keyword paragraph and its corresponding preset weight as the weight sum;
[0140] The data determination unit is used to arrange each keyword paragraph according to its weight and in descending order, and to concatenate the keyword paragraphs to form product data that matches the text information.
[0141] In one embodiment, the product data matching device further includes:
[0142] The knowledge base construction module is used to build a knowledge base for product data documents and store all product data documents in the knowledge base;
[0143] The identifier construction module is used to preprocess all product data documents, divide the product data documents into at least one document category according to the pre-stored question tendency, and establish a document category identifier for each document category.
[0144] The product data matching device provided in this embodiment of the invention can execute the product data matching method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0145] Example 5
[0146] Figure 4 This is a schematic diagram of the structure of an electronic device implementing the product data matching method of embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0147] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0148] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0149] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the matching method for commodity data.
[0150] In some embodiments, the product data matching method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the product data matching method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the product data matching method by any other suitable means (e.g., by means of firmware).
[0151] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0152] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0153] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0154] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0155] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0156] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0157] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0158] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for matching commodity data, characterized in that, include: Obtain text information input by the user, and determine the reasonableness of the text information and the questioning tendency; The document category identifier to be matched is determined based on the reasonableness and the questioning tendency. Based on the document category identifier to be matched and the preset matching threshold, the product data document associated with the text information is determined as the first type of document, and a first product data document set is generated; Based on the keywords contained in the text information, a second set of product data documents is generated by matching the second type of documents associated with the text information in all product data documents; The first type of document in the first product data document set and the second type of document in the second product data document set are merged to obtain product data that matches the text information.
2. The method according to claim 1, characterized in that, The process of obtaining text information input by the user and determining the reasonableness of the text information and the questioning tendency includes: Semantic analysis is performed on the text information to determine the questioning tendency of the text information, wherein the categories of the questioning tendency include at least: product-oriented questions, product recommendation questions, product application method questions, and technical tutorial-oriented questions; The text information is segmented into word groups, and the word groups are matched with a preset prohibited word database. If it is determined that the word groups do not belong to the preset prohibited word database, the legality of the text information is determined to be legal. The semantic information of the segmented word groups is determined, and the segmented word groups existing in the preset lexicon, as well as the word groups existing in the preset lexicon that have the same semantic information as the segmented word groups, are combined as keywords of the text information; The keywords are vectorized into text to generate keyword vectors, and the word vectors of phrases in the preset product thesaurus are determined. The first cosine similarity between the keyword vector and each of the phrase vectors is determined. If the first cosine similarity is greater than or equal to the preset first matching threshold, the text information is determined to be within the scope of business operations. If the legality of the keyword is legal and the keyword is within the scope of business operations, the reasonableness of the text information is determined to be reasonable. If the legality of the keyword is deemed illegal or the keyword does not conform to the business scope of the product, the reasonableness of the text information is determined to be unreasonable.
3. The method according to claim 1, characterized in that, The step of determining the document category identifier to be matched based on the reasonableness and the question tendency includes: If the reasonableness level is deemed unreasonable, it is determined that there is no matching document category identifier that matches the text information. If the reasonableness level is reasonable, the identification information associated with the category of question tendency is extracted from the preset configuration information as the category identifier of the document to be matched.
4. The method according to claim 1, characterized in that, The step of determining the product data document associated with the text information as a first type of document based on the document category identifier to be matched and a preset matching threshold, and generating a first product data document set, includes: The product data document with the specified document category identifier is selected as the document to be matched. The text information is vectorized to generate a text information vector, and the document to be matched is vectorized to generate a document vector; Determine the second cosine similarity between the text information vector and each of the document vectors; Product data documents with a second cosine similarity greater than or equal to the preset second matching threshold are designated as first type documents, and the first type documents are stored as a first product data document set.
5. The method according to claim 1, characterized in that, The step of matching second-type documents associated with the text information in all product data documents based on the keywords contained in the text information to generate a second product data document set includes: Extract keywords from the text information, and match the keywords with all the product data documents to obtain the target product data document; The target product data document is used as a second type of document, and the second type of document is stored as a second product data document set.
6. The method according to claim 1, characterized in that, The step of merging the first type of documents in the first product data document set and the second type of documents in the second product data document set to obtain product data matching the text information includes: Identify identical documents of the first type and the second type, and delete duplicate documents of the first type or the second type. Extract the keyword paragraphs to which the keywords belong from the first type of document and the second product data, respectively; Determine the preset weight of each keyword in the first type of document and the second product data document, and determine the number of times each keyword appears in the keyword paragraph; The sum of the products of the frequency of occurrence of each keyword in each keyword paragraph and the corresponding preset weight is determined as the weight sum; The keyword paragraphs are arranged according to their weights and in descending order, and then concatenated to form product data that matches the text information.
7. The method according to claim 6, characterized in that, After concatenating the aforementioned keyword paragraphs to obtain product data that matches the text information, the process also includes: The system generates and outputs questions related to the product data based on the product data.
8. The method according to claim 1, characterized in that, Before acquiring user-inputted text information and determining the reasonableness of the text information and the question's intent, the process also includes: Construct a knowledge base for the product data documents and store all the product data documents in the knowledge base; All the product data documents are preprocessed with text, and the product data documents are divided into at least one document category according to the pre-stored question tendency, and a document category identifier is established for each document category.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the matching method for commodity data according to any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the matching method for commodity data according to any one of claims 1-8.