A method for intelligently supplementing product information based on a large language model
By combining image feature analysis with user feedback, and utilizing a multimodal pre-training model to generate product information, the problem of existing technologies being unable to automatically generate product information and determine product authenticity is resolved. This enables efficient and accurate product information generation and authenticity determination, adapting to the needs of different scenarios.
Patent Information
- Application Number
- CN202510758740.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-09
AI Technical Summary
Existing intelligent product information supplementation methods based on large language models cannot automatically generate preliminary product information from pictures uploaded by sellers, cannot automatically determine whether the product is official and authentic, and cannot automatically generate final information based on user feedback, resulting in data bias, reduced efficiency and user experience.
By obtaining a collection of target product images, image feature extraction and analysis are performed, combined with product feature comparison and user feedback, and a multimodal pre-training model is used to generate product information, including image detail analysis, product feature judgment and user feedback fusion, to optimize model parameters to adapt to different scenarios.
It realizes automated product information generation and authenticity judgment, reduces manual intervention, improves efficiency and accuracy, adapts to different products and scenarios, improves user experience and platform product quality, and enhances the robustness and fault tolerance of the algorithm.
Smart Images

Figure CN120338829B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method for intelligently supplementing product information based on a large language model. Background Art
[0002] With the rapid development of e-commerce, corporate websites, and multi-channel product releases, the richness and accuracy of product information have become key factors in improving user experience and conversion rates. However, manually maintaining large amounts of product information is time-consuming and labor-intensive, and is prone to incomplete or inconsistent information. In recent years, large pre-trained language models (such as the GPT series and BERT, etc.) have provided new solutions for automatically generating and supplementing product information with their powerful natural language understanding and generation capabilities. They can automatically generate large amounts of high-quality product information, saving manual time, and can adjust generated content according to different product categories and market demands. They have semantic understanding and contextual association capabilities, generating coherent and rich content. Leveraging the powerful text generation and comprehension capabilities of large language models, they can automatically supplement and improve product information, effectively improving the company's content management efficiency and user experience.
[0003] Existing intelligent product information supplementation methods based on large language models cannot automatically generate preliminary product information through pictures uploaded by sellers, cannot automatically determine whether the seller's products are official genuine products, and cannot automatically generate final product information based on user feedback. This makes data prone to deviations and requires manual intervention, which reduces efficiency and cannot adapt to different products and scenarios. It reduces the quality of authenticity judgment and information generation, reduces user experience and platform product quality, and is prone to disputes due to counterfeit and shoddy products, thereby disrupting market order. Summary of the Invention
[0004] The present invention provides a method for intelligently supplementing product information based on a large language model, which is used to promote the solution of the problems raised in the background technology.
[0005] The present invention provides the following technical solution: a method for intelligently supplementing product information based on a large language model, comprising:
[0006] Get the image collection of the target product , automatically generate preliminary product information;
[0007] Get image collection The authenticity results of each image in and ;
[0008] Extract image collection Elements corresponding to all real pictures in:
[0009] ;
[0010] in, Indicates traversing the image collection Each picture in ;
[0011] Set up a product judgment function to determine whether the seller's product is official and authentic:
[0012] ;
[0013] in, is the set threshold;
[0014] like , the target product is initially determined to be genuine, and the final product information is generated according to the official product information;
[0015] like , the target product is initially analyzed to be a counterfeit product, and all user feedback related to the target product is read from the database or feedback storage system to form a feedback set :
[0016] ;
[0017] in, is a function that retrieves all user feedback related to a specific product from a database, file system, or other storage medium;
[0018] Set a feedback judgment function to determine the user feedback of the target product:
[0019] ;
[0020] like , then it is determined that there is user feedback, and the final product information is generated based on the user feedback;
[0021] like , it is determined that there is no user feedback, and the final product information is assisted in generating through the fusion of knowledge graph and cross-modal retrieval and generation.
[0022] As an optional solution of the method for intelligently supplementing product information based on a large language model of the present invention, preliminary product information is automatically generated, specifically:
[0023] Get all the pictures of the target product and generate a picture collection, recorded as ;
[0024] The key features of the target product are extracted from the image collection through the feature extraction function, denoted as :
[0025] ;
[0026] in, is the input image set, is the feature extraction function, is the extracted feature vector;
[0027] Get a set of predefined prompt templates, denoted as ;
[0028] For the extracted features, a selection strategy function is used to select a suitable hint template from the hint template set:
[0029] ;
[0030] in, Is a set of predefined prompt templates, each template is a string containing placeholders for inserting feature values. is a selection strategy function to select the most appropriate template based on the content of the feature and the applicability of the prompt template, From the prompt template collection A specific template selected from
[0031] Set up a prompt generation function , combine the extracted features with the selected prompt template to generate specific prompt information:
[0032] ;
[0033] in, Is the generated prompt information;
[0034] Obtaining a multimodal pre-trained model ;
[0035] Through multimodal pre-training model Image feature extraction component in , from the picture collection Extract image features from :
[0036] ;
[0037] in, Pre-training models for multimodal The image feature extraction component in For pictures from the collection The image features extracted from
[0038] Through multimodal pre-training model Text feature extraction component in , from the prompt information Extract text features from :
[0039] ;
[0040] in, It is the text feature extraction component in the multimodal pre-training model. For prompt information The text features extracted from
[0041] The image features and text features Input to the multimodal feature fusion component In the fusion, we get the multimodal features, which are recorded as :
[0042] ;
[0043] in, It is the feature fusion component in the multimodal pre-training model. is the fused multimodal feature;
[0044] The fused multimodal features Generative components that are input to the model In the example, the initial product information is generated and recorded as :
[0045] ;
[0046] in, Generative components in multimodal pre-trained models.
[0047] As an optional solution to the method for intelligently supplementing product information based on a large language model described in the present invention, a preliminary authenticity analysis of the target product is performed, including image feature analysis, specifically:
[0048] Get the image collection of the target product ;
[0049] Gray-level co-occurrence matrix , extract each picture Texture features:
[0050] ;
[0051] in, For picture collection The images analyzed in
[0052] Extract each picture through template matching Pattern characteristics :
[0053] ;
[0054] Extract each image through Canny edge detection Edge features :
[0055] ;
[0056] Calculate each image The Laplace operator response is used to evaluate the blurriness of the image. :
[0057] ;
[0058] Detect each image through local consistency check Is there any distortion? :
[0059] ;
[0060] Fusion of texture, pattern, edge, blur and distortion features into image detail features :
[0061] ;
[0062] For each pair of images and , calculate their image detail features and Similarity :
[0063] ;
[0064] Calculate the average similarity of all image pairs as the multi- Figure 1 Consistency score :
[0065] ;
[0066] Extract each image Metadata characteristics of :
[0067] ;
[0068] Calculate the consistency score of metadata features of all images :
[0069] ;
[0070] in, is the reference metadata, is the number of fields matched, is the total number of fields;
[0071] Use image forensics algorithms to detect whether images have been cropped or spliced :
[0072] ;
[0073] Combine metadata consistency and cropping and splicing detection results to calculate image integrity and originality scores :
[0074] ;
[0075] in, and It is a weight parameter used to balance the contribution of different parts to the final result, satisfying ;
[0076] Calculate the average score of image detail features for all images :
[0077] ;
[0078] in, Indicates the Picture No. Detailed features;
[0079] Combined with image detail score, multiple Figure 1 The consistency score and the image integrity and originality score are used to calculate the final authenticity score. :
[0080] ;
[0081] in, 、 and It is a weight parameter used to balance the contribution of different parts to the final result, satisfying ;
[0082] According to the authenticity score Perform threshold judgment to determine the authenticity of the results :
[0083] ;
[0084] in, is a predefined threshold.
[0085] As an optional solution to the method for intelligently supplementing product information based on a large language model of the present invention, the preliminary authenticity analysis of the target product also includes product characteristics analysis, specifically:
[0086] Get the image collection of the target product ;
[0087] From each picture Extract product features :
[0088] ;
[0089] in, It refers to extracting product-related features from images;
[0090] Calculate product features for each image Compared with official product features Similarity :
[0091] ;
[0092] in, For the Product features extracted from the image, Official product features are standard product information obtained from the brand's official website, officially authorized channels, or other trusted sources. For the The similarity between the product features in the image and the official product features, and Represents the product feature vector and the official product feature vector The norm of
[0093] Calculate the average of the product feature similarities of all images as the product detail comparison score :
[0094] ;
[0095] in, Indicates the total number of images in the image collection uploaded by the seller;
[0096] Detect each image Check whether there are brand logos, trademarks and anti-counterfeiting marks in the :
[0097] ;
[0098] in, For the The brand logo detection result in the picture indicates whether the brand logo, trademark and anti-counterfeiting mark are detected. Refers to detecting specific patterns of brand logos, trademarks, and anti-counterfeiting marks in images;
[0099] Check the clarity, accuracy and completeness of the detected brand logo and obtain the inspection results :
[0100] ;
[0101] in, For the Check results for clarity, accuracy, and completeness of brand logos in images. Refers to the quality assessment of detected brand logos;
[0102] Calculate the brand logo and logo inspection score by combining the brand logo detection results and inspection results :
[0103] ;
[0104] The final authenticity score is calculated by combining the product details comparison score and the brand logo and logo inspection score. :
[0105] ;
[0106] in, and It is a weight parameter used to balance the contribution of different parts to the final result, satisfying ;
[0107] According to the authenticity score Perform threshold judgment to determine the authenticity of the results :
[0108] ;
[0109] in, is a predefined threshold.
[0110] As an optional solution to the method for intelligently supplementing product information based on a large language model described in the present invention, if the target product is genuine, the final product information is generated according to the official product information, specifically:
[0111] Get image collection ;
[0112] Extract all official information of the target product from the database and integrate it to generate an official information set, denoted as ;
[0113] Obtaining a multimodal pre-trained model ;
[0114] Through multimodal pre-training model Image feature extraction component in , from the picture collection Extract image features from :
[0115] ;
[0116] Through multimodal pre-training model Text feature extraction component in , from official information Extract text features from :
[0117] ;
[0118] The image features and text features Input to the multimodal feature fusion component In the fusion, we get the multimodal features, which are recorded as :
[0119] ;
[0120] The fused multimodal features Generative components that feed into the updated model In the process, new product information is generated and recorded as :
[0121] ;
[0122] Output the generated new product information .
[0123] As an optional solution to the method for intelligently supplementing product information based on a large language model described in the present invention, if the target product is initially analyzed to be a counterfeit and there is user feedback, the final product information is generated based on the user feedback, including model optimization, specifically:
[0124] Get feedback collection ;
[0125] For feedback collection Every user feedback in , calculate its relationship with the current product information Cosine similarity of:
[0126] ;
[0127] in, For the The norm of the user feedback vector, The norm of the product information vector generated for the model;
[0128] Based on similarity score , design a reward function to convert the similarity score into a reward signal:
[0129] ;
[0130] in, is the similarity threshold;
[0131] All reward signals Combined into a reward signal vector :
[0132] ;
[0133] in, is the number of user feedback;
[0134] Through the loss function , measure the product information generated by the model and user feedback The differences between:
[0135] ;
[0136] ;
[0137] Among them, the loss function Specifically, it is the mean square error, The product information generated for the model A quantity, For the User feedback Quantity
[0138] Calculating the loss function Model parameters Gradient:
[0139] ;
[0140] Using learning rate And the calculated gradient, update the model parameters:
[0141] .
[0142] As an optional solution to the method for intelligently supplementing product information based on a large language model according to the present invention, the method of generating final product information by referring to user feedback further includes generating final product information, specifically:
[0143] Get the model parameters after model optimization ;
[0144] Get image collection and prompt information ;
[0145] Obtaining a multimodal pre-trained model ;
[0146] Extracting multimodal pre-trained models The current model parameters in are denoted as ;
[0147] The current model parameters Update to the optimized model parameters ;
[0148] Through multimodal pre-training model Image feature extraction component in , from the picture collection Extract image features from :
[0149] ;
[0150] Through multimodal pre-training model Text feature extraction component in , from the prompt information Extract text features from :
[0151] ;
[0152] The image features and text features Input to the multimodal feature fusion component In the fusion, we get the multimodal features, which are recorded as :
[0153] ;
[0154] The fused multimodal features Generative components that feed into the updated model In the process, new product information is generated and recorded as :
[0155] ;
[0156] From new product information Extract key product feature vectors from:
[0157] ;
[0158] in, Indicates the The value of the key feature;
[0159] Assign weights to each key characteristic:
[0160] ;
[0161] in, Indicates the The weight of the features, and ;
[0162] Calculate the total weighted score of the product's authenticity :
[0163] ;
[0164] The weighted total score Converted into the final authenticity judgment result :
[0165] ;
[0166] in, is the preset threshold;
[0167] like , then the generated new product information is output , and mark the product as authentic;
[0168] like , then the generated new product information is output , and mark the product as a counterfeit.
[0169] As an optional solution to the large language model-based intelligent product information supplementation method of the present invention, if the target product is initially analyzed to be a counterfeit and there is no user feedback, the final product information is assisted in generating by integrating knowledge graph and cross-modal retrieval and generation, specifically:
[0170] Generate a knowledge graph for the target product:
[0171] ;
[0172] in, is a collection of entities, is a collection of relations;
[0173] In the knowledge graph Search and product categories Related entity collections :
[0174] ;
[0175] in, Representing an entity Category label of
[0176] From each entity Extract feature vectors from :
[0177] ;
[0178] Get feature vector The weight of :
[0179] ;
[0180] Multiply the feature vector of each entity by its corresponding weight, and then add these weighted feature vectors to get the category feature vector :
[0181] ;
[0182] Normalize the weighted summed eigenvectors:
[0183] ;
[0184] in, represents the L2 norm of the feature vector.
[0185] As an optional solution to the method for intelligently supplementing product information based on a large language model described in the present invention, the method of assisting in generating final product information by integrating knowledge graph and cross-modal retrieval and generation also includes:
[0186] In the product image database In the input image Perform cross-modal retrieval to obtain a set of similar products :
[0187] ;
[0188] in, Represents a function that determines similar products by calculating the similarity between the image features and the features of images in the database;
[0189] From each similar product Extract feature vectors from :
[0190] ;
[0191] Aggregate the feature vectors of similar products to obtain the feature vectors of similar products :
[0192] ;
[0193] in, Similar products The specific formula is:
[0194] ;
[0195] in, Represents a collection of similar products For each product image in For each similar product With input picture The cosine similarity is:
[0196] ;
[0197] in, Is the input image The eigenvector of Similar products The eigenvector of Represents the L2 norm of the vector;
[0198] Knowledge graph features and similar product features Fusion is performed to obtain the fusion feature vector :
[0199] ;
[0200] in, and is the weight parameter;
[0201] The fused feature vector Generative components that feed into the updated model In the process, new product information is generated and recorded as :
[0202] ;
[0203] From new product information Extract key product feature vectors from:
[0204] ;
[0205] in, Indicates the The value of the key feature;
[0206] Assign weights to each key characteristic:
[0207] ;
[0208] in, Indicates the The weight of the features, and ;
[0209] Calculate the total weighted score of the product's authenticity :
[0210] ;
[0211] The weighted total score Converted into the final authenticity judgment result :
[0212] ;
[0213] in, is the preset threshold;
[0214] like , then the generated new product information is output , and mark the product as authentic;
[0215] like , then the generated new product information is output , and mark the product as a counterfeit.
[0216] The present invention has the following beneficial effects:
[0217] 1. This intelligent product information supplementation method based on a large language model automatically generates preliminary product information from pictures uploaded by sellers, and combines image analysis and product feature comparison to determine whether the seller's products are official and authentic. It automates the entire process from image feature extraction to product information generation and authenticity judgment, reducing manual intervention and improving efficiency. It can more comprehensively assess product authenticity and is more accurate and reliable than a single method. It effectively utilizes multimodal information such as images and text to more fully understand product characteristics and improve the quality of authenticity judgment and information generation. When image quality is poor or information is incomplete, the fusion of multimodality and multiple methods enhances the robustness and fault tolerance of the algorithm.
[0218] 2. This intelligent product information supplementation method based on a large language model generates final product information based on official product information if the product is authentic. If the product is a counterfeit, it extracts user feedback data from the database to determine whether there is any user feedback for the product. In scenarios with complex product information such as e-commerce and second-hand trading platforms, this method helps combat counterfeiting and maintain market order, uses user feedback to form a closed-loop optimization, and improves user experience and platform product quality.
[0219] 3. This intelligent product information supplementation method based on a large language model generates the final product information based on user feedback if there is any. If there is no user feedback, the final product information is assisted in being generated through the fusion of knowledge graph and cross-modal retrieval and generation. The authenticity of the final product information is then judged again to determine whether the product is genuine. The model parameters are optimized through reinforcement learning, so that the algorithm can continuously improve based on user feedback, adapt to different products and scenarios, and use user feedback to form a closed-loop optimization to improve user experience and the quality of platform products. BRIEF DESCRIPTION OF THE DRAWINGS
[0220] Figure 1 This is a flow chart of the method for intelligently supplementing product information based on a large language model of the present invention. DETAILED DESCRIPTION
[0221] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0222] Example 1: A method for intelligently supplementing product information based on a large language model. Figure 1 ,include:
[0223] Get the image collection of the target product , automatically generate preliminary product information;
[0224] Get image collection The authenticity results of each image in and ;
[0225] Extract image collection Elements corresponding to all real pictures in:
[0226] ;
[0227] in, Indicates traversing the image collection Each picture in ;
[0228] Set up a product judgment function to determine whether the seller's product is official and authentic:
[0229] ;
[0230] in, The threshold is set to determine whether the target product is an official product. If the number of real pictures in the picture collection is If the proportion of exceeds the set threshold, it means that the target product is genuine;
[0231] like , the target product is initially determined to be genuine, and the final product information is generated according to the official product information;
[0232] like , the target product is initially analyzed to be a counterfeit product, and all user feedback related to the target product is read from the database or feedback storage system to form a feedback set :
[0233] ;
[0234] in, This function retrieves all user feedback related to a specific product from a database, file system, or other storage medium. It primarily collects user reviews, opinions, or suggestions for a product for further analysis and processing. It retrieves all feedback related to a specific product from the location where user feedback is stored. It typically requires a parameter to specify the product to retrieve, such as a product ID or name. The result is typically a collection or list of multiple user feedback items.
[0235] Set a feedback judgment function to determine the user feedback of the target product:
[0236] ;
[0237] like , then it is determined that there is user feedback, and the final product information is generated based on the user feedback;
[0238] like , it is determined that there is no user feedback, and the final product information is assisted in generating through the fusion of knowledge graph and cross-modal retrieval and generation.
[0239] This embodiment also provides that if the target product is genuine, the final product information is generated according to the official product information, specifically, wherein the method of determining that the target product is genuine includes determining that the target product is genuine through preliminary analysis and when Output the new product information generated , and mark the product as authentic:
[0240] Get image collection ;
[0241] Extract all official information of the target product from the database and integrate it to generate an official information set, denoted as ;
[0242] Obtaining a multimodal pre-trained model ;
[0243] Through multimodal pre-training model Image feature extraction component in , from the picture collection Extract image features from :
[0244] ;
[0245] Through multimodal pre-training model Text feature extraction component in , from official information Extract text features from :
[0246] ;
[0247] The image features and text features Input to the multimodal feature fusion component In the fusion, we get the multimodal features, which are recorded as :
[0248] ;
[0249] The fused multimodal features Generative components that feed into the updated model In the process, new product information is generated and recorded as :
[0250] ;
[0251] Output the generated new product information ;
[0252] In this embodiment, after determining that the target product is genuine, the official information of the brand corresponding to the target product is extracted from the database. , this official information Substitute into the multimodal pre-training model , through multimodal pre-training model Components in 、 、 as well as , generate final, more complete product information that can be displayed to buyers .
[0253] This embodiment also provides that if the target product is initially analyzed to be a counterfeit and there is user feedback, the final product information is generated with reference to the user feedback, including model optimization, specifically:
[0254] Get feedback collection ;
[0255] For feedback collection Every user feedback in , calculate its relationship with the current product information The cosine similarity reflects the degree of correlation between the two:
[0256] ;
[0257] in, For the The norm of the user feedback vector, used to normalize the vector, The norm of the product information vector generated by the model, used to normalize the vector;
[0258] Based on similarity score , a reward function is designed to convert the similarity score into a reward signal to guide the update of model parameters. High similarity corresponds to high reward:
[0259] ;
[0260] in, is the similarity threshold, which is used to filter out feedback with low similarity to ensure that only valuable feedback affects the model update;
[0261] All reward signals Combined into a reward signal vector , summarizing the reward signals corresponding to all user feedback and using them to reinforce the learning process to optimize the model:
[0262] ;
[0263] in, is the amount of user feedback, which determines the dimension of the reward signal vector;
[0264] Through the loss function , measure the product information generated by the model and user feedback The differences between:
[0265] ;
[0266] ;
[0267] Among them, the loss function Specifically, it is the mean square error MSE, The product information generated for the model A quantity, For the User feedback Quantity
[0268] Calculating the loss function Model parameters The gradient of is used to guide the direction of parameter update:
[0269] ;
[0270] Using learning rate And the calculated gradient, update the model parameters:
[0271] ;
[0272] Among them, the learning rate Used to control the step size of parameter updates.
[0273] The generating of final product information by referring to user feedback also includes generating final product information, specifically:
[0274] Get the model parameters after model optimization ;
[0275] Get image collection and prompt information ;
[0276] Obtaining a multimodal pre-trained model ;
[0277] Extracting multimodal pre-trained models The current model parameters in are denoted as ;
[0278] The current model parameters Update to the optimized model parameters ;
[0279] Through multimodal pre-training model Image feature extraction component in , from the picture collection Extract image features from :
[0280] ;
[0281] Through multimodal pre-training model Text feature extraction component in , from the prompt information Extract text features from :
[0282] ;
[0283] The image features and text features Input to the multimodal feature fusion component In the fusion, we get the multimodal features, which are recorded as :
[0284] ;
[0285] The fused multimodal features Generative components that feed into the updated model In the process, new product information is generated and recorded as :
[0286] ;
[0287] From new product information Extract key product feature vectors from:
[0288] ;
[0289] in, Indicates the The value of the key feature;
[0290] Assign a weight to each key feature to determine its relative importance in authenticity judgment:
[0291] ;
[0292] in, Indicates the The weight of the features, and ;
[0293] Calculate the total weighted score of the product's authenticity :
[0294] ;
[0295] The weighted total score Converted into the final authenticity judgment result :
[0296] ;
[0297] in, is a preset threshold used to convert the weighted total score into a true or false judgment result;
[0298] like , then the generated new product information is output , and mark the product as authentic;
[0299] like , then the generated new product information is output , and mark the product as a counterfeit;
[0300] In this embodiment, after the target product is initially determined to be a counterfeit, the multimodal pre-training model is updated based on the user's feedback on the product. The model parameters are set, and the pictures uploaded by the seller and the product description text are substituted into the multimodal pre-training model , through multimodal pre-training model Components in 、 、 as well as , generate a new and more complete product information , and analyze the new product information generated, judge the authenticity of the product again, and based on the results of the second judgment, determine the product information that is finally displayed to the buyer and mark the authenticity of the product.
[0301] This embodiment also provides that if the target product is initially analyzed to be a counterfeit and there is no user feedback, the final product information is assisted in generating by integrating knowledge graph and cross-modal retrieval and generation, specifically:
[0302] Generate a knowledge graph for the target product:
[0303] ;
[0304] in, It is a collection of entities, including products, brands, functions, etc. It is a collection of relationships, such as brand-product, product-feature, etc.
[0305] In the knowledge graph Search and product categories Related entity collections :
[0306] ;
[0307] in, Representing an entity Category label of
[0308] From each entity Extract feature vectors from , these features can include attributes of the entity, such as brand name, function description, etc.:
[0309] ;
[0310] Get feature vector The weight of :
[0311] ;
[0312] Multiply the feature vector of each entity by its corresponding weight, and then add these weighted feature vectors to get the category feature vector :
[0313] ;
[0314] Normalize the weighted summed eigenvectors:
[0315] ;
[0316] in, represents the L2 norm of the feature vector.
[0317] The method of assisting in generating the final product information by integrating knowledge graph and cross-modal retrieval and generation also includes:
[0318] In the product image database In the input image Perform cross-modal retrieval to obtain a set of similar products :
[0319] ;
[0320] in, Represents a function that determines similar products by calculating the similarity between the image features and the features of images in the database;
[0321] From each similar product Extract feature vectors from , such as appearance description, functional features, etc.:
[0322] ;
[0323] in, It is a function that extracts feature vectors from images or other data types. Feature vectors can contain a variety of image information, such as color, texture, shape, edges, etc. In image processing and machine learning tasks, images are converted into feature vectors to enable computers to understand and process image data. These feature vectors can be used for tasks such as classification, retrieval, and clustering.
[0324] Aggregate the feature vectors of similar products to obtain the feature vectors of similar products :
[0325] ;
[0326] in, Similar products The specific formula is:
[0327] ;
[0328] in, Represents a collection of similar products Each product image in is used to traverse all product images in the similar product collection to calculate their similarity with the input image The similarity of For each similar product With input picture The cosine similarity is:
[0329] ;
[0330] in, Is the input image The eigenvector of Similar products The eigenvector of Represents the L2 norm of the vector;
[0331] Knowledge graph features and similar product features Fusion is performed to obtain the fusion feature vector :
[0332] ;
[0333] in, and is a weight parameter used to balance the contribution of the two features, that is, to determine the relative importance of each feature to the fusion feature vector;
[0334] The fused feature vector Generative components that feed into the updated model In the process, new product information is generated and recorded as :
[0335] ;
[0336] From new product information Extract key product feature vectors from:
[0337] ;
[0338] in, Indicates the The value of the key feature;
[0339] Assign a weight to each key feature to determine its relative importance in authenticity judgment:
[0340] ;
[0341] in, Indicates the The weight of the features, and ;
[0342] Calculate the total weighted score of the product's authenticity :
[0343] ;
[0344] The weighted total score Converted into the final authenticity judgment result :
[0345] ;
[0346] in, is a preset threshold used to convert the weighted total score into a true or false judgment result;
[0347] like , then the generated new product information is output , and mark the product as authentic;
[0348] like , then the generated new product information is output , and mark the product as a counterfeit.
[0349] Through the above method, preliminary product information is automatically generated through the pictures uploaded by the seller, and it is judged whether the seller's product is an official genuine product. If the product is genuine, the final product information is generated according to the official product information. If the product is a counterfeit, it is judged whether there is user feedback for the product. If there is user feedback, the final product information is generated with reference to the user feedback. If there is no user feedback, the final product information is assisted in generating the final product information through the fusion of knowledge graph and cross-modal retrieval and generation. By combining multiple means such as image analysis, product feature comparison and user feedback, the authenticity of the product is comprehensively evaluated, which is more accurate and reliable than a single method. The automated processing is based on the information extracted from image features. The entire process of product information generation and authenticity judgment is obtained, which reduces manual intervention and improves efficiency. Through reinforcement learning, the model parameters are optimized so that the algorithm can continuously improve according to user feedback, adapt to different products and scenarios, and effectively utilize multimodal information such as images and text to more fully understand product characteristics, improve the authenticity judgment and information generation quality. When the image quality is poor or the information is incomplete, the fusion of multimodality and multi-method enhances the robustness and fault tolerance of the algorithm, uses user feedback to form a closed-loop optimization, improves user experience and platform product quality, and reduces disputes caused by counterfeit and shoddy products in scenarios with complex product information such as e-commerce and second-hand trading platforms, and maintains market order.
[0350] Example 2: This example is an improvement made on the basis of Example 1. The intelligent product information supplement method based on the large language model automatically generates preliminary product information, specifically:
[0351] Get all the pictures of the target product and generate a picture collection, recorded as ;
[0352] The key features of the target product are extracted from the image collection through the feature extraction function, denoted as :
[0353] ;
[0354] in, For the input image set, each image can be represented as a matrix or tensor, containing pixel values and color channel information. For feature extraction functions, a variety of computer vision techniques can be used, such as color histogram calculation, texture analysis algorithms such as gray-level co-occurrence matrix, edge detection operators such as Canny edge detection, object detection models such as YOLO or Faster R-CNN, etc., to extract features describing the product appearance and content from the image. It is the extracted feature vector, which is used in the subsequent steps to select the appropriate prompt template and generate specific prompt information;
[0355] Get a set of predefined prompt templates, denoted as ;
[0356] For the extracted features, a selection strategy function is used to select a suitable hint template from the hint template set:
[0357] ;
[0358] in, It is a set of predefined prompt templates. Each template is a string containing placeholders for inserting feature values. It is designed to guide the model to generate expected product information, such as describing the product's appearance, functions, etc. It is a selection strategy function that selects the most appropriate template based on the content of the feature and the applicability of the prompt template. For example, it can select the corresponding template based on the type of object detected in the image, such as mobile phones, clothes, etc., or select a more detailed or more concise template based on the richness of the feature. From the prompt template collection A specific template selected from is used to generate prompt information;
[0359] Set up a prompt generation function , used to fill the feature value into the placeholder of the prompt template to generate specific prompt information, and combine the extracted features with the selected prompt template to generate specific prompt information:
[0360] ;
[0361] in, The generated prompt information is used to guide the multimodal pre-trained model to generate the expected product information. The prompt information is clear and specific, and can guide the model to focus on the key product features in the image.
[0362] Obtaining a multimodal pre-trained model ;
[0363] Through multimodal pre-training model Image feature extraction component in , from the picture collection Extract image features from :
[0364] ;
[0365] in, Pre-training models for multimodal The image feature extraction component in the image is usually based on the convolutional neural network (CNN) and is used to extract image features from images. The image feature extraction component is responsible for extracting useful visual features from images. It can identify information such as color distribution, texture pattern, edge contour, and position and shape of objects in the image, and convert the pixel data of the image into a higher-level feature representation, laying the foundation for subsequent feature fusion and product information generation. For pictures from the collection The image features extracted from the image contain visual information in the image, such as color, texture, shape, object position, etc.
[0366] Through multimodal pre-training model Text feature extraction component in , from the prompt information Extract text features from :
[0367] ;
[0368] in, It is a text feature extraction component in a multimodal pre-training model, usually based on Transformer or its variants, used to extract text features from text. The text feature extraction component converts prompt information into text features that the model can understand. It can capture information such as keywords, phrases, and semantic relationships in the text, and convert the semantic connotation of the text into a feature representation in the form of a vector, so that the text information can be effectively integrated with the image information in subsequent steps. For prompt information The text features extracted from the prompt include semantic information in the prompt information, such as keywords, semantic relationships, etc.
[0369] The image features and text features Input to the multimodal feature fusion component In the fusion, we get the multimodal features, which are recorded as :
[0370] ;
[0371] in, It is a feature fusion component in the multimodal pre-training model, used to fuse image features and text features, such as by using attention mechanisms, fully connected networks, etc. The feature fusion component organically integrates image features and text features, so that the information of the two modalities complements and strengthens each other. Through feature fusion, the model can comprehensively utilize visual and text information to generate more comprehensive and accurate product information. For example, the feature fusion component can combine the product color mentioned in the prompt information with the color features extracted from the image to generate a more specific color description. The fused multimodal features combine information from image and text features to generate product information. The fused multimodal features integrate information from both image and text modalities and are a comprehensive feature representation that includes both visual information such as the product's appearance and shape, as well as the semantic requirements of the prompt information. This provides a richer and more complete information foundation for generating components, enabling the generated product information to meet both visual and semantic requirements.
[0372] The fused multimodal features Generative components that are input to the model In the example, the initial product information is generated and recorded as :
[0373] ;
[0374] in, It is the generation component in the multimodal pre-training model, usually based on the Transformer decoder structure, and is used to generate product information based on the fused features. The generation component generates the final product information based on the fused features, can convert multimodal features into natural language descriptions, and output product appearance descriptions, function speculations, and other content that meet user needs. The generation component usually uses technologies such as the attention mechanism to ensure that the generated text is closely related to the input features and has good language fluency and logic. The generated initial product information includes basic content such as appearance description and function speculation. It describes the product's appearance, functions and other basic content in natural language. The initial product information is generated based on the pictures and prompt information uploaded by the seller, and is intended to provide users with accurate and detailed product introductions to help users better understand the product features;
[0375] Among them, the multimodal pre-training model The specific generation process is as follows:
[0376] Multimodal pre-training model Initialize all components in:
[0377] ;
[0378] Among them, the multimodal pre-training model The components in include image feature extraction components , text feature extraction component , feature fusion component And generate components ;
[0379] Get the image-text pair for training, denoted as ;
[0380] Integrate all image-text pairs to generate a pre-training dataset, denoted as :
[0381] ;
[0382] use Extracting images Features:
[0383] ;
[0384] in, It is a parameter set of the image feature extraction component, including convolution layer parameters, normalization layer parameters and fully connected layer parameters. The convolution layer parameters include the weights and biases of the convolution kernels. The convolution kernels are used to slide on the image to extract local features such as edges and textures. For example, a convolution layer may have multiple convolution kernels, each of which is responsible for extracting a specific feature. The normalization layer parameters include the parameters of the batch normalization BatchNorm layer, such as the scaling factor and offset factor, which are used to normalize the output of the convolution layer and stabilize the training process. The fully connected layer parameters are the parameters of the fully connected layer. The fully connected layer is usually in the last few layers of the convolutional neural network CNN. Its parameters include the weight matrix and bias vector, which are used to map the extracted features to a higher-level semantic space.
[0385] use Extract text Features:
[0386] ;
[0387] in, This is a parameter set for the text feature extraction component, including word embedding layer parameters, self-attention layer parameters, and feedforward neural network layer parameters. The word embedding layer maps words or subwords in the text to a vector space of fixed dimension. Its parameter is the embedding matrix, where each row corresponds to the embedding vector of a word or subword. The self-attention layer parameters are the parameters of the self-attention layer. In the Transformer model, the self-attention layer is used to capture the dependencies between words in the text. Its parameters include the weights and biases of the query, key, and value matrices. The feedforward neural network layer parameters are the parameters of the feedforward neural network layer. The feedforward neural network in the Transformer is used to perform nonlinear transformations on the output of the self-attention layer. Its parameters include the weight matrix and bias vector.
[0388] Will and Input In the fusion feature, we get:
[0389] ;
[0390] in, It is a parameter set of the feature fusion component, including the fully connected layer parameters and the attention mechanism parameters. If the feature fusion adopts a simple splicing followed by a fully connected layer, the weight matrix and bias vector of the fully connected layer are used to map the spliced features to the fusion feature space. If the attention mechanism is used for feature fusion, the parameters of the attention layer include the weight matrix and bias vector, which are used to calculate the attention weights between image features and text features, and determine the importance of each modal feature in the fusion process.
[0391] use ,according to , generating text:
[0392] ;
[0393] in, It is a parameter set of the generation component, including word embedding layer parameters, self-attention layer parameters, feedforward neural network layer parameters, and output layer parameters. The word embedding layer parameters are similar to those of the text feature extraction component. The word embedding layer in the generation component maps words or subwords to the vector space. Its parameters are the embedding matrix. The self-attention layer parameters are the parameters of the self-attention layer. During the generation process, the self-attention layer is used to capture the dependencies between generated words. Its parameters include the weights and biases of the query, key, and value matrices. The feedforward neural network layer parameters are used to perform nonlinear transformations on the output of the self-attention layer. Its parameters include the weight matrix and bias vector. The output layer parameters include the weight matrix and bias vector, which are used to map the output of the generation component to the vocabulary space and calculate the generation probability of each word.
[0394] Through the loss function , calculate and generate text With real text The loss between :
[0395] ;
[0396] ;
[0397] Among them, the loss function Specifically, the cross entropy loss function is one of the most commonly used loss functions in natural language processing tasks, especially for text generation tasks. It measures the difference between the word distribution of the generated text and the word distribution of the real text. is the size of the vocabulary, i.e. the number of possible words, For real text Middle One-hot encoding, if If the word is a real word, , otherwise it is 0. For example, if the vocabulary size is 10,000 and the real word is the 100th word in the vocabulary, then , the rest of the positions are 0, To generate text Middle The probability distribution of words, which is the probability of each word at that position predicted by the model, and the probability distribution of generated text It is the result of the softmax layer output by the model, indicating the probability of each word at that position predicted by the model. The softmax layer is a commonly used activation function in neural networks, mainly used in multi-classification problems, converting any real-valued logits output by the model into a probability distribution. The softmax layer is usually used together with the cross-entropy loss function, which measures the difference between the predicted probability distribution and the true probability distribution, which is usually one-hot encoded. This combination can effectively guide the learning of the model during training, making the model's predicted probability distribution as close to the true distribution as possible.
[0398] According to the loss Compute the gradient:
[0399] ;
[0400] Use the optimizer to update the model parameters:
[0401] ;
[0402] in, is the parameter set of the model, including 、 、 and These parameters are continuously updated and optimized during the training process through back propagation and optimization algorithms, so that the model can gradually learn the association between images and texts, improve the quality and accuracy of generated texts, and generate product information that meets expectations. The learning rate controls the step size of parameter updates and determines the magnitude of parameter updates in each iteration.
[0403] Based on the pre-training dataset Each image-text pair in Corresponding generated text after training , integrated to form a multimodal pre-training model .
[0404] Example 3: This example is an improvement made on the basis of Example 2. In this example, initial information is generated for the target product, specifically:
[0405] Get all the pictures of the target product and generate a picture collection, recorded as ;
[0406] The key features of the target product are extracted from the image collection through the feature extraction function, denoted as :
[0407] ;
[0408] in, For the input image set, each image can be represented as a matrix or tensor, containing pixel values and color channel information. For feature extraction functions, a variety of computer vision techniques can be used, such as color histogram calculation, texture analysis algorithms such as gray-level co-occurrence matrix, edge detection operators such as Canny edge detection, object detection models such as YOLO or Faster R-CNN, etc., to extract features describing the product appearance and content from the image. It is the extracted feature vector, which is used in the subsequent steps to select the appropriate prompt template and generate specific prompt information;
[0409] Get a set of predefined prompt templates, denoted as ;
[0410] For the extracted features, a selection strategy function is used to select a suitable hint template from the hint template set:
[0411] ;
[0412] in, It is a set of predefined prompt templates. Each template is a string containing placeholders for inserting feature values. It is designed to guide the model to generate expected product information, such as describing the product's appearance, functions, etc. It is a selection strategy function that selects the most appropriate template based on the content of the feature and the applicability of the prompt template. For example, it can select the corresponding template based on the type of object detected in the image, such as mobile phones, clothes, etc., or select a more detailed or more concise template based on the richness of the feature. From the prompt template collection A specific template selected from is used to generate prompt information;
[0413] Set up a prompt generation function , used to fill the feature value into the placeholder of the prompt template to generate specific prompt information, and combine the extracted features with the selected prompt template to generate specific prompt information:
[0414] ;
[0415] in, The generated prompt information is used to guide the multimodal pre-trained model to generate the expected product information. The prompt information is clear and specific, and can guide the model to focus on the key product features in the image.
[0416] Obtaining a multimodal pre-trained model ;
[0417] Through multimodal pre-training model Image feature extraction component in , from the picture collection Extract image features from :
[0418] ;
[0419] in, Pre-training models for multimodal The image feature extraction component in the image is usually based on the convolutional neural network (CNN) and is used to extract image features from images. The image feature extraction component is responsible for extracting useful visual features from images. It can identify information such as color distribution, texture pattern, edge contour, and position and shape of objects in the image, and convert the pixel data of the image into a higher-level feature representation, laying the foundation for subsequent feature fusion and product information generation. For pictures from the collection The image features extracted from the image contain visual information in the image, such as color, texture, shape, object position, etc.
[0420] Through multimodal pre-training model Text feature extraction component in , from the prompt information Extract text features from :
[0421] ;
[0422] in, It is a text feature extraction component in a multimodal pre-training model, usually based on Transformer or its variants, used to extract text features from text. The text feature extraction component converts prompt information into text features that the model can understand. It can capture information such as keywords, phrases, and semantic relationships in the text, and convert the semantic connotation of the text into a feature representation in the form of a vector, so that the text information can be effectively integrated with the image information in subsequent steps. For prompt information The text features extracted from the prompt include semantic information in the prompt information, such as keywords, semantic relationships, etc.
[0423] The image features and text features Input to the multimodal feature fusion component In the fusion, we get the multimodal features, which are recorded as :
[0424] ;
[0425] in, It is a feature fusion component in the multimodal pre-training model, used to fuse image features and text features, such as by using attention mechanisms, fully connected networks, etc. The feature fusion component organically integrates image features and text features, so that the information of the two modalities complements and strengthens each other. Through feature fusion, the model can comprehensively utilize visual and text information to generate more comprehensive and accurate product information. For example, the feature fusion component can combine the product color mentioned in the prompt information with the color features extracted from the image to generate a more specific color description. The fused multimodal features combine information from image and text features to generate product information. The fused multimodal features integrate information from both image and text modalities and are a comprehensive feature representation that includes both visual information such as the product's appearance and shape, as well as the semantic requirements of the prompt information. This provides a richer and more complete information foundation for generating components, enabling the generated product information to meet both visual and semantic requirements.
[0426] The fused multimodal features Generative components that are input to the model In the example, the initial product information is generated and recorded as :
[0427] ;
[0428] in, It is the generation component in the multimodal pre-training model, usually based on the Transformer decoder structure, and is used to generate product information based on the fused features. The generation component generates the final product information based on the fused features, can convert multimodal features into natural language descriptions, and output product appearance descriptions, function speculations, and other content that meet user needs. The generation component usually uses technologies such as the attention mechanism to ensure that the generated text is closely related to the input features and has good language fluency and logic. The generated initial product information includes basic content such as appearance description and function speculation. It describes the product's appearance, functions and other basic content in natural language. The initial product information is generated based on the pictures and prompt information uploaded by the seller, and is intended to provide users with accurate and detailed product introductions to help users better understand the product features;
[0429] Multimodal pre-trained models The specific generation process is as follows:
[0430] Multimodal pre-training model Initialize all components in:
[0431] ;
[0432] Among them, the multimodal pre-training model The components in include image feature extraction components , text feature extraction component , feature fusion component And generate components ;
[0433] Get the image-text pair for training, denoted as ;
[0434] Integrate all image-text pairs to generate a pre-training dataset, denoted as :
[0435] ;
[0436] use Extracting images Features:
[0437] ;
[0438] in, It is a parameter set of the image feature extraction component, including convolution layer parameters, normalization layer parameters and fully connected layer parameters. The convolution layer parameters include the weights and biases of the convolution kernels. The convolution kernels are used to slide on the image to extract local features such as edges and textures. For example, a convolution layer may have multiple convolution kernels, each of which is responsible for extracting a specific feature. The normalization layer parameters include the parameters of the batch normalization BatchNorm layer, such as the scaling factor and offset factor, which are used to normalize the output of the convolution layer and stabilize the training process. The fully connected layer parameters are the parameters of the fully connected layer. The fully connected layer is usually in the last few layers of the convolutional neural network CNN. Its parameters include the weight matrix and bias vector, which are used to map the extracted features to a higher-level semantic space.
[0439] use Extract text Features:
[0440] ;
[0441] in, This is a parameter set for the text feature extraction component, including word embedding layer parameters, self-attention layer parameters, and feedforward neural network layer parameters. The word embedding layer maps words or subwords in the text to a vector space of fixed dimension. Its parameter is the embedding matrix, where each row corresponds to the embedding vector of a word or subword. The self-attention layer parameters are the parameters of the self-attention layer. In the Transformer model, the self-attention layer is used to capture the dependencies between words in the text. Its parameters include the weights and biases of the query, key, and value matrices. The feedforward neural network layer parameters are the parameters of the feedforward neural network layer. The feedforward neural network in the Transformer is used to perform nonlinear transformations on the output of the self-attention layer. Its parameters include the weight matrix and bias vector.
[0442] Will and Input In the fusion feature, we get:
[0443] ;
[0444] in, It is a parameter set of the feature fusion component, including the fully connected layer parameters and the attention mechanism parameters. If the feature fusion adopts a simple splicing followed by a fully connected layer, the weight matrix and bias vector of the fully connected layer are used to map the spliced features to the fusion feature space. If the attention mechanism is used for feature fusion, the parameters of the attention layer include the weight matrix and bias vector, which are used to calculate the attention weights between image features and text features, and determine the importance of each modal feature in the fusion process.
[0445] use ,according to , generating text:
[0446] ;
[0447] in, It is a parameter set of the generation component, including word embedding layer parameters, self-attention layer parameters, feedforward neural network layer parameters, and output layer parameters. The word embedding layer parameters are similar to those of the text feature extraction component. The word embedding layer in the generation component maps words or subwords to the vector space. Its parameters are the embedding matrix. The self-attention layer parameters are the parameters of the self-attention layer. During the generation process, the self-attention layer is used to capture the dependencies between generated words. Its parameters include the weights and biases of the query, key, and value matrices. The feedforward neural network layer parameters are used to perform nonlinear transformations on the output of the self-attention layer. Its parameters include the weight matrix and bias vector. The output layer parameters include the weight matrix and bias vector, which are used to map the output of the generation component to the vocabulary space and calculate the generation probability of each word.
[0448] Through the loss function , calculate and generate text With real text The loss between :
[0449] ;
[0450] ;
[0451] Among them, the loss function Specifically, the cross entropy loss function is one of the most commonly used loss functions in natural language processing tasks, especially for text generation tasks. It measures the difference between the word distribution of the generated text and the word distribution of the real text. is the size of the vocabulary, i.e. the number of possible words, For real text Middle One-hot encoding, if If the word is a real word, , otherwise it is 0. For example, if the vocabulary size is 10,000 and the real word is the 100th word in the vocabulary, then , the rest of the positions are 0, To generate text Middle The probability distribution of words, which is the probability of each word at that position predicted by the model, and the probability distribution of generated text It is the result of the softmax layer output by the model, indicating the probability of each word at that position predicted by the model. The softmax layer is a commonly used activation function in neural networks, mainly used in multi-classification problems, converting any real-valued logits output by the model into a probability distribution. The softmax layer is usually used together with the cross-entropy loss function, which measures the difference between the predicted probability distribution and the true probability distribution, which is usually one-hot encoded. This combination can effectively guide the learning of the model during training, making the model's predicted probability distribution as close to the true distribution as possible.
[0452] According to the loss Compute the gradient:
[0453] ;
[0454] Use the optimizer to update the model parameters:
[0455] ;
[0456] in, is the parameter set of the model, including 、 、 and These parameters are continuously updated and optimized during the training process through back propagation and optimization algorithms, so that the model can gradually learn the association between images and texts, improve the quality and accuracy of generated texts, and generate product information that meets expectations. The learning rate controls the step size of parameter updates and determines the magnitude of parameter updates in each iteration.
[0457] Based on the pre-training dataset Each image-text pair in Corresponding generated text after training , integrated to form a multimodal pre-training model ;
[0458] By processing the product images and product descriptions provided by the seller, and substituting the processed data into the multimodal pre-training model , through multimodal pre-training model Components in 、 、 as well as , generate an initial product information .
[0459] The preliminary authenticity analysis of the target product also includes product characteristics analysis, specifically:
[0460] Get the image collection of the target product ;
[0461] From each picture Extract product features , including appearance, function, specifications, materials, etc.:
[0462] ;
[0463] in, Extracting product-related features from images, such as appearance, function, specifications, and materials, helps compare the product with official information and determine whether the product meets expectations.
[0464] Calculate product features for each image Compared with official product features Similarity :
[0465] ;
[0466] in, For the Product features extracted from the image, including appearance, function, specifications, materials, etc., are used for comparison with official product features. Official product features are standard product information obtained from the brand's official website, officially authorized channels, or other trusted sources, and are used as a benchmark for comparison. For the The similarity between the product features in the picture and the official product features measures the degree of match between the product in the picture and the official product. and Represents the product feature vector and the official product feature vector The norm of , usually refers to the L2 norm, that is, the Euclidean norm, that is, the square root of the sum of the squares of the elements in the vector;
[0467] Calculate the average of the product feature similarities of all images as the product detail comparison score , reflecting the average similarity between the product details in all images and the official product details. A higher score indicates that the product details are closer to the official standards:
[0468] ;
[0469] in, Indicates the total number of images in the image collection uploaded by the seller;
[0470] Detect each image Check whether there are brand logos, trademarks and anti-counterfeiting marks in the , where 1 indicates existence and 0 indicates non-existence:
[0471] ;
[0472] in, For the The brand logo detection result in the picture indicates whether key elements such as brand logos, trademarks, and anti-counterfeiting marks are detected, which is the basis for judging brand authenticity. It refers to the detection of specific patterns such as brand logos, trademarks, anti-counterfeiting marks, etc. in images to determine the presence and location of the logo;
[0473] Check the clarity, accuracy and completeness of the detected brand logo and obtain the inspection results , check the results A value between 0 and 1, where 1 indicates complete clarity, accuracy, and completeness:
[0474] ;
[0475] in, For the The results of the clarity, accuracy, and completeness inspection of brand logos in the images are used to assess the quality of brand logos and prevent blurred, incomplete, or forged logos from passing through inspection. This involves assessing the quality of detected brand logos, including clarity, accuracy, and completeness, to ensure they conform to official designs and specifications;
[0476] Calculate the brand logo and logo inspection score by combining the brand logo detection results and inspection results , used to determine the authenticity of a product in terms of brand identity:
[0477] ;
[0478] The final authenticity score is calculated by combining the product details comparison score and the brand logo and logo inspection score. :
[0479] ;
[0480] in, and It is a weight parameter used to balance the contribution of different parts to the final result, satisfying When multiple factors are combined to obtain the final result, different factors may have different influences on the final conclusion. By setting the weight parameter and , the proportion of each factor in the overall evaluation can be adjusted to make the model more in line with actual business needs and data characteristics;
[0481] According to the authenticity score Perform threshold judgment to determine the authenticity of the results :
[0482] ;
[0483] in, Is a predefined threshold used to determine whether the image is real. If the authenticity score Above threshold , it means the picture is highly realistic.
[0484] In this embodiment, preliminary product information is automatically generated through the pictures uploaded by the seller, and it is determined whether the seller's product is an official genuine product. If the product is genuine, the final product information is generated according to the official product information. If the product is a counterfeit, it is determined whether there is user feedback for the product. If there is user feedback, the final product information is generated with reference to the user feedback. If there is no user feedback, the final product information is assisted in generating the final product information through the fusion of knowledge graph and cross-modal retrieval and generation. In combination with multiple means such as image analysis, product feature comparison and user feedback, the authenticity of the product is comprehensively evaluated, which is more accurate and reliable than a single method. The automatic processing extracts the features from the image. It covers the entire process of product information generation and authenticity judgment, reducing manual intervention and improving efficiency. By optimizing model parameters through reinforcement learning, the algorithm can continuously improve based on user feedback, adapt to different products and scenarios, and effectively utilize multimodal information such as images and text to more fully understand product characteristics, improve the quality of authenticity judgment and information generation. When the image quality is poor or the information is incomplete, the fusion of multimodality and multi-methods enhances the robustness and fault tolerance of the algorithm, uses user feedback to form a closed-loop optimization, improves user experience and platform product quality, and reduces disputes caused by counterfeit and shoddy products in scenarios with complex product information such as e-commerce and second-hand trading platforms, thereby maintaining market order.
Claims
1. A method for intelligently supplementing product information based on a large language model, characterized by: include: Get the image collection of the target product , automatically generate preliminary product information; Get image collection The authenticity results of each image in and ; Extract image collection Elements corresponding to all real pictures in: ; in, Indicates traversing the image collection Each picture in ; Set up a product judgment function to determine whether the seller's product is official and authentic: ; in, The threshold is set to determine whether the target product is an official product. If the number of real pictures in the picture collection is If the proportion of exceeds the set threshold, it means that the target product is genuine; like , the target product is initially determined to be genuine, and the final product information is generated according to the official product information; like , the target product is initially analyzed to be a counterfeit product, and all user feedback related to the target product is read from the database or feedback storage system to form a feedback set : ; in, is a function that retrieves all user feedback related to a specific product from a database, file system, or other storage medium; Set a feedback judgment function to determine the user feedback of the target product: ; like , then it is determined that there is user feedback, and the final product information is generated based on the user feedback; like , it is determined that there is no user feedback, and the final product information is assisted in generating through the fusion of knowledge graph and cross-modal retrieval and generation.
2. The method for intelligently supplementing product information based on a large language model according to claim 1, characterized in that: Automatically generate preliminary product information, specifically: Get all the pictures of the target product and generate a picture collection, recorded as ; The key features of the target product are extracted from the image collection through the feature extraction function, denoted as : ; in, is the input image set, is the feature extraction function, is the extracted feature vector; Get a set of predefined prompt templates, denoted as ; For the extracted features, a selection strategy function is used to select a suitable hint template from the hint template set: ; in, Is a set of predefined prompt templates, each template is a string containing placeholders for inserting feature values. is a selection strategy function to select the most appropriate template based on the content of the feature and the applicability of the prompt template, From the prompt template collection A specific template selected from Set up a prompt generation function , combine the extracted features with the selected prompt template to generate specific prompt information: ; in, Is the generated prompt information; Obtaining a multimodal pre-trained model ; Through multimodal pre-training model Image feature extraction component in , from the picture collection Extract image features from : ; in, Pre-training models for multimodal The image feature extraction component in For pictures from the collection The image features extracted from Through multimodal pre-training model Text feature extraction component in , from the prompt information Extract text features from : ; in, It is the text feature extraction component in the multimodal pre-training model. For prompt information The text features extracted from The image features and text features Input to the multimodal feature fusion component In the fusion, we get the multimodal features, which are recorded as : ; in, It is the feature fusion component in the multimodal pre-training model. is the fused multimodal feature; The fused multimodal features Generative components that are input to the model In the example, the initial product information is generated and recorded as : ; in, Generative components in multimodal pre-trained models.
3. The method for intelligently supplementing product information based on a large language model according to claim 1, characterized in that: Conduct preliminary authenticity analysis on the target product, including image feature analysis, specifically: Get the image collection of the target product ; Gray-level co-occurrence matrix , extract each picture Texture features: ; in, For picture collection The images analyzed in Extract each picture through template matching Pattern characteristics : ; Extract each image through Canny edge detection Edge features : ; Calculate each image The Laplace operator response is used to evaluate the blurriness of the image. : ; Detect each image through local consistency check Is there any distortion? : ; Fusion of texture, pattern, edge, blur and distortion features into image detail features : ; For each pair of images and , calculate their image detail features and Similarity : ; Calculate the average similarity of all image pairs as the multi-image consistency score : ; Extract each image Metadata characteristics of : ; Calculate the consistency score of metadata features of all images : ; in, is the reference metadata, is the number of fields matched, is the total number of fields; Use image forensics algorithms to detect whether images have been cropped or spliced : ; Combine metadata consistency and cropping and splicing detection results to calculate image integrity and originality scores : ; in, and It is a weight parameter used to balance the contribution of different parts to the final result, satisfying ; Calculate the average score of image detail features for all images : ; in, Indicates the Picture No. Detailed features; Combine the image detail score, multi-image consistency score, and image integrity and originality score to calculate the final authenticity score : ; in, 、 and It is a weight parameter used to balance the contribution of different parts to the final result, satisfying ; According to the authenticity score Perform threshold judgment to determine the authenticity of the results : ; in, is a predefined threshold.
4. The method for intelligently supplementing product information based on a large language model according to claim 3, characterized in that: The preliminary authenticity analysis of the target product also includes product characteristics analysis, specifically: Get the image collection of the target product ; From each picture Extract product features : ; in, It refers to extracting product-related features from images; Calculate product features for each image Compared with official product features Similarity : ; in, For the Product features extracted from the image, Official product features are standard product information obtained from the brand's official website, officially authorized channels, or other trusted sources. For the The similarity between the product features in the image and the official product features, and Represents the product feature vector and the official product feature vector The norm of Calculate the average of the product feature similarities of all images as the product detail comparison score : ; in, Indicates the total number of images in the image collection uploaded by the seller; Detect each image Check whether there are brand logos, trademarks and anti-counterfeiting marks in the : ; in, For the The brand logo detection result in the picture indicates whether the brand logo, trademark and anti-counterfeiting mark are detected. Refers to detecting specific patterns of brand logos, trademarks, and anti-counterfeiting marks in images; Check the clarity, accuracy and completeness of the detected brand logo and obtain the inspection results : ; in, For the Check results for clarity, accuracy, and completeness of brand logos in images. Refers to the quality assessment of detected brand logos; Calculate the brand logo and logo inspection score by combining the brand logo detection results and inspection results : ; The final authenticity score is calculated by combining the product details comparison score and the brand logo and logo inspection score. : ; in, and It is a weight parameter used to balance the contribution of different parts to the final result, satisfying ; According to the authenticity score Perform threshold judgment to determine the authenticity of the results : ; in, is a predefined threshold.
5. The method for intelligently supplementing product information based on a large language model according to claim 2, characterized in that: If the target product is genuine, the final product information is generated according to the official product information, specifically: Get image collection ; Extract all official information of the target product from the database and integrate it to generate an official information set, denoted as ; Obtaining a multimodal pre-trained model ; Through multimodal pre-training model Image feature extraction component in , from the picture collection Extract image features from : ; Through multimodal pre-training model Text feature extraction component in , from official information Extract text features from : ; The image features and text features Input to the multimodal feature fusion component In the fusion, we get the multimodal features, which are recorded as : ; The fused multimodal features Generative components that feed into the updated model In the process, new product information is generated and recorded as : ; Output the generated new product information .
6. The method for intelligently supplementing product information based on a large language model according to claim 5, characterized in that: If the target product is initially analyzed as a counterfeit and there is user feedback, the final product information is generated based on the user feedback, including model optimization, specifically: Get feedback collection ; For feedback collection Every user feedback in , calculate its relationship with the current product information Cosine similarity of: ; in, For the The norm of the user feedback vector, The norm of the product information vector generated for the model; Based on similarity score , design a reward function to convert the similarity score into a reward signal: ; in, is the similarity threshold; All reward signals Combined into a reward signal vector : ; in, is the number of user feedback; Through the loss function , measure the product information generated by the model and user feedback The differences between: ; ; Among them, the loss function Specifically, it is the mean square error, The product information generated for the model A quantity, For the User feedback Quantity Calculating the loss function Model parameters Gradient: ; Using learning rate And the calculated gradient, update the model parameters: 。 7. The method for intelligently supplementing product information based on a large language model according to claim 6, characterized in that: The generating of final product information by referring to user feedback also includes generating final product information, specifically: Get the model parameters after model optimization ; Get image collection and prompt information ; Obtaining a multimodal pre-trained model ; Extracting multimodal pre-trained models The current model parameters in are denoted as ; The current model parameters Update to the optimized model parameters ; Through multimodal pre-training model Image feature extraction component in , from the picture collection Extract image features from : ; Through multimodal pre-training model Text feature extraction component in , from the prompt information Extract text features from : ; The image features and text features Input to the multimodal feature fusion component In the fusion, we get the multimodal features, which are recorded as : ; The fused multimodal features Generative components that feed into the updated model In the process, new product information is generated and recorded as : ; From new product information Extract key product feature vectors from: ; in, Indicates the The value of the key feature; Assign weights to each key characteristic: ; in, Indicates the The weight of the features, and ; Calculate the total weighted score of the product's authenticity : ; The weighted total score Converted into the final authenticity judgment result : ; in, is the preset threshold; like , then the generated new product information is output , and mark the product as authentic; like , then the generated new product information is output , and mark the product as a counterfeit.
8. The method for intelligently supplementing product information based on a large language model according to claim 7, characterized in that: If the target product is initially analyzed to be a counterfeit and there is no user feedback, the final product information is generated by integrating knowledge graph and cross-modal retrieval and generation. Specifically: Generate a knowledge graph for the target product: ; in, is a collection of entities, is a collection of relations; In the knowledge graph Search and product categories Related entity collections : ; in, Representing an entity Category label of From each entity Extract feature vectors from : ; Get feature vector The weight of : ; Multiply the feature vector of each entity by its corresponding weight, and then add these weighted feature vectors to get the category feature vector : ; Normalize the weighted summed eigenvectors: ; in, represents the L2 norm of the feature vector.
9. The method for intelligently supplementing product information based on a large language model according to claim 8, characterized in that: The method of assisting in generating the final product information by integrating knowledge graph and cross-modal retrieval and generation also includes: In the product image database In the input image Perform cross-modal retrieval to obtain a set of similar products : ; in, Represents a function that determines similar products by calculating the similarity between the image features and the features of images in the database; From each similar product Extract feature vectors from : ; Aggregate the feature vectors of similar products to obtain the feature vectors of similar products : ; in, Similar products The specific formula is: ; in, Represents a collection of similar products For each product image in For each similar product With input picture The cosine similarity is: ; in, Is the input image The eigenvector of Similar products The eigenvector of Represents the L2 norm of the vector; Knowledge graph features and similar product features Fusion is performed to obtain the fusion feature vector : ; in, and is the weight parameter; The fused feature vector Generative components that feed into the updated model In the process, new product information is generated and recorded as : ; From new product information Extract key product feature vectors from: ; in, Indicates the The value of the key feature; Assign weights to each key characteristic: ; in, Indicates the The weight of the features, and ; Calculate the total weighted score of the product's authenticity : ; The weighted total score Converted into the final authenticity judgment result : ; in, is the preset threshold; like , then the generated new product information is output , and mark the product as authentic; like , then the generated new product information is output , and mark the product as a counterfeit.
Citation Information
Patent Citations
Method and system for generating commodity selling point information from e-commerce website link
CN118446776A
Multimodal large language model counterfeit information detection method introducing expert knowledge
CN118606892A