Product information intelligent supplement method based on large language model
Through image feature extraction and multimodal pre-training model combined with user feedback, we automatically judge the authenticity of the product and generate information, solving the problem of not being able to automatically generate product information and judge the authenticity of the product in the existing technology, and achieving efficient and accurate product information generation and authenticity judgment, adapting to different scenarios and reducing market disputes.
Patent Information
- Application Number
- CN202510758740.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-09
AI Technical Summary
The existing intelligent supplementary method of product information based on large language models cannot automatically generate preliminary product information through pictures uploaded by sellers, cannot automatically determine whether the product is an official genuine product, and cannot automatically generate the final product information based on user feedback, resulting in data deviations, reducing efficiency and user experience, and easily causing market disputes.
Through image feature extraction and multimodal pre-training models, combined with user feedback and knowledge graphs, the authenticity of the product is automatically judged and the final information is generated, including image analysis, product feature comparison and user feedback judgment. Multimodal information is used to enhance robustness and fault tolerance, and model parameters are optimized to adapt to different products and scenarios.
It realizes automated product information generation and authenticity judgment, reduces manual intervention, improves efficiency and accuracy, adapts to different products and scenarios, enhances the robustness and fault tolerance of algorithms, reduces disputes between counterfeit and shoddy products, and maintains market order.
Smart Images

Figure CN120338829A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and specifically to an intelligent product information supplement method based on a large language model. Background Art
[0002] With the rapid development of e-commerce, corporate websites, and multi-channel product releases, the richness and accuracy of product information have become key factors in enhancing the user experience and conversion rate. However, manually maintaining a large amount of product information is time-consuming and laborious, and it is prone to problems such as incomplete or inconsistent information. In recent years, large pre-trained language models (such as the GPT series, BERT, etc.) have provided new solutions for automatically generating and supplementing product information with their powerful natural language understanding and generation capabilities. They can automatically generate a large amount of high-quality product information, saving manual time, and can adjust the generated content according to different product categories and market demands. They have semantic understanding and context association capabilities, and the generated content is coherent and rich. With the powerful text generation and understanding capabilities of large language models, it is possible to automatically supplement and improve product information, effectively enhancing the content management efficiency of enterprises and the user experience; Existing intelligent product information supplement methods based on large language models cannot automatically generate preliminary product information from the pictures uploaded by sellers, cannot automatically determine whether the products of sellers are official genuine products, and cannot automatically generate final product information based on user feedback, making the data prone to deviation and requiring manual intervention, reducing efficiency. They cannot adapt to different products and scenarios, reducing the authenticity judgment and information generation quality, reducing the user experience and the quality of platform products, and are prone to disputes caused by counterfeit and shoddy products, thus disrupting the market order. Summary of the Invention
[0003] The present invention provides an intelligent product information supplement method based on a large language model to help solve the problems mentioned in the background art.
[0004] The present invention provides the following technical solution: An intelligent product information supplement method based on a large language model, including: Obtaining a set of pictures of the target product ; Obtaining the authenticity result of each picture in the set of pictures ; and ; Extracting the elements corresponding to all real pictures in the set of pictures ; ; Wherein, represents traversing each picture in the set of pictures ; ; Set up a product judgment function to determine the authenticity of the target product: ; Among them, is the set threshold; If , then the target product is initially judged as genuine; If , then the target product is initially judged as a counterfeit, and all user feedback related to the target product is read from the database or feedback storage system to form a feedback set : ; Among them, is a function to obtain all user feedback related to a specific product from the database, file system, or other storage media; Set up a feedback judgment function to judge the user feedback situation of the target product: ; If , then it is determined that there is user feedback; If , then it is determined that there is no user feedback.
[0005] As an alternative solution of the product information intelligent supplement method based on the large language model described in the present invention, among them: Generate initial information for the target product, specifically: Obtain all pictures of the target product to generate a picture set, denoted as ; Extract the key features of the target product from the picture set through a feature extraction function, denoted as : ; Among them, is the input picture set, is the feature extraction function, is the extracted feature vector; Obtain the predefined prompt template set, denoted as ; For the extracted features, use a selection strategy function to select a suitable prompt template from the prompt template set: ; Among them, is a set of predefined prompt templates, each template is a string containing placeholders for inserting feature values, is the selection strategy function to select the most suitable template according to the content of the feature and the applicability of the prompt template, is from the prompt template set A specific template selected from Set a prompt generation function , combine the extracted features with the selected prompt template to generate specific prompt information: ; Among them, is the generated prompt information; Obtain a multi-modal pre-trained model ; Through the image feature extraction component in the multi-modal pre-trained model , extract image features from the image set , denoted as : ; Among them, is the image feature extraction component in the multi-modal pre-trained model , is the image feature extracted from the image set ; Through the text feature extraction component in the multi-modal pre-trained model , extract text features from the prompt information , denoted as : ; Among them, is the text feature extraction component in the multi-modal pre-trained model, is the text feature extracted from the prompt information ; Input the image feature and the text feature into the multi-modal feature fusion component , and obtain the fused multi-modal feature, denoted as : ; Among them, is the feature fusion component in the multi-modal pre-trained model, is the fused multi-modal feature; Input the fused multi-modal feature into the generation component of the model to generate initial product information, denoted as : ; Among them, is the generation component in the multi-modal pre-trained model.
[0006] As an alternative solution to the method for intelligent supplementation of product information based on large language models according to the present invention, wherein: a preliminary authenticity analysis of the target product is performed, including image feature analysis, specifically: Obtain a set of pictures of the target product ; Through the gray-level co-occurrence matrix , extract the texture features of each picture : ; Among them, is the picture analyzed in the set of pictures ; Through template matching, extract the pattern features of each picture : : ; Through Canny edge detection, extract the edge features of each picture : : ; Calculate the Laplacian operator response of each picture and evaluate the blurriness of the picture : ; Through local consistency check, detect whether there is distortion in each picture : : ; Fuse the texture, pattern, edge, blur, and distortion features into image detail features : ; For each pair of pictures and , calculate the similarity of their image detail features and : : ; Calculate the average value of the similarities of all picture pairs as the multi Figure 1 -consistency score : ; Extract the metadata features of each picture : : ; Calculate the consistency score of the metadata features of all images : ; wherein, is the reference metadata, is the number of matching fields, is the total number of fields; Use an image forensics algorithm to detect whether the image has been cropped or spliced : ; Combine the metadata consistency and the cropping and splicing detection results to calculate the image integrity and originality scores : ; wherein, and are weight parameters used to balance the contributions of different parts to the final result, satisfying ; Calculate the average score of the image detail features of all images : ; wherein, represents the th detail feature of the th Figure 1 image; : ; wherein, , and are weight parameters used to balance the contributions of different parts to the final result, satisfying ; Perform threshold judgment based on the authenticity score to determine the authenticity result : ; wherein, is the predefined threshold.
[0007] As an alternative solution of the product information intelligent supplement method based on the large language model described in the present invention, wherein: the preliminary authenticity analysis of the target product further includes product characteristic analysis, specifically: Obtain a set of images of the target product ; From each image Extract product features : ; Among them, refers to extracting features related to the product from the image; Calculate the product features of each picture and the official product features similarity : ; Among them, is the product feature extracted from the th picture, is the official product feature, which is the standard product information obtained from the brand official website, official authorized channels or other reliable sources, is the similarity between the product feature of the th picture and the official product feature, and represent the norms of the product feature vector and the official product feature vector ; Calculate the average value of the product feature similarities of all pictures as the product detail comparison score : ; Among them, represents the total number of pictures in the set of pictures uploaded by the seller; Detect each picture to check whether there are brand logos, trademarks and anti-counterfeiting marks, and obtain the detection result : ; Among them, is the brand logo detection result in the th picture, indicating whether the brand logo, trademark and anti-counterfeiting marks are detected, refers to the specific pattern for detecting the brand logo, trademark and anti-counterfeiting marks in the image; Check the clarity, accuracy and integrity of the detected brand logo, and obtain the inspection result : ; Among them, is the clarity, accuracy and integrity inspection result of the brand logo in the th picture, refers to the quality assessment of the detected brand logo; Calculate the brand logo and identification inspection score by combining the brand logo detection result and the inspection result : ; Calculate the final authenticity score by combining the product detail comparison score and the brand logo and identification inspection score : ; wherein and are weight parameters used to balance the contributions of different parts to the final result, satisfying ; Perform a threshold judgment based on the authenticity score to determine the authenticity result : ; wherein is a predefined threshold
[0008] As an alternative solution of the product information intelligent supplementation method based on the large language model described in the present invention, wherein: if the target product is genuine, the final product information of the target product is generated accordingly, specifically: Obtain a set of pictures ; Extract all official information of the target product from the database, integrate and generate an official information set, denoted as ; Obtain a multi-modal pre-trained model ; Through the image feature extraction component in the multi-modal pre-trained model , extract image features from the set of pictures , denoted as : ; Through the text feature extraction component in the multi-modal pre-trained model , extract text features from the official information , denoted as : ; Input the image features and the text features into the multi-modal feature fusion component to obtain the fused multi-modal features, denoted as : ; Input the fused multi-modal features The generation component input into the updated model generates new product information, denoted as : ; Outputs the generated new product information .
[0009] As an alternative solution of the product information intelligent supplementation method based on the large language model described in the present invention, wherein: if the target product is initially analyzed as a counterfeit and there is user feedback, then the final product information of the target product is generated with reference to the user feedback, including model optimization, specifically: Obtain the feedback set ; For each piece of user feedback in the feedback set , calculate its cosine similarity with the current product information : ; wherein, is the norm of the th user feedback vector, is the norm of the product information vector generated by the model; According to the similarity score , design a reward function to convert the similarity score into a reward signal: ; wherein, is the similarity threshold; Combine all the reward signals into a reward signal vector : ; wherein, is the number of user feedbacks; Through the loss function , measure the difference between the product information generated by the model and the user feedback : ; ; wherein, the loss function is specifically the mean square error, is the th component in the product information generated by the model, is the th component in the th user feedback; Calculate the loss function The gradients of the model parameters are as follows: ; Using the learning rate and the calculated gradients, update the model parameters: .
[0010] As an alternative solution of the intelligent product information supplementation method based on the large language model according to the present invention, wherein: generating the final product information of the target product based on the reference user feedback further includes generating the final product information, specifically: Obtain the model parameters after model optimization ; Obtain the image set and the prompt information ; Obtain the multi-modal pre-trained model ; Extract the current model parameters in the multi-modal pre-trained model , denoted as ; Update the current model parameters to the optimized model parameters ; Through the image feature extraction component in the multi-modal pre-trained model , extract the image features from the image set , denoted as : ; Through the text feature extraction component in the multi-modal pre-trained model , extract the text features from the prompt information , denoted as : ; Input the image features and the text features into the multi-modal feature fusion component to obtain the fused multi-modal features, denoted as : ; Input the fused multi-modal features into the generation component of the updated model to generate new product information, denoted as : ; Extract the key product feature vectors from the new product information: ; ; Among them, represents the value of the th key feature; Assign weights to each key feature: ; Among them, represents the weight of the th feature, and ; Calculate the weighted total score of the authenticity of the product : ; Convert the weighted total score into the final authenticity judgment result : ; Among them, is the preset threshold; If , then output the generated new product information , and label the product as genuine; If , then output the generated new product information , and label the product as a counterfeit.
[0011] As an alternative solution of the product information intelligent supplement method based on the large language model described in the present invention, wherein: If the target product is initially analyzed as a counterfeit and there is no user feedback, then assist in generating the final product information of the target product, specifically: Generate a knowledge graph for the target product: ; Among them, is a set of entities, is a set of relationships; In the knowledge graph , retrieve the set of entities related to the product category : ; Among them, represents the category label of the entity ; Extract the feature vectors from each entity : ; Obtain the feature vector The weight of, marked as : ; Multiply the feature vector of each entity by its corresponding weight, and then add these weighted feature vectors to obtain the category feature vector : ; Normalize the feature vector after weighted summation: ; Among them, Represents the L2 norm of the feature vector.
[0012] As an alternative solution of the intelligent supplementary method for product information based on the large language model described in the present invention, wherein: the auxiliary generation of the final product information of the target product further includes: In the product picture database Perform cross-modal retrieval on the input picture P to obtain a set of similar products : ; Among them, Represents a function for determining similar products by calculating the feature similarity between the picture features and the pictures in the database; Extract the feature vector From each similar product : ; Aggregate the feature vectors of the similar products to obtain the similar product feature vector : ; Among them, Is the weight of the similar product , and the specific formula is: ; Among them, Represents each product picture in the product picture database , Is each product picture in the database And the input picture The cosine similarity of, specifically: ; Among them, Is the feature vector of the input picture , Is the product picture in the database The eigenvector, represents the L2 norm of the vector; Fuse the knowledge graph features and the similar product features to obtain a fused feature vector : ; wherein, and are weight parameters; Input the fused feature vector into the generation component of the updated model to generate new product information, denoted as : ; Extract the key product feature vector from the new product information : ; wherein, represents the value of the th key feature; Assign weights to each key feature: ; wherein, represents the weight of the th feature, and ; Calculate the weighted total score of the authenticity of the product : ; Convert the weighted total score into the final authenticity judgment result : ; wherein, is a preset threshold; If , then output the generated new product information , and label the product as genuine; If , then output the generated new product information , and label the product as a counterfeit.
[0013] The present invention has the following beneficial effects: 1. The intelligent product information supplementation method based on large language models automatically generates preliminary product information through the pictures uploaded by sellers, and combines image analysis and product feature comparison to determine whether the sellers' products are official genuine products. It automates the entire process from image feature extraction to product information generation and authenticity judgment, reduces manual intervention, improves efficiency, can more comprehensively evaluate product authenticity, is more accurate and reliable than single methods, effectively utilizes multi-modal information such as images and texts, more fully understands product characteristics, improves the quality of authenticity judgment and information generation. When the image quality is poor or the information is incomplete, the fusion of multi-modal and multi-methods enhances the robustness and fault tolerance of the algorithm.
[0014] 2. For the intelligent product information supplementation method based on large language models, if the product is genuine, the final product information is generated according to the official product information. If the product is a counterfeit, user feedback data of the product is extracted from the database to determine whether there is user feedback. In scenarios with complex product information such as e-commerce and second-hand trading platforms, it helps to combat counterfeits and maintain market order, forms a closed-loop optimization using user feedback, and improves user experience and the quality of platform commodities.
[0015] 3. For the intelligent product information supplementation method based on large language models, if there is user feedback, the final product information is generated with reference to the situation of user feedback. If there is no user feedback, the final product information is assisted to be generated through the fusion of knowledge graph and cross-modal retrieval and generation, and the authenticity of the finally generated product information is judged again to determine whether the product is genuine. The model parameters are optimized through reinforcement learning, enabling the algorithm to continuously improve according to user feedback, adapt to different products and scenarios, form a closed-loop optimization using user feedback, and improve user experience and the quality of platform commodities. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a flowchart of the intelligent product information supplementation method based on large language models of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0018] Embodiment 1. An intelligent product information supplementation method based on large language models, referring to Figure 1 , includes: Obtain a set of pictures of the target product ; Get the image collection The authenticity result of each image in and ; Extract image collection Elements corresponding to all real pictures in: ; in, Represents traversal of a collection of images Each picture in ; Set up a product judgment function to determine the authenticity of the target product: ; in, is the set threshold used to determine whether the target product is an official product. If the number of real pictures is in the picture collection If the percentage of exceeds the set threshold, it means that the target product is genuine; like , then the target product is preliminarily determined to be genuine; like , the target product is initially determined to be a counterfeit product, and all user feedback related to the target product is read from the database or feedback storage system to form a feedback set : ; in, It is a function that obtains all user feedback related to a specific product from a database, file system or other storage medium. It mainly collects user evaluations, opinions or suggestions on the product for further analysis and processing. The function is to retrieve all feedback related to a specific product from the place where user feedback is stored. Usually a parameter is required to specify the product to be retrieved, such as product ID or product name. The returned result is usually a collection or list containing multiple user feedbacks; Set a feedback judgment function to determine the user feedback of the target product: ; like , it is determined that there is user feedback; like , it is determined that there is no user feedback.
[0019] This embodiment also provides that if the target product is genuine, the final product information of the target product is generated accordingly, specifically, wherein the way to determine that the target product is genuine includes determining that the target product is initially analyzed to be genuine and when Output the new product information generated , and label the product as genuine: Obtain an image set ; Extract all official information of the target product from the database, integrate and generate an official information set, denoted as ; Obtain a multi-modal pre-trained model ; Through the image feature extraction component in the multi-modal pre-trained model , extract image features from the image set , denoted as : ; Through the text feature extraction component in the multi-modal pre-trained model , extract text features from the official information , denoted as : ; Input the image features and the text features into the multi-modal feature fusion component , and obtain the fused multi-modal features, denoted as : ; Input the fused multi-modal features into the generation component of the updated model to generate new product information, denoted as : ; Output the generated new product information ; In this embodiment, after determining that the target product is genuine, by extracting the official information of the brand corresponding to the target product in the database , substituting the official information into the multi-modal pre-trained model , through the components in the multi-modal pre-trained model , , and , generate the final, more complete product information that can be shown to the buyer .
[0020] This embodiment also provides that if the target product is initially analyzed as a counterfeit and there is user feedback, the final product information of the target product is generated with reference to the user feedback, including model optimization, specifically: Obtain the feedback set ; For each piece of user feedback in the feedback set , calculate its cosine similarity with the current product information , reflecting the degree of correlation between the two: ; Among them, is the norm of the th user feedback vector, used to normalize the vector, is the norm of the product information vector generated by the model, used to normalize the vector; According to the similarity score , design a reward function to convert the similarity score into a reward signal, used to guide the update of model parameters, where high similarity corresponds to high reward: ; Among them, is the similarity threshold, used to filter out feedback with low similarity, ensuring that only valuable feedback affects model update; Combine all the reward signals into a reward signal vector , summarize the reward signals corresponding to all user feedback, and use them in the reinforcement learning process to optimize the model: ; Among them, is the number of user feedbacks, used to determine the dimension of the reward signal vector; Through the loss function , measure the difference between the product information generated by the model and the user feedback : ; ; Among them, the loss function is specifically the mean squared error MSE, is the th component in the product information generated by the model, is the th th component in the th user feedback; Calculate the gradient of the loss function with respect to the model parameters, used to guide the parameter update direction: ; Utilize the learning rate and the calculated gradient to update the model parameters: ; wherein, the learning rate is used to control the step size of parameter update.
[0021] Among them, generating the final product information of the target product based on the reference user feedback further includes generating the final product information, specifically: Obtain the model parameters after model optimization ; Obtain the image set and the prompt information ; Obtain the multi-modal pre-trained model ; Extract the current model parameters in the multi-modal pre-trained model and denote them as ; Update the current model parameters to the optimized model parameters ; Through the image feature extraction component in the multi-modal pre-trained model , extract the image features from the image set and denote them as : ; Through the text feature extraction component in the multi-modal pre-trained model , extract the text features from the prompt information and denote them as : ; Input the image features and the text features into the multi-modal feature fusion component to obtain the fused multi-modal features, denoted as : ; Input the fused multi-modal features into the generation component of the updated model to generate new product information, denoted as : ; From the new product information Extract the key product feature vectors: ; Among them, represents the value of the th key feature; Assign weights to each key feature to determine the relative importance of each key feature in authenticity judgment: ; Among them, represents the weight of the th feature, and ; Calculate the weighted total score of the authenticity of the product : ; Convert the weighted total score into the final authenticity judgment result : ; Among them, is a preset threshold for converting the weighted total score into an authenticity judgment result; If , then output the generated new product information , and label the product as genuine; If , then output the generated new product information , and label the product as a counterfeit; In this embodiment, after initially determining that the target product is a counterfeit, combined with the user's feedback on the product, update the model parameters of the multi-modal pre-training model , and substitute the pictures uploaded by the seller and the description text of the product into the multi-modal pre-training model , through the components in the multi-modal pre-training model , , and , generate a new and more perfect product information , and analyze the generated new product information, judge the authenticity of the product again, determine the final product information to be shown to the buyer according to the result of the re-judgment, and label the authenticity of the product.
[0022] This embodiment also provides that if the target product is initially analyzed as a counterfeit and there is no user feedback, assist in generating the final product information of the target product, specifically: Generate a knowledge graph for the target product: ; Among them, is a collection of entities, including products, brands, functions, etc., is a collection of relationships, such as brand - product, product - function, etc.; In the knowledge graph retrieve the collection of entities related to the product category : : ; Among them, represents the category label of the entity ; Extract the feature vector from each entity , and these features can include the attributes of the entity, such as brand name, function description, etc.: ; Obtain the weight of the feature vector , marked as : ; Multiply the feature vector of each entity by its corresponding weight, and then add these weighted feature vectors to obtain the category feature vector : ; Normalize the feature vector after weighted summation: ; Among them, represents the L2 norm of the feature vector.
[0023] Among them, the auxiliary generation of the final product information of the target product further includes: In the product image database perform cross - modal retrieval on the input image P to obtain the set of similar products : ; Among them, represents the function of determining similar products by calculating the feature similarity between the image features and the images in the database; Extract the feature vector from each similar product , such as appearance description, function characteristics, etc.: ; Among them, A function that extracts feature vectors from images or other data types. The feature vectors can contain various information about the image, such as color, texture, shape, edges, etc. In image processing and machine learning tasks, converting an image into a feature vector enables a computer to understand and process the image data. These feature vectors can be used for tasks such as classification, retrieval, and clustering; Aggregate the feature vectors of similar products to obtain the feature vectors of similar products : ; Among them, is the weight of the similar product , and the specific formula is: ; Among them, represents each product image in the product image database , which is used to traverse all product images in the database in order to calculate their similarity with the input image , is the cosine similarity between each product image in the database and the input image , specifically: ; Among them, is the feature vector of the input image , is the feature vector of the product image in the database, represents the L2 norm of the vector; Fuse the knowledge graph features and the similar product features to obtain the fused feature vector : ; Among them, and are weight parameters used to balance the contributions of the two features, that is, to determine the relative importance of each feature for the fused feature vector; Input the fused feature vector into the generation component of the updated model to generate new product information, denoted as : ; Extract the key product feature vector from the new product information : ; Among them, Represents the value of the th key feature; Assign weights to each key feature to determine the relative importance of each key feature in authenticity judgment: ; Among them, Represents the weight of the th feature, and ; Calculate the weighted total score of the authenticity of the product : ; Convert the weighted total score into the final authenticity judgment result : ; Among them, is a preset threshold for converting the weighted total score into an authenticity judgment result; If , then output the generated new product information , and label the product as genuine; If , then output the generated new product information , and label the product as a counterfeit.
[0024] Through the above method, based on the pictures uploaded by the seller, automatically generate preliminary product information, and judge whether the seller's product is an official genuine product. If the product is genuine, generate the final product information according to the official product information. If the product is a counterfeit, judge whether there is user feedback. If there is user feedback, refer to the user feedback to generate the final product information. If there is no user feedback, assist in generating the final product information through the integration of knowledge graph and cross-modal retrieval and generation, combining multiple means such as image analysis, product feature comparison, and user feedback to comprehensively evaluate the authenticity of the product, which is more accurate and reliable than a single method. Automatically process the entire process from image feature extraction to product information generation and authenticity judgment, reduce manual intervention, improve efficiency, optimize the model parameters through reinforcement learning, enable the algorithm to continuously improve according to user feedback, adapt to different products and scenarios, effectively utilize multi-modal information such as images and texts, more fully understand the product characteristics, improve the authenticity judgment and information generation quality. When the image quality is poor or the information is incomplete, the multi-modal and multi-method fusion enhances the robustness and fault tolerance of the algorithm, uses user feedback to form a closed-loop optimization, improves the user experience and the quality of platform commodities, reduces disputes caused by counterfeit and shoddy products in scenarios with complex product information such as e-commerce and second-hand trading platforms, and maintains market order.
[0025] Example 2. This example is an improvement based on Example 1. For the intelligent product information supplementation method based on a large language model, initial information is generated for the target product, specifically as follows: Obtain all the pictures of the target product, generate a picture set, denoted as ; Extract the key features of the target product from the picture set through a feature extraction function, denoted as : ; Among them, is the input picture set. Each picture can be represented as a matrix or tensor, containing pixel values and color channel information. is the feature extraction function. Multiple computer vision techniques can be used, such as color histogram calculation, texture analysis algorithms such as gray-level co-occurrence matrix, edge detection operators such as Canny edge detection, object detection models such as YOLO or Faster R-CNN, etc., to extract the features describing the product appearance and content from the pictures. is the extracted feature vector, which is used to select a suitable prompt template and generate specific prompt information in the subsequent steps; Obtain the predefined set of prompt templates, denoted as ; For the extracted features, use a selection strategy function to select a suitable prompt template from the set of prompt templates: ; Among them, is a set of predefined prompt templates. Each template is a string, containing placeholders for inserting feature values, aiming to guide the model to generate product information that meets expectations, such as describing the product appearance, function, etc. is the selection strategy function, which selects the most suitable template according to the content of the features and the applicability of the prompt templates. For example, it can select the corresponding template according to the object types detected in the pictures, such as mobile phones, clothes, etc., or select a more detailed or more concise template according to the richness of the features. is a specific template selected from the set of prompt templates for generating prompt information; Set a prompt generation function , which is used to fill the feature values into the placeholders of the prompt template to generate specific prompt information, combine the extracted features with the selected prompt template to generate specific prompt information: ; Among them, It is the generated prompt information used to guide the multi-modal pre-trained model to generate product information that meets expectations. Among them, the prompt information is clear and specific, and can guide the model to focus on the key product features in the picture; Obtain the multi-modal pre-trained model ; Through the image feature extraction component in the multi-modal pre-trained model , extract image features from the picture set , denoted as : ; Among them, is the image feature extraction component in the multi-modal pre-trained model, usually based on the convolutional neural network CNN, used to extract image features from pictures. The image feature extraction component is responsible for extracting useful visual features from pictures. It can identify information such as color distribution, texture pattern, edge contour, and the position and shape of objects in the picture, and convert the pixel data of the picture into a higher-level feature representation, laying a foundation for subsequent feature fusion and product information generation, is the image feature extraction component in the multi-modal pre-trained model, usually based on the convolutional neural network CNN, used to extract image features from pictures. The image feature extraction component is responsible for extracting useful visual features from pictures. It can identify information such as color distribution, texture pattern, edge contour, and the position and shape of objects in the picture, and convert the pixel data of the picture into a higher-level feature representation, laying a foundation for subsequent feature fusion and product information generation, is the image feature extracted from the picture set , containing visual information in the picture, such as color, texture, shape, object position, etc.; Through the text feature extraction component in the multi-modal pre-trained model , extract text features from the prompt information , denoted as : ; Among them, is the text feature extraction component in the multi-modal pre-trained model, usually based on Transformer or its variants, used to extract text features from text. The text feature extraction component converts the prompt information into text features that the model can understand. It can capture information such as keywords, phrases, and semantic relationships in the text, and convert the semantic connotation of the text into a feature representation in vector form, enabling the text information to be effectively fused with the image information in subsequent steps, is the text feature extracted from the prompt information , containing semantic information in the prompt information, such as keywords, semantic relationships, etc.; Input the image feature and the text feature into the multi-modal feature fusion component , and obtain the fused multi-modal feature, denoted as : ; Among them, is a feature fusion component in the multi-modal pre-training model, which is used to fuse image features and text features, such as by using attention mechanisms, fully connected networks, etc. The feature fusion component organically fuses image features and text features, enabling the information of the two modalities to complement and reinforce each other. Through feature fusion, the model can comprehensively utilize visual and text information to generate more comprehensive and accurate product information. For example, the feature fusion component can combine the product color mentioned in the prompt information with the color features extracted from the picture to generate a more specific color description. is the fused multi-modal feature, which synthesizes the information of image features and text features and is used to generate product information. The fused multi-modal feature integrates the information of the image and text modalities and is a comprehensive feature representation that contains both visual information such as the appearance and shape of the product and the semantic requirements in the prompt information, providing a richer and more complete information basis for the generation component, enabling the generated product information to meet both visual and semantic requirements simultaneously; Input the fused multi-modal feature into the generation component of the model to generate the initial product information, denoted as : ; Among them, is the generation component in the multi-modal pre-training model, usually based on the decoder structure of Transformer, which is used to generate product information according to the fused features. The generation component generates the final product information according to the fused features, can transform the multi-modal features into natural language descriptions, and outputs content such as product appearance descriptions and function speculations that meet the user's needs. The generation component usually uses technologies such as attention mechanisms to ensure that the generated text is closely related to the input features and has good language fluency and logic. is the generated initial product information, including basic content such as appearance descriptions and function speculations, which describes the appearance, function, etc. of the product in the form of natural language. The initial product information is generated based on the pictures uploaded by the seller and the prompt information, aiming to provide users with accurate and detailed product introductions to help users better understand the product features; Among them, the generation process of the multi-modal pre-training model is specifically as follows: Initialize all components in the multi-modal pre-training model : ; Among them, the components in the multi-modal pre-training model include the image feature extraction component , a text feature extraction component , a feature fusion component and a generation component ; Obtain image-text pairs for training, denoted as ; Integrate all image-text pairs to generate a pre-training dataset, denoted as : ; Use to extract the features of the image : ; Among them, is the parameter set of the image feature extraction component, including convolutional layer parameters, normalization layer parameters, and fully connected layer parameters. Among them, the convolutional layer parameters include the weights and biases of the convolutional kernels. The convolutional kernels are used to slide on the image to extract local features such as edges and textures. For example, a convolutional layer may have multiple convolutional kernels, and each convolutional kernel is responsible for extracting a specific feature. The normalization layer parameters include the parameters of the BatchNorm layer, such as the scale factor and the offset factor, which are used to normalize the output of the convolutional layer and stabilize the training process. The fully connected layer parameters are the parameters of the fully connected layer. The fully connected layer is usually in the last few layers of the convolutional neural network (CNN), and its parameters include the weight matrix and the bias vector, which are used to map the extracted features to a higher-level semantic space; Use to extract the features of the text : ; Among them, is the parameter set of the text feature extraction component, including word embedding layer parameters, self-attention layer parameters, and feed-forward neural network layer parameters. The word embedding layer maps words or sub-words in the text to a vector space of a fixed dimension, and its parameter is the embedding matrix, where each row corresponds to the embedding vector of a word or sub-word. The self-attention layer parameters are the parameters of the self-attention layer. In the Transformer model, the self-attention layer is used to capture the dependencies between words in the text, and its parameters include the weights and biases of the query, key, and value matrices. The feed-forward neural network layer parameters are the parameters of the feed-forward neural network layer. The feed-forward neural network in the Transformer is used to perform a non-linear transformation on the output of the self-attention layer, and its parameters include the weight matrix and the bias vector; Input and into to obtain the fused features: ; Among them, is the parameter set of the feature fusion component, including the fully connected layer parameters and the attention mechanism parameters. If the feature fusion adopts the method of simple concatenation followed by a fully connected layer, the weight matrix and bias vector of the fully connected layer are used to map the concatenated features to the fusion feature space. If the attention mechanism is used for feature fusion, the parameters of the attention layer include the weight matrix and bias vector, which are used to calculate the attention weights between the image features and the text features, and determine the importance of each modality feature in the fusion process; Use , according to , generate text: ; Among them, is the parameter set of the generation component, including the word embedding layer parameters, the self-attention layer parameters, the feed-forward neural network layer parameters, and the output layer parameters. The word embedding layer parameters are similar to those of the text feature extraction component. The word embedding layer in the generation component maps words or sub-words to the vector space, and its parameter is the embedding matrix. The self-attention layer parameters are the parameters of the self-attention layer. During the generation process, the self-attention layer is used to capture the dependencies between the generated words, and its parameters include the weights and biases of the query, key, and value matrices. The feed-forward neural network layer parameters are used to perform non-linear transformation on the output of the self-attention layer, and its parameters include the weight matrix and bias vector. The output layer parameters include the weight matrix and bias vector, which are used to map the output of the generation component to the vocabulary space and calculate the generation probability of each word; Through the loss function , calculate the loss between the generated text and the real text , denoted as : ; ; Among them, the loss function is specifically the cross-entropy loss function. Cross-entropy loss is one of the most commonly used loss functions in natural language processing tasks, especially for text generation tasks. It measures the difference between the word distribution of the generated text and the word distribution of the real text. is the size of the vocabulary, that is, the number of possible words. is the one-hot encoding of the th word in the real text . If the th word is the real word, then , otherwise it is 0. For example, if the vocabulary size is 10000 and the real word is the 100th word in the vocabulary, then , and the rest of the positions are 0. To generate text The probability distribution of the th word in , which is the probability of each word at this position predicted by the model, and the probability distribution of the generated text is the result of the softmax layer output by the model, indicating the probability of each word at this position predicted by the model. Among them, the softmax layer is a commonly used activation function in neural networks, mainly used in multi-classification problems to convert any real-valued logits output by the model into a probability distribution. The Softmax layer is usually used together with the cross-entropy loss function. The cross-entropy loss measures the difference between the predicted probability distribution and the true probability distribution, usually in one-hot encoding. This combination can effectively guide the learning of the model during the training process, making the predicted probability distribution of the model as close as possible to the true distribution; Calculate the gradient according to the loss : ; Update the model parameters using an optimizer: ; Among them, is the set of model parameters, including , , and . These parameters are continuously updated and optimized through backpropagation and optimization algorithms during the training process, enabling the model to gradually learn the association between images and texts, improving the quality and accuracy of the generated text, and generating product information that meets expectations. is the learning rate, which controls the step size of parameter updates and determines the magnitude of parameter updates in each iteration of the model; According to each image-text pair in the pre-training dataset corresponding to the generated text after training, integrate and form a multi-modal pre-training model . .
[0026] Example 3. This example is an improvement based on Example 2. In this example, for the target product, initial information is generated as follows: Obtain all the pictures of the target product, generate a picture set, denoted as ; Extract the key features of the target product from the picture set through a feature extraction function, denoted as : ; Among them, is the input picture set. Each picture can be represented as a matrix or tensor, containing pixel values and color channel information. As a feature extraction function, various computer vision techniques can be used, such as color histogram calculation, texture analysis algorithms like gray-level co-occurrence matrix, edge detection operators like Canny edge detection, object detection models like YOLO or Faster R-CNN, etc., to extract features describing the appearance and content of the product from the picture. is the extracted feature vector, which is used to select a suitable prompt template and generate specific prompt information in the subsequent steps; Obtain the predefined set of prompt templates, denoted as ; For the extracted features, use a selection strategy function to select a suitable prompt template from the set of prompt templates: ; Among them, is a set of predefined prompt templates, and each template is a string containing placeholders for inserting feature values, aiming to guide the model to generate product information that meets expectations, such as describing the appearance, function, etc. of the product. is the selection strategy function, which selects the most suitable template according to the content of the features and the applicability of the prompt templates. For example, it can select the corresponding template according to the object types detected in the picture, such as mobile phones, clothes, etc., or select a more detailed or more concise template according to the richness of the features. is a specific template selected from the set of prompt templates and is used to generate prompt information; Set a prompt generation function , which is used to fill the feature values into the placeholders of the prompt template to generate specific prompt information, combining the extracted features with the selected prompt template to generate specific prompt information: ; Among them, is the generated prompt information, which is used to guide the multi-modal pre-trained model to generate product information that meets expectations. Among them, the prompt information is clear and specific, and can guide the model to focus on the key product features in the picture. Obtain the multi-modal pre-trained model ; Through the image feature extraction component in the multi-modal pre-trained model , extract the image features from the picture set , denoted as : ; Among them, is the multi-modal pre-trained model The image feature extraction component in it, usually based on the convolutional neural network CNN, is used to extract image features from pictures. The image feature extraction component is responsible for extracting useful visual features from pictures. It can identify information such as color distribution, texture patterns, edge contours, and the position and shape of objects in the pictures, converting the pixel data of the pictures into a higher-level feature representation, laying a foundation for subsequent feature fusion and product information generation. The image features extracted from the picture set contain visual information in the pictures, such as color, texture, shape, object position, etc.; Through the text feature extraction component in the multi-modal pre-trained model extract text features from the prompt information , denoted as : ; Among them, is the text feature extraction component in the multi-modal pre-trained model, usually based on Transformer or its variants, and is used to extract text features from text. The text feature extraction component converts the prompt information into text features that the model can understand. It can capture information such as keywords, phrases, and semantic relationships in the text, converting the semantic connotation of the text into a feature representation in vector form, enabling the text information to be effectively fused with the image information in subsequent steps. is the text feature extracted from the prompt information , containing semantic information in the prompt information, such as keywords, semantic relationships, etc.; Input the image features and text features into the multi-modal feature fusion component to obtain the fused multi-modal features, denoted as : ; Among them, is the feature fusion component in the multi-modal pre-trained model, which is used to fuse image features and text features, such as by using attention mechanisms, fully connected networks, etc. The feature fusion component organically fuses image features and text features, making the information of the two modalities complement and reinforce each other. Through feature fusion, the model can comprehensively utilize visual and text information to generate more comprehensive and accurate product information. For example, the feature fusion component can combine the product color mentioned in the prompt information with the color features extracted from the picture to generate a more specific color description. The fused multi-modal features integrate the information of image features and text features and are used to generate product information. The fused multi-modal features integrate the information of both image and text modalities and are a comprehensive feature representation. They contain both visual information such as the appearance and shape of the product and the semantic requirements in the prompt information, providing a richer and more complete information basis for the generation component, enabling the generated product information to meet both visual and semantic requirements simultaneously; Input the fused multi-modal features into the generation component of the model to generate the initial product information, denoted as : ; Among them, is the generation component in the multi-modal pre-training model, usually based on the decoder structure of Transformer, and is used to generate product information according to the fused features. The generation component generates the final product information according to the fused features, can transform the multi-modal features into natural language descriptions, and output content such as product appearance descriptions and function speculations that meet user needs. The generation component usually uses techniques such as attention mechanisms to ensure that the generated text is closely related to the input features and has good language fluency and logic, is the generated initial product information, including basic content such as appearance descriptions and function speculations, which describes the appearance, function, etc. of the product in the form of natural language. The initial product information is generated based on the pictures uploaded by the seller and the prompt information, aiming to provide users with accurate and detailed product introductions to help users better understand the product features; The multi-modal pre-training model The specific generation process is as follows: Initialize all components in the multi-modal pre-training model : ; Among them, the components in the multi-modal pre-training model include the image feature extraction component , the text feature extraction component , the feature fusion component and the generation component ; Obtain the image-text pairs for training, denoted as ; Integrate all the image-text pairs to generate the pre-training dataset, denoted as : ; Use to extract the features of the image : ; Among them, is the parameter set of the image feature extraction component, including convolutional layer parameters, normalization layer parameters, and fully connected layer parameters. Among them, the convolutional layer parameters include the weights and biases of the convolutional kernels. The convolutional kernels are used to slide on the image to extract local features such as edges and textures. For example, a convolutional layer may have multiple convolutional kernels, and each convolutional kernel is responsible for extracting a specific feature. The normalization layer parameters include the parameters of the batch normalization (BatchNorm) layer, such as the scale factor and the offset factor, which are used to normalize the output of the convolutional layer and stabilize the training process. The fully connected layer parameters are the parameters of the fully connected layer. The fully connected layer is usually in the last few layers of the convolutional neural network (CNN), and its parameters include the weight matrix and the bias vector, which are used to map the extracted features to a higher-level semantic space; Use to extract the features of the text : ; Among them, is the parameter set of the text feature extraction component, including word embedding layer parameters, self-attention layer parameters, and feed-forward neural network layer parameters. The word embedding layer maps words or sub-words in the text to a vector space of a fixed dimension, and its parameter is the embedding matrix, where each row corresponds to the embedding vector of a word or sub-word. The self-attention layer parameters are the parameters of the self-attention layer. In the Transformer model, the self-attention layer is used to capture the dependencies between words in the text, and its parameters include the weights and biases of the query, key, and value matrices. The feed-forward neural network layer parameters are the parameters of the feed-forward neural network layer. The feed-forward neural network in the Transformer is used to perform a non-linear transformation on the output of the self-attention layer, and its parameters include the weight matrix and the bias vector; Input and into to obtain the fused features: ; Among them, is the parameter set of the feature fusion component, including fully connected layer parameters and attention mechanism parameters. If the feature fusion adopts the method of simple concatenation followed by a fully connected layer, the weight matrix and the bias vector of the fully connected layer are used to map the concatenated features to the fused feature space. If the attention mechanism is used for feature fusion, the parameters of the attention layer include the weight matrix and the bias vector, which are used to calculate the attention weights between the image features and the text features and determine the importance of each modality feature in the fusion process; Use , according to , to generate the text: ; Among them, is the parameter set of the generation component, including word embedding layer parameters, self-attention layer parameters, feed-forward neural network layer parameters, and output layer parameters. The word embedding layer parameters are similar to those of the text feature extraction component. The word embedding layer in the generation component maps words or sub-words to a vector space, and its parameter is the embedding matrix. The self-attention layer parameters are the parameters of the self-attention layer. During the generation process, the self-attention layer is used to capture the dependencies between the generated words, and its parameters include the weights and biases of the query, key, and value matrices. The feed-forward neural network layer parameters are used to perform a non-linear transformation on the output of the self-attention layer, and its parameters include the weight matrix and bias vector. The output layer parameters include the weight matrix and bias vector, which are used to map the output of the generation component to the vocabulary space and calculate the generation probability of each word; Through the loss function , calculate the loss between the generated text and the real text , denoted as : ; ; Among them, the loss function is specifically the cross-entropy loss function. Cross-entropy loss is one of the most commonly used loss functions in natural language processing tasks, especially for text generation tasks. It measures the difference between the word distribution of the generated text and the word distribution of the real text. is the size of the vocabulary, that is, the number of possible words. is the one-hot encoding of the st word in the real text . If the th word is the real word, then , otherwise it is 0. For example, if the vocabulary size is 10000 and the real word is the 100th word in the vocabulary, then , and the rest of the positions are 0. is the probability distribution of the th word in the generated text . This is the probability of each word predicted by the model at this position. The probability distribution of the generated text It is the result of the softmax layer of the model output, representing the probability of each word at that position predicted by the model. Among them, the softmax layer is a commonly used activation function in neural networks, mainly used in multi-classification problems to convert any real-valued logits output by the model into a probability distribution. The Softmax layer is usually used together with the cross-entropy loss function. The cross-entropy loss measures the difference between the predicted probability distribution and the true probability distribution, usually in one-hot encoding. This combination can effectively guide the learning of the model during the training process, making the predicted probability distribution of the model as close as possible to the true distribution; According to the loss Calculate the gradient: ; Use the optimizer to update the model parameters: ; Among them, is the set of model parameters, including , , and . These parameters are continuously updated and optimized through backpropagation and optimization algorithms during the training process, enabling the model to gradually learn the association between images and texts, improve the quality and accuracy of the generated texts, and generate product information that meets expectations. is the learning rate, which controls the step size of parameter updates and determines the amplitude of parameter updates in each iteration of the model; According to each image-text pair in the pre-training dataset corresponding to the generated text after training, integrate and form a multi-modal pre-training model ; By processing the product pictures and product descriptions provided by the seller and substituting the processed data into the multi-modal pre-training model , through the components in the multi-modal pre-training model , , and , generate an initial product information .
[0027] Among them, the preliminary authenticity analysis of the target product also includes product feature analysis, specifically: Obtain the picture set of the target product ; Extract product features from each picture , including appearance, function, specifications, materials, etc.: ; Among them, refers to extracting product-related features from images, such as appearance, function, specifications, materials, etc. These features can help compare the product with official information and determine whether the product meets expectations; Calculating the product features of each image and the official product features similarity : ; Among them, is the product feature extracted from the th image, including appearance, function, specifications, materials, etc., and is used for comparison with the official product features. is the official product feature, which is the standard product information obtained from the brand's official website, official authorized channels or other reliable sources and is used as the benchmark for comparison. is the similarity between the product feature of the th image and the official product feature, which measures the matching degree between the product in the image and the official product. and represent the norms of the product feature vector and the official product feature vector , usually referring to the L2 norm, that is, the Euclidean norm, which is the square root of the sum of the squares of the elements in the vector; Calculating the average value of the product feature similarities of all images as the product detail comparison score , reflecting the average similarity between the product details in all images and the official product details. The higher the score, the closer the product details are to the official standard: ; Among them, represents the total number of images in the set of images uploaded by the seller; Detecting whether there are brand logos, trademarks and anti-counterfeiting marks in each image to obtain the detection result , among which, 1 means present and 0 means absent: ; Among them, is the brand logo detection result of the th image, indicating whether key elements such as brand logos, trademarks, and anti-counterfeiting marks are detected, and is the basis for judging brand authenticity. refers to detecting specific patterns such as brand logos, trademarks, and anti-counterfeiting marks in the image to determine whether the logo is present and its location; Check the clarity, accuracy, and integrity of the detected brand logo to obtain the inspection results , the inspection results are values between 0 and 1, where 1 indicates completely clear, accurate, and complete: ; Among them, is the inspection result of the clarity, accuracy, and integrity of the brand logo in the th picture, which is used to evaluate the quality of the brand logo and prevent blurred, incomplete, or forged logos from passing through refers to the quality assessment of the detected brand logo, including clarity, accuracy, and integrity, to ensure that the logo meets the official design and specifications; Combine the brand logo detection results and inspection results to calculate the brand logo and identification inspection scores , which is used to judge the authenticity of the product in terms of brand identification: ; Combine the product detail comparison score and the brand logo and identification inspection score to calculate the final authenticity score : ; Among them, and are weight parameters used to balance the contributions of different parts to the final result, satisfying . When obtaining the final result by integrating multiple factors, the influence degrees of different factors on the final conclusion may be different. By setting the weight parameters and , the proportion of each factor in the overall evaluation can be adjusted to make the model more in line with the actual business requirements and data characteristics; Make a threshold judgment based on the authenticity score to determine the authenticity result : ; Among them, is a predefined threshold used to judge whether the picture is authentic. If the authenticity score is higher than the threshold , it indicates that the picture has a high authenticity degree.
[0028] In this embodiment, based on the pictures uploaded by the seller, preliminary product information is automatically generated, and it is determined whether the seller's product is an official genuine product. If the product is genuine, the final product information is generated according to the official product information. If the product is a counterfeit, it is determined whether there is user feedback for the product. If there is user feedback, the user feedback is referred to generate the final product information. If there is no user feedback, the final product information is assisted to be generated through the integration of knowledge graph and cross-modal retrieval and generation. By combining multiple means such as image analysis, product feature comparison, and user feedback, the authenticity of the product is comprehensively evaluated, which is more accurate and reliable than a single method. The whole process from image feature extraction to product information generation and authenticity judgment is automatically processed, reducing manual intervention and improving efficiency. By optimizing the model parameters through reinforcement learning, the algorithm can continuously improve according to user feedback and adapt to different products and scenarios. The multi-modal information such as images and texts is effectively utilized to more fully understand the product characteristics, improving the authenticity judgment and information generation quality. When the image quality is poor or the information is incomplete, the multi-modal and multi-method integration enhances the robustness and fault tolerance of the algorithm. Using user feedback to form a closed-loop optimization improves the user experience and the quality of platform products. In scenarios with complex product information such as e-commerce and second-hand trading platforms, disputes caused by counterfeit and shoddy products are reduced, and the market order is maintained.
Claims
1. An intelligent product information supplement method based on large language models, characterized in that: Including: Obtain a set of pictures of the target product ; Obtain the image set The authenticity result of each image in and ; Extract the elements corresponding to all real pictures in the picture set in: ; Among them, represents traversing each image in the image set ; Set up a product judgment function to judge the authenticity of the target product: ; wherein, is a set threshold value; If , it is determined that the target product is preliminarily analyzed as genuine; If , it is determined that the target product is preliminarily analyzed as a counterfeit, and all user feedback related to the target product is read from the database or feedback storage system to form a feedback set : ; Among them, is a function that retrieves all user feedback related to a specific product from a database, file system, or other storage medium; Set up a feedback judgment function to judge the user feedback situation of the target product: ; If , it is determined that there is user feedback; If , it is determined that there is no user feedback.
2. The intelligent supplementary method for product information based on large language model according to claim 1, wherein: Generate initial information for the target product, specifically: Obtain all the pictures of the target product and generate a picture set, denoted as ; Extract the key features of the target product from the image set through the feature extraction function, denoted as : ; Among them, is the input image set, is the feature extraction function, is the extracted feature vector; Obtain a set of predefined prompt templates, denoted as ; For the extracted features, use a selection strategy function to select a suitable prompt template from the set of prompt templates: ; Among them, is a set of predefined hint templates, and each template is a string containing placeholders for inserting feature values, is a selection strategy function for selecting the most suitable template according to the content of the feature and the applicability of the hint template, is selected from the set of hint templates and is a specific template selected therefrom; Set up a prompt generation function , combine the extracted features with the selected prompt template to generate specific prompt information: ; Among them, is the generated prompt message; Obtain a multi-modal pre-trained model ; Through a multi-modal pre-trained model the image feature extraction component in extracts image features from the picture set which are denoted as : ; Among them, is an image feature extraction component in the multi-modal pre-training model , and is the image feature extracted from the picture set ; Through the multi-modal pre-training model The text feature extraction component in , extract text features from the prompt information , denoted as : ; Among them, is the text feature extraction component in the multi-modal pre-training model, is the text feature extracted from the prompt information ; Input the image features and text features into the multimodal feature fusion component to obtain the fused multimodal features, denoted as : ; Among them, is a feature fusion component in the multi-modal pre-training model, is the fused multi-modal feature; The fused multi-modal features are input into the generation component of the model to generate initial product information, denoted as : ; Among them, is the generation component in the multi-modal pre-training model.
3. The intelligent supplementary method for product information based on large language model according to claim 1, characterized in that: Conduct a preliminary authenticity analysis of the target product, including image feature analysis, specifically: Obtain a set of pictures of the target product ; Using the gray-level co-occurrence matrix , extract the texture features of each image : ; Among them, is the image analyzed in the set of images; Extract the pattern features of each image through template matching : ; Extract the edge features of each image through Canny edge detection : ; Calculate the Laplacian response of each image and evaluate the blurriness of the image : ; Detect each image through local consistency check Whether there is distortion : ; Fuse texture, pattern, edge, blur, and distortion features into image detail features : ; For each pair of images and , calculate the similarity of their image detail features and : : ; Calculate the average similarity of all image pairs as the multi-image consistency score : ; Extract the metadata features of each image and : ; Calculate the consistency score of the metadata features of all images : ; Among them, is the reference metadata, is the number of matching fields, is the total number of fields; Detect whether an image has been cropped or spliced using an image forensics algorithm : ; Calculate the picture integrity and originality scores by combining the metadata consistency and the cropping and splicing detection results : ; Among them, and are weight parameters used to balance the contributions of different parts to the final result, satisfying ; Calculate the average score of the image detail features of all images : ; Among them, represents the th picture's th detailed feature; Calculate the final authenticity score by combining the image detail score, multi-image consistency score, and picture integrity and originality score : ; Among them, , and are weight parameters used to balance the contributions of different parts to the final result, satisfying ; Based on the authenticity score Perform a threshold judgment to determine the authenticity result : ; Among them, is a predefined threshold value.
4. The intelligent supplementary method for product information based on large language model according to claim 3, wherein: The preliminary authenticity analysis of the target product also includes product characteristic analysis, specifically: Obtain a set of pictures of the target product ; Extract product features from each image : ; Among them, refers to extracting product-related features from an image; Calculate the product features of each picture with the official product features similarity : ; Among them, is the product feature extracted from the th picture, is the official product feature, which is the standard product information obtained from the brand official website, official authorized channels or other reliable sources, is the similarity between the product feature of the th picture and the official product feature, and represent the norms of the product feature vector and the official product feature vector ; Calculate the average of the product feature similarities of all images as the product detail comparison score : ; Among them, represents the total number of images in the set of images uploaded by the seller; Detect each image to check whether there are brand logos, trademarks, and anti-counterfeiting marks, and obtain the detection results : ; Among them, is the detection result of the brand logo in the th picture, indicating whether the brand logo, trademark, and anti-counterfeiting label are detected. It refers to the specific pattern for detecting the brand logo, trademark, and anti-counterfeiting label in the image; Check the clarity, accuracy, and integrity of the detected brand logo to obtain the inspection results : ; Among them, is the inspection result of the clarity, accuracy, and integrity of the brand logo in the th picture, which refers to the quality assessment of the detected brand logo; Calculate the brand logo and identification inspection scores by combining the brand logo detection results and the inspection results : ; Compare the score with product details, brand logos and identification marks, and calculate the final authenticity score : ; Among them, and are weight parameters used to balance the contributions of different parts to the final result, satisfying ; Based on the authenticity score Perform threshold judgment to determine the authenticity result : ; Among them, is a predefined threshold value.
5. The intelligent product information supplement method based on a large language model according to claim 1, characterized in that: If the target product is genuine, correspondingly generate the final product information of the target product, specifically: Obtain a collection of images ; Extract all official information of the target product from the database, integrate and generate an official information set, denoted as ; Obtain a multi-modal pre-trained model ; Through the multi-modal pre-training model The image feature extraction component in Extract image features from the picture collection And denote it as : ; Through the multi-modal pre-training model in the text feature extraction component , extract text features from the official information , and denote it as : ; Input the image features and text features into the multi-modal feature fusion component to obtain the fused multi-modal features, denoted as : ; The fused multi-modal features are input into the generation component of the updated model to generate new product information, denoted as : ; The newly generated product information of the output .
6. The intelligent supplementary method for product information based on a large language model according to claim 1 or 5, characterized in that: If the target product is preliminarily analyzed as a counterfeit and there is user feedback, generate the final product information of the target product with reference to the user feedback, including model optimization, specifically: Obtain a feedback set ; For the feedback set For each piece of user feedback in it, calculate its cosine similarity with the current product information as follows: ; Among them, is the norm of the th user feedback vector, and is the norm of the product information vector generated by the model; According to the similarity score , design a reward function to convert the similarity score into a reward signal: ; Among them, is the similarity threshold; Combine all the reward signals into a reward signal vector : ; Among them, is the quantity of user feedback; Through the loss function , measure the product information generated by the model and user feedback The difference between: ; ; Among them, the loss function is specifically the mean squared error, is the th component in the product information generated by the model, is the th component in the th user feedback; Calculate the loss function For the model parameters Gradient of: ; Using the learning rate and the calculated gradient, update the model parameters: 。 7. The intelligent product information supplement method based on a large language model according to claim 6, characterized in that: The generation of the final product information of the target product with reference to the user feedback also includes generating the final product information, specifically: Obtain the model parameters after model optimization ; Obtain a collection of images and prompt information ; Obtain a multi-modal pre-trained model ; Extract the current model parameters in the multi-modal pre-training model and denote them as ; ; Update the current model parameters to the optimized model parameters ; Through the multi-modal pre-training model The image feature extraction component in , extract image features from the picture set , denoted as : ; Through the multi-modal pre-training model The text feature extraction component in , extract text features from the prompt information , denoted as : ; Input the image features and text features into the multi-modal feature fusion component to obtain the fused multi-modal features, denoted as : ; The fused multi-modal features are input into the generation component of the updated model to generate new product information, denoted as : ; Extract key product feature vectors from new product information : ; Among them, represents the value of the key feature; Assign weights to each key feature: ; Among them, represents the weight of the th feature, and ; Calculate the weighted total score of the authenticity of the product : ; Convert the weighted total score into the final true / false judgment result : ; wherein, is a preset threshold value; If , output the generated new product information , and label this product as genuine; If , output the generated new product information , and label the product as a counterfeit.
8. The intelligent product information supplementation method based on a large language model according to claim 7, characterized in that: If the target product is preliminarily analyzed as a counterfeit and there is no user feedback, assist in generating the final product information of the target product, specifically: Generate a knowledge graph for the target product: ; Among them, is a set of entities, is a set of relationships; In the knowledge graph retrieve the entity set related to the product category : ; Among them, represents the category label of the entity Extract feature vectors from each entity : ; Obtain the feature vector weights, marked as : ; Multiply the feature vector of each entity by its corresponding weight, and then add these weighted feature vectors together to obtain the class feature vector : ; Normalize the weighted sum of the feature vectors: ; Among them, represents the L2 norm of the feature vector.
9. The intelligent product information supplementation method based on a large language model according to claim 8, wherein: The assistance in generating the final product information of the target product also includes: In the product image database perform cross-modal retrieval on the input image P to obtain a set of similar products : ; Among them, denotes a function for determining similar products by calculating the similarity of image features with those of images in the database; Extract feature vectors from each similar product : ; Aggregate the feature vectors of similar products to obtain the feature vectors of similar products : ; Among them, is the weight of similar products and the specific formula is: ; Among them, represents each product image in the product image database, for each product image in the database and the input image is the cosine similarity, specifically: ; Among them, is the feature vector of the input image , is the feature vector of the product image in the database , represents the L2 norm of the vector. Fuse the knowledge graph features and the features of similar products to obtain a fused feature vector : ; Among them, and are weight parameters; Input the fused feature vector into the generation component of the updated model to generate new product information, denoted as : ; Extract key product feature vectors from new product information : ; Among them, represents the value of the key feature No. Assign weights to each key feature: ; Among them, represents the weight of the th feature, and ; Calculate the weighted total score of the authenticity of the product : ; Convert the weighted total score into the final authenticity judgment result : ; wherein, is a preset threshold value; If , output the newly generated product information , and label the product as genuine; If , output the generated new product information , and label this product as a counterfeit.
Citation Information
Patent Citations
Method and equipment for completing input information complementation by utilizing picture attribute extraction
CN107729900A
Illegal commodity identification method and device, computer equipment and storage medium
CN116051132A
Automatic monitoring system and method for infringement and counterfeit commodities
CN118154988A
Method and system for generating commodity selling point information from e-commerce website link
CN118446776A
Multimodal large language model counterfeit information detection method introducing expert knowledge
CN118606892A