Commodity selling point extraction and generation method and system based on RAG and multi-modal model
Through the product selling point extraction and generation method and system based on RAG and multimodal models, the problems of low efficiency and different quality of product selling point extraction in the existing technology are solved, automated processing and intelligent integration are realized, and the quality and efficiency of product explanation copy are improved.
Patent Information
- Application Number
- CN202411871072.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The product selling points of existing product description files are usually extracted manually, with low efficiency and varying quality, making it difficult to achieve intelligent integration and quality control of selling point information and product information.
Using a product selling point extraction and generation method and system based on RAG and multimodal models, users upload product description information, perform document analysis, text vector processing and information fusion to generate product explanation copy.
It realizes automated processing of product description files in multiple formats, improves information processing efficiency and accuracy, and intelligently integrates selling point information, improves the quality and efficiency of product explanation copy and reduces manual writing costs.
Smart Images

Figure CN119940304A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and computer science technology, and more specifically, to a method and system for extracting and generating commodity selling points based on RAG and multimodal models. Background Art
[0002] Multimodal big models are a cutting-edge direction in the field of artificial intelligence. They achieve in-depth understanding of complex scenarios by integrating and analyzing information from different data sources, such as text, images, audio, and video. The core advantage of this model is that it can cross the boundaries of traditional single-modal data processing to achieve comprehensive capture and comprehensive analysis of information; in terms of application, multimodal big models have shown broad potential. For example, in social media analysis, it can understand the pictures, videos, and texts posted by users, and provide sentiment analysis and trend prediction. In addition, multimodal big models can also play a role in many fields such as education, customer service, and security monitoring.
[0003] RAG, or Retrieval-Augmented Generation, is a machine learning framework that combines retrieval and generation. It retrieves information from an external knowledge base and combines it with the model's generation capabilities to improve the ability to handle complex tasks. The core of the RAG model is its ability to leverage massive amounts of external data to enhance its own understanding and generation capabilities, which enables it to excel in handling tasks that require a broad knowledge background; the RAG model is widely used, and it has received particular attention in the field of natural language processing. For example, in a question-and-answer system, RAG can retrieve relevant documents or paragraphs, and then generate accurate and comprehensive answers based on this information. In text summarization tasks, it is able to understand large amounts of text content and generate refined summaries. In addition, RAG is also suitable for dialogue systems, capable of generating coherent and relevant responses based on context.
[0004] The selling points of existing product description files are generally extracted manually, and the product explanation copy is manually written. This places high demands on the number and efficiency of manual labor, and the quality of manually written product explanation copy varies, making it inconvenient to achieve intelligent integration of selling point information and product information, and it is also inconvenient to accurately control the quality of the product explanation copy.
[0005] In view of this, the present invention proposes a method and system for extracting and generating product selling points based on RAG and multimodal models to solve the above problems. Summary of the invention
[0006] In order to overcome the above-mentioned defects of the prior art and to achieve the above-mentioned purpose, the present invention provides the following technical solution: a method and system for extracting and generating product selling points based on RAG and multimodal model, comprising the following steps:
[0007] S1. The user uploads product description information, which includes a product description file and product pictures;
[0008] S2, parsing the product description file and determining whether it is a PDF. If so, performing OCR typesetting optimization on the parsed product description file. If not, segmenting the parsed product description file to form a Chinese corpus of product information.
[0009] S3, performing text vectorization processing on the Chinese corpus of product information to generate vector data and store it in a vector database, using the vector database and a search engine to perform search, and re-arranging the retrieved preliminary search results to obtain product information;
[0010] S4. Extract selling point information from product images through a multimodal big model, and use a text big model to fuse product information with selling point information to obtain product explanation copy.
[0011] Furthermore, the product description file includes document formats such as Excel, WORD, and PDF, and the product image includes image formats such as PNG and JPEG.
[0012] Furthermore, the method of parsing the product description file and determining whether it is a PDF, if so, performing OCR typesetting optimization on the parsed product description file, and if not, performing segmentation processing on the parsed product description file to form a Chinese corpus of product information includes:
[0013] After the product description file is uploaded, the document format of the product description file can be automatically identified;
[0014] If the product description file is in PDF format, use Tesseract-OCR technology to perform optical character recognition on it, that is, convert the product description file in PDF format into editable text, and then optimize the layout of the text and table content read in the editable text;
[0015] If the product description file is not in PDF format, use the corresponding parsing method according to the different document formats;
[0016] Use the NLP library NLTK to process the recognized text, including word segmentation, part-of-speech tagging, and named entity recognition;
[0017] The identified text and table contents are segmented to form Chinese corpus of product information.
[0018] Furthermore, the typesetting optimization method is: using regular expressions, text cleaning algorithms and template matching algorithms to identify and adjust the format of the text until it meets the preset standardization requirements of the product information.
[0019] Furthermore, the segmentation processing method is to segment the text and table content by chunk segmentation, thereby obtaining chunk corpus, and the chunk corpus is the Chinese corpus of product information.
[0020] Furthermore, the method of performing text vectorization processing on the Chinese corpus of the product information to generate vector data and store it in a vector database, using the vector database and a search engine to search, and rearranging the retrieved preliminary search results to obtain the product information includes:
[0021] Use a pre-trained language model (BGE-M3) based on the BERT architecture for text vectorization;
[0022] Read the Chinese corpus of product information, input the Chinese corpus of product information into the pre-trained language model, and obtain the corresponding vector data;
[0023] The generated vector data is stored in a vector database;
[0024] A vector data index is constructed in the vector database. Users input product information fragments and use the search engine to search for the most similar vector data in the vector database.
[0025] According to the retrieved vector data, a corresponding preliminary search result is obtained from the vector database;
[0026] The reranker model is used to sort the preliminary search results, and the preliminary search result ranked first is regarded as the most relevant product information for the user.
[0027] Furthermore, the method of extracting selling point information from the product image by using the multimodal macro model and fusing the product information with the selling point information by using the text macro model to obtain the product explanation copy includes:
[0028] Input the product image into the multimodal big model to obtain selling point information, including color, material and design details;
[0029] Then, the product information is associated with the selling point information;
[0030] The extracted selling point information is then integrated with the product information using the large text model to obtain the product explanation copy.
[0031] Furthermore, the method of integrating the extracted selling point information with the product information by using the text macro model includes:
[0032] Collect and pre-process product information, including product descriptions, user reviews, and expert comments;
[0033] Use natural language processing technology to extract prompt words of selling point information and establish the correlation between the prompt words of selling point information and product information;
[0034] Construct a prompt word guidance model, embed, replace and supplement the prompt words of the selling point information into the product information through the text big model, and then obtain the explanation copy;
[0035] Evaluate the explanation copy, optimize and adjust the explanation copy based on the evaluation results, and use the optimized and adjusted explanation copy as the product explanation copy.
[0036] The product selling point extraction and generation system based on RAG and multimodal model includes:
[0037] Information upload module, where users upload product description information, including product description files and product images;
[0038] The document parsing module parses the product description file and determines whether it is a PDF. If so, the parsed product description file is optimized for OCR typesetting. If not, the parsed product description file is segmented to form the Chinese corpus of product information.
[0039] The information retrieval and rearrangement module performs text vectorization processing on the Chinese corpus of product information to generate vector data and store it in the vector database. It uses the vector database and the search engine to search and rearrange the retrieved preliminary search results to obtain product information.
[0040] The information fusion module extracts the selling point information from the product image through the multimodal big model, and uses the text big model to fuse the product information with the selling point information to obtain the product explanation copy.
[0041] The technical effects and advantages of the method and system for extracting and generating product selling points based on RAG and multimodal models of the present invention are as follows:
[0042] 1. Through the automatic document format recognition and processing mechanism, it can quickly process product description files in various formats; through the use of Tesseract-OCR technology and NLP processing methods, it ensures the accurate extraction and standardized processing of text information; through the chunk segmentation technology, it can achieve accurate segmentation of text, improve the efficiency of subsequent processing, and improve the information processing efficiency and accuracy of product description files; through the use of the BERT architecture pre-trained language model for text vectorization, it can improve the accuracy of text representation, and through the dual retrieval mechanism of combining vector database and retrieval engine, it can improve the efficiency of information retrieval; through the introduction of the reranker rearrangement model, it can optimize the relevance ranking of preliminary search results and enhance the accuracy of information retrieval;
[0043] 2. The multimodal large model can accurately extract visual selling point information from product images, realize the intelligent integration of selling point information and product information, and improve and ensure the quality of product explanation copy by establishing a complete copy evaluation and optimization mechanism; by adjusting the copy style and expression method according to the characteristics of the target audience and the quantitative management of evaluation indicators, the quality of product explanation copy can be accurately controlled, and the continuous optimization and iterative improvement of product explanation copy can be supported; and, through the automated document processing process, manual intervention is reduced, the cost of manual writing is reduced, and the workload of manual review is reduced, thereby improving the efficiency of generating product explanation copy. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a flowchart of the method for extracting and generating commodity selling points based on RAG and multimodal model of the present invention;
[0045] Figure 2 It is a schematic diagram of the framework of the system for extracting and generating commodity selling points based on RAG and multimodal model of the present invention;
[0046] Figure 3 It is a flow chart of the method for extracting and generating commodity selling points based on RAG and multimodal model according to the present invention. DETAILED DESCRIPTION
[0047] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0048] Example 1
[0049] See also Figure 1 and Figure 3As shown, the method for extracting and generating commodity selling points based on RAG and multimodal model described in this embodiment includes the following steps:
[0050] S1. The user uploads product description information, which includes a product description file and product pictures;
[0051] Specifically, the product description file includes document formats such as Excel, WORD, and PDF, and the product image includes image formats such as PNG and JPEG.
[0052] S2, parsing the product description file and determining whether it is a PDF. If so, performing OCR typesetting optimization on the parsed product description file. If not, segmenting the parsed product description file to form a Chinese corpus of product information.
[0053] Specifically, the product description file is parsed and it is determined whether it is a PDF. If so, the parsed product description file is subjected to OCR typesetting optimization. If not, the parsed product description file is segmented to form a Chinese corpus of product information, including:
[0054] After the product description file is uploaded, the document format of the product description file can be automatically identified;
[0055] If the product description file is in PDF format, use Tesseract-OCR technology to perform optical character recognition on it, that is, convert the product description file in PDF format into editable text, and then optimize the layout of the text and table content read in the editable text;
[0056] If the product description file is not in PDF format, use the corresponding parsing method according to the different document formats;
[0057] Use the NLP library NLTK to process the recognized text, including word segmentation, part-of-speech tagging, and named entity recognition;
[0058] The identified text and table contents are segmented to form Chinese corpus of product information.
[0059] The method of typesetting optimization is: using regular expressions, text cleaning algorithms, and template matching algorithms to identify and adjust the format of the text until it meets the preset standardization requirements of product information;
[0060] The segmentation processing method is to segment the text and table content by chunk segmentation, thereby obtaining chunk corpus, which is the Chinese corpus of product information;
[0061] It should be noted that after the user uploads the product description file in PDF format, the system automatically recognizes its format as PDF. Then, using Tesseract-OCR technology, the system converts the PDF document into an editable text format. During the conversion process, the system recognizes the text and table content in the document. Then, the system optimizes the layout of the recognized text and table. For example, regular expressions are used to adjust the paragraph format to ensure that there is a clear separation between the title, text and table. Text cleaning algorithms are applied to remove extra spaces and line breaks. Template matching algorithms are used to adjust the column width and row height of the table to meet the standardized requirements of product information. After typesetting optimization, the information structure is ensured to be clear.
[0062] The processing flow of non-PDF format (taking WORD as an example) is as follows: after the user uploads the product description file in WORD format, the system automatically recognizes that its format is non-PDF. The system uses a parsing method suitable for WORD documents to convert the document content into an editable text format; then, using the NLP library NLTK, the system performs word segmentation, part-of-speech tagging, and named entity recognition on the recognized text; then, the system uses chunk segmentation technology to divide the text and table content into multiple paragraphs or blocks (chunks corpus), each chunk corpus contains a relatively complete information unit, such as product features, specifications, instructions for use, etc.; these chunk corpora are the Chinese corpora of product information, which provide the basis for subsequent product information extraction, indexing, and display.
[0063] As a result, after OCR recognition and typesetting optimization, the product information in the PDF format product description file is presented in a more standardized and beautiful form, which is easy for users to read and understand. CR technology converts the text in the PDF / image into editable text to ensure the accurate extraction of information; for product description files in non-PDF format (such as WORD), through NLP processing and segmentation technology, the system can accurately extract the product information and organize it into a corpus that is easy to retrieve and display. This step ensures the logic and readability of the information to facilitate subsequent processing and analysis.
[0064] S3, performing text vectorization processing on the Chinese corpus of product information to generate vector data and store it in a vector database, using the vector database and a search engine to perform search, and re-arranging the retrieved preliminary search results to obtain product information;
[0065] Specifically, the method of performing text vectorization processing on the Chinese corpus of product information to generate vector data and store it in a vector database, using the vector database and a search engine to search, and rearranging the retrieved preliminary search results to obtain the product information includes:
[0066] Use a pre-trained language model (BGE-M3) based on the BERT architecture for text vectorization;
[0067] Read the Chinese corpus of product information, input the Chinese corpus of product information into the pre-trained language model, and obtain the corresponding vector data;
[0068] The generated vector data is stored in a vector database;
[0069] A vector data index is constructed in the vector database. Users input product information fragments and use the search engine to search for the most similar vector data in the vector database.
[0070] According to the retrieved vector data, a corresponding preliminary search result is obtained from the vector database;
[0071] The reranker model is used to sort the preliminary search results, and the preliminary search result ranked first is regarded as the most relevant product information for the user.
[0072] It should be noted that in S3, you can use the pre-trained language model BGE-M3 based on the BERT architecture to perform embedding text vectorization processing, generate vector data and store it in a vector database, such as Milvus, and use search engines such as Elasticsearch to provide fast query of product information;
[0073] Among them, when retrieving product information fragments, the system conducts preliminary searches through vector databases (such as Milvus) and search engines (such as Elasticsearch) to obtain the product information fragments most relevant to the user's query; the reranker rearrangement model (such as BERT or RoBERTa) is used to sort the preliminary search results. This model re-evaluates the priority of each search result by analyzing the context and relevance; compared with traditional retrieval methods, the introduction of the rearrangement model significantly improves the accuracy of information retrieval; through the context understanding ability of the deep learning model, the system can more accurately select the most relevant information for the product and reduce the interference of irrelevant information; the system will select the most relevant information for the product and the information that best reflects the characteristics of the product to generate explanation copy; by combining the user's query intention and product characteristics, the generated copy is ensured to be more targeted and attractive; after each search, the system will record the user's feedback and use this data to continuously optimize the parameters of the rearrangement model to improve future retrieval results.
[0074] S4. Extract selling point information from product images through a multimodal big model, and use a text big model to fuse product information with selling point information to obtain product explanation copy.
[0075] Specifically, the selling point information in the product image is extracted through the multimodal big model, and the product information and the selling point information are integrated using the text big model to obtain the product explanation copy, including:
[0076] Input the product image into the multimodal big model to obtain selling point information, including color, material and design details;
[0077] Then, the product information is associated with the selling point information;
[0078] The extracted selling point information is then integrated with the product information using the large text model to obtain the product explanation copy.
[0079] Specifically, the method of integrating the extracted selling point information with the product information by using the text macro model includes:
[0080] Collect and pre-process product information, including product descriptions, user reviews, and expert comments;
[0081] Use natural language processing technology (NLP technology) to extract prompt words of selling point information and establish the correlation between the prompt words of selling point information and product information;
[0082] Construct a prompt word guidance model, embed, replace and supplement the prompt words of the selling point information into the product information through the text big model, and then obtain the explanation copy;
[0083] Evaluate the explanation copy, optimize and adjust the explanation copy according to the evaluation results, and use the optimized and adjusted explanation copy as the product explanation copy;
[0084] It should be noted that the evaluation indicators for the explanation copy include:
[0085] Copywriting quality dimensions: grammatical accuracy (0-1), fluency of expression (0-1), structural integrity (0-1);
[0086] Marketing effect dimensions: selling point prominence (0-1), persuasiveness (0-1), attractiveness index (0-1);
[0087] Brand consistency dimensions: tonality matching (0-1), value alignment (0-1);
[0088] Optimize and adjust the explanation copy based on the evaluation results:
[0089] When the score is lower than the preset threshold (0.7), the optimization mechanism of the explanation copy is triggered. The optimization mechanism includes:
[0090] Correct grammatical errors and improve sentence structure; strengthen selling point information and adjust information priority; unify brand terms and adjust expressions;
[0091] Optimization strategies include: placing the selling point information of the explanation copy in a prominent position, which can be set by relevant personnel according to actual conditions;
[0092] Sorting out the logical relationship of the explanation copy: ensuring that the copy hierarchy is arranged in the form of major title-medium title / subtitle-content;
[0093] Adjust the language style according to the target audience (for example, if the target audience is young women, choose a language style that is light-hearted, humorous, fashionable, etc.; if the target audience is middle-aged and elderly men, choose a language style that is concise and easy to understand, such as vernacular and colloquial);
[0094] Moreover, an iterative approach is adopted in the optimization process: re-evaluation is carried out after each round of optimization, and the optimization history is recorded for continuous improvement until the preset quality standards are achieved.
[0095] In this embodiment, through the automatic document format recognition and processing mechanism, product description files in various formats can be quickly processed; by adopting Tesseract-OCR technology and NLP processing methods, accurate extraction and standardized processing of text information are ensured; through chunk segmentation technology, accurate segmentation of text can be achieved, the efficiency of subsequent processing is improved, and the information processing efficiency and accuracy of product description files can be improved; by using the pre-trained language model of the BERT architecture for text vectorization, the accuracy of text representation can be improved, and by combining the dual retrieval mechanism of the vector database and the retrieval engine, the efficiency of information retrieval can be improved; by introducing the reranker rearrangement model, the optimization The relevance ranking of preliminary search results can also enhance the accuracy of information retrieval; the multimodal large model can accurately extract visual selling point information from product images, and realize the intelligent integration of selling point information and product information. By establishing a complete copy evaluation and optimization mechanism, the quality of product explanation copy can be improved and ensured; by adjusting the copy style and expression method according to the characteristics of the target audience and the quantitative management of evaluation indicators, the quality of product explanation copy can be accurately controlled, and the continuous optimization and iterative improvement of product explanation copy can be supported; and, through the automated document processing process, human intervention is reduced, the cost of manual writing is reduced, and the workload of manual review is reduced, thereby improving the efficiency of generating product explanation copy.
[0096] Example 2
[0097] See also Figure 2 As shown, the system for extracting and generating product selling points based on RAG and multimodal models described in this embodiment includes:
[0098] Information upload module, where users upload product description information, including product description files and product images;
[0099] The document parsing module parses the product description file and determines whether it is a PDF. If so, the parsed product description file is optimized for OCR typesetting. If not, the parsed product description file is segmented to form the Chinese corpus of product information.
[0100] The information retrieval and rearrangement module performs text vectorization processing on the Chinese corpus of product information to generate vector data and store it in the vector database. It uses the vector database and the search engine to search and rearrange the retrieved preliminary search results to obtain product information.
[0101] The information fusion module extracts the selling point information from the product image through the multimodal big model, and uses the text big model to fuse the product information with the selling point information to obtain the product explanation copy.
[0102] In this embodiment, the present invention can be put into practice in the live broadcast scene of e-commerce digital people, and can improve the generation efficiency of product explanation copy, reducing the original generation time of 5 hours in a live broadcast room to 2 hours; it can improve the generation quality and achieve the automatic generation effect of product explanation copy, and its quality is close to the level of manual writing.
[0103] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the present invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0104] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the units is only one, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0105] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
[0106] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for extracting and generating commodity selling points based on RAG and multimodal models, characterized in that: The following steps are involved: S1. The user uploads product description information, which includes a product description file and product pictures; S2, parsing the product description file and determining whether it is a PDF. If so, performing OCR typesetting optimization on the parsed product description file. If not, segmenting the parsed product description file to form a Chinese corpus of product information. S3, performing text vectorization processing on the Chinese corpus of product information to generate vector data and store it in a vector database, using the vector database and a search engine to perform search, and re-arranging the retrieved preliminary search results to obtain product information; S4. Extract selling point information from product images through a multimodal big model, and use a text big model to fuse product information with selling point information to obtain product explanation copy.
2. The method for extracting and generating commodity selling points based on RAG and multimodal model according to claim 1 is characterized in that: The product description file includes document formats such as Excel, WORD, and PDF, and the product image includes image formats such as PNG and JPEG.
3. The method for extracting and generating commodity selling points based on RAG and multimodal model according to claim 1 is characterized in that: The method of parsing the product description file and determining whether it is a PDF, and if so, performing OCR typesetting optimization on the parsed product description file, and if not, segmenting the parsed product description file to form a Chinese corpus of product information includes: After the product description file is uploaded, the document format of the product description file can be automatically identified; If the product description file is in PDF format, use Tesseract-OCR technology to perform optical character recognition on it, that is, convert the product description file in PDF format into editable text, and then optimize the layout of the text and table content read in the editable text; If the product description file is not in PDF format, use the corresponding parsing method according to the different document formats; Use the NLP library NLTK to process the recognized text, including word segmentation, part-of-speech tagging, and named entity recognition; The identified text and table contents are segmented to form Chinese corpus of product information.
4. The method for extracting and generating commodity selling points based on RAG and multimodal model according to claim 3 is characterized in that: The typesetting optimization method is: using regular expressions, text cleaning algorithms and template matching algorithms to identify and adjust the format of the text until it meets the preset standardization requirements of the product information.
5. The method for extracting and generating commodity selling points based on RAG and multimodal model according to claim 3 is characterized in that: The segmentation processing method is to segment the text and table content by chunk segmentation, thereby obtaining chunk corpus, which is the Chinese corpus of product information.
6. The method for extracting and generating commodity selling points based on RAG and multimodal model according to claim 1 is characterized in that: The method of performing text vectorization processing on the Chinese corpus of product information to generate vector data and storing it in a vector database, performing retrieval using the vector database and a retrieval engine, and reordering the retrieved preliminary retrieval results to obtain the product information includes: Use pre-trained language models for text vectorization; Read the Chinese corpus of product information, input the Chinese corpus of product information into the pre-trained language model, and obtain the corresponding vector data; The generated vector data is stored in a vector database; A vector data index is constructed in the vector database. Users input product information fragments and use the search engine to search for the most similar vector data in the vector database. According to the retrieved vector data, a corresponding preliminary search result is obtained from the vector database; The reranker model is used to sort the preliminary search results, and the preliminary search result ranked first is regarded as the most relevant product information for the user.
7. The method for extracting and generating commodity selling points based on RAG and multimodal model according to claim 6 is characterized in that: The method of extracting the selling point information from the product image by using the multimodal big model, and fusing the product information with the selling point information by using the text big model, and then obtaining the product explanation copy includes: Input the product image into the multimodal big model to obtain selling point information, including color, material and design details; Then, the product information is associated with the selling point information; The extracted selling point information is then integrated with the product information using the large text model to obtain the product explanation copy.
8. The method for extracting and generating commodity selling points based on RAG and multimodal model according to claim 7 is characterized in that: The method of fusing the extracted selling point information with the product information by using the text macro model includes: Collect and pre-process product information, including product descriptions, user reviews, and expert comments; Use natural language processing technology to extract prompt words of selling point information and establish the correlation between the prompt words of selling point information and product information; Construct a prompt word guidance model, embed, replace and supplement the prompt words of the selling point information into the product information through the text big model, and then obtain the explanation copy; Evaluate the explanation copy, optimize and adjust the explanation copy based on the evaluation results, and use the optimized and adjusted explanation copy as the product explanation copy.
9. A system for extracting and generating selling points of products based on RAG and multimodal models, used to execute the method for extracting and generating selling points of products based on RAG and multimodal models as claimed in any one of claims 1 to 8, characterized in that: include: Information upload module, where users upload product description information, including product description files and product images; The document parsing module parses the product description file and determines whether it is a PDF. If so, the parsed product description file is optimized for OCR typesetting. If not, the parsed product description file is segmented to form the Chinese corpus of product information. The information retrieval and rearrangement module performs text vectorization processing on the Chinese corpus of product information to generate vector data and store it in the vector database. The vector database and the retrieval engine are used for retrieval, and the preliminary retrieval results are rearranged to obtain product information. The information fusion module extracts the selling point information from the product image through the multimodal big model, and uses the text big model to fuse the product information with the selling point information to obtain the product explanation copy.
Citation Information
Patent Citations
Cboth generation method, device and equipment and storage medium
CN115983227A
Cboth generation method and device, electronic equipment and storage medium
CN118333017A
Method and system for generating commodity selling point information from e-commerce website link
CN118446776A
Marketing commodity short video generation method and device based on AI, medium and product
CN119130502A
Cited By
Massive network live broadcast batch data acquisition method and system
CN120769077A
Generation method and device of live broadcast commodity explanation text, and computer readable medium
CN121174019A
Live commodity explanation text generation method and device and computer readable medium
CN121174019B