Tobacco question-answering method and device based on multi-modal large model and medium
Through image recognition and multimodal big model technology, combined with user identity information and tobacco knowledge vector library, target tobacco display suggestions are generated, which solves the problem that the existing technology cannot adapt to the dynamic market and different user needs, and achieves efficient and personalized tobacco Q&A and display suggestions generation.
Patent Information
- Application Number
- CN202510192955.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-21
AI Technical Summary
In the prior art, when dealing with multimodal large models, the tobacco question and answer system cannot effectively adapt to the dynamic and changeable real-time tobacco market and the question and answer needs of users of different identities.
The cigarette and placement description information is obtained through image recognition technology, and input it into the multimodal model to generate supplementary description information and general description information. Combining user identity information and tobacco knowledge vector library, target tobacco display suggestions are generated.
It realizes accurate extraction, enrichment display and efficient query of tobacco product information, provides personalized product recommendations and display suggestions, and improves user experience and the operational efficiency of tobacco companies.
Smart Images

Figure CN120086391A_ABST
Abstract
Description
[0001] The present invention relates to the technical field of tobacco intelligent question - answering systems, and particularly relates to a tobacco question - answering method, device, and medium based on a multimodal large model. Background Art
[0002] In the era of the rapid development of artificial intelligence, the development models of all walks of life have gradually moved towards intelligence. In the field of intelligent question - answering, the most concerned is the current multimodal data question - answering. It has gradually developed from the traditional text - to - text dialogue into data dialogues in different forms such as pictures, documents, and web pages. Among the dialogues of multimodal large models, picture question - answering is the most popular. This is not only because pictures are more intuitive for humans, but also for large models.
[0003] In the tobacco field, especially from the perspective of tobacco retailers, it would be more desirable for business operators if a multimodal large model can accurately identify retail display pictures and give good suggestions on the display placement rules. However, currently, the picture question - answering ability of multimodal large models for tobacco scenarios is weak, and a large amount of detailed semantic information contained in pictures is lost. Especially when making display recommendations, the knowledge points in the answers are relatively fixed, unable to adapt to the dynamic and changeable real - time tobacco market, and the answers for different - identity users are the same, unable to adapt to the questions and answers of different - identity users. Summary of the Invention
[0004] The purpose of the present invention is to provide a tobacco question - answering method, device, and medium based on a multimodal large model to solve the technical problems that the results of tobacco question - answering in related scenarios cannot adapt to the dynamic and changeable real - time tobacco market and cannot adapt to the questions and answers of different - identity users.
[0005] The purpose of the present invention can be achieved through the following technical solutions: In a first aspect, a tobacco question - answering method based on a multimodal large model is provided, including: Performing image recognition on the acquired tobacco image to be recognized to obtain the cigarettes and placement description information in the tobacco image to be recognized, and inputting the cigarettes and placement description information into the multimodal large model to obtain supplementary description information of the multimodal large model for the cigarettes and placement description information, and generating general description information; Responding to the received question vector retrieval, generating a question database table structure, and based on the question vector retrieval and the question database table structure, performing vector retrieval from the vector library to obtain key database structure information, converting the key database structure information from a vector to text, and combining the question vector retrieval to generate a structured query language statement; According to the structured query language statement, use the cigarette and placement description information and the general description information as query information, and combine with the tobacco knowledge vector library to supplement the prompt words to generate an initial tobacco display suggestion; Obtain the user identity information corresponding to the problem vector retrieval and the display schemes pre-stored in the display vector library, and generate a target tobacco display suggestion according to the user identity information, the display schemes pre-stored in the display vector library, and the initial display suggestion.
[0006] In a possible implementation manner, the inputting the cigarette and placement description information into the multimodal large model to obtain the supplementary description information of the multimodal large model for the cigarette and placement description information, and generating the general description information includes: Input the cigarette and placement description information into the multimodal large model, and the multimodal large model retrieves multiple target tobacco images according to the cigarette and placement description information; Generate description information for the image vectors corresponding to the multiple target tobacco images to obtain the supplementary description information of the multimodal large model for the cigarette and placement description information; Add a first positional encoding to the first vector corresponding to the supplementary description information, and add a second positional encoding to the text vector corresponding to the cigarette and placement description information, where the first positional encoding and the second positional encoding are used to represent the input order information; According to the first positional encoding and the second positional encoding, splice the first vector corresponding to the supplementary description information and the text vector corresponding to the cigarette and placement description information to generate the general description information.
[0007] In a possible implementation manner, the method further includes: Determine the compliance and length of the text descriptions corresponding to the multiple target tobacco images and determine the symbols and length of the text description corresponding to the cigarette and placement description information; According to the preset length and preset symbol rules, perform chunking processing on the text descriptions corresponding to the multiple target tobacco images and the text description corresponding to the cigarette and placement description information, where the preset symbol rules include line break chunking and paragraph identifier chunking; Vectorize the text descriptions corresponding to the multiple target tobacco images after chunking processing to generate the image vectors, and vectorize the text description corresponding to the cigarette and placement description information after chunking processing to generate the text vector corresponding to the cigarette and placement description information.
[0008] In a possible implementation manner, retrieving according to the problem vector and the problem database table structure, performing vector retrieval from the vector library to obtain key database structure information, converting the key database structure information from a vector to text, and combining the problem vector retrieval to generate a Structured Query Language statement, including: Generating a problem vector according to the problem vector retrieval, retrieving, in the vector library, a table structure information vector most similar to the problem vector according to the problem vector and the problem database table structure, and selecting a target structure table according to the similarity sorting of the table structure information vector; Filling key data information in the target structure table to obtain key database structure information; Converting the extracted key database structure information from a vector form to a text form; Based on template matching and / or grammar rules, combining the key database structure information in the text form with the problem vector retrieval, and generating a Structured Query Language statement when at least one of selecting an optimal query path, adding necessary indexes, and optimizing query conditions is satisfied.
[0009] In a possible implementation manner, taking the cigarette and placement description information and the general description information as query information according to the Structured Query Language statement, supplementing prompt words in combination with a tobacco knowledge vector library, and generating an initial tobacco display suggestion, including: Performing preprocessing of word segmentation and stop word removal on the cigarette and placement description information and the general description information as query information to generate a query text; Generating an SQL query statement according to the query text and the SQL statement template corresponding to the Structured Query Language statement; Executing the SQL query statement to retrieve target data corresponding to the query text from the database, where the target data includes at least one of sales data and consumer preferences; Converting the query text into a vector representation and performing prompt word retrieval in the tobacco knowledge vector library to generate prompt words; Supplementing the original query information according to the prompt words, and combining the retrieved target data and the query information supplemented with prompt words to generate an initial tobacco display suggestion, where the initial tobacco display suggestion includes at least one of a brand, a type, a placement position, and a display method.
[0010] In a possible implementation manner, generating a target tobacco display suggestion according to the user identity information, the display schemes pre-stored in the display vector library, and the initial display suggestion, including: Determining a retail grade according to the problem vector retrieval and the user identity information; Determine the satisfied preset conditions according to the retail grade, the display plan pre-stored in the display vector library, and the initial display suggestion, where the preset conditions include at least one of the following: whether there are continuous packs and continuous strips, the orderliness of display colors, price orderliness, and the rationality of the cigarette display positions of key product specifications; Generate a target tobacco display suggestion according to the satisfied preset conditions.
[0011] In a possible implementation manner, the display plan pre-stored in the display vector library is generated through the following method: Construct a professional knowledge base in the form of a document according to the professional terms and professional knowledge of the tobacco industry, where the professional terms and professional knowledge include at least one of order instructions, sales instructions, store grade instructions, content ratio meaning explanations and calculations, and tobacco law and regulation instructions; Utilize the retrieval and learning capabilities of the large model and RAG to self-organize and digest the professional knowledge base, and after completing the self-organization and digestion, generate a vectorized display vector library; According to the pre-constructed display template, perform mold entry in the display vector library to generate the display plan pre-stored in the display vector library.
[0012] In a possible implementation manner, the cigarette and placement description information includes at least one of the following: The placement position of the cigarette, the product specification of the cigarette, the placement direction of the cigarette, the front cabinet information of the cigarette placement, the back cabinet information of the cigarette placement, and the back cabinet grid information of the cigarette placement.
[0013] In a second aspect, an electronic device is provided, including: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to execute the executable instructions stored in the memory to implement the method described in any item of the first aspect.
[0014] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any item of the first aspect are implemented.
[0015] Compared with the existing solutions, the beneficial effects achieved by the present invention: Through image recognition technology, information such as the brand, type, quantity, and placement status of cigarettes can be accurately recognized, reducing the error of manual recognition. The automated recognition process is faster than manual recognition, significantly shortening the recognition time and improving work efficiency. Inputting the recognized cigarette and placement description information into a multi-modal large model to generate supplementary description information and general description information. The multi-modal large model can combine various information such as the brand, taste, and origin of cigarettes to generate a more comprehensive and detailed product description, enhancing the user's understanding of the product. By combining user identity information and consumption habits, more personalized product recommendations and display suggestions can be provided to users, enhancing the user experience.
[0016] Furthermore, in response to the received question vector retrieval, generate the structure of the question database table, and perform vector retrieval from the vector library to obtain the key database structure information, which is then converted into text and a structured query language statement is generated. Vector retrieval technology can efficiently retrieve key information matching the question vector from a large amount of data, shortening the query time. Through the structured query language statement, the required information can be accurately extracted from the database, reducing query errors. According to the structured query language statement and the tobacco knowledge vector library, combined with user identity information and the pre-stored display plan, generate the target tobacco display suggestion. By combining multiple information sources and algorithm optimization, tobacco display suggestions that meet user needs and market trends can be generated, enhancing the attractiveness and purchase rate of product display. The generated display suggestions can provide strong support for the operation decision-making of tobacco enterprises, helping enterprises optimize product layout and improve sales performance.
[0017] In summary, through key technical means such as image recognition, multi-modal large model, vector retrieval, database query, and intelligent display suggestion generation, the accurate extraction, rich display, and efficient query of tobacco product information, as well as the generation of intelligent display suggestions, are achieved. These technical effects together enhance the operation efficiency of tobacco enterprises and the user experience, providing strong support for the digital transformation and intelligent upgrade of enterprises. It realizes the evaluation of each picture and the display suggestion response under different identities and templates by making a standard measurement of the current picture and the template based on its own recognition results and traditional detection and recognition results. Using a mature database management tool, the rapid and convenient management of personnel identity, permissions, and display standards can be realized, dynamically modifying the standard configuration and real-time updating the Q&A results. Thus, the effect of the multi-modal large model for picture Q&A and display suggestions in the tobacco field is improved, and an intelligent, real-time, and efficient tobacco Q&A system is realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The present invention will be further described below with reference to the accompanying drawings.
[0019] Figure 1It is a flowchart of a tobacco Q&A method based on a multi-modal large model shown according to the embodiments of the specification.
[0020] Figure 2 It is a schematic diagram of a tobacco Q&A based on a multi-modal large model shown according to the embodiments of the specification.
[0021] Figure 3 It is a block diagram of a tobacco Q&A device based on a multi-modal large model shown according to the embodiments of the specification. Detailed implementation manners
[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0023] As Figure 1 shown, it is a flowchart of a tobacco Q&A method based on a multi-modal large model of the present invention, including: In step S11, perform image recognition on the acquired tobacco image to be recognized, obtain the cigarettes and placement description information in the tobacco image to be recognized, and input the cigarettes and placement description information into the multi-modal large model to obtain supplementary description information of the multi-modal large model for the cigarettes and placement description information, and generate general description information; Among them, image recognition refers to the process of using computer vision technology to extract useful information or features from the input image, and then identifying the objects, scenes or other contents in the image.
[0024] In the embodiments of the present disclosure, it relies on deep learning algorithms, especially convolutional neural networks (CNNs). These networks extract features in the image layer by layer through multiple convolutional layers and pooling layers, and finally perform classification or regression through fully connected layers. In tobacco image recognition, CNNs can learn features such as the shape, color, and texture of cigarettes, so as to accurately identify the cigarettes and their placement states in the image. Detect the position of the cigarettes in the picture, identify the product specifications of the cigarettes, and then judge whether they are ordered and whether they are placed in the correct orientation (text direction) according to the recognition results of the cigarettes. Then detect information such as the front cabinet, back cabinet, and back cabinet compartments, and save all this information to the database according to formatted fields. Different tables save the detection and recognition results of different algorithms, and record the id of the picture. For example: Suppose there is an image containing multiple brands of cigarettes, and the image recognition system can identify the brand, quantity of each cigarette, and their relative positions in the image.
[0025] Among them, a multimodal large model refers to a deep learning model that can process and integrate information from multiple different modalities (such as text, images, audio, etc.).
[0026] In the embodiments of the present disclosure, the multimodal large model is usually based on the Transformer architecture and captures the associations between different modality data through the self-attention mechanism. In the tobacco field, the multimodal large model can accept cigarette and placement description information (text modality) as input, and combine its internal knowledge base to generate richer supplementary description information, such as the brand story and taste evaluation of cigarettes.
[0027] In step S12, in response to the received question vector retrieval, a question database table structure is generated, and based on the question vector retrieval and the question database table structure, vector retrieval is performed from the vector library to obtain key database structure information. The key database structure information is converted from a vector to text, and combined with the question vector retrieval, a structured query language statement is generated; Among them, the structured query language statement SQL is a programming language specifically used to communicate with a database, and is used to store, retrieve, update, and manage data in the database.
[0028] In the embodiments of the present disclosure, the SQL statement allows users to perform various data operations by defining the structure, type, and relationships between data. During the vector retrieval and text conversion process, the generated SQL statement is used to retrieve key information matching the question vector from the database.
[0029] In step S13, based on the structured query language statement, the cigarette and placement description information and the general description information are used as query information, and combined with the tobacco knowledge vector library for prompt word supplementation to generate an initial tobacco display suggestion; Among them, the tobacco knowledge vector library refers to a database that stores a large amount of tobacco-related knowledge information and represents it in the form of vectors.
[0030] In the embodiments of the present disclosure, vectorization is the process of converting text or other types of data into numerical vectors for easy processing by machine learning models. Each vector in the tobacco knowledge vector library represents a specific tobacco knowledge or attribute, such as brand, taste, origin, etc.
[0031] In the embodiments of the present disclosure, accurate SQL statements are generated through the excellent text question and answer ability of the large language model; and in the form of a vector library, professional knowledge points in the tobacco industry can be effectively supplemented, realizing the knowledge sorting and analysis of excellent display templates, and being able to effectively give suggestions for the user's display pictures.
[0032] In step S14, obtain the user identity information corresponding to the problem vector retrieval and the display schemes pre-stored in the display vector library, and generate a target tobacco display suggestion according to the user identity information, the display schemes pre-stored in the display vector library, and the initial display suggestion.
[0033] Among them, the target tobacco display suggestion refers to the tobacco product placement suggestion generated according to user needs, product information, and predefined display rules.
[0034] In the embodiments of the present disclosure, by combining user identity information (such as consumption habits, preferences), preset schemes in the display vector library (such as the best display angle, color matching, etc.), and the initial display suggestion, the system optimizes through algorithms to generate a tobacco display scheme that meets specific goals.
[0035] For example: For users who prefer to buy high-end cigarettes, the system may suggest placing Yellow Crane Tower cigarettes in a prominent position in the store, and matching them with a delicate display stand and lighting effects to enhance the attractiveness and purchase rate of the products.
[0036] Starting from image recognition, enrich the description information through a multimodal large model, then generate SQL statements based on the problem vector retrieval to retrieve key information from the database, and finally generate a target tobacco display suggestion by combining the tobacco knowledge vector library and user information, realizing the complete conversion from images to intelligent display suggestions.
[0037] The above technical solution can accurately identify information such as the brand, type, quantity, and placement status of cigarettes through image recognition technology, reducing the error of manual recognition. The automated recognition process is faster than manual recognition, can significantly shorten the recognition time, and improve work efficiency. Input the recognized cigarette and placement description information into the multimodal large model to generate supplementary description information and general description information. The multimodal large model can combine various information such as the brand, taste, and origin of the cigarettes to generate a more comprehensive and detailed product description, enhancing the user's understanding of the product. By combining user identity information and consumption habits, more personalized product recommendations and display suggestions can be provided for users, enhancing the user experience.
[0038] Further, in response to the received question vector retrieval, a question database table structure is generated, and vector retrieval is performed from the vector library to obtain key database structure information, which is then converted into text and a Structured Query Language (SQL) statement is generated. Vector retrieval technology can efficiently retrieve key information matching the question vector from massive data, shortening the query time. Through the SQL statement, the required information can be accurately extracted from the database, reducing query errors. According to the SQL statement and the tobacco knowledge vector library, combined with the user identity information and the pre-stored display plan, a target tobacco display recommendation is generated. By combining multiple information sources and algorithm optimization, a tobacco display recommendation that meets the user's needs and market trends can be generated, enhancing the attractiveness and purchase rate of product display. The generated display recommendation can provide strong support for the operation decision-making of tobacco enterprises, helping the enterprises optimize the product layout and improve the sales performance.
[0039] In summary, through key technical means such as image recognition, multi-modal large model, vector retrieval, database query, and intelligent display recommendation generation, the accurate extraction, rich display, and efficient query of tobacco product information are achieved, as well as the generation of intelligent display recommendations. These technical effects together improve the operation efficiency and user experience of tobacco enterprises, providing strong support for the digital transformation and intelligent upgrade of the enterprises. It is realized to make a standard measurement of the current picture and the template according to its own recognition result and the traditional detection and recognition result, and evaluate each picture to generate a display recommendation reply under different identities and different templates. Using a mature database management tool, the rapid and convenient management of personnel identity, permissions, and display standards can be realized, dynamically modifying the standard configuration and real-time updating the Q&A results. Thereby improving the effect of the multi-modal large model for picture Q&A and display recommendations in the tobacco field, and realizing an intelligent, real-time, and efficient tobacco Q&A system.
[0040] In a possible implementation manner, in step S11, the inputting the cigarette and placement description information into the multi-modal large model to obtain supplementary description information of the multi-modal large model for the cigarette and placement description information and generating general description information includes: In step S111, the cigarette and placement description information is input into the multi-modal large model, and the multi-modal large model retrieves multiple target tobacco images according to the cigarette and placement description information; In the embodiment of the present disclosure, the cigarette and placement description information is input into the multi-modal large model. According to these description information, the multi-modal large model uses its internal knowledge and algorithms to retrieve multiple target tobacco images that are most matched with the description information from the pre-trained image library. These images are highly semantically related to the description information.
[0041] In step S112, description information is generated for the image vectors corresponding to the multiple target tobacco images, and supplementary description information of the multi-modal large model for the cigarette and placement description information is obtained; In the embodiments of the present disclosure, for multiple retrieved target tobacco images, the multi-modal large model extracts their image vectors. Then, using these image vectors and the internal generation mechanism of the model, supplementary description information corresponding to the image content is generated. These information may be detailed descriptions of features such as cigarette brands, appearances, and packaging in the images.
[0042] In step S113, a first positional encoding is added to the first vector corresponding to the supplementary description information, and a second positional encoding is added to the text vector corresponding to the cigarette and placement description information. The first positional encoding and the second positional encoding are used to represent the order information of the input; Among them, the positional encoding is used to introduce positional information in the sequence data and represent the order information of the input.
[0043] In the embodiments of the present disclosure, in order to maintain the sequentiality of the input information, a first positional encoding is added to the first vector corresponding to the supplementary description information, and a second positional encoding is added to the text vector corresponding to the original cigarette and placement description information. The introduction of the positional encoding enables the model to perceive the order relationship when processing this information, which is crucial for subsequent splicing and generating general description information.
[0044] In step S114, according to the first positional encoding and the second positional encoding, the first vector corresponding to the supplementary description information is spliced with the text vector corresponding to the cigarette and placement description information to generate the general description information.
[0045] In the embodiments of the present disclosure, according to the first positional encoding and the second positional encoding, the multi-modal large model splices the first vector of the supplementary description information with the text vector of the cigarette and placement description information. The spliced vector contains the original text description and the supplementary description information generated by the image, jointly constituting the general description information. This step realizes the conversion from multi-modal input to unified text output and provides a rich information basis for subsequent applications.
[0046] Through the multi-modal large model, the cigarette and placement description information and related image data can be fully utilized to generate general description information that includes both the original text description and the image supplementary description. This information processing method not only improves the accuracy and richness of the description but also provides more extensive possibilities for subsequent applications.
[0047] In a possible implementation manner, the method further includes: In step S115, determine the compliance and length of the text descriptions corresponding to multiple target tobacco images, and determine the symbols and length of the text descriptions corresponding to the cigarette and placement description information; In the embodiments of the present disclosure, determine the symbols in the text descriptions corresponding to multiple target tobacco images and the text descriptions corresponding to the cigarette and placement description information, such as line break characters, paragraph identifiers, etc. These symbols are used to mark the structure and format of the text descriptions. Then, calculate the number of characters or the length of these text descriptions for subsequent chunking and vectorization.
[0048] In step S116, according to the preset length and preset symbol rules, perform chunking on the text descriptions corresponding to multiple target tobacco images and the text descriptions corresponding to the cigarette and placement description information, where the preset symbol rules include line break character chunking and paragraph identifier chunking; Among them, chunking is the process of splitting a longer text description into multiple shorter text chunks. This processing method helps to reduce the processing difficulty, improve the processing efficiency, and also helps to maintain the coherence and readability of the text description.
[0049] In the embodiments of the present disclosure, according to the preset length (such as the number of characters included in each chunk) and preset symbol rules (such as line break character chunking, paragraph identifier chunking, etc.), perform chunking on the text descriptions corresponding to multiple target tobacco images and the text descriptions corresponding to the cigarette and placement description information. The specific operations of chunking may include identifying symbol positions, cutting text chunks, and maintaining the order and coherence of text chunks. Through chunking, the original text description is split into multiple shorter text chunks.
[0050] In step S117, vectorize the text descriptions corresponding to multiple target tobacco images after chunking to generate the image vectors, and vectorize the text descriptions corresponding to the cigarette and placement description information after chunking to generate the text vectors corresponding to the cigarette and placement description information.
[0051] In the embodiments of the present disclosure, vectorize the text descriptions corresponding to multiple target tobacco images after chunking and the text descriptions corresponding to the cigarette and placement description information. The process of vectorization may include text preprocessing (such as removing stop words, stemming, etc.), word embedding (such as Word2Vec, BERT, etc.) or sentence embedding (such as Sentence-BERT). These steps convert text chunks into vector representations to capture their semantic features.
[0052] Through vectorization, image vectors and text vectors that can be used in subsequent machine learning or deep learning models are obtained. These vectors represent the semantic content of the original text description and provide a basis for subsequent processing and analysis.
[0053] In this way, it is possible to process multi-modal input information (including images and text descriptions) and convert it into a unified vector representation form. This processing method not only improves the information processing efficiency but also provides a rich feature basis for subsequent applications (such as text generation, information retrieval, recommendation systems, etc.).
[0054] In a possible implementation manner, in step S12, according to the problem vector retrieval and the problem database table structure, vector retrieval is performed from the vector library to obtain key database structure information, the key database structure information is converted from a vector to text, and in combination with the problem vector retrieval, a structured query language statement is generated, including: In step S121, according to the problem vector retrieval, a problem vector is generated, and according to the problem vector and the problem database table structure, the table structure information vector most similar to the problem vector is retrieved in the vector library, and the target structure table is selected according to the similarity sorting of the table structure information vector; Among them, problem vector retrieval refers to converting the user's problem or query into a vector representation and retrieving in the vector library to find the vector most similar to it. In the content of this technology, problem vector retrieval is used to understand the user's intention and find the database table structure information related to the problem.
[0055] The table structure information vector refers to converting the structure information of the database table (such as table name, field name, field type, etc.) into a vector representation. This representation method can capture the similarities and differences between table structures, facilitating subsequent retrieval and matching.
[0056] Similarity sorting refers to the process of sorting the retrieval results according to the similarity between vectors. In this technology, similarity sorting is used to select the table structure information vector most similar to the target problem vector.
[0057] In the embodiments of the present disclosure, a problem vector is generated according to the problem or query input by the user. This usually involves converting the problem text into a vector representation, possibly using word embedding or sentence embedding techniques. Then, according to the problem vector and the problem database table structure information, retrieval is performed in the vector library to find the table structure information vector most similar to the problem vector. The retrieval results are sorted according to the similarity of the table structure information vector, and the table with the highest similarity is selected as the target structure table.
[0058] In step S122, key data information is filled in the target structure table to obtain key database structure information; Among them, the key data information refers to the data fields and information that are crucial for solving user problems or queries in the database table structure. In this technology, the key data information is filled into the target structure table to generate a complete database query condition.
[0059] In the embodiments of the present disclosure, the key data information is filled according to the user's problem or query and the field information of the target structure table. This may involve extracting keyword field values from the user input or inferring field values based on the context. The filled target structure table contains the key database structure information.
[0060] In step S123, the extracted key database structure information is converted from vector form to text form; In the embodiments of the present disclosure, the extracted key database structure information is converted from vector form to text form. This generally involves reversing the vector representation into a readable text representation, such as converting a field name vector into an actual field name.
[0061] In step S124, based on template matching and / or grammar rules, the key database structure information in text form is combined with the problem vector retrieval, and a structured query language statement is generated when at least one of the following conditions is met: selecting the optimal query path, adding necessary indexes, and optimizing query conditions.
[0062] Among them: Template matching refers to matching the input information with a preset template or rule to generate a result that conforms to a specific format or requirement. In this technology, template matching is used to combine the key database structure information in text form with the problem vector to generate a structured query language (SQL) statement.
[0063] Grammar rules refer to the rules and conventions that define the language structure. In this technology, grammar rules are used to ensure that the generated SQL statement conforms to the grammar requirements of the SQL language and can correctly execute the database query.
[0064] In the embodiments of the present disclosure, based on template matching and / or grammar rules, the key database structure information in text form is combined with the problem vector to generate an SQL statement. During the process of generating the SQL statement, multiple factors may be considered to optimize the query, such as selecting the optimal query path, adding necessary indexes, and optimizing query conditions. The finally generated SQL statement conforms to the grammar requirements of the SQL language, can correctly execute the database query, and returns the results required by the user.
[0065] Through this series of steps, the user's question or query can be converted into a vector representation, and the database table structure information related to the question can be found in the vector library. Then, the system generates structured SQL statements based on this information to execute database queries and return results. This technical method improves the efficiency and accuracy of database queries.
[0066] In a possible implementation, in step S13, according to the structured query language statement, using the cigarette and placement description information and the general description information as query information, and combining with the tobacco knowledge vector library for prompt word supplementation to generate an initial tobacco display suggestion, including: In step S131, preprocess the cigarette and placement description information and the general description information as query information by word segmentation and stop word removal to generate a query text; Among them, word segmentation refers to the process of cutting a continuous text string into individual words or phrases. In Chinese text processing, word segmentation is an important step in text preprocessing, which helps subsequent tasks such as word frequency statistics, keyword extraction, and text classification. Stop words refer to words that appear frequently in the text but contribute little to the text semantics, such as "de", "le", "zai", etc. Removing stop words can reduce the redundancy of text data and improve the efficiency and accuracy of subsequent text processing.
[0067] In the embodiments of the present disclosure, using the cigarette and placement description information and the general description information as query information, perform preprocessing of word segmentation and stop word removal. The word segmentation operation cuts the text into independent words or phrases, and the stop word removal operation reduces the redundancy of text data. The preprocessed text is used as the query text.
[0068] In step S132, generate an SQL query statement according to the query text and the SQL statement template corresponding to the structured query language statement; Among them, the SQL statement template refers to a preset SQL statement structure containing placeholders. By replacing the placeholders with actual values, an SQL query statement that meets the requirements can be quickly generated.
[0069] In the embodiments of the present disclosure, generate an SQL query statement according to the query text and the SQL statement template corresponding to the structured query language statement. The SQL statement template contains placeholders, and the system replaces the placeholders with the actual values in the query text to generate an SQL query statement that meets the requirements.
[0070] In step S133, execute the SQL query statement to retrieve target data corresponding to the query text from the database, where the target data includes at least one of sales data and consumer preferences; Among them, sales data refers to the data recording the sales situation of goods, including sales volume, sales amount, sales time, sales channels, etc. Consumer preference refers to the degree of preference of consumers for goods or services.
[0071] In step S134, the query text is converted into a vector representation, and a prompt word is retrieved and generated in the tobacco knowledge vector library. In the embodiment of the present disclosure, the query text is converted into a vector representation and retrieved in the tobacco knowledge vector library. Through vector retrieval, the system finds the tobacco knowledge vector most similar to the query text and extracts a prompt word from it. In step S135, the original query information is supplemented according to the prompt word, and an initial tobacco display suggestion is generated by combining the retrieved target data and the query information supplemented with the prompt word. The initial tobacco display suggestion includes at least one of brand, type, placement position, and display method.
[0072] In the embodiment of the present disclosure, the original query information is supplemented according to the prompt word, and an initial tobacco display suggestion is generated by combining the retrieved target data and the query information supplemented with the prompt word. The initial tobacco display suggestion includes at least one of brand, type, placement position, and display method. These suggestions are based on sales data, consumer preferences, and information in the tobacco knowledge vector library, aiming to improve the display effect and sales performance of tobacco products.
[0073] Through this series of steps, the system can generate an initial tobacco display suggestion according to the query information and structured query language statements input by the user, in combination with information such as the tobacco knowledge vector library and sales data, which helps to optimize the display method and position of tobacco products and improve consumers' purchase intention and sales performance.
[0074] In a possible implementation manner, in step S14, generating a target tobacco display suggestion according to the user identity information, the display scheme pre-stored in the display vector library, and the initial display suggestion includes: In step S141, the retail grade is determined according to the problem vector retrieval and the user identity information. Among them, the retail grade refers to the hierarchical classification for the hierarchical management of retail households according to factors such as the business scale, sales ability, and geographical location of retail households. Different retail grades enjoy different policy treatments and resource allocations, such as supply quantity, display support, etc.
[0075] In the embodiments of the present disclosure, the system determines the retail level of the user based on the problem vector retrieval (which may include information such as the problem description input by the user, historical purchase records, etc.) and the user identity information (such as the retailer number, business address, etc.). The retail level is an important reference factor for subsequent generation of display suggestions because retailers at different levels have differences in aspects such as supply quantity and display requirements.
[0076] In step S142, according to the retail level, the display schemes pre-stored in the display vector library, and the initial display suggestion, determine the satisfied preset conditions, where the preset conditions include at least one of the following: whether there are consecutive packs and strips, the orderliness of display colors, price orderliness, and the rationality of the display position of key product specifications of cigarettes; Among them, consecutive packs and strips mean that when tobacco products are displayed, multiple cigarette packs or strips of the same brand are closely arranged together to form a neat and uniform visual effect. This display method helps to highlight the brand image and attract consumers' attention.
[0077] The orderliness of display colors means that when tobacco products are displayed, they are arranged in an orderly manner according to visual factors such as color depth and warm and cold color tones to enhance the overall aesthetics and attractiveness of the display.
[0078] Price orderliness means that when tobacco products are displayed, they are arranged in an orderly manner from low to high or from high to low so that consumers can quickly find products that meet their budgets.
[0079] Key product specifications refer to tobacco product specifications with large sales volume, high popularity, and high profit contribution. These products should occupy prominent positions when displayed so that consumers can see and purchase them at a glance. The rationality of the display position directly affects the sales performance of key product specifications.
[0080] In the embodiments of the present disclosure, according to the retail level, the display schemes pre-stored in the display vector library, and the initial display suggestion, determine the satisfied preset conditions. The preset conditions include whether there are consecutive packs and strips, the orderliness of display colors, price orderliness, and the rationality of the display position of key product specifications of cigarettes, etc. These conditions are used to evaluate whether the display scheme meets the actual needs of the retailer and the display requirements of tobacco products.
[0081] Specifically, the system may compare the initial display suggestion with the schemes in the display vector library to check whether the initial suggestion meets the requirements of consecutive packs and strips (such as whether products of the same brand are closely arranged); whether they are arranged in an orderly manner according to color depth or warm and cold color tones; whether the prices are arranged in ascending or descending order; and whether the key product specifications occupy prominent positions, etc.
[0082] In step S143, generate the target tobacco display suggestion according to the satisfied preset conditions.
[0083] In the embodiments of the present disclosure, according to the satisfied preset conditions, the system generates target tobacco display suggestions. These suggestions may include adjusting the display method to meet the requirements of continuous packs and continuous strips, optimizing the color arrangement to enhance the visual effect, adjusting the price order to facilitate consumers' comparison, and rearranging the display positions of key product specifications to increase sales volume, etc.
[0084] The target tobacco display suggestions are designed to comprehensively consider the retail grade, the display plan library, and the initial suggestions, so as to provide tobacco product display plans that meet the actual needs and are attractive to retail households. These suggestions help improve the display effect and sales performance of tobacco products, while enhancing consumers' shopping experience and satisfaction.
[0085] Through this series of steps, the system can generate target tobacco display suggestions that meet the actual needs of retail households and the display requirements of tobacco products based on information such as user identity information, the display vector library, and the initial display suggestions. These suggestions provide guidance and support for retail households to optimize the display of tobacco products, helping to improve sales performance and consumer satisfaction.
[0086] In a possible implementation manner, the display plans pre-stored in the display vector library are generated through the following method: According to the professional terms and professional knowledge of the tobacco industry, a professional knowledge base in the form of a document is constructed, where the professional terms and professional knowledge include at least one of order instructions, sales instructions, store grade instructions, content ratio meaning instructions and calculations, and tobacco laws and regulations instructions; In the embodiments of the present disclosure, according to the professional terms and professional knowledge of the tobacco industry, a professional knowledge base in the form of a document is constructed. These documents cover multiple aspects such as order instructions, sales instructions, store grade instructions, content ratio meaning instructions and calculations, and tobacco laws and regulations instructions, ensuring the comprehensiveness and accuracy of the knowledge base.
[0087] Using the retrieval and learning capabilities of the large model and RAG, the professional knowledge base is self-organized and digested, and after completing the self-organization and digestion, a vectorized display vector library is generated; In the embodiments of the present disclosure, the retrieval and learning capabilities of the large model and RAG are used to self-organize and digest the professional knowledge base. The large model has powerful text understanding and reasoning capabilities, and can deeply understand the terms and instructions in the professional knowledge base. The RAG framework allows the model to retrieve external knowledge bases as needed during the organization process (although in this example the external knowledge base is the professional knowledge base itself, the introduction of the RAG framework emphasizes the combination of retrieval and generation capabilities), thereby enhancing the understanding and absorption of professional knowledge. After completing the self-organization and digestion, vectorization technology is used to convert the text information in the professional knowledge base into vector representations. These vector representations retain the semantic relationships and similarities between texts in the high-dimensional space.
[0088] Perform modulo entry in the display vector library according to the pre-constructed display template to generate the display scheme pre-stored in the display vector library.
[0089] In the embodiments of the present disclosure, preset display templates are constructed according to the display standards of the tobacco industry and consumer demands. These templates contain a series of rules and suggestions on how to place and display tobacco products, aiming to ensure that the display effect meets industry standards and consumer aesthetics.
[0090] Perform modulo entry in the display vector library. Retrieve and match relevant information from the vectorized display vector library according to the rules and suggestions in the display template to generate the pre-stored display scheme. These schemes combine the industry knowledge in the professional knowledge base and the practical experience in the display template, providing comprehensive and accurate guidance for the display of tobacco products.
[0091] Through the above steps, a display vector library containing rich industry knowledge and practical experience, as well as pre-stored display schemes, can be generated.
[0092] In a possible implementation manner, the cigarette and placement description information includes at least one of the following: The placement position of the cigarette, the product specification of the cigarette, the placement direction of the cigarette, the front cabinet information of the cigarette placement, the back cabinet information of the cigarette placement, the back cabinet grid information of the cigarette placement.
[0093] Further, referring to Figure 2 As shown in the block diagram, when the user uploads a picture, use the multi-modal large model to generate a general description for each picture and save the description information to the database. Here, mainly use the capabilities of the multi-modal large model itself to obtain general picture detail information. Secondly, the detection and recognition of the deep learning module further obtains more detail information about the picture and saves all relevant information of the pictures in the database.
[0094] Further, when the user asks a question related to the picture content, the large model first makes a self-analysis of the question and the user identity information, and then combines the original question and the analysis result to convert it into an SQL query statement to query the corresponding picture detection and recognition result. In the process of converting to SQL, it is necessary to combine the knowledge of the vector library, discover information related to the database from the question, and supplement the prompt words for the proper names through the knowledge vector library in the tobacco industry field to improve the quality of SQL generation.
[0095] Further, based on the image detection and recognition results obtained from the SQL query, the general description of the image is supplemented with details. For example, in the display image, how many cigarettes, brands, industrial company names, colors, row and column numbers, etc. exist. The large model then combines the question and the relevant description of the image to give a reply, realizing the ability to reply to questions in detail.
[0096] Further, in terms of display suggestions, the large model compares the results of the current image with excellent display templates and makes reasonable display suggestions for retail customers at different levels. Calculate whether there are consecutive packs and strips, the orderliness of display colors, price orderliness, and the rationality of the display positions of cigarettes of key product specifications. Conduct vector retrieval from the knowledge base in different dimensions and compare the display effects of the current image. The large model then gives a suggested answer based on the analysis results.
[0097] In this way, through the multi-modal large model combined with the conversion of different forms of professional field documents into vector libraries, and using the database and traditional detection and recognition algorithms, the implemented multi-modal tobacco intelligent question-answering system can ensure the timeliness and accuracy of data, and enhance the recognition ability of the display images uploaded by users, and effectively answer user questions and display suggestions.
[0098] The embodiments of the present disclosure also provide an electronic device, including: A processor; A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the executable instructions stored in the memory to implement the method described in any one of the above embodiments.
[0099] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the foregoing embodiments are implemented.
[0100] Figure 3 The tobacco question-answering device 100 based on the multi-modal large model shown includes a processor 1001 and a memory 1003. Among them, the processor 1001 and the memory 1003 are connected, such as connected through a bus 1002. Optionally, the tobacco question-answering device 100 based on the multi-modal large model may further include a communication component 1004. The communication component 1004 can be used for data interaction between the device 100 and other devices, such as sending and / or receiving data, etc. It should be noted that in actual scheduling, the communication component 1004 is not limited to one, and the structure of the tobacco question-answering device 100 based on the multi-modal large model does not constitute a limitation to the embodiments of the present application.
[0101] The processor 1001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 1001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0102] The bus 1002 may include a path for transmitting information between the above components. The bus 1002 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 1002 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 3 only a thick line is used to represent it herein, but it does not mean that there is only one bus or one type of bus.
[0103] The memory 1003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store program code and can be read by a computer, which is not limited herein.
[0104] The memory 1003 is used to store program codes for executing the embodiments of the present disclosure, and the execution is controlled by the processor 1001. The processor 1001 is used to execute the program codes stored in the memory 1003 to implement the steps shown in the embodiment of the tobacco question-answering method based on the multimodal large model.
[0105] The embodiments of the present disclosure also provide a computer-readable storage medium having program codes stored thereon. When the program codes are executed by a processor, the steps and corresponding contents of the aforementioned tobacco question-and-answer method embodiment based on a multimodal large model can be implemented.
[0106] The preferred embodiments of the present disclosure are described in detail above in conjunction with the accompanying drawings; however, the present disclosure is not limited to the specific details in the above embodiments; within the technical concept of the present disclosure, various changes, modifications, substitutions and variations may be made to these embodiments, and these changes, modifications, substitutions and variations all fall within the protection scope of the present disclosure.
[0107] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any appropriate manner without contradiction, and they should also be regarded as the contents disclosed in this disclosure. In order to avoid unnecessary repetition, this disclosure will not further describe various possible combinations. The technical scope of this application is not limited to the contents in the specification, and its technical scope must be determined according to the scope of the claims.
Claims
1. A tobacco question-answering method based on a multimodal large model, characterized in that: include: Performing image recognition on the acquired tobacco image to be identified, obtaining cigarette and placement description information in the tobacco image to be identified, and inputting the cigarette and placement description information into the multimodal large model to obtain supplementary description information of the multimodal large model for the cigarette and placement description information, and generating general description information; In response to the received question vector search, a question database table structure is generated, and according to the question vector search and the question database table structure, a vector search is performed from a vector library to obtain key database structure information, the key database structure information is converted from a vector to a text, and combined with the question vector search, a structured query language statement is generated; According to the structured query language statement, the cigarette and placement description information and the general description information are used as query information, and the prompt words are supplemented in combination with the tobacco knowledge vector library to generate an initial tobacco display suggestion; The user identity information corresponding to the question vector retrieval and the display scheme pre-stored in the display vector library are obtained, and the target tobacco display suggestion is generated according to the user identity information, the display scheme pre-stored in the display vector library and the initial display suggestion.
2. According to the tobacco question-answering method based on the multimodal big model of claim 1, the step of inputting the cigarette and placement description information into the multimodal big model, obtaining the supplementary description information of the multimodal big model for the cigarette and placement description information, and generating the general description information comprises: Inputting the cigarette and placement description information into a multimodal large model, wherein the multimodal large model retrieves a plurality of target tobacco images according to the cigarette and placement description information; Generate description information for image vectors corresponding to a plurality of target tobacco images to obtain supplementary description information of the multimodal large model for the cigarette and placement description information; Adding a first position code to the first vector corresponding to the supplementary description information, and adding a second position code to the text vector corresponding to the cigarette and placement description information, wherein the first position code and the second position code are used to represent input sequence information; According to the first position code and the second position code, the first vector corresponding to the supplementary description information and the text vector corresponding to the cigarette and placement description information are concatenated to generate the general description information.
3. The tobacco question-answering method based on a multimodal large model according to claim 2, further comprising: Determine the text descriptions corresponding to the plurality of target tobacco images for conformance and length, and determine the symbols and length of the text descriptions corresponding to the cigarette and placement description information; According to a preset length and a preset symbol rule, the text descriptions corresponding to the plurality of target tobacco images and the text descriptions corresponding to the cigarette and placement description information are divided into blocks, wherein the preset symbol rule includes line break blocks and paragraph identifier blocks; The text descriptions corresponding to the multiple target tobacco images after block processing are vectorized to generate the image vectors, and the text descriptions corresponding to the cigarettes and placement description information after block processing are vectorized to generate text vectors corresponding to the cigarettes and placement description information.
4. The tobacco question-answering method based on a multimodal large model according to claim 1, wherein the vector search is performed from a vector library according to the question vector search and the question database table structure to obtain key database structure information, the key database structure information is converted from a vector to a text, and combined with the question vector search, a structured query language statement is generated, including: Retrieve the question vector and generate the question vector, and retrieve the table structure information vector most similar to the question vector in the vector database according to the question vector and the table structure of the question database, and select the target structure table according to the similarity sorting of the table structure information vectors; Fill key data information in the target structure table to obtain key database structure information; Converting the extracted key database structure information from vector form to text form; Based on template matching and / or grammatical rules, the key database structure information in text form is combined with the question vector for retrieval, and a structured query language statement is generated while satisfying at least one of selecting an optimal query path, adding necessary indexes, and optimizing query conditions.
5. The tobacco question-answering method based on a multimodal large model according to claim 1, characterized in that: The method of using the cigarette and placement description information and the general description information as query information according to the structured query language statement and supplementing prompt words with the tobacco knowledge vector library to generate an initial tobacco display suggestion includes: The cigarette and placement description information and the general description information are used as query information for word segmentation and preprocessing of removing stop words to generate a query text; Generate an SQL query statement according to the query text and the SQL statement template corresponding to the structured query language statement; Execute an SQL query statement to retrieve target data corresponding to the query text from a database, wherein the target data includes at least one of sales data and consumer preferences; Convert the query text into a vector representation, and perform prompt word search in the tobacco knowledge vector database to generate prompt words; The original query information is supplemented according to the prompt word, and an initial tobacco display suggestion is generated by combining the retrieved target data and the query information supplemented by the prompt word. The initial tobacco display suggestion includes at least one of brand, type, placement location, and display method.
6. The tobacco question-answering method based on a multimodal large model according to claim 1, characterized in that: The step of generating a target tobacco display suggestion according to the user identity information, the display scheme pre-stored in the display vector library, and the initial display suggestion comprises: Determine a retail stall according to the question vector retrieval and the user identity information; Determine the preset conditions that are satisfied according to the retail stall, the display scheme pre-stored in the display vector library, and the initial display suggestion, wherein the preset conditions include at least one of the following: whether the packs and strips are connected, the orderliness of the display colors, the orderliness of the prices, and the rationality of the display positions of the key cigarettes; Generate target tobacco display recommendations based on the preset conditions that are met.
7. The tobacco question-answering method based on a multimodal large model according to claim 6, characterized in that: The display scheme pre-stored in the display vector library is generated in the following way: Based on the professional terms and professional knowledge of the tobacco industry, a professional knowledge base in the form of documents is constructed, wherein the professional terms and professional knowledge include at least one of ordering instructions, sales instructions, store stall instructions, content ratio meaning instructions and calculations, and tobacco laws and regulations instructions; Using the large model and the retrieval learning capabilities of RAG, the professional knowledge base is self-sorted and digested, and after the self-sorting and digestion are completed, a vectorized display vector library is generated; According to the pre-built display template, the template is entered into the display vector library to generate the display plan pre-stored in the display vector library.
8. The tobacco question-answering method based on a multimodal large model according to any one of claims 1 to 7, characterized in that: The cigarette and placement description information includes at least one of the following: The placement position of the cigarettes, the specifications of the cigarettes, the placement direction of the cigarettes, the front cabinet information where the cigarettes are placed, the back cabinet information where the cigarettes are placed, and the back cabinet compartment information where the cigarettes are placed.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Cigarette display method based on LDA model
CN116882685A
Visitor display guidance suggestion generation method and system based on AIGC technology
CN117197643A
Fine-grained multi-modal large model training method
CN118072128A
Method and system for automatically generating sql and replying questions
CN118093634A
Text abstract generation method and system for cigarette retail marketing public opinion monitoring
CN118673134A
Cited By
Image-text question answering method and device based on target detection and rule enhancement and electronic equipment
CN120892590A