Tobacco question and answer method, device and medium based on multi-modal large model

By combining image recognition and multimodal large models with vector retrieval to generate target tobacco display recommendations, the problem of insufficient question-answering capabilities of multimodal large models in the tobacco field in existing technologies is solved, and accurate and personalized display recommendations are achieved, improving the operational efficiency and user experience of tobacco companies.

CN120086391BActive Publication Date: 2025-10-10HEFEI CO OF ANHUI TOBACCO CO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510192955.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-10-10
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

The existing multimodal large model is insufficient in image question answering capabilities in the tobacco field and cannot adapt to the dynamic and ever-changing real-time tobacco market and the question-answering needs of users with different identities, resulting in inaccurate and in personalized display recommendations.

Method used

Cigarettes and their placement information are identified through image recognition technology, supplementary description information is generated by combining a multimodal large model, target tobacco display suggestions are generated using vector retrieval and structured query language, and display plans are optimized by combining user identity information and a display vector library.

Benefits of technology

It achieves accurate extraction and personalized recommendation of cigarette information, improves the operational efficiency and user experience of tobacco companies, and provides an intelligent, real-time and efficient tobacco question-and-answer system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086391B_ABST
    Figure CN120086391B_ABST
Patent Text Reader

Abstract

The application discloses a tobacco question and answer method and device based on a multi-modal large model, and a medium, which comprises the following steps: identifying a tobacco image to be identified to obtain a cigarette and placement description information, inputting the cigarette and placement description information into a multi-modal large model to obtain supplementary description information and generate general description information; in response to a question vector retrieval, generating a question database table structure, according to the question vector retrieval and the question database table structure, performing vector retrieval from a vector library to obtain key database structure information, converting the key database structure information from a vector into text, and combining the question vector retrieval to generate a structured query language statement; according to the structured query language statement, taking the cigarette and placement description information and the general description information as query information, combining a tobacco knowledge vector library to supplement a prompt word, and generating an initial tobacco display suggestion; and according to user identity information, a pre-stored display scheme in a display vector library and the initial display suggestion, generating a target tobacco display suggestion.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to the technical field of tobacco intelligent question-answering systems, and in particular to a tobacco question-answering method, device, and medium based on a multimodal large model. Background Art

[0002] In an era of rapid advancements in artificial intelligence, the development models of various industries are gradually moving towards intelligentization. In the field of intelligent question answering, the most noteworthy area is the multimodal data question answering. This has evolved from traditional text-to-text conversations to conversations involving data in various forms, such as images, documents, and web pages. Among multimodal large-scale model conversations, image question answering is undoubtedly the most popular. This is not only because images are more intuitive for humans, but also for large models.

[0003] In the tobacco industry, especially from the perspective of tobacco retailers, managers might prefer a multimodal large model that can accurately identify retail display images and provide sound recommendations for display placement. However, current multimodal large models are limited in their ability to answer questions about tobacco images. They lose significant semantic information from detailed images, especially when making display recommendations. The answers often contain fixed knowledge points, making them incapable of adapting to the dynamic and ever-changing tobacco market. Furthermore, the answers are often the same for users with different identities, making them incapable of adapting to different user identities. Summary of the Invention

[0004] The purpose of the present invention is to provide a tobacco question-and-answer method, device, and medium based on a multimodal large model to solve the technical problems in related scenarios where the results of tobacco question-and-answer cannot adapt to the dynamic and changeable real-time tobacco market and cannot adapt to questions and answers of users with different identities.

[0005] The purpose of the present invention can be achieved through the following technical solutions:

[0006] A first aspect provides a tobacco question-answering method based on a multimodal large model, comprising:

[0007] Performing image recognition on the acquired tobacco image to be identified to obtain description information of the cigarettes and their placement in the tobacco image to be identified, inputting the description information of the cigarettes and their placement into a multimodal macro model to obtain supplementary description information of the multimodal macro model for the description information of the cigarettes and their placement, and generating general description information;

[0008] In response to the received question vector search, a question database table structure is generated, and based on the question vector search and the question database table structure, a vector search is performed from a vector library to obtain key database structure information, the key database structure information is converted from a vector to text, and combined with the question vector search, a structured query language statement is generated;

[0009] According to the structured query language statement, the cigarette and placement description information and the general description information are used as query information, and the prompt words are supplemented in combination with the tobacco knowledge vector library to generate initial tobacco display suggestions;

[0010] Obtain user identity information corresponding to the question vector retrieval and a display scheme pre-stored in a display vector library, and generate a target tobacco display suggestion based on the user identity information, the display scheme pre-stored in the display vector library, and the initial display suggestion.

[0011] In one possible implementation, inputting the cigarette and placement description information into a multimodal macro model, obtaining supplementary description information of the multimodal macro model for the cigarette and placement description information, and generating general description information includes:

[0012] Inputting the cigarette and placement description information into a multimodal large model, wherein the multimodal large model retrieves a plurality of target tobacco images based on the cigarette and placement description information;

[0013] Generating description information for image vectors corresponding to the plurality of target tobacco images to obtain supplementary description information of the multimodal large model for the cigarettes and placement description information;

[0014] Adding a first position code to the first vector corresponding to the supplementary description information, and adding a second position code to the text vector corresponding to the cigarette and placement description information, wherein the first position code and the second position code are used to represent input sequence information;

[0015] According to the first position code and the second position code, the first vector corresponding to the supplementary description information and the text vector corresponding to the cigarette and placement description information are spliced ​​to generate the general description information.

[0016] In a possible implementation, the method further includes:

[0017] Determining the correspondence and length of the text descriptions corresponding to the plurality of target tobacco images and determining the symbols and length of the text descriptions corresponding to the cigarette and placement description information;

[0018] Segmenting the text descriptions corresponding to the plurality of target tobacco images and the text descriptions corresponding to the cigarette and placement description information into blocks according to a preset length and a preset symbol rule, wherein the preset symbol rule includes line break segmentation and paragraph identifier segmentation;

[0019] The text descriptions corresponding to the multiple target tobacco images after block processing are vectorized to generate the image vectors, and the text descriptions corresponding to the cigarettes and placement description information after block processing are vectorized to generate the text vectors corresponding to the cigarettes and placement description information.

[0020] In one possible implementation, performing a vector search from a vector library based on the question vector search and the question database table structure to obtain key database structure information, converting the key database structure information from a vector into text, and generating a structured query language statement in combination with the question vector search includes:

[0021] Retrieving the question vector, generating a question vector, and searching the vector database for a table structure information vector most similar to the question vector based on the question vector and the table structure of the question database, and sorting the table structure information vectors based on their similarity to select a target structure table;

[0022] Fill key data information in the target structure table to obtain key database structure information;

[0023] Converting the extracted key database structure information from vector form to text form;

[0024] Based on template matching and / or grammatical rules, the key database structure information in text form is combined with the question vector for retrieval, and a structured query language statement is generated while satisfying at least one of selecting an optimal query path, adding necessary indexes, and optimizing query conditions.

[0025] In a possible implementation, the structured query language statement is used to take the cigarette and placement description information and the general description information as query information, and the prompt words are supplemented in combination with the tobacco knowledge vector library to generate initial tobacco display suggestions, including:

[0026] The cigarette and placement description information and the general description information are used as query information for word segmentation and stop word removal preprocessing to generate a query text;

[0027] Generate an SQL query statement according to the query text and the SQL statement template corresponding to the structured query language statement;

[0028] executing an SQL query statement to retrieve target data corresponding to the query text from a database, wherein the target data includes at least one of sales data and consumer preferences;

[0029] Convert the query text into a vector representation, and perform prompt word search in the tobacco knowledge vector database to generate prompt words;

[0030] The original query information is supplemented according to the prompt word, and an initial tobacco display suggestion is generated by combining the retrieved target data and the query information supplemented with the prompt word. The initial tobacco display suggestion includes at least one of brand, type, placement location, and display method.

[0031] In a possible implementation, generating a target tobacco display suggestion based on the user identity information, the display scheme pre-stored in the display vector library, and the initial display suggestion includes:

[0032] Determining a retail stall based on the question vector retrieval and the user identity information;

[0033] Determining, based on the retail stall, the display scheme pre-stored in the display vector library, and the initial display suggestion, whether preset conditions are satisfied, wherein the preset conditions include at least one of the following: whether the packs and strips are connected, the orderliness of the display colors, the orderliness of the prices, and the rationality of the display locations of key cigarette specifications;

[0034] Generate target tobacco display recommendations based on the preset conditions that are met.

[0035] In a possible implementation, the display scheme pre-stored in the display vector library is generated in the following manner:

[0036] Construct a professional knowledge base in the form of documents based on professional terminology and professional knowledge of the tobacco industry, wherein the professional terminology and professional knowledge include at least one of ordering instructions, sales instructions, store stall instructions, content ratio meaning and calculation, and tobacco laws and regulations.

[0037] Utilize the large model and RAG's retrieval and learning capabilities to self-sort and digest the professional knowledge base, and after completing the self-sorting and digestion, generate a vectorized display vector library;

[0038] According to the pre-built display template, the template is entered into the display vector library to generate the display plan pre-stored in the display vector library.

[0039] In one possible implementation, the cigarette and placement description information includes at least one of the following:

[0040] The placement position of the cigarettes, the specifications of the cigarettes, the placement direction of the cigarettes, the front cabinet information where the cigarettes are placed, the back cabinet information where the cigarettes are placed, and the back cabinet compartment information where the cigarettes are placed.

[0041] A second aspect provides an electronic device, including:

[0042] processor;

[0043] a memory for storing processor-executable instructions;

[0044] The processor is configured to execute the executable instructions stored in the memory to implement any method described in the first aspect.

[0045] A third aspect provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any one of the methods described in the first aspect.

[0046] Compared with the existing solutions, the present invention achieves the following beneficial effects:

[0047] Image recognition technology can accurately identify information such as the brand, type, quantity, and placement of cigarettes, reducing errors in manual recognition. The automated recognition process is faster than manual recognition, significantly shortening recognition time and improving work efficiency. The identified cigarettes and placement description information are input into a multimodal large model to generate supplementary description information and general description information. The multimodal large model can combine various information such as the cigarette brand, taste, and origin to generate a more comprehensive and detailed product description, thereby improving users' understanding of the product. By combining user identity information and consumption habits, users can be provided with more personalized product recommendations and display suggestions, enhancing the user experience.

[0048] Furthermore, in response to the received question vector retrieval, a question database table structure is generated, and a vector search is performed from the vector library to obtain key database structure information, which is then converted into text and a structured query language statement is generated. Vector retrieval technology can efficiently retrieve key information that matches the question vector from massive data, shortening query time. Through structured query language statements, the required information can be accurately extracted from the database, reducing query errors. Based on the structured query language statement and the tobacco knowledge vector library, combined with user identity information and pre-stored display plans, target tobacco display suggestions are generated. By combining multiple information sources and algorithm optimization, tobacco display suggestions that meet user needs and market trends can be generated, thereby improving the attractiveness and purchase rate of product displays. The generated display suggestions can provide strong support for the operational decisions of tobacco companies, helping them optimize product layout and improve sales performance.

[0049] In summary, through key technologies such as image recognition, multimodal large models, vector retrieval and database query, and intelligent display suggestion generation, we have achieved accurate extraction, enriched display, efficient query, and intelligent display suggestion generation for tobacco product information. These technical achievements have collectively improved the operational efficiency and user experience of tobacco companies, providing strong support for their digital transformation and intelligent upgrades. Based on its own recognition results and traditional detection and recognition results, it benchmarks the current image against the template, evaluates each image, and responds to display suggestions based on different identities and templates. Leveraging sophisticated database management tools, it enables fast and convenient management of personnel identities, permissions, and display standards, dynamically modifying standard configurations, and updating Q&A results in real time. This improves the effectiveness of the multimodal large model for image Q&A and display suggestion generation in the tobacco sector, resulting in an intelligent, real-time, and efficient tobacco Q&A system. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The present invention will be further described below with reference to the accompanying drawings.

[0051] Figure 1 This is a flowchart of a tobacco question-answering method based on a multimodal large model according to an embodiment of the specification.

[0052] Figure 2 This is a schematic diagram of a tobacco question-answering system based on a multimodal large model according to an embodiment of the specification.

[0053] Figure 3 This is a block diagram of a tobacco question-and-answer device based on a multimodal large model according to an embodiment of the specification. DETAILED DESCRIPTION

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0055] like Figure 1 FIG. 1 is a flow chart of a tobacco question-answering method based on a multimodal large model according to the present invention, comprising:

[0056] In step S11, image recognition is performed on the acquired tobacco image to be identified to obtain description information of the cigarettes and their placement in the tobacco image to be identified, and the description information of the cigarettes and their placement is input into a multimodal macro model to obtain supplementary description information of the multimodal macro model for the description information of the cigarettes and their placement, thereby generating general description information;

[0057] Among them, image recognition refers to the process of using computer vision technology to extract useful information or features from the input image, and then identify objects, scenes or other content in the image.

[0058] The disclosed embodiments rely on deep learning algorithms, particularly convolutional neural networks (CNNs). These networks extract features from images layer by layer through multiple convolutional and pooling layers, ultimately performing classification or regression through fully connected layers. In tobacco image recognition, CNNs can learn features such as cigarette shape, color, and texture, accurately identifying cigarettes in the image and their placement. They detect the position of cigarettes in the image and identify their specifications. Based on the identification results, they determine whether the cigarettes are in order and positioned correctly (in the text direction). They then detect information such as the front cabinet, back cabinet, and back cabinet compartments, and store all of this information in a database using formatted fields. Different tables store the detection and recognition results of different algorithms, and record the image ID. For example, consider an image containing multiple brands of cigarettes. The image recognition system can identify the brand, quantity, and relative position of each cigarette in the image.

[0059] Among them, a multimodal large model refers to a deep learning model that can process and integrate information from multiple different modalities (such as text, images, audio, etc.).

[0060] In the disclosed embodiments, the multimodal large model is typically based on the Transformer architecture, using a self-attention mechanism to capture the associations between data from different modalities. In the tobacco field, the multimodal large model can accept cigarettes and their placement descriptions (textual modality) as input and, combined with its internal knowledge base, generate richer supplementary descriptive information, such as the cigarette brand story and taste evaluation.

[0061] In step S12, in response to the received question vector search, a question database table structure is generated, and based on the question vector search and the question database table structure, a vector search is performed from the vector library to obtain key database structure information, the key database structure information is converted from a vector to text, and combined with the question vector search, a structured query language statement is generated;

[0062] Among them, the Structured Query Language statement SQL is a programming language specifically used to communicate with the database, and is used to store, retrieve, update and manage data in the database.

[0063] In the disclosed embodiment, SQL statements allow users to perform various data operations by defining the structure, type, and relationship between data. During the vector retrieval and text conversion process, the generated SQL statements are used to retrieve key information matching the question vector from the database.

[0064] In step S13, according to the structured query language statement, the cigarette and placement description information and the general description information are used as query information, and the prompt words are supplemented in combination with the tobacco knowledge vector library to generate initial tobacco display suggestions;

[0065] Among them, the tobacco knowledge vector library refers to a database that stores a large amount of tobacco-related knowledge information and is expressed in the form of vectors.

[0066] In the disclosed embodiments, vectorization is the process of converting text or other types of data into numerical vectors, making them easier to process with machine learning models. Each vector in the tobacco knowledge vector library represents specific tobacco knowledge or attributes, such as brand, taste, and origin.

[0067] In the disclosed embodiment, accurate SQL statements are generated through the excellent text question-answering capabilities of the large language model. Furthermore, the use of a vector library effectively supplements professional knowledge points of the tobacco industry, enables knowledge sorting and analysis of excellent display templates, and can effectively provide suggestions for display images for users.

[0068] In step S14, the user identity information corresponding to the question vector retrieval and the display scheme pre-stored in the display vector library are obtained, and the target tobacco display suggestion is generated according to the user identity information, the display scheme pre-stored in the display vector library and the initial display suggestion.

[0069] Among them, target tobacco display suggestions refer to tobacco product placement suggestions generated based on user needs, product information and pre-defined display rules.

[0070] In the disclosed embodiment, the system generates a tobacco display plan that meets specific goals through algorithm optimization by combining user identity information (such as consumption habits and preferences), preset plans in the display vector library (such as optimal display angles, color matching, etc.), and initial display suggestions.

[0071] For example: For users who prefer to buy high-end cigarettes, the system may recommend placing Yellow Crane Tower cigarettes in a prominent location in the store, and pairing them with exquisite display racks and lighting effects to increase the attractiveness and purchase rate of the products.

[0072] Starting from image recognition, the descriptive information is enriched through a large multimodal model, and then SQL statements are generated based on question vector retrieval to retrieve key information from the database. Finally, the tobacco knowledge vector library and user information are combined to generate target tobacco display recommendations, realizing the complete transformation from image to intelligent display recommendations.

[0073] The above technical solution uses image recognition technology to accurately identify information such as the brand, type, quantity, and placement of cigarettes, reducing errors in manual recognition. The automated recognition process is faster than manual recognition, which can significantly shorten recognition time and improve work efficiency. The identified cigarettes and placement description information are input into a multimodal large model to generate supplementary description information and general description information. The multimodal large model can combine various information such as the brand, taste, and origin of the cigarettes to generate a more comprehensive and detailed product description, thereby improving users' understanding of the product. By combining user identity information and consumption habits, users can be provided with more personalized product recommendations and display suggestions, enhancing the user experience.

[0074] Furthermore, in response to the received question vector retrieval, a question database table structure is generated, and a vector search is performed from the vector library to obtain key database structure information, which is then converted into text and a structured query language statement is generated. Vector retrieval technology can efficiently retrieve key information that matches the question vector from massive data, shortening query time. Through structured query language statements, the required information can be accurately extracted from the database, reducing query errors. Based on the structured query language statement and the tobacco knowledge vector library, combined with user identity information and pre-stored display plans, target tobacco display suggestions are generated. By combining multiple information sources and algorithm optimization, tobacco display suggestions that meet user needs and market trends can be generated, thereby improving the attractiveness and purchase rate of product displays. The generated display suggestions can provide strong support for the operational decisions of tobacco companies, helping them optimize product layout and improve sales performance.

[0075] In summary, through key technologies such as image recognition, multimodal large models, vector retrieval and database query, and intelligent display suggestion generation, we have achieved accurate extraction, enriched display, efficient query, and intelligent display suggestion generation for tobacco product information. These technical achievements have collectively improved the operational efficiency and user experience of tobacco companies, providing strong support for their digital transformation and intelligent upgrades. Based on its own recognition results and traditional detection and recognition results, it benchmarks the current image against the template, evaluates each image, and responds to display suggestions based on different identities and templates. Leveraging sophisticated database management tools, it enables fast and convenient management of personnel identities, permissions, and display standards, dynamically modifying standard configurations, and updating Q&A results in real time. This improves the effectiveness of the multimodal large model for image Q&A and display suggestion generation in the tobacco sector, resulting in an intelligent, real-time, and efficient tobacco Q&A system.

[0076] In a possible implementation, in step S11, the cigarette and placement description information is input into the multimodal macro model, and the multimodal macro model obtains supplementary description information for the cigarette and placement description information, thereby generating general description information, including:

[0077] In step S111, the cigarette and placement description information are input into a multimodal large model, and the multimodal large model retrieves a plurality of target tobacco images based on the cigarette and placement description information;

[0078] In this disclosed embodiment, cigarette and placement descriptions are fed into a multimodal macromodel. Based on these descriptions, the multimodal macromodel leverages its internal knowledge and algorithms to retrieve multiple target cigarette images from a pre-trained image library that best match the descriptions. These images are highly semantically correlated with the descriptions.

[0079] In step S112, description information is generated for the image vectors corresponding to the plurality of target tobacco images to obtain supplementary description information of the multimodal large model for the cigarette and placement description information;

[0080] In this disclosed embodiment, the multimodal large model extracts image vectors from multiple retrieved target tobacco images. These image vectors and the model's internal generation mechanism are then used to generate supplementary descriptive information corresponding to the image content. This information may include a detailed description of features such as the cigarette brand, appearance, and packaging.

[0081] In step S113, a first position code is added to the first vector corresponding to the supplementary description information, and a second position code is added to the text vector corresponding to the cigarette and placement description information, wherein the first position code and the second position code are used to represent the input sequence information;

[0082] Among them, position encoding is used to introduce position information in sequence data to represent the order information of the input.

[0083] In this disclosed embodiment, to maintain the order of the input information, a first positional encoding is added to the first vector corresponding to the supplementary description information, and a second positional encoding is added to the text vector corresponding to the original cigarette and placement description information. The introduction of positional encoding enables the model to perceive the order of these information when processing them, which is crucial for subsequent splicing and generating general description information.

[0084] In step S114, based on the first position code and the second position code, the first vector corresponding to the supplementary description information and the text vector corresponding to the cigarette and placement description information are concatenated to generate the general description information.

[0085] In this disclosed embodiment, based on the first and second position codes, the multimodal macro model concatenates the first vector of the supplementary description information with the text vectors of the cigarette and placement description information. The concatenated vector contains both the original text description and the supplementary description information generated from the image, together forming the universal description information. This step achieves the conversion from multimodal input to unified text output, providing a rich information foundation for subsequent applications.

[0086] The multimodal large model can fully utilize the cigarette and placement description information and related image data to generate a universal description that combines both the original text description and the image supplementary description. This information processing method not only improves the accuracy and richness of the description, but also opens up a wider range of possibilities for subsequent applications.

[0087] In a possible implementation, the method further includes:

[0088] In step S115, the text descriptions corresponding to the plurality of target tobacco images are determined to be consistent and have their lengths determined, and the text descriptions corresponding to the cigarette and placement description information are determined to have their symbols and lengths determined;

[0089] In this disclosed embodiment, symbols, such as line breaks and paragraph identifiers, are determined within the text descriptions corresponding to multiple target tobacco images and the text descriptions of the cigarettes and their placement. These symbols are used to mark the structure and format of the text descriptions. The number of characters or length of these text descriptions is then calculated to facilitate subsequent block processing and vectorization.

[0090] In step S116, the text descriptions corresponding to the plurality of target tobacco images and the text descriptions corresponding to the cigarette and placement description information are segmented according to a preset length and a preset symbol rule, wherein the preset symbol rule includes segmentation by line break and segmentation by paragraph identifier;

[0091] Chunking is the process of breaking a long text description into multiple shorter text blocks. This processing method helps reduce processing difficulty and improve processing efficiency, while also helping to maintain the coherence and readability of the text description.

[0092] In the disclosed embodiment, the text descriptions corresponding to multiple target tobacco images, as well as the text descriptions corresponding to the cigarette and placement descriptions, are segmented into blocks based on a preset length (e.g., the number of characters per block) and preset symbol rules (e.g., segmentation by line break, segmentation by paragraph identifier, etc.). Specific operations of the segmentation process may include identifying symbol positions, segmenting text blocks, and maintaining the order and coherence of text blocks. Through segmentation, the original text description is divided into multiple shorter text blocks.

[0093] In step S117, the text descriptions corresponding to the multiple target tobacco images after block processing are vectorized to generate the image vectors, and the text descriptions corresponding to the cigarettes and placement description information after block processing are vectorized to generate the text vectors corresponding to the cigarettes and placement description information.

[0094] In this disclosed embodiment, the text descriptions corresponding to the multiple block-processed target tobacco images, as well as the text descriptions corresponding to the cigarettes and their placement, are vectorized. The vectorization process may include steps such as text preprocessing (such as stop word removal and stemming), word embedding (such as Word2Vec and BERT), or sentence embedding (such as Sentence-BERT). These steps convert the text blocks into vector representations, capturing their semantic features.

[0095] Through vectorization, we obtain image and text vectors that can be used in subsequent machine learning or deep learning models. These vectors represent the semantic content of the original text description and provide a basis for subsequent processing and analysis.

[0096] This allows for processing multimodal input information (including images and text descriptions) and converting it into a unified vector representation. This approach not only improves information processing efficiency but also provides a rich feature base for subsequent applications (such as text generation, information retrieval, and recommendation systems).

[0097] In one possible implementation, in step S12, performing a vector search from a vector library based on the question vector search and the question database table structure to obtain key database structure information, converting the key database structure information from a vector into text, and generating a structured query language statement in combination with the question vector search includes:

[0098] In step S121, a question vector is generated based on the question vector search, and based on the question vector and the question database table structure, a table structure information vector that is most similar to the question vector is searched in the vector database, and the table structure information vectors are sorted according to their similarity to select a target structure table;

[0099] Question vector retrieval involves converting a user's question or query into a vector representation and searching a vector database to find the most similar vector. In this technology, question vector retrieval is used to understand user intent and locate database table structure information related to the question.

[0100] Table structure information vectorization involves converting the structural information of a database table (such as table name, field name, and field type) into a vector representation. This representation method can capture the similarities and differences between table structures, facilitating subsequent retrieval and matching.

[0101] Similarity sorting is the process of sorting retrieval results based on the similarity between vectors. In this technology, similarity sorting is used to select the table structure information vector that is most similar to the target question vector.

[0102] In the disclosed embodiments, a question vector is generated based on the question or query input by the user. This typically involves converting the question text into a vector representation, possibly using word embedding or sentence embedding techniques. Next, based on the question vector and the table structure information in the question database, a search is performed in the vector library to find the table structure information vector that is most similar to the question vector. The search results are sorted based on the similarity of the table structure information vectors, and the table with the highest similarity is selected as the target structure table.

[0103] In step S122, key data information is filled in the target structure table to obtain key database structure information;

[0104] Key data refers to the data fields and information in the database table structure that are crucial for solving user questions or queries. In this technology, key data is filled into the target structure table to generate complete database query conditions.

[0105] In the disclosed embodiment, key data information is populated based on the user's question or query and the field information of the target structure table. This may involve extracting key field values ​​from the user input or inferring field values ​​based on the context. The populated target structure table contains key database structure information.

[0106] In step S123, the extracted key database structure information is converted from vector form to text form;

[0107] In the embodiment of the present disclosure, the extracted key database structure information is converted from vector form to text form, which generally involves reversing the vector representation into a readable text representation, such as converting a field name vector into an actual field name.

[0108] In step S124, based on template matching and / or grammatical rules, the key database structure information in text form is combined with the question vector retrieval, and a structured query language statement is generated while satisfying at least one of selecting the optimal query path, adding necessary indexes, and optimizing query conditions.

[0109] Template matching involves matching input information with a template based on a preset template or rules to generate results that conform to a specific format or requirement. In this technology, template matching is used to combine key database structure information in text form with the question vector to generate a Structured Query Language (SQL) statement.

[0110] Syntax rules are the rules and conventions that define the structure of a language. In this technology, grammar rules are used to ensure that the generated SQL statements conform to the grammatical requirements of the SQL language and can correctly execute database queries.

[0111] In the disclosed embodiments, key database structure information in text form is combined with the question vector based on template matching and / or grammatical rules to generate SQL statements. During the SQL statement generation process, multiple factors may be considered to optimize the query, such as selecting the optimal query path, adding necessary indexes, and optimizing query conditions. The resulting SQL statement complies with the grammatical requirements of the SQL language, correctly executes the database query, and returns the results desired by the user.

[0112] Through this series of steps, the user's question or query is converted into a vector representation and the database table structure information related to the question is found in the vector library. The system then generates a structured SQL statement based on this information to execute the database query and return the results. This technical approach improves the efficiency and accuracy of database queries.

[0113] In a possible implementation, in step S13, the cigarette and placement description information and the general description information are used as query information based on the structured query language statement, and the prompt words are supplemented in combination with the tobacco knowledge vector library to generate initial tobacco display suggestions, including:

[0114] In step S131, the cigarette and placement description information and the general description information are used as query information for word segmentation and stop word removal to generate a query text;

[0115] Word segmentation is the process of breaking a continuous text string into individual words or phrases. In Chinese text processing, word segmentation is a crucial step in text preprocessing, facilitating subsequent tasks such as word frequency statistics, keyword extraction, and text classification. Stop words, such as "的," "了," and "在," appear frequently in text but contribute little to the text's semantics. Removing stop words reduces redundancy in text data and improves the efficiency and accuracy of subsequent text processing.

[0116] In this disclosed embodiment, the cigarette and placement descriptions, along with the general description, are used as query information, and preprocessed by word segmentation and stop word removal. Word segmentation breaks the text into individual words or phrases, while stop word removal reduces redundancy in the text data. The preprocessed text serves as the query text.

[0117] In step S132, an SQL query statement is generated according to the query text and the SQL statement template corresponding to the structured query language statement;

[0118] The SQL statement template refers to a preset SQL statement structure containing placeholders. By replacing the placeholders with actual values, a required SQL query statement can be quickly generated.

[0119] In the embodiments of the present disclosure, the SQL query statement is generated according to the query text and the SQL statement template corresponding to the structured query language statement. The SQL statement template contains placeholders, and the system generates a required SQL query statement by replacing the placeholders with actual values in the query text.

[0120] In step S133, the SQL query statement is executed to retrieve target data corresponding to the query text from the database, wherein the target data at least includes one of sales data and consumer preferences;

[0121] The sales data refers to data recording the sales of goods, including sales volume, sales amount, sales time, sales channel, etc. The consumer preference refers to the degree of preference of consumers for goods or services.

[0122] In step S134, the query text is converted into a vector representation, and a prompt word is retrieved in the tobacco knowledge vector library;

[0123] In the embodiments of the present disclosure, the query text is converted into a vector representation, and retrieval is performed in the tobacco knowledge vector library. Through vector retrieval, the system finds the most similar tobacco knowledge vector to the query text and extracts the prompt word therefrom

[0124] In step S135, the original query information is supplemented according to the prompt word, and an initial tobacco display suggestion is generated based on the retrieved target data and the query information supplemented by the prompt word, wherein the initial tobacco display suggestion includes at least one of brand, type, placement position, and display method.

[0125] In the embodiments of the present disclosure, the original query information is supplemented according to the prompt word, and an initial tobacco display suggestion is generated based on the retrieved target data and the query information supplemented by the prompt word. The initial tobacco display suggestion includes at least one of brand, type, placement position, and display method, which is based on information in the sales data, consumer preferences, and tobacco knowledge vector library, aiming to improve the display effect and sales performance of tobacco products.

[0126] Through the series of steps, the system can generate an initial tobacco display suggestion based on the query information input by the user and the structured query language statement, in combination with information in the tobacco knowledge vector library and sales data, which helps to optimize the display method and position of tobacco products, and improve the purchase willingness of consumers and sales performance.

[0127] In a possible implementation, in step S14, generating a target tobacco display suggestion based on the user identity information, the display scheme pre-stored in the display vector library, and the initial display suggestion includes:

[0128] In step S141, a retail stall is determined based on the question vector retrieval and the user identity information;

[0129] Retail tiers refer to the hierarchical management of retailers based on factors such as their scale of operation, sales capacity, and geographic location. Different retail tiers enjoy different policy benefits and resource allocations, such as supply volume and display support.

[0130] In this disclosed embodiment, the system determines the user's retail location based on a question vector search (which may include information such as the user's input question description and past purchase history) and user identity information (such as retailer ID and business address). Retail location is an important factor in generating subsequent display suggestions, as retailers in different locations may have different supply volumes and display requirements.

[0131] In step S142, based on the retail location, the display schemes pre-stored in the display vector library, and the initial display suggestion, the preset conditions are determined to be satisfied, wherein the preset conditions include at least one of the following: whether the packs and strips are connected, the orderliness of the display colors, the orderliness of the prices, and the rationality of the display locations of key cigarettes;

[0132] Among them, "pack and strip" refers to the display of tobacco products by closely arranging multiple packs or strips of the same brand together to create a uniform visual effect. This display method helps to highlight the brand image and attract consumers' attention.

[0133] Display color orderliness means that tobacco products should be arranged in an orderly manner according to visual factors such as color depth, warm and cool tones, etc., when displayed, so as to enhance the overall beauty and appeal of the display.

[0134] Price orderliness means that tobacco products are displayed in an orderly manner from low to high or from high to low according to price, so that consumers can quickly find products that meet their budget.

[0135] Key product specifications are those with high sales volume, high visibility, and high profit contribution. These products should be displayed in a prominent position so that consumers can immediately see and purchase them. The rationality of display placement directly affects the sales performance of key product specifications.

[0136] In the disclosed embodiment, the preset conditions are determined based on the retail location, pre-stored display plans in the display vector library, and initial display suggestions. These conditions include whether the packs and strips are connected, the orderliness of the display colors and prices, and the rationality of the display locations for key cigarette types. These conditions are used to assess whether the display plan meets the actual needs of retailers and the display requirements of tobacco products.

[0137] Specifically, the system may compare the initial display suggestions with the solutions in the display vector library to check whether the initial suggestions meet the requirements of continuous packages and strips (for example, whether products of the same brand are arranged closely); whether they are arranged in order according to color depth or warm and cool tones; whether prices are arranged in order from low to high or from high to low; and whether key product specifications occupy a prominent position.

[0138] In step S143, a target tobacco display suggestion is generated based on the satisfied preset conditions.

[0139] In the disclosed embodiment, the system generates targeted tobacco display suggestions based on the pre-set conditions met. These suggestions may include adjusting the display layout to meet the requirements of continuous packs and strips, optimizing color arrangement to enhance visual appeal, adjusting price order to facilitate consumer comparison, and rearranging the display location of key product specifications to increase sales.

[0140] Targeted tobacco display recommendations are designed to comprehensively consider retail locations, display solution libraries, and initial recommendations to provide retailers with both practical and attractive tobacco product display solutions. These recommendations help improve tobacco product display effectiveness and sales performance, while also enhancing consumer shopping experience and satisfaction.

[0141] Through this series of steps, the system generates targeted tobacco display recommendations that meet retailers' actual needs and tobacco product display requirements based on user identity information, a display vector library, and initial display recommendations. These recommendations provide retailers with guidance and support for optimizing tobacco product displays, helping to improve sales performance and consumer satisfaction.

[0142] In a possible implementation, the display scheme pre-stored in the display vector library is generated in the following manner:

[0143] Construct a professional knowledge base in the form of documents based on professional terminology and professional knowledge of the tobacco industry, wherein the professional terminology and professional knowledge include at least one of ordering instructions, sales instructions, store stall instructions, content ratio meaning and calculation, and tobacco laws and regulations.

[0144] In the disclosed embodiments, a knowledge base in the form of documents is constructed based on tobacco industry terminology and expertise. These documents cover ordering instructions, sales instructions, store location descriptions, content ratio explanations and calculations, tobacco laws and regulations, and other aspects, ensuring the comprehensiveness and accuracy of the knowledge base.

[0145] Utilize the large model and RAG's retrieval and learning capabilities to self-sort and digest the professional knowledge base, and after completing the self-sorting and digestion, generate a vectorized display vector library;

[0146] In the disclosed embodiment, the large model and the retrieval and learning capabilities of RAG are used to self-sort and digest the professional knowledge base. The large model has powerful text comprehension and reasoning capabilities, and can deeply understand the terms and descriptions in the professional knowledge base. The RAG framework allows the model to retrieve external knowledge bases as needed during the sorting process (although the external knowledge base in this case is the professional knowledge base itself, the introduction of the RAG framework emphasizes the combination of retrieval and generation capabilities), thereby enhancing the understanding and absorption of professional knowledge. After completing the self-sorting and digestion, the text information in the professional knowledge base is converted into vector representations using vectorization technology. These vector representations retain the semantic relationships and similarities between texts in high-dimensional space.

[0147] According to the pre-built display template, the template is entered into the display vector library to generate the display plan pre-stored in the display vector library.

[0148] In the disclosed embodiments, preset display templates are constructed based on tobacco industry display standards and consumer demand. These templates contain a series of rules and suggestions on how to place and display tobacco products, aiming to ensure that the display effect meets industry standards and consumer aesthetics.

[0149] The template is entered into the display vector library. Based on the rules and recommendations in the display template, relevant information is retrieved and matched from the vectorized display vector library to generate pre-stored display plans. These plans combine industry knowledge from the professional knowledge library with practical experience from the display templates to provide comprehensive and accurate guidance for tobacco product display.

[0150] Through the above steps, a display vector library containing rich industry knowledge and practical experience, as well as pre-stored display plans, can be generated.

[0151] In one possible implementation, the cigarette and placement description information includes at least one of the following:

[0152] The placement position of the cigarettes, the specifications of the cigarettes, the placement direction of the cigarettes, the front cabinet information where the cigarettes are placed, the back cabinet information where the cigarettes are placed, and the back cabinet compartment information where the cigarettes are placed.

[0153] Further, see Figure 2 As shown in the block diagram, after a user uploads an image, the multimodal large model generates a general description for each image and saves this description in the database. This primarily utilizes the inherent capabilities of the multimodal large model to obtain a general overview of the image's details. Next, the deep learning module performs detection and recognition to further obtain more detailed information about the image, and all relevant information about the images is stored in the database.

[0154] Furthermore, when a user asks a question related to the image content, the large model first analyzes the question and the user's identity information. It then combines the original question and the analysis results into a SQL query to retrieve the corresponding image detection and recognition results. This SQL conversion process leverages knowledge from the vector library to extract database-related information from the question. Furthermore, the knowledge vector library within the tobacco industry is used to supplement proper nouns with prompts to improve SQL generation quality.

[0155] Furthermore, based on the image detection and recognition results obtained from the SQL query, the general description of the image is supplemented with details. For example, the number of cigarettes, brand, industrial company name, color, row and column numbers, etc., are displayed in the image. The large model then combines the question and the relevant description of the image to respond in detail.

[0156] Furthermore, for display recommendations, the large-scale model compares the current image's results with excellent display templates, making reasonable display suggestions for retailers at different price points. It calculates whether the packs and strips are connected, the orderliness of the display colors and prices, and the rationality of the display positions for key cigarette sizes. It then searches the knowledge base for vectors from different dimensions and compares the current image's display effects. The large-scale model then makes recommendations based on the analysis results.

[0157] In this way, by combining a large multimodal model with different forms of professional documents to transform them into a vector library, and using databases and traditional detection and recognition algorithms, a multimodal tobacco intelligent question-and-answer system is realized. This system can ensure the real-time and accuracy of the data, enhance the recognition capability of display images uploaded by users, and provide effective responses to user questions and display suggestions.

[0158] The present disclosure also provides an electronic device, including:

[0159] processor;

[0160] a memory for storing processor-executable instructions;

[0161] The processor is configured to execute the executable instructions stored in the memory to implement the method described in any one of the above embodiments.

[0162] An embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method described in any one of the aforementioned embodiments are implemented.

[0163] Figure 3 The tobacco question-and-answer device 100 based on a multimodal large model shown includes: a processor 1001 and a memory 1003. The processor 1001 and the memory 1003 are connected, such as via a bus 1002. Optionally, the tobacco question-and-answer device 100 based on a multimodal large model may further include a communication component 1004, which may be used for data exchange between the device 100 and other devices, such as data transmission and / or data reception. It should be noted that in actual scheduling, the number of communication components 1004 is not limited to one, and the structure of the tobacco question-and-answer device 100 based on a multimodal large model does not constitute a limitation on the embodiments of the present application.

[0164] Processor 1001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 1001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, or a combination of a DSP and a microprocessor.

[0165] Bus 1002 may include a path for transmitting information between the above components. Bus 1002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. Bus 1002 may be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, Figure 3 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0166] The memory 1003 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store program code and can be read by a computer, without limitation herein.

[0167] The memory 1003 is used to store program codes for executing the embodiments of the present disclosure, and the execution is controlled by the processor 1001. The processor 1001 is used to execute the program codes stored in the memory 1003 to implement the steps shown in the embodiment of the tobacco question-answering method based on the multimodal large model.

[0168] The embodiment of the present disclosure also provides a computer-readable storage medium, which stores program code. When the program code is executed by a processor, it can implement the steps and corresponding contents of the aforementioned tobacco question-and-answer method embodiment based on a multimodal large model.

[0169] The preferred embodiments of the present disclosure are described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the technical concept of the present disclosure, various changes, modifications, replacements and variations can be made to these embodiments, and these changes, modifications, replacements and variations all fall within the scope of protection of the present disclosure.

[0170] It should also be noted that the various specific technical features described in the above specific embodiments may be combined in any suitable manner, unless there is any contradiction, and these combinations shall also be considered as the contents disclosed in this disclosure. To avoid unnecessary repetition, this disclosure will not further describe various possible combinations. The technical scope of this application is not limited to the contents of the specification and must be determined based on the scope of the claims.

Claims

1. A tobacco question-answering method based on a multimodal large model, characterized in that: include: Performing image recognition on the acquired tobacco image to be identified to obtain description information of the cigarettes and their placement in the tobacco image to be identified, inputting the description information of the cigarettes and their placement into a multimodal macro model to obtain supplementary description information of the multimodal macro model for the description information of the cigarettes and their placement, and generating general description information; In response to the received question vector search, a question database table structure is generated, and based on the question vector search and the question database table structure, a vector search is performed from a vector library to obtain key database structure information, the key database structure information is converted from a vector to text, and combined with the question vector search, a structured query language statement is generated; According to the structured query language statement, the cigarette and placement description information and the general description information are used as query information, and the prompt words are supplemented in combination with the tobacco knowledge vector library to generate initial tobacco display suggestions; Obtain user identity information corresponding to the question vector retrieval and a display scheme pre-stored in a display vector library, and generate a target tobacco display suggestion based on the user identity information, the display scheme pre-stored in the display vector library, and the initial tobacco display suggestion.

2. The tobacco question-answering method based on a multimodal large model according to claim 1, wherein the step of inputting the cigarette and placement description information into the multimodal large model, obtaining supplementary description information of the cigarette and placement description information by the multimodal large model, and generating general description information comprises: Inputting the cigarette and placement description information into a multimodal large model, wherein the multimodal large model retrieves a plurality of target tobacco images based on the cigarette and placement description information; Generating description information for image vectors corresponding to the plurality of target tobacco images to obtain supplementary description information of the multimodal large model for the cigarettes and placement description information; Adding a first position code to the first vector corresponding to the supplementary description information, and adding a second position code to the text vector corresponding to the cigarette and placement description information, wherein the first position code and the second position code are used to represent input sequence information; According to the first position code and the second position code, the first vector corresponding to the supplementary description information and the text vector corresponding to the cigarette and placement description information are spliced ​​to generate the general description information.

3. The tobacco question-answering method based on a multimodal large model according to claim 2, further comprising: Determining the symbols and lengths of the text descriptions corresponding to the plurality of target tobacco images and determining the symbols and lengths of the text descriptions corresponding to the cigarette and placement description information; Segmenting the text descriptions corresponding to the plurality of target tobacco images and the text descriptions corresponding to the cigarette and placement description information into blocks according to a preset length and a preset symbol rule, wherein the preset symbol rule includes line break segmentation and paragraph identifier segmentation; The text descriptions corresponding to the multiple target tobacco images after block processing are vectorized to generate the image vectors, and the text descriptions corresponding to the cigarettes and placement description information after block processing are vectorized to generate the text vectors corresponding to the cigarettes and placement description information.

4. The tobacco question-answering method based on a multimodal large model according to claim 1, wherein the step of performing a vector search from a vector library based on the question vector search and the question database table structure to obtain key database structure information, converting the key database structure information from vectors into text, and combining the question vector search to generate a structured query language statement comprises: Retrieving the question vector and generating a question vector, and searching the vector library for a table structure information vector that is most similar to the question vector based on the question vector and the table structure of the question database, and sorting the table structure information vectors according to their similarity to select a target structure table; Filling key data information in the target structure table to obtain key database structure information; converting the extracted key database structure information from vector form to text form; Based on template matching and / or grammatical rules, the key database structure information in text form is combined with the question vector for retrieval, and a structured query language statement is generated while satisfying at least one of selecting an optimal query path, adding necessary indexes, and optimizing query conditions.

5. The tobacco question-answering method based on a multimodal large model according to claim 1, characterized in that: The method of generating an initial tobacco display suggestion by using the cigarette and placement description information and the general description information as query information based on the structured query language statement and supplementing prompt words with a tobacco knowledge vector library in combination includes: The cigarette and placement description information and the general description information are used as query information for word segmentation and stop word removal preprocessing to generate a query text; Generate an SQL query statement according to the query text and the SQL statement template corresponding to the structured query language statement; executing an SQL query statement to retrieve target data corresponding to the query text from a database, wherein the target data includes at least one of sales data and consumer preferences; Convert the query text into a vector representation, and perform prompt word search in the tobacco knowledge vector database to generate prompt words; The original query information is supplemented according to the prompt word, and an initial tobacco display suggestion is generated by combining the retrieved target data and the query information supplemented with the prompt word. The initial tobacco display suggestion includes at least one of brand, type, placement location, and display method.

6. The tobacco question-answering method based on a multimodal large model according to claim 1, characterized in that: Generating a target tobacco display suggestion based on the user identity information, the display scheme pre-stored in the display vector library, and the initial tobacco display suggestion includes: Determining a retail stall based on the question vector retrieval and the user identity information; Based on the retail stall, the display plans pre-stored in the display vector library, and the initial tobacco display suggestion, determine whether preset conditions are met, wherein the preset conditions include at least one of the following: whether the packs and strips are connected, the orderliness of the display colors, the orderliness of the prices, and the rationality of the display location of key cigarette specifications; based on the satisfied preset conditions, generate a target tobacco display suggestion.

7. The tobacco question-answering method based on a multimodal large model according to claim 6, characterized in that: The display schemes pre-stored in the display vector library are generated in the following manner: Construct a professional knowledge base in the form of documents based on professional terminology and professional knowledge of the tobacco industry, wherein the professional terminology and professional knowledge include at least one of ordering instructions, sales instructions, store stall instructions, content ratio meaning and calculation, and tobacco laws and regulations. Utilize the large model and RAG's retrieval and learning capabilities to self-sort and digest the professional knowledge base, and after completing the self-sorting and digestion, generate a vectorized display vector library; According to the pre-built display template, the template is entered into the display vector library to generate the display plan pre-stored in the display vector library.

8. The tobacco question-answering method based on a multimodal large model according to any one of claims 1 to 7, characterized in that: The cigarette and placement description information includes at least one of the following: The placement position of the cigarettes, the specifications of the cigarettes, the placement direction of the cigarettes, the front cabinet information where the cigarettes are placed, the back cabinet information where the cigarettes are placed, and the back cabinet compartment information where the cigarettes are placed.

9. An electronic device, characterized in that: include: processor; A memory for storing processor-executable instructions; wherein the processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Visitor display guidance suggestion generation method and system based on AIGC technology

    CN117197643A

  • Fine-grained multi-modal large model training method

    CN118072128A