Method for generating commodity label on the basis of large language model, and electronic device
By obtaining product description information from multiple sources, and using large language models and OCR technology to generate product labels, solving the problems of low accuracy and batch mounts of product labels, achieving more attractive and covert product label generation.
Patent Information
- Application Number
- PCT/CN2024/115448
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-08-29
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, the generation of product selling point labels on the product display page depends on the product title, resulting in low label accuracy and lack of attractiveness, and inconsistent labels between different merchants or platforms, so that batch mounts of the same product cannot be achieved.
By obtaining product description information from multiple sources, using fine-tuned large language models for structured processing and language prompt thesaurus processing, objective attribute features are generated, and product labels are generated in combination with FABE rules, including product characteristics, advantages and benefits, and OCR is used to identify graphic details to realize batch mount of product labels.
A more reliable, highly confident and traceable product label was generated, which can highlight the characteristics and advantages of the product, improve the attractiveness and shopping experience of the product, and realize batch mount of the same product on different platforms.
Smart Images

Figure CN2024115448_03072025_PF_FP_ABST
Abstract
Description
Method and electronic device for generating product labels based on large language model Technical Field
[0001] The present application relates to the field of electronic technology, and in particular to a method and electronic device for generating product labels based on a large language model. Background Art
[0002] With the development of the e-commerce industry, users are increasingly shopping online. Users can browse different stores and different types of goods on e-commerce platforms or food delivery platforms and purchase the goods they need according to their needs.
[0003] When browsing products, the product selling point label is one of the important elements on the product display page. It can concisely summarize the characteristics and advantages of the product, attract the attention of potential consumers, and allow consumers to see the highlights and advantages of the product at a glance, thereby promoting sales.
[0004] Currently, the generation of product selling point labels presented on product display pages depends on the product title. The information obtained for the product selling point labels is very limited, the generated product selling point labels have low accuracy, and lack appeal to consumers.
[0005] In addition, the same product may be sold by different merchants or platforms. The current label generation method mainly relies on the product title or product description entered by different merchants. Therefore, different product labels may be generated for the same product, and batch mounting of the same product cannot be achieved.
[0006] Summary of the Invention
[0007] The present application provides a method and electronic device for generating product labels based on a large language model. The method can generate product selling point labels through product description information obtained from multiple sources and multiple channels, expand the sources of product description information, and generate more reliable, higher confidence, and traceable selling point labels.
[0008] In a first aspect, a method for generating product labels based on a large language model is provided, comprising:
[0009] Obtain product description information related to the target product;
[0010] Structuring the product description information to obtain objective attribute characteristics of the target product;
[0011] The objective attribute features are processed using a fine-tuned large language model based on a language prompt vocabulary to obtain one or more product labels for the target product.
[0012] In a second aspect, a device for generating product labels based on a large language model is provided, comprising:
[0013] An acquisition unit, used to acquire product description information related to a target product;
[0014] A first processing unit is configured to perform structured processing on the product description information to obtain objective attribute characteristics of the target product;
[0015] The second processing unit is configured to process the objective attribute features using the fine-tuned large language model based on the language prompt vocabulary to obtain one or more product labels for the target product.
[0016] In combination with the second aspect, in certain implementations of the second aspect, the acquisition unit is also used to obtain a detail picture associated with the target product; the device for generating product labels also includes a third processing unit, which is used to perform character recognition on the detail picture to obtain one or more text fragments in the detail picture; and the one or more text fragments are spliced to obtain the text information.
[0017] In a third aspect, a server is provided, comprising a memory and a processor. The memory is configured to store executable program code, and the processor is configured to call and execute the executable program code from the memory, so that the device executes the method of the first aspect or any possible implementation of the first aspect.
[0018] In a fourth aspect, a computer program product is provided, comprising: a computer program code, which, when executed on a computer, enables the computer to execute the method in the first aspect or any possible implementation of the first aspect.
[0019] In a fifth aspect, a computer-readable storage medium is provided, which stores a computer program code. When the computer program code runs on a computer, the computer executes the method in the above-mentioned first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] FIG1 is a schematic diagram of a graphical user interface started by a user running application A. FIG.
[0021] FIG2 is a schematic flowchart of a method for generating a product label provided in an embodiment of the present application.
[0022] FIG3 is a general flow chart of an implementation process of generating a product label provided in an embodiment of the present application.
[0023] FIG4 is a flowchart of another example of an implementation process for generating a product label provided in an embodiment of the present application.
[0024] FIG5 is a flowchart of another example of an implementation process for generating a product label provided in an embodiment of the present application.
[0025] FIG6 is a flowchart of another example of an implementation process for generating a product label provided in an embodiment of the present application.
[0026] FIG7 is a flowchart of another example of an implementation process for generating a product label provided in an embodiment of the present application.
[0027] FIG8 is a flowchart of another example of an implementation process for generating a product label provided in an embodiment of the present application.
[0028] FIG9 is a schematic diagram of an apparatus for generating product labels based on a large language model according to an embodiment of the present application.
[0029] FIG10 is a schematic structural diagram of a computer device 1000 provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] For ease of understanding, the following will take the first application (such as application A) installed on the mobile phone as an example, and combine with the accompanying drawings to specifically explain the scenario in which the running interface of application A is displayed to the user in the first application (such as application A) installed on the mobile phone.
[0031] FIG1 is a schematic diagram of a graphical user interface (GUI) started by a user running application A. FIG.
[0032] For example, Figure 1 (a) shows an interface 101 displayed on a mobile phone in unlocked mode. This interface 101 displays a weather clock component and multiple applications (apps). Applications may include phone, messages, settings, and application A. It should be understood that this interface 101 may also include other applications, and this embodiment of the application is not limited to this.
[0033] As shown in Figure 1 (a), a user clicks the icon of Application A. In response to the user's click, the mobile phone displays the main interface 102 of Application A, as shown in Figure 1 (b). This interface 102 can also be referred to as the "homepage of Application A." This main interface 102 of Application A can display multiple category menus, operable controls or buttons, images, and other interface content to meet the user's usage needs.
[0034] For example, as shown in Figure 1 (b), the main interface 102 of the application A displays the current delivery address (for example, "XX District XX Street"), a search box, and different category menus such as food takeout, supermarkets, fruits, buying medicine, desserts, hamburgers, lobsters and barbecue, as well as a list of merchants, etc. The embodiment of the present application does not limit the display content, the size of the display area, etc. on the main interface 102 of the application A.
[0035] Currently, on the main interface 102 of application A, one or more merchants can be displayed for the user in the merchant list, and the user can view more merchant information by sliding up and down. Optionally, each merchant display area can display one or more contents such as the merchant's user rating, monthly sales, delivery time, delivery distance, minimum delivery price, delivery fee, promotional information, etc. For example, as shown in Figure 1 (b), the display area 10 of supermarket B displays promotional information such as user rating 4.9 points, monthly sales 443, delivery time 23 minutes, delivery distance from the current device 1.2km, minimum delivery price 20 yuan, delivery fee 6 yuan, and 8 yuan no threshold; the display area 20 of bakery C displays promotional information such as user rating 4.9 points, monthly sales 100, delivery time 30 minutes, delivery distance from the current device 3.5km, minimum delivery price 20 yuan, delivery fee 6 yuan, 26 minus 1, 38 minus 2, etc.
[0036] When the user clicks anywhere on the B supermarket display area 10, in response to the user's click operation, the mobile phone displays the B supermarket product display interface 103 as shown in Figure 1 (c). The interface 103 can display the merchant information of the B supermarket and display a product list for the user. The product list can categorize and display one or more products sold by the B supermarket. For example, as shown in Figure 1 (c), the interface 103 displays different product categories such as "meat, eggs and poultry", "home appliances and digital products", "snacks", "prepared food and frozen products", "wine and beverages", "aquatic products and seafood", and "imported food". When the user clicks on the "snacks" category, the product display area of the interface 103 can display one or more snack food products such as A potato chips, B melon seeds, and C ham sausage. The user can also view more products by sliding up and down, etc. For simplicity, it is not repeated here.
[0037] When the user clicks anywhere in the product display area 30 of potato chips A, in response to the user's click operation, the mobile phone displays the product purchase interface 104 corresponding to potato chips A as shown in Figure (d) in Figure 1. The interface 104 can display product information related to potato chips A, such as the taste characteristics of potato chips A, monthly sales, price information, red envelope discounts, promotional information, delivery information, as well as one or more contents such as product details, product reviews and related recommended products. This embodiment of the present application is not limited to this.
[0038] In the above-mentioned process of browsing products through application A, the product display interface 103 of supermarket B of application A displays one or more product contents of supermarket B, but the product label displayed in the product display area 30 of potato chips A is single, for example, only the product flavor is cucumber flavor, monthly sales volume, etc., which fails to highlight the characteristics and advantages of the product, fails to let consumers see the highlights and advantages of the product at a glance, and fails to attract customers to buy.
[0039] In addition, the product information of the potato chips A displayed on the product purchase interface 104 generally relies only on the product title to generate selling points. The product-related information that the model can obtain in this way is very limited, resulting in a single product label displayed to the user. For example, it only shows that the product flavor is cucumber flavor, monthly sales volume, etc., which fails to highlight the characteristics and advantages of the product and lacks purchasing appeal.
[0040] In addition, the same product may be sold by different merchants or platforms. The current label generation method mainly relies on the product title or product description entered by different merchants. Therefore, different product labels may be generated for the same product, and batch mounting of the same product cannot be achieved.
[0041] Therefore, in response to the above problems, this application will provide a method for generating product labels, which can realize batch mounting of product labels, and at the same time generate product labels for each product that can highlight the advantages and characteristics of the product, increase the attractiveness of the product to consumers, and improve the user's shopping experience.
[0042] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0043] To facilitate understanding of the embodiments of the present application, the following explains the professional terms involved in the embodiments of the present application:
[0044] 1. Large Language Model (LLM)
[0045] LLM is a natural language processing model based on deep learning. Its goal is to generate text that is similar to human language. Trained with large amounts of text, LLM can generate language prompts or understand the meaning of text. It can handle a variety of natural language tasks, including conversational question-answering, information extraction, and text classification, and has demonstrated great potential in many of these tasks.
[0046] Large language models (LLMs) typically consist of billions of parameters and are trained using large amounts of training data. The core idea behind large language models is to automatically capture the semantic and grammatical relationships between words, phrases, sentences, and paragraphs by learning from large amounts of language data. These models can generate coherent text from a given context and generate new sentences with a certain degree of logic and rationality. They can be applied to natural language processing tasks such as automatic text generation, machine translation, question-answering systems, and text summarization.
[0047] 2. Prompt Model
[0048] Refers to the input text paragraph or phrase, which is added before the task text to be solved and passed to the LLM to achieve the expected task. It has the meaning of instructions and prompts, usually in the form of questions, dialogues, descriptions, etc. The prompt input enables the LLM to adapt to various downstream applications.
[0049] A prompt typically provides input and guidance to an artificial intelligence (AI) model or system to produce specific output text. It can be a brief instruction, question, hint, or description that guides the AI model to generate a corresponding answer, article, code, or other textual output. Prompts are widely used in natural language processing and generative modeling. By providing a specific prompt, AI models can be guided to generate text more accurately and specifically. For example, in an automated question-answering system, a user's question can serve as a prompt to guide the system in generating the appropriate answer. In text generation tasks, a starting prompt or context can be provided, and the model will generate coherent text based on this prompt. The design and selection of the prompt has a significant impact on the generated output. A clear, specific, and reasonable prompt can help the model understand the task requirements and generate accurate and expected text. Furthermore, the flexibility of the prompt allows users to adjust and optimize the prompt to achieve better output results.
[0050] 3. ChatGPT
[0051] ChatGPT is a conversational generative model developed by OpenAI. It's built on the Generative Pre-trained Transformer (GPT) architecture and boasts powerful natural language processing and generation capabilities. ChatGPT focuses on generating conversational text, simulating natural language conversations and interacting with users in a human-like manner. ChatGPT learns language patterns and semantic relationships through pre-training on large-scale text data, and can generate coherent responses based on context. It can be applied to a variety of conversational tasks, such as question answering, text summarization, and guidance prompts. Users can interact with ChatGPT by providing an initial conversation text or question and receiving a system-generated response.
[0052] 4. ChatGLM
[0053] ChatGLM is an AI-based language model developed by Tsinghua University's KEG Lab and Zhipu AI. By learning from large amounts of text data, it can identify users' emotions and needs, providing helpful advice and guidance. ChatGLM can communicate with users in real time, answering their questions and requests, and even conducting voice conversations. It can help users solve a variety of problems and needs, such as emotional counseling, health consultation, education consultation, and financial consultation.
[0054] 5. FABE system
[0055] The FABE principle is commonly used in product marketing copywriting. The FABE principle stands for Features (F), Advantages (A), Benefits (B), and Evidence (E). It's a traditional product sales model that attracts consumers by describing the product's features, advantages, and benefits, and fosters trust through emotion and evidence.
[0056] 6.CPV
[0057] The CPV system is a method used in the e-commerce industry to characterize and manage products. Specifically, it consists of categories (C), properties (P), and values (V). In the e-commerce industry, some channels are often divided into front-end and back-end systems. Back-end categories are more professional and stable, used for back-end product management such as product releases, and are suitable for sellers and groups. Front-end categories are more flexible and adaptable, meeting the needs of marketing operations and allowing for rapid product definition and display, and are suitable for users.
[0058] 7. Optical Character Recognition (OCR)
[0059] OCR is a technology that converts text in images into editable and searchable text. It can recognize various forms of text, including printed text, handwritten text, and text in printed materials, and convert them into a computer-processable text format. OCR technology has a wide range of applications, such as document scanning, automated data entry, license plate recognition, ID card recognition, and bill recognition. Its emergence has greatly improved text processing efficiency, reduced manual data entry errors, and facilitated a variety of text-related applications.
[0060] Figure 2 is a schematic flow chart of a method for generating a product label provided in an embodiment of the present application. It should be understood that the method 200 can be applied to electronic devices with data processing capabilities, or cloud servers, etc. The embodiment of the present application does not limit the device form of the execution subject. The method 200 includes the following steps:
[0061] S201, obtaining product description information related to the target product.
[0062] It should be understood that before S201, you can first circle the target product that needs to be mounted with a product label. After circled, you can generate the product label of the target product in different ways, and finally mount the target product according to the product label.
[0063] In one possible implementation, the product selection process can determine the target product by item granularity, which can be called "item selection." For example, the "item selection" method can select a certain product in a store.
[0064] In another possible implementation, the product selection process can identify the target product by barcode granularity, which can be called "barcode selection." For example, the item ID of the same product in different stores may be different. Therefore, item selection may only select product 1 from Supermarket A, but may miss product 1 from Supermarket B. "Barcode selection" can select all products of the same type from a product perspective. In other words, you can select product 1 from Supermarket A, as well as product 1 from Supermarket B, product 1 from Supermarket C, and so on.
[0065] The above-mentioned method of circling the target product selects the same product in different stores. After the product label is generated according to the method provided in the embodiment of the present application, batch mounting can be achieved for the same product, that is, the generated label can be applied to products 1 in different stores and different platforms, thereby improving the coverage rate of the product label.
[0066] It should also be understood that in the embodiment of the present application, "product description information" may include information associated with the target product from one or more different sources.
[0067] Optionally, the "product description information" may be information derived from the product title of the target product. For example, when the same product is sold on different platforms or in different stores, the merchant may enter different product titles. In other words, the product description information may be derived from different product titles.
[0068] Alternatively, the "product description information" may also be derived from the promotional content of the target product's graphic details interface. For example, the graphic details interface of a Taobao store may include pictures of the target product, promotional selling points, and other information.
[0069] Alternatively, the "product description information" may also be information obtained from product detail pages obtained from other platforms, mini-programs, etc. For example, label information of the target product on some other mini-programs may be obtained.
[0070] Alternatively, the "product description information" may also come from group synchronization information, for example, the product data source of a preset database of a certain product sales group.
[0071] S202: Structural processing is performed on the product description information to obtain objective attribute characteristics of the target product.
[0072] S203 : Using the fine-tuned large language model to process the objective attribute features based on the language prompt vocabulary, obtain one or more product labels for the target product.
[0073] The FABE principle is often used when creating product marketing copy. The FABE principle, which stands for Features, Advantages, Benefits, and Evidence, is a traditional product sales model that attracts consumers by describing the product's features, advantages, and benefits, and builds trust through emotion and evidence.
[0074] Product Features lists product features and functions, such as size, color, material, and brand. Advantages describes the product's advantages, namely, the benefits of its features to consumers, such as performance, durability, and ease of use. Benefits describes the product's benefits, namely, the actual benefits to consumers, such as improved production efficiency, time savings, and cost savings. Evidence provides product verification, such as consumer reviews, certifications from relevant organizations, and use cases, to enhance consumer trust.
[0075] In the embodiment of the present application, the AB class labels introduced above are referred to as subjective labels; in contrast, the FE class introduced above can be referred to as objective labels.
[0076] In S202 , the obtained product description information from different sources is structured to obtain objective attribute characteristics of the target product, such as the size, color, material, brand, etc. of the target product.
[0077] With the emergence of ChatGPT, LLM can handle a variety of natural language tasks, such as conversational question-answering, information extraction, and text classification. The process of generating selling point labels is essentially a natural language generation (NLG) task. Through the process of S203, LLM's capabilities can be used to process product description information from multiple different sources, thereby generating unique and specific product labels that reflect the characteristics and advantages of the target product and enhance the confidence and richness of the labels.
[0078] FIG3 is a flowchart of an example of the implementation process of generating product labels provided by the embodiment of the present application. As shown in FIG3, the embodiment of the present application divides the process of generating product labels described above into four different stages, including:
[0079] Phase 1: The process of product selection.
[0080] The first stage can refer to the implementation process of the aforementioned S201, and for the sake of simplicity, it will not be repeated here.
[0081] Phase 2: The process of generating product labels.
[0082] In the second stage, different sources of product description information can be used to generate different product labels.
[0083] For example, when the "product description information" comes from the product title of the target product, corresponding to the first implementation path in Figure 3, the product title of the target product is structured to obtain the objective attribute characteristics of the target product, and then the fine-tuned large language model is used to summarize the corresponding product labels or promotional selling points.
[0084] For example, when the "product description information" comes from the promotional content of the graphic and text details interface of the target product, corresponding to the second implementation path in Figure 3, the graphic and text details corresponding to the product are extracted, and OCR recognition is performed to obtain the product description information. Then, the fine-tuned large language model is used to summarize the corresponding product labels or promotional selling points.
[0085] Exemplarily, when the "product description information" is information from a product details page obtained from other platforms, mini-programs, etc., corresponding to the third implementation path in Figure 3, the first data source is obtained, and the first data source and the target product are aligned, that is, the product information contained in the obtained first data source is aligned to the target product in the product library, which is a "same product normalization" process; after obtaining the product description information of the target product from the first data source, the fine-tuned large language model is used to summarize the corresponding product labels or promotional selling points.
[0086] Exemplarily, the "product description information" can also be derived from information different from the first data source, for example, the "product description information" can also be derived from information from a second data source, corresponding to the fourth implementation path in Figure 3. For example, the second data source can be product information from a product database of a certain product sales group, also known as "group synchronization information." It is necessary to determine the relevance between the product information in the product database and the target product, for example, by correlating them based on barcode product granularity. After aligning the data source to the product database to achieve "same product normalization," the product description information of the target product is obtained from the product information in the product database (group synchronization information), and then the fine-tuned large language model is used to summarize the corresponding product label or promotional selling point.
[0087] It should be understood that in the embodiments of the present application, “first data source” and “second data source” are used to represent commodity information from different sources.
[0088] For example, when the "product description information" can also be derived from preset rules set by business personnel, it corresponds to the fifth implementation path in Figure 3. For example, the business personnel set association patterns for the target product in different scenarios, and can perform pattern matching on the target product, matching different categories under different patterns, and thus determining the corresponding product label or promotional selling point.
[0089] The third stage: post-processing process.
[0090] The third stage can process the product labels generated in the second stage, such as performing blacklist filtering, so as to improve the accuracy and confidence of the generated product labels.
[0091] In one possible implementation, after obtaining one or more product tags of the target product, the post-processing process may include one or more of the following:
[0092] Filtering the one or more product tags, removing product tags containing risky words from the one or more product tags, and mounting the target product based on the filtered product tags;
[0093] Alternatively, calculating the similarity between any two of the one or more product tags, classifying the two product tags whose similarity is greater than or equal to a preset threshold, and listing the target product based on the classified product tags;
[0094] Alternatively, a correlation between each of the one or more product tags and the objective attribute characteristics of the target product is calculated, the correlations are ranked from highest to lowest as a priority order for each product tag, and product tags that meet a preset priority requirement are attached to the target product according to the priority order;
[0095] Alternatively, according to a preset rule, a scene keyword associated with each of the one or more product tags is determined, and a product tag is selected to mount the target product according to the scene keyword.
[0096] The fourth stage: the product mounting process.
[0097] The fourth stage can be understood as the process of attaching the target product based on the product label obtained in the third stage, that is, displaying product labels that can highlight the characteristics and advantages of the product and attract users.
[0098] The following details the different implementation processes of generating product tags based on different sources of product description information.
[0099] FIG4 is a flowchart of another example of the implementation process of generating a product label provided by an embodiment of the present application. When the "product description information" is derived from the product title of the target product, as shown in FIG4 , the specific implementation process of the four stages of the process of generating a product label described above may include:
[0100] Phase 1: The process of product selection.
[0101] Phase 2: The process of generating product labels.
[0102] When the "product description information" originates from the target product's title, after obtaining the target product's title, structured processing can be performed to obtain the target product's objective attribute characteristics. Specifically, the obtained product title can be a natural language text, which may have many descriptions, and different merchants may have different ways of writing it. "Structural processing" involves extracting the required information from a natural language text, such as brand information, entity information, color information, flavor information, etc., and standardizing it into a structured representation.
[0103] After obtaining the objective attributes of the product, we can combine them into different combinations. This process can be performed using LLM, specifically including: 1) generating more attractive expressions based on the objective attributes; 2) converting them into deeper user benefits based on the specific product entity.
[0104] For example, descriptions of different colors may be normalized to the same value.
[0105] For example, the taste of a product A is described as bland. Through structural processing, the taste description can be transformed into a more attractive expression and into a point of interest that users care more about.
[0106] For example, Product A is a 200g spicy rabbit head. After structured processing, objective attributes extracted from this product may include the flavor: Spicy Rabbit Head 200g, resulting in a product label of "Fragrant and Spicy." Product B is clothing. Objective attributes extracted from this product may include the material: pure cotton, resulting in a product label of "Soft and Skin-Friendly and / or Anti-static." Product C is fruit. Origin information extracted from this product may include "Fruit Origin Characteristics," etc. Examples are not provided here.
[0107] The above-mentioned method of generating product tags has a clear advantage: the product's selling point tags and objective tags can be directly linked; and the objective tags are linked to the item-level products. Therefore, the selling point tags can be white-boxed and batch-mounted on the product side for downstream delivery applications.
[0108] However, ChatGLM test results show that its output does not fully cover the instructions and cannot be used directly, requiring post-processing. Regarding label quality, manual review is required during the cold start phase to accumulate high-quality samples. This not only ensures the correctness of batch mounting, but also serves as training samples for subsequent model fine-tuning.
[0109] In Figure 4, "attribute combinations" can be understood as different combinations of objective attribute features after structured processing. For example, in the example of "Material: Pure Cotton - Soft and Skin-Friendly / No Static Electricity," a pure cotton product can be either clothing or a face towel, while the selling point of "No Static Electricity" is clearly not applicable to face towels. Therefore, for certain attributes, the selling point labels can be divided into general labels and entity / category-specific labels, which requires the production of unique labels through attribute combinations.
[0110] In order to improve the accuracy of product labels, in this embodiment of the application, objective labels can be roughly divided into three categories:
[0111] Category: Level 1, Level 2, Level 3
[0112] Entity: Structured understanding of core entities
[0113] Attributes: brand / flavor / color / material…
[0114] Attribute combinations include: category, entity, category + attribute, category + entity, entity + attribute, and so on. Based on the different attribute combinations, a fine-tuned large language model is used to process the objective attribute features based on a language prompt lexicon. Specifically, prompts from different language prompt lexicons are selected to obtain one or more product labels for the target product, thereby ensuring that LLM generates more targeted selling point labels. Table 1 shows some examples of white-box mounting rules and their corresponding labels.
[0115] Table 1
[0116] The third stage: post-processing process.
[0117] After one or more product tags are generated, products can be batch-mounted based on them. However, the one or more product tags generated in the first and second stages may contain tags with similar or similar meanings, or tags that are relatively general. To improve the accuracy of the product tags, the one or more tags generated in the third stage can be post-processed.
[0118] Optionally, the operations of the post-processing process may include one or more of the following processes: word count control of product labels, subtitle extraction, blacklist filtering, pointwise mutual information (PMI) sorting, and label item classification filtering.
[0119] It should be understood that the post-processing process will add some statistical information to the generated product labels. "Blacklist filtering" can be understood as removing risky banned words from product labels, such as "lowest price ever" and "health benefits."
[0120] "Tag item classification filtering" can divide product tags into "tag items" and "tag values". "Tag items" are generally called "selling point tags", which are further divided into those describing materials, describing trial scenarios, and describing taste. The selling points can be subdivided and classified in the post-processing stage. During classification, "classification items" or irrelevant tags are added. In this way, tags generated by LLM that do not conform to the facts and are irrelevant to the current product can be filtered out through "tag item classification filtering". This involves the "classification model", and the classification model is used to filter product tags. It will not be repeated here.
[0121] For example, the product labels output by the black-box model may contain risky or banned words, which need to be filtered out during post-processing to avoid risks at the product display level. Alternatively, the LLM may only output product labels with four characters, requiring post-processing to control the number of characters. Alternatively, the post-processing process may rank and score the product labels output by the LLM based on some algorithms, and the ranking and reference scores can be used for downstream display.
[0122] Optionally, when mounting product labels in batches, the following questions are also required:
[0123] (1) Label normalization or deduplication
[0124] When native ChatGLM is generated, semantically similar tags appear under the same combination rule. However, when launching the application, we hope to display them from different dimensions to improve the efficiency of the placement. Therefore, similar tags need to be normalized and deduplicated.
[0125] To address the above problems, we can use the graph structure to build edges based on the jaccard score and normalize the labels of connected subgraphs.
[0126] (2) Selection of targeted and generalized labels
[0127] Some tags have broad applicability, such as "delicious," which can be applied to any food-related category; while others are very targeted, such as "rich in anthocyanins" and "promotes digestion." Obviously, more specific tags should be prioritized.
[0128] To address the above issues, we can use Pointwise Mutual Information (PMI) to calculate the degree of match between labels and categories, entities, and attribute dimensions, and arrange them in descending order to obtain priorities. In this way, during the label mounting process, we can select targeted labels or generalized labels based on the priority.
[0129] Through the label processing in the above post-processing process, the quality of LLM generated labels can be controlled, thereby improving the accuracy of LLM label generation.
[0130] The fourth stage: the product mounting process.
[0131] In one possible implementation, generated product labels can be reviewed and graded during the product placement process. Optionally, this review and grading process can be performed manually, with scores assigned, for example by sales personnel, to assess the quality and accuracy of the product labels. The review and evaluation results are then fed back to the algorithm to assist with subsequent iterations.
[0132] Optionally, you can also let LLM automatically evaluate product labels based on manual labeling scores, automatically score product labels, and generate another PMI score, which can be used as a reference score when displaying to downstream.
[0133] For example, during the review and grading process, one or more product tags can be marked as strongly relevant tags, weakly relevant tags, or irrelevant tags based on the correlation between the products and the tags, which can also be used as a reference when displaying downstream.
[0134] "Rule storage" can be understood as a batch mounting process, which clarifies the rules for generating objective label escape combinations. For example, among all the circled products, after structured understanding of "category + flavor" = "snacks + cucumber flavor", all products with the same product label can be attached.
[0135] When the product description information is derived from the target product's title, LLM can leverage the capabilities of the aforementioned processing stages to generate unique and specific product tags that reflect the product's characteristics and advantages, enhancing both the confidence and richness of the tags. Furthermore, this approach enables batch tagging, improving product tag coverage.
[0136] In another possible implementation, the product description information includes text information obtained from a detail picture associated with the target product, and obtaining the product description information related to the target product includes: obtaining the detail picture associated with the target product; performing character recognition on the detail picture to obtain one or more text fragments in the detail picture; and splicing the one or more text fragments to obtain the text information.
[0137] In other words, when the "product description information" comes from the graphic and text details of the target product, the corresponding promotional selling points can be summarized through the product's corresponding graphic and text details page. The advantages of the product labels generated by this method are traceability and high confidence.
[0138] FIG5 is a flowchart of another example of the implementation process of generating product labels provided by an embodiment of the present application. As shown in FIG5, the four stages of the process of generating product labels described above include:
[0139] Phase 1: The process of product selection.
[0140] Phase 2: The process of generating product labels.
[0141] When the product description information is derived from the target product's image and text details, the system first obtains the details image and performs character recognition on it to obtain one or more text segments within the details image. These one or more text segments are then concatenated to obtain the text information. The text information is then processed using a fine-tuned LLM to obtain one or more product tags for the target product.
[0142] It should be understood that this solution relies on pre-processed OCR recognition. For example, when you click on a product and enter its product details page, you'll see a long description or a picture-based introduction. The picture-based introduction may include "selling points" text, which can be extracted to generate a more reliable label.
[0143] OCR only recognizes text and can extract text information from images. ORC extracts text from different areas, such as Area 1 and Area 2. This text can be in artistic fonts or arranged horizontally or vertically, but this is not limited in this embodiment of the application. "Text splicing processing" can splice the text extracted from different areas to obtain complete text information.
[0144] Optionally, after receiving the recognized text, multiple text paragraphs from the same image need to be spliced together. A single product may correspond to multiple images, and the text from different images also needs to be spliced together. Complex text splicing strategies can involve factors such as font size, background color, and proximity. Simpler ones can use length filters (4-16) and splice in a left-right or top-down order. Use spaces between paragraphs in a single image, and line breaks between multiple images.
[0145] After concatenation, we obtain the complete text information, from which we need to extract the selling point tags. Text summarization is a task that ChatGLM excels at, so we consider using a large model.
[0146] In one possible implementation process, the method further includes: obtaining an original large language model used to generate product labels; using the text information and the one or more product labels as training samples, adjusting the original large language model to obtain the fine-tuned large language model.
[0147] Specifically, as shown in the dotted-line steps in Figure 5, by manually analyzing the images and pairing the selling point labels generated by business engineers with the model's OCR results, we can generate a batch of over 2,000 training samples. This constitutes the training sample, which allows for LLM fine-tuning. Business engineers can perform manual OCR over an extended period, continuously accumulating samples and iteratively optimizing the model.
[0148] Through the above method, the labels obtained from the real pictures in the product's graphic details interface have higher confidence and reliability. Similarly, by using these samples to iteratively train the LLM, the LLM can also be optimized to improve the accuracy of the LLM-generated labels.
[0149] Considering the low image and text detail ratio in the current product library, the same product may have images in store A but not in store B. In this case, the product normalization capability can be used to improve the coverage and utilization of labels.
[0150] The third stage: post-processing process.
[0151] The fourth stage: the product mounting process.
[0152] The processes of the third and fourth stages can be referred to the introduction of FIG. 3 and FIG. 4 , and for the sake of brevity, they will not be described again here.
[0153] In another possible implementation, the product description information includes all tag information of the target product when it is sold on other shopping platforms, and obtaining the product description information related to the target product includes: obtaining all tag information of the target product when it is sold on other shopping platforms; and the method also includes: using the fine-tuned large language model to analyze the tag information, determine a first language prompt word, and update the language prompt word library in the fine-tuned large language model according to the first language prompt word.
[0154] In other words, when the "product description information" can be derived from information of the first data source, such as obtaining tag information mounted under the store detail page or product detail page of other platforms, ready-made selling point tags can be obtained.
[0155] FIG6 is a flowchart of another example of the implementation process of generating product labels provided by an embodiment of the present application. As shown in FIG6, the four stages of the process of generating product labels described above include:
[0156] Phase 1: The process of product selection.
[0157] The second stage: the process of generating product labels, which specifically includes structured understanding and normalization of similar products.
[0158] In this process, information from the first data source is obtained and preprocessed to obtain tag information mounted on store detail pages and product detail pages of other platforms. Classification filtering of tag items is then performed, and the filtered tag information is structured and understood.
[0159] Optionally, there are two key steps to utilize the information from the first data source:
[0160] 1) Label cleaning, which is the label item classification and filtering process shown in Figure 6. Information from the first data source is obtained, and ready-made selling point labels are obtained. Some short labels can be used directly, while some are irrelevant descriptive labels and need to be removed. Optionally, the label cleaning process can label samples based on a classification model (such as the BERT classification model), classify the samples, and then perform structured processing based on the large model.
[0161] 2) Label mounting: This requires product alignment, which involves aligning the product information from the first data source to the target product database, and then performing label alignment and mounting.
[0162] Label mounting can include mounting at the product level, that is, product alignment. For example, aligning acquired products to the Ele.me product library requires the use of the same-product normalization capability.
[0163] Optionally, product normalization currently supports CPV-level normalization. This involves normalizing products with the same category, entity, and attributes based on the structured understanding of the products (objective product information, category, entity, and attributes). Products in different categories may have different key attributes. If a certain attribute is considered important for distinguishing between products, it must be identified during the structured understanding; otherwise, it will be ignored. Therefore, CPV-level product normalization relies on the identification of F-type attribute tags, and the determination of key attributes requires the assistance of business personnel.
[0164] Alternatively, an upgraded solution for product alignment can be used. This solution leverages the large model's text and vector search capabilities to establish an online link, eliminating the need for relying solely on structured processing results. To improve product alignment, LLM can be used to enhance product understanding. This process involves product-level alignment, mapping products from the primary data source to the target products. It does not involve batch loading at the product level.
[0165] Alternatively, tagging can be achieved at the objective tag level. After a structured understanding of the products in the first data source, LLM can be used to discover the relationship between subjective and objective tags. Specifically, LLM can be used to discover tags and determine how these tags are derived based on attribute combinations. After discovering patterns, the rules are extracted and used directly as tagging rules for product tags. For example, in a 4bc product listing, the product title is "Freego Disposable Towels 75*35cm (2-pack)" and the selling point tag is "Natural Plant Fiber | Skin-friendly and Comfortable." It can be found that this selling point is derived from the product's "category" and "material." This selling point can then be expanded to other products with the same "category" and "material."
[0166] It's important to understand that objective labels don't necessarily appear in product titles. For example, "Freego Disposable Towels 75*35cm (Pack of 2)"—(Material) Natural Plant Fiber / Skin-Friendly and Comfortable." Furthermore, some labels are unique to specific products, raising questions about whether the model can understand brand characteristics. For example, "Ghetto Pulled Cheese (Original) 25g*2"—6x Milk Protein / Low Carb. These special cases can be tested and fine-tuned within the LLM, enhancing its structured output capabilities.
[0167] The third stage: post-processing process.
[0168] The fourth stage: the product mounting process.
[0169] The processes of the third and fourth stages can be referred to the introduction of FIG. 3 and FIG. 4 , and for the sake of brevity, they will not be described again here.
[0170] In another possible implementation, the product description information may also include information from a preset database, which includes information on one or more products from multiple different shopping platforms. The obtaining of product description information related to the target product includes: obtaining the barcode of the target product; and searching the preset database for product information associated with the product whose barcode is consistent with the target product based on the barcode of the target product.
[0171] It should be understood that the "preset database information" here can be the information of the product library, or it can be called group synchronization information. For example, the data source (second data source) and data return information of the product library of a certain product sales group. A certain product sales group has many platforms, such as Hema, Ele.me, Taobao, etc. The labels obtained from different retail channels can be used on products on other platforms. Then, in the process of matching the information in the product library with the target product information, mapping can be performed through barcodes.
[0172] For example, the barcode of the target product is obtained. The barcode provided by the group matches the barcode of the current product directly, and the corresponding relationship between the label and the barcode can be obtained. Then, the label is mounted according to the barcode. However, if the group information is incorrect, there may be irrelevant issues. The label and the product may not match. For example, the barcode may be incorrectly typed or the label may be incorrectly provided, resulting in the label and product not matching.
[0173] In order to solve the above problems, relevance filtering can be performed to remove irrelevant tags.
[0174] Optionally, after searching the preset database for product information associated with a product having the same barcode as the target product, the method further includes: calculating the correlation between the product information and the objective attribute characteristics of the target product, and retaining the product information whose correlation meets the preset requirements.
[0175] Optionally, after obtaining the barcode of the target product, the method further includes: classifying the target product according to a classification model to determine the category to which the target product belongs; and searching the preset database for product information associated with the product that is consistent with the barcode of the target product based on the barcode of the target product, including: searching the preset database for product information associated with the product that is consistent with the barcode of the target product and belongs to the same category as the target product based on the barcode of the target product and the category to which the target product belongs.
[0176] Optionally, finding the product information associated with the product whose barcode is the same as that of the target product from the preset database according to the barcode of the target product includes: according to the barcode of the target product, using the text retrieval and / or vector retrieval function of the fine-tuned large language model to find the product information associated with the product whose barcode is the same as that of the target product from the preset database.
[0177] FIG. 7 is a flowchart of the implementation process of another example of generating a product label provided by an embodiment of the present application. As shown in FIG. 7, after preprocessing the obtained group information, perform label item classification and filtering processing, perform relevance filtering on the obtained product labels, and then map the information in the product library and the target product information through a barcode.
[0178] Specifically, perform association according to the barcode standard product granularity. The label needs to be additionally cleaned and filtered, and a classification model can be used. It should be noted that this process requires manual discovery and manual intervention. For example, for spelling mistakes, such as "kouwei soft" being adjusted to "taste soft". Another example is the barcode mis-association problem, the problem of normalizing the same product between the barcode and the title, etc. Although the barcode inherently has the ability to integrate the same product normalization, considering the accuracy of the association, the self-built same product normalization ability can still be used for extended coverage.
[0179] It should also be noted that there is a correlation problem between the barcode and the label. After statistically analyzing the distribution relationship between the category and the label, it is found that there are irrelevant label descriptions under the category. For example, there are meat-related descriptions under the wine category. This part needs to be solved by a model.
[0180] In addition, it is also found that many of the group-synchronized labels are the product highlights of a certain brand's product for external publicity, which belong to professional knowledge and are not directly reflected in the title. For example, for "Xiaolangjiu (Classic Sharing Pack) 45 degrees, 100 ml * 6 bottles / box", one of the selling point labels is "pure grain solid-state fermentation". On the one hand, this increases the difficulty of judging the correlation between the barcode and the label. On the other hand, it also inspires us to extend such labels to all items with the same category + product name / brand category (that is, the relationship discovery between the subjective label and the objective label in 4bc). In the attempt, compared with general category words ("flavored milk"), the confidence of direct generalization and coverage of the labels under the product name ("Mengniu") is higher.
[0181] Similarly, the same relationship discovery and extension can be performed on the standard product part in the graphic and text details introduced in FIG. 5. For the sake of simplicity, it will not be elaborated here.
[0182] In another possible implementation, after obtaining one or more product tags of the target product, the method further includes: determining the scene keywords associated with each of the one or more product tags according to preset rules, and selecting the product tag to mount the target product according to the scene keywords.
[0183] Figure 8 is a flowchart of another example of the process for generating product tags, as provided in an embodiment of the present application. As shown in Figure 8, after selecting a product, an automated process can be designed based on business rules for automated pattern matching. The rules are presented in Table 2 below.
[0184] Table 2
[0185] In the rule matching process, the "Trial Category" in Table 2 is used as a constraint. If the title contains keywords, that is, if the keywords are hit, a label value is assigned. Optionally, the rules can be business rules given by business personnel, or rules given by operations personnel based on industry experience and operational experience. This embodiment of the application does not limit this.
[0186] Through the above method, this application mainly generates product selling point labels from product description information obtained from multiple sources and multiple channels. Conventional methods rely solely on product titles to generate selling point labels, resulting in very limited product-related information. The method provided by this application expands the sources of product description information, can identify graphic details, and obtain data information from different sources, etc., which can help the model better understand the product and generate more reliable, higher-confidence, and traceable selling point labels.
[0187] Specifically, the embodiment of the present application combines the current trend of big models and fully utilizes the advantages of the big model's capabilities to generate labels for multi-source data. Specifically, in the extraction of graphic and text details, the embodiment of the present application fully utilizes the overview capability of the big model to extract concise, informative selling point labels from a large amount of OCR text; in the objective label escape, it fully utilizes the semantic understanding capability of the big model to rewrite the description of related entities and attributes to generate more attractive selling point labels; in the first data source (4bc source) part, it utilizes the deeper understanding and trial capabilities of the big model to discover the relationship between objective labels and subjective labels, thereby improving label utilization and product coverage. At the same time, thanks to the diversity of the big model's generation capabilities, the richness of labels has also been rapidly accumulated.
[0188] Furthermore, this application fully utilizes structured understanding and product normalization capabilities to achieve efficient batch labeling of product tags and improve label coverage. Product normalization is useful across multiple sources and, in practice, can increase label coverage by approximately 10%.
[0189] FIG9 is a schematic diagram of an apparatus for generating product labels based on a large language model according to an embodiment of the present application. The apparatus 900 for generating product labels includes an acquisition unit 910 , a first processing unit 920 , and a second processing unit 930 .
[0190] An acquisition unit 910 is configured to acquire product description information related to a target product;
[0191] A first processing unit 920 is configured to perform structured processing on the product description information to obtain objective attribute characteristics of the target product;
[0192] The second processing unit 930 is configured to process the objective attribute features using the fine-tuned large language model based on the language prompt vocabulary to obtain one or more product labels for the target product.
[0193] 10 is a schematic diagram of the structure of a computer device 1000 provided in an embodiment of the present application. Optionally, the computer device 1000 may be a device with computing functions, or a server, which is not limited in the embodiment of the present application.
[0194] Exemplarily, as shown in FIG10 , the computer device 1000 includes: a memory 1001, a processor 1002, and a computer program 1003 stored in the memory 1001 and running on the processor 1002, wherein when the processor 1002 executes the computer program 1003, the computer device can execute any of the aforementioned methods for generating product labels based on a large language model.
[0195] The implementation of each module in the server or computer device provided in the embodiments of the present application may be in the form of a computer program. The computer program can be run on the server or computer device. The program modules comprising the computer program can be stored in the memory of the server or computer device. When the computer program is executed by a processor, all or part of the steps of the method described in the embodiments of the present application are implemented.
[0196] It should also be understood that in the various embodiments of the present application, the size of the serial number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0197] It is understandable that, in order to implement the above functions, the computer device includes hardware and / or software modules that perform the corresponding functions. In combination with the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to be beyond the scope of this application.
[0198] In this embodiment, the computer device can be divided into functional modules according to the above-mentioned method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into a single processing module. The above-mentioned integrated modules can be implemented in the form of hardware. It should be noted that the module division in this embodiment is illustrative and only represents a logical functional division. In actual implementation, other division methods may be used.
[0199] In the case of dividing the functional modules according to their respective functions, a possible schematic diagram of the computer device involved in the above embodiments is shown, in which the computer device may include: a display unit, a detection unit, and a processing unit. The display unit, the detection unit, and the processing unit cooperate with each other to support the computer device in executing the above steps, etc., and / or other processes of the technology described herein.
[0200] It should be noted that all relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module and will not be repeated here.
[0201] It should be noted that the information and data involved in this application (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0202] The computer device provided in this embodiment is used to execute the above-mentioned method of generating product labels based on a large language model, and thus can achieve the same effect as the above-mentioned implementation method.
[0203] When integrated, the computer device may include a processing module, a storage module, and a communication module. The processing module can be used to control and manage the computer device's operations. For example, it can be used to support the computer device in executing the steps performed by the display unit, detection unit, and processing unit. The storage module can be used to support the computer device in executing and storing program code and data. The communication module can be used to support communication between the computer device and other devices.
[0204] The processing module may be a processor or a controller. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, and so on. The storage module may be a memory. The communication module may specifically be a device that interacts with other computer devices, such as a radio frequency circuit, a Bluetooth chip, or a Wi-Fi chip.
[0205] This embodiment also provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed on a computer device, the computer device executes the above-mentioned related method steps to implement the method of generating product labels based on a large language model in the above-mentioned embodiment.
[0206] This embodiment also provides a computer program product. When the computer program product is run on a computer, it enables the computer to execute the above-mentioned related steps to implement the method for generating product labels based on a large language model in the above-mentioned embodiment.
[0207] In addition, an embodiment of the present application also provides a device, which can be a chip, component or module. The device may include a connected processor and memory; wherein the memory is used to store computer-executable instructions. When the device is running, the processor can execute the computer-executable instructions stored in the memory to enable the chip to execute the method of generating product labels based on a large language model in the above-mentioned method embodiments.
[0208] Among them, the computer device, computer-readable storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0209] Through the description of the above implementation methods, technical personnel in the relevant field can understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0210] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0211] Units described as separate components may or may not be physically separate, and components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0212] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0213] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0214] The above content is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for generating product labels based on large language models, characterized in that, Including: Obtain product description information related to the target product; Perform structured processing on the product description information to obtain the objective attribute features of the target product; Use the fine-tuned large language model to process the objective attribute features based on the language prompt library to obtain one or more product labels of the target product.
2. The method according to claim 1, wherein The product description information includes text information obtained from the detail pictures associated with the target product. The obtaining of the product description information related to the target product includes: Obtain the detail pictures associated with the target product; Perform character recognition on the detail pictures to obtain one or more text segments in the detail pictures; Perform splicing processing on the one or more text segments to obtain the text information.
3. The method according to claim 2, wherein The method further includes: Obtain the original large language model for generating product labels; Use the text information and the one or more product labels as training samples to adjust the original large language model to obtain the fine-tuned large language model.
4. The method according to claim 1, characterized in that, The product description information includes information from a preset database, and the preset database includes information on one or more products from multiple different shopping platforms. The obtaining of the product description information related to the target product includes: Obtain the barcode of the target product; According to the barcode of the target product, search in the preset database for the product information associated with the product whose barcode is the same as that of the target product.
5. The method according to claim 4, wherein After searching in the preset database for the product information associated with the product whose barcode is the same as that of the target product, the method further includes: Calculate the correlation between the product information and the objective attribute features of the target product, and retain the product information whose correlation meets the preset requirements.
6. The method according to claim 4, characterized in that After obtaining the barcode of the target product, the method further includes: Classify the target product according to a classification model to determine the category to which the target product belongs; And, the searching in the preset database for the product information associated with the product whose barcode is the same as that of the target product includes: According to the barcode of the target product and the category to which the target product belongs, search in the preset database for the product information associated with the product whose barcode is the same as that of the target product and whose category is the same as that of the target product.
7. The method according to claim 4, wherein The searching in the preset database for the product information associated with the product whose barcode is the same as that of the target product includes: According to the barcode of the target product, use the text retrieval and / or vector retrieval function of the fine-tuned large language model to search in the preset database for the product information associated with the product whose barcode is the same as that of the target product.
8. The method according to any one of claims 1 to 7, characterized in that, After obtaining one or more product labels of the target product, the method further includes: Perform filtering processing on the one or more product labels to remove the product labels containing risk vocabulary, and mount the target product according to the filtered product labels; and / or, Calculate the similarity between any two of the one or more product labels, classify the two product labels with a similarity greater than or equal to a preset threshold, and mount the target product according to the classified product labels; and / or, Calculate the correlation between each of the one or more product labels and the objective attribute features of the target product, determine the priority order of each product label in descending order of the correlation, and mount the product labels that meet the preset priority requirements on the target product according to the priority order.
9. The method according to any one of claims 1 to 7, characterized in that, After obtaining one or more product labels of the target product, the method further includes: According to a preset rule, determine the scenario keywords associated with each of the one or more product labels, and select product labels to mount the target product according to the scenario keywords.
10. An apparatus for generating product labels based on a large language model, characterized in that, Includes: An acquisition unit for acquiring product description information related to the target product; A first processing unit for performing structured processing on the product description information to obtain the objective attribute features of the target product; A second processing unit for processing the objective attribute features by using a fine-tuned large language model based on a language prompt library to obtain one or more product labels of the target product.
11. A server, characterized in that, Includes: A memory for storing executable program code; A processor for calling and running the executable program code from the memory, so that the server executes the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions run on an electronic device, the electronic device executes the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Commodity and tag automatic correlation method, system and equipment and storage medium
CN108416403A
Recommendation tag generation method, recommendation tag display method, corresponding device and electronic equipment
CN114579896A
Data processing method, text display method, data processing system and equipment
CN115828862A
Text analysis method and device for commodity information
CN116644735A
Advertisement copywriting generation method and device, equipment and medium
CN116797280A
Cited By
Commodity selling point knowledge graph construction method and electronic equipment
CN120671797A
Multi-source data-based selling point generation method for AI intelligent marketing
CN120875963A
Question and answer interaction method, device and equipment based on large model and storage medium
CN121301547A
Mass data automatic label generation method based on large model and rule engine
CN121350628A
Commodity label generation method and system for multi-modal commodity information analysis
CN121616381A