Commodity information processing method and device, storage medium and electronic equipment

By pre-training a lightweight coding model based on LLM to perform multimodal encoding of product information on e-commerce platforms, the real-time and accuracy issues of product information encoding on e-commerce platforms are solved, and real-time and accurate product feature extraction is achieved.

CN120596885APending Publication Date: 2025-09-05ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510562477.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

How to obtain accurate product features while ensuring the real-time encoding of product information, especially in the context of processing a wide variety of product information on e-commerce platforms.

Method used

A lightweight encoding model pre-trained based on a large language model (LLM) is used to encode multimodal product information. The product information is encoded into product features through the lightweight encoding model, and the encoding features of each modality are fused using an adaptation network to ensure real-time and accurate processing.

Benefits of technology

While ensuring the real-time processing of product information, it achieves accuracy comparable to that of a large language model, thereby improving the accuracy and efficiency of product features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596885A_ABST
    Figure CN120596885A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a commodity information processing method, and the method comprises the steps: training a lightweight encoding model based on a heavy LLM in advance, inputting the multi-modal commodity information of a target commodity into an encoder of a corresponding mode in the encoding model when the commodity information is processed, and carrying out the processing of the multi-modal commodity information. According to the commodity information processing method based on the lightweight encoding model, the encoding features of the corresponding modes output by the encoders are obtained, then the commodity features of the target commodity are obtained according to the encoding features of the modes, and finally the target commodity is processed according to the commodity features. According to the invention, the coding mode is obtained based on LLM training in advance, so that the real-time performance of commodity information processing can be ensured, the coding mode is obtained based on LLM training in advance, the accuracy of the commodity features obtained by the coding model can be equivalent to that of the LLM, and the accurate commodity features can be obtained on the premise of ensuring the real-time performance of commodity information coding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to a method, apparatus, storage medium, and device for processing commodity information. Background Art

[0002] At present, more and more merchants choose to sell their products on e-commerce platforms, and the products on e-commerce platforms are becoming more and more abundant. This has brought convenience to merchants and users, but also brought huge challenges to e-commerce platforms.

[0003] With the development of AI technology, e-commerce platforms can directly use machine learning models to encode product information in various product information processing scenarios, such as recommending products to users and automatically classifying products, in order to obtain the product features corresponding to the product information.

[0004] However, due to the wide variety of products on e-commerce platforms, how to obtain accurate product features while ensuring the real-time encoding of product information has become an urgent problem to be solved. Summary of the Invention

[0005] The embodiments of this specification provide a method, device, storage medium, and electronic device for processing product information to partially solve the problems existing in the above-mentioned prior art.

[0006] The embodiments of this specification adopt the following technical solutions:

[0007] This specification provides a method for processing product information, the method comprising:

[0008] Obtain multimodal product information of the target product;

[0009] For each modality, the product information of the target product in that modality is input into the encoder corresponding to that modality in a pre-trained encoding model, and the encoding features of that modality output by the encoder corresponding to that modality are obtained; wherein the encoding model is pre-trained based on the Large Language Model (LLM);

[0010] Determining product features of the target product based on the coding features of each modality of the target product;

[0011] The target product is processed according to the product characteristics of the target product.

[0012] This specification provides a device for processing commodity information, the device comprising:

[0013] The acquisition module is used to obtain multimodal product information of the target product;

[0014] An encoding module is configured to input the product information of the target product in the modality into an encoder corresponding to the modality in a pre-trained encoding model for each modality, and obtain encoding features of the modality output by the encoder corresponding to the modality; wherein the encoding model is pre-trained based on a large language model (LLM);

[0015] a determination module, configured to determine the commodity features of the target commodity based on the coding features of each modality of the target commodity;

[0016] The processing module is used to process the target product according to the product characteristics of the target product.

[0017] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned method for processing product information.

[0018] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the above-mentioned method for processing product information is implemented.

[0019] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects:

[0020] The embodiments of this specification disclose a method for processing product information. The method pre-trains a lightweight encoding model based on a heavyweight LLM. When processing the product information, the multimodal product information of the target product is respectively input into the encoders of the corresponding modes in the encoding model to obtain the encoding features of the corresponding modes output by each encoder, and then the product features of the target product are obtained based on the encoding features of each mode. Finally, the target product is processed based on the product features. Since the above method uses a lightweight encoding model when processing product information, the real-time processing of product information can be guaranteed, and the encoding mode is pre-trained based on the LLM. Therefore, the accuracy of the product features obtained by the encoding model can be comparable to that of the LLM, thereby obtaining accurate product features while ensuring the real-time encoding of product information. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The exemplary embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings:

[0022] Figure 1 A flowchart of a method for processing commodity information provided in an embodiment of this specification;

[0023] Figure 2 Schematic diagram of the coding model structure in the embodiment of this specification;

[0024] Figure 3 Flowchart of the method for training the coding model provided in the embodiment of this specification

[0025] Figure 4 A schematic diagram of a device for processing commodity information provided in an embodiment of this specification;

[0026] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION

[0027] In actual application scenarios, there are a large number of scenarios that require real-time processing of product information. For example, for a new product, it is necessary to quickly and in real time classify the new product into an existing product category, or, for a new product category, it is necessary to quickly and in real time classify existing products into the new product category, or, for a user, when the user's demand information has been obtained through big data and other methods, it is necessary to quickly and in real time screen out products that meet the user's needs in the product recall pool and push the product to the user.

[0028] In all of the above application scenarios, it is necessary to encode the product information of a product into product features. The Large Language Model (LLM) can indeed cope with a wide variety of product information encoding scenarios and can obtain accurate product features. However, due to the excessive weight of the LLM and the excessive number of parameters, it cannot guarantee the real-time processing of product information. Based on this, the embodiments of this specification pre-train a lightweight encoding model based on the heavyweight LLM. When processing product information, the lightweight encoding model is used to encode the product information into product features, so as to obtain accurate product features while ensuring the real-time encoding of the product information.

[0029] To make the objectives, technical solutions, and advantages of this specification more clear, the following will clearly and completely describe the technical solutions of this specification in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.

[0030] The technical solutions provided by the embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0031] Figure 1 A flowchart of a method for processing product information provided in an embodiment of this specification includes the following steps:

[0032] S100: Acquire multimodal product information of a target product.

[0033] In the embodiment of this specification, the Figure 1 The device that processes product information using the method shown can be a server of an e-commerce platform, or of course other electronic devices. The following description will only take the server as an example.

[0034] Since most product information is currently multimodal, the server needs to first obtain the multimodal product information of the target product when processing the product information. The multimodal product information described in the embodiments of this specification includes, but is not limited to: product information in image mode (such as photos and pictures of the product) and product information in text mode (such as product name, product description, etc.).

[0035] It should be noted that the target products described in the embodiments of this specification may vary depending on the application scenario. For example, in a scenario where newly added products are automatically classified, the target products may be newly added products. In a scenario where product recommendations are made to users, the target products may be every product in the product recall pool. Specific application scenarios will be described below.

[0036] S102: For each modality, the product information of the modality of the target product is input into the encoding network corresponding to the modality in the pre-trained encoding model to obtain the encoding features of the modality output by the encoding network corresponding to the modality.

[0037] In the embodiments of this specification, the coding model may be pre-trained based on a heavyweight LLM using a comparative learning method to obtain a trained coding model, wherein the number of parameters of the coding model is much smaller than that of the LLM.

[0038] Because the product information in the embodiments of this specification is multimodal, the encoding model in these embodiments also includes an encoder corresponding to each modality. That is, the encoding model includes at least an image encoder corresponding to product information in image modality and a text encoder corresponding to product information in text modality. Therefore, after obtaining the multimodal product information of the target product, the server can input the product information for each modality into the encoder corresponding to that modality in the encoding model to obtain the corresponding encoding features for that modality.

[0039] S104: Determine the product features of the target product according to the coding features of each modality of the target product.

[0040] After obtaining the encoding features of the corresponding modalities output by the encoders corresponding to each modality in the encoding model, the server can fuse the encoding features of each modality of the target product and use the fused features as the product features of the target product. When fusing the encoding features of each modality of the target product, methods such as splicing and pooling can be used to fuse the encoding features of each modality. In order to further improve the accuracy of encoding product information, the encoding model in the embodiment of this specification may also include an adaptation network, such as Figure 2 shown.

[0041] Figure 2 This is a schematic diagram of the coding model structure in the embodiment of this specification. Figure 2 In the

[15] , the encoding model includes not only the image encoder and the text encoder, but also an adaptation network. After obtaining the encoding features of each modality of the target product, the server can input the encoding features of each modality into the adaptation network, so that the adaptation network can fuse the encoding features of each modality to obtain the product features of the target product.

[0042] Specifically, such as Figure 2 As shown, the adaptation network includes an image adapter, a text adapter, and a fusion adapter. If the product information of the target product includes both image product information and text product information, then in step S102, the server can obtain the image encoding features output by the image encoder and the text encoding features output by the text encoder. In step S104, since the encoding features of the target product have more than two modalities, the encoding features of each modality (i.e., image encoding features and text encoding features) can be input into the fusion adapter in the adaptation network to obtain the product features of the target product output by the fusion adapter.

[0043] If the target product's product information only has one modality, either image product information or text product information, then in step S102, the server can only obtain one of the image encoding features and the text encoding features. Therefore, in step S104, the encoding features of this single modality can be input into the adapter corresponding to that modality in the adaptation network to obtain the product features of the target product output by the adapter corresponding to that modality. That is, if only the image encoding features are available, the image encoding features are input into the image adapter to obtain the product features of the target product output by the image adapter; if only the text encoding features are available, the text encoding features are input into the text adapter to obtain the product features of the target product output by the text adapter.

[0044] S106: Process the target product according to its characteristics.

[0045] After obtaining the product features of the target product, the server can process the target product according to the product features.

[0046] Since the above method adopts a lightweight encoding model when processing product information, the real-time processing of product information can be guaranteed. Moreover, the encoding mode is obtained in advance based on LLM training. Therefore, the accuracy of the product features obtained by the encoding model can be comparable to that of LLM, thereby obtaining accurate product features while ensuring the real-time encoding of product information.

[0047] In the embodiment of this specification, a comparative learning method can be used to obtain the following LLM training based on pre-training: Figure 2 The encoding model shown is Figure 3 shown.

[0048] Figure 3 The flowchart of the method for training the coding model provided in the embodiment of this specification includes the following steps:

[0049] S300: Obtain multimodal product information of a sample product.

[0050] Since the encoding model has far fewer parameters than the LLM and is a lightweight machine learning model, when training the encoding model, the server can use any new type of product as a sample product and use the sample products to train the encoding model in real time online.

[0051] S302: Input the multimodal product information of the sample product into the LLM, and obtain the product features of the sample product output by the LLM as the annotation features.

[0052] Since the encoding model is trained by contrastive learning in the embodiment of this specification, the server can input the multimodal product information of the sample product into the LLM and use the LLM to obtain the product features of the sample product as the annotation features.

[0053] S304: For each modality, the product information of the modality of the sample product is input into the encoder corresponding to the modality in the encoding model to be trained, and the encoding features of the modality output by the encoder corresponding to the modality are obtained. Based on the encoding features of each modality of the sample product, the product features of the sample product are determined as comparison features.

[0054] While using LLM to obtain the labeled features of the sample product, the server can input the multimodal product information of the sample product into the encoding model to be trained to obtain the product features of the sample product output by the encoding model to be trained as comparison features.

[0055] It should be noted that the process in which the encoding model to be trained encodes the multimodal product information of the sample product to obtain the product features of the sample product is similar to the process in which the encoding model to be trained encodes the multimodal product information of the sample product to obtain the product features of the sample product. Figure 1Steps S102 to S104 are the same and will not be described in detail here. Figure 3 The execution order of steps S302 and S304 is shown in no particular order.

[0056] S306: Determine a loss value based on the similarity between the comparison feature and the annotation feature.

[0057] The loss value is negatively correlated with the similarity, that is, the higher the similarity between the comparison feature and the annotation feature, the smaller the loss value, and vice versa.

[0058] S308: Taking reducing the loss value as a training goal, adjust the model parameters in the coding model to be trained.

[0059] In the embodiments of this specification, Figure 2 The model parameters of the image encoder and text encoder in the coding model, as well as the image adapter, text adapter, and fusion adapter in the adaptation network, can all be used as model parameters that need to be adjusted when training the coding model. The server can use minimizing the loss value obtained in step S306 as the training objective and adjust the above model parameters until the coding model converges.

[0060] The LLM mentioned in the embodiment of this specification can be a knowledge graph-based LLM, including a knowledge graph-based retrieval-augmented generation (RAG) large model. Figure 3 If the sample product is a new type of product (i.e., a type of product that has never appeared in history), the new type of product can be added offline to the knowledge graph on which the LLM is based before training the coding model, so that the LLM can directly obtain knowledge about the new type of product based on the updated knowledge graph, and then train the coding model. This eliminates the need to change the model structure of the coding model or retrain the LLM, which can effectively improve the efficiency of product information processing.

[0061] Specifically, when expanding the LLM's knowledge graph with a new product type, the attributes and attribute values ​​of the new product type can be determined first. Other products and / or other attributes related to the new product type can be identified in the current knowledge graph. Based on the attributes and attribute values ​​of the new product type, as well as other products and / or other attributes related to the new product type, the new product type can be added to the current knowledge graph. For example, the attributes and attribute values ​​of a product type may include: if the attribute is size, the corresponding attribute value is XL; if the attribute is color, the corresponding attribute value is blue.

[0062] After training the encoding model, you can Figure 1The method shown processes the product information of the target product. The following describes the specific processing process for different application scenarios.

[0063] For the scenario of automatically classifying newly added products, the server can input the category information of each product category into the pre-trained LLM in advance, obtain the category features of each product category output by the LLM, and save them. When automatically classifying newly added products, in step S100, the server can use the newly added product as the target product, and continue to execute subsequent steps S102 to S106 to obtain the product features of the target product, and then determine the product category to which the target product belongs based on the product features of the target product and the category features of each pre-saved product category, as the target category, and classify the target product under the target category. Specifically, when determining the product category to which the target product belongs, the server can first determine the similarity between the product features of the target product and the category features of each pre-saved product category, and determine the product category with the greatest similarity as the target category to which the target product belongs.

[0064] Correspondingly, if a new product category appears, the category information of the new product category can also be input into the above-mentioned LLM in advance, so that the LLM outputs the category characteristics of the new product category and saves it.

[0065] In the scenario of recommending products to users, the server may first obtain the user's demand information. The method for obtaining the demand information may include: determining the demand information through the user's user information and / or behavior information. After obtaining the user's demand information, the demand information may be input into the above-mentioned LLM to encode the user's demand information through the LLM to obtain the user's demand characteristics, and then all the products in the product recall pool are used as Figure 1 The target products shown are obtained, and the product characteristics of each target product are obtained. Finally, according to the product characteristics of the target product and the user's demand characteristics, it is determined whether the target product is a product to be recalled. If so, the target product is recalled for the user.

[0066] It can be seen from the above method that this manual uses the powerful encoding ability of LLM to train the encoding model, so that the encoding features of the encoding model regardless of the modality are aligned to the feature space of LLM. The trained encoding model can accurately encode product information whether in a single modality or in multimodal fusion, and can obtain accurate product features while ensuring the real-time encoding of product information.

[0067] The above is a robot-based service provision method provided in an embodiment of this specification. Based on the same idea, this specification also provides corresponding devices, storage media and electronic devices.

[0068] Figure 4 This is a schematic diagram of a device for processing product information provided in an embodiment of this specification, the device comprising:

[0069] Acquisition module 401, used to acquire multimodal product information of a target product;

[0070] Encoding module 402 is configured to input, for each modality, the product information of the target product in that modality into an encoder corresponding to that modality in a pre-trained encoding model, and obtain encoding features of that modality output by the encoder corresponding to that modality; wherein the encoding model is pre-trained based on a large language model (LLM), and the number of parameters of the encoding model is smaller than that of the LLM;

[0071] A determination module 403 is configured to determine the product features of the target product based on the coding features of each modality of the target product;

[0072] The processing module 404 is configured to process the target product according to the product characteristics of the target product.

[0073] Optionally, the multimodal product information includes product information in image mode and product information in text mode;

[0074] The coding model includes an image encoder and a text encoder.

[0075] Optionally, the coding model further includes an adaptation network;

[0076] The determining module 403 is specifically configured to input the encoding features of each modality of the target product into the adaptation network to obtain product features of the target product output by the adaptation network.

[0077] Optionally, the adaptation network includes an image adapter, a fusion adapter, and a text adapter;

[0078] The determination module 403 is specifically used to, when the coding features of the target product have only one modality, input the coding features of the modality into the adapter corresponding to the modality in the adaptation network, and obtain the product features of the target product output by the adapter corresponding to the modality; when the coding features of the target product have more than two modalities, input the coding features of each modality into the fusion adapter in the adaptation network, and obtain the product features of the target product output by the fusion adapter.

[0079] Optionally, the device further comprises:

[0080] The training module 405 is used to pre-acquire the multimodal product information of the sample product; input the multimodal product information of the sample product into the LLM, and obtain the product features of the sample product output by the LLM as the annotation features; and, for each modality, input the product information of the modality of the sample product into the encoder corresponding to the modality in the encoding model to be trained, and obtain the encoding features of the modality output by the encoder corresponding to the modality, and determine the product features of the sample product according to the encoding features of each modality of the sample product as the comparison features; determine the loss value according to the similarity between the comparison features and the annotation features, wherein the loss value is negatively correlated with the similarity; and adjust the model parameters in the encoding model to be trained with reducing the loss value as the training goal.

[0081] Optionally, the LLM is a knowledge graph-based LLM;

[0082] The training module 405 is further configured to, when the sample product is a new type of product, add the sample product to the knowledge graph before inputting the multimodal product information of the sample product into the LLM.

[0083] Optionally, the processing module 404 is specifically used to determine the product category to which the target product belongs as the target category based on the product characteristics of the target product and the category characteristics of each pre-saved product category; wherein the category characteristics of each product category are the category characteristics output by the LLM obtained by inputting the category information of each product category into the LLM in advance; and classify the target product into the target category.

[0084] Optionally, the processing module 404 is specifically used to input the user's demand information into the LLM to obtain the user's demand characteristics; determine whether the target product is a recalled product based on the product characteristics of the target product and the user's demand characteristics; if so, recall the target product for the user.

[0085] This specification also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can be used to execute the above-mentioned method for processing commodity information.

[0086] based on Figure 1 The method for processing the commodity information shown in the embodiment of this specification also provides Figure 5 The structural diagram of the electronic device shown in FIG. Figure 5At the hardware level, the electronic device includes a processor, an internal bus, a network interface, memory, and non-volatile storage, and may also include other hardware required for its operations. The processor reads the corresponding computer program from the non-volatile storage into the memory and then runs it to implement the aforementioned method for processing product information.

[0087] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.

Claims

1. A method for processing product information, the method comprising: Obtain multimodal product information of the target product; For each modality, the product information of the target product in that modality is input into the encoder corresponding to that modality in a pre-trained encoding model, and the encoding features of that modality output by the encoder corresponding to that modality are obtained; wherein the encoding model is pre-trained based on a large language model (LLM), and the number of parameters of the encoding model is smaller than that of the LLM; Determining product features of the target product based on the coding features of each modality of the target product; The target product is processed according to the product characteristics of the target product.

2. The method according to claim 1, wherein the multimodal product information includes product information in an image mode and product information in a text mode; The coding model includes an image encoder and a text encoder.

3. The method of claim 2, wherein the coding model further comprises an adaptation network; Determining the product features of the target product based on the coding features of each modality of the target product specifically includes: The encoding features of each modality of the target product are input into the adaptation network to obtain the product features of the target product output by the adaptation network.

4. The method of claim 3, wherein the adaptation network comprises an image adapter, a fusion adapter, and a text adapter; Inputting the encoding features of each modality of the target product into the adaptation network to obtain the product features of the target product output by the adaptation network specifically includes: When the coding feature of the target product has only one mode, the coding feature of the mode is input into the adapter corresponding to the mode in the adaptation network to obtain the product feature of the target product output by the adapter corresponding to the mode; When the coding features of the target product have more than two modalities, the coding features of each modality are input into the fusion adapter in the adaptation network to obtain the product features of the target product output by the fusion adapter.

5. The method according to claim 3, wherein the encoding model is pre-trained based on a large language model (LLM), specifically comprising: Obtain multimodal product information for sample products; Input the multimodal product information of the sample product into the LLM, and obtain the product features of the sample product output by the LLM as the annotation features; Furthermore, for each modality, the product information of the sample product in that modality is input into the encoder corresponding to that modality in the encoding model to be trained, and the encoding features of that modality output by the encoder corresponding to that modality are obtained. Based on the encoding features of each modality of the sample product, the product features of the sample product are determined as comparison features; Determining a loss value according to a similarity between the comparison feature and the annotation feature, wherein the loss value is negatively correlated with the similarity; Taking reducing the loss value as a training goal, the model parameters in the encoding model to be trained are adjusted.

6. The method according to claim 5, wherein the LLM is a knowledge graph-based LLM; Before inputting the multimodal product information of the sample product into the LLM, the method further includes: When the sample product is a new type of product, the sample product is added to the knowledge graph.

7. The method according to claim 1, wherein processing the target product according to the product characteristics of the target product specifically comprises: Determining the product category to which the target product belongs as the target category based on the product characteristics of the target product and the pre-stored category characteristics of each product category; wherein the category characteristics of each product category are obtained by pre-inputting the category information of each product category into the LLM and outputting the category characteristics of the LLM; Classify the target product into the target category.

8. The method according to claim 1, wherein processing the target product according to the product characteristics of the target product specifically comprises: Inputting the user's demand information into the LLM to obtain the user's demand characteristics; Determining whether the target product is a product to be recalled based on the product characteristics of the target product and the user's demand characteristics; If so, the target product is recalled for the user.

9. A device for processing commodity information, comprising: The acquisition module is used to obtain multimodal product information of the target product; An encoding module is configured to input, for each modality, the product information of the target product in that modality into an encoder corresponding to that modality in a pre-trained encoding model, and obtain encoding features of that modality output by the encoder corresponding to that modality; wherein the encoding model is pre-trained based on a large language model (LLM), and the number of parameters of the encoding model is smaller than that of the LLM; a determination module, configured to determine the commodity features of the target commodity based on the coding features of each modality of the target commodity; The processing module is used to process the target product according to the product characteristics of the target product.

10. A computer-readable storage medium storing a computer program, wherein the computer program implements the method according to any one of claims 1 to 8 when executed by a processor.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 8 when executing the program.