Training method, application method and training device of commodity processing model
By processing sample images containing multiple sub-regions and training a multimodal large language model, the problem of insufficient performance of multimodal large models in the prior art during fine-grained area tasks is solved, and more efficient product-level details are achieved.
Patent Information
- Application Number
- CN202411946355.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-26
AI Technical Summary
Existing multimodal large models have insufficient performance when dealing with fine-grained area or pixel-level tasks and cannot meet the usage needs of retail scenarios.
By obtaining sample images containing multiple sub-regions, processing the image and text description information of each sub-region, and training using a pre-trained multimodal large language model, enhancing the model's ability to understand product-level details.
The model's understanding and execution effect of fine-grained tasks is improved, and product-level details in retail scenarios can be captured and processed more accurately.
Smart Images

Figure CN120014381A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a training method, an application method and a training device for a commodity processing model. Background Art
[0002] With the rapid development of artificial intelligence and deep learning technologies, multimodal large models have been widely used in various fields. In retail scenarios, for example, the pre-trained model of contrasting language images usually focuses on the overall relationship between text and image, but lacks accurate capture of details. At present, the performance of multimodal large models is insufficient for fine-grained areas or pixel-level tasks, and cannot meet the needs of business use. Summary of the invention
[0003] This application provides a training method, application method and training device for a commodity processing model, which helps to solve the problem that the current multi-modal large model has insufficient performance when processing fine-grained regional tasks. The following introduces various aspects involved in this application.
[0004] In a first aspect, the present application provides a method for training a product processing model, comprising: obtaining a first sample image, the first sample image being any sample image in a labeled training sample set, the first sample image having first sample image information and first sample text description information, and the first sample image having multiple sub-regions, the multiple sub-regions corresponding one-to-one to multiple product sub-images, the first product sub-image having first product image information and first product text description information, the first product sub-image being any product sub-image among the multiple product sub-images, and the first product text description information at least including the location and brand information of the first product sub-image; calling a pre-trained multimodal large language model to process the first sample image information and the first sample text description information of the first sample image, Obtain a first sample image feature vector and a first sample text feature vector of the first sample image, process the first product image information and the first product text description information of the first product sub-image, and obtain a first product image feature vector and a first product text feature vector of the first product sub-image; obtain generated text description information of the first sample image based on the first sample image feature vector and the instruction information of the first sample image; train the pre-trained multimodal large language model based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, and the total loss function to obtain the multimodal large language model.
[0005] In a second aspect, the present application provides an application method of a commodity processing model, comprising: obtaining a first image to be identified and instruction information of the first image; inputting the first image to be identified and the instruction information of the first image into a pre-trained commodity processing model for identification to obtain a target output; wherein the commodity processing model is a multimodal large language model trained using the training method described in the first aspect.
[0006] In the third aspect, the present application provides a training device for a product processing model, including: a first acquisition module, used to acquire a first sample image, the first sample image is any sample image in a labeled training sample set, the first sample image has first sample image information and first sample text description information, and the first sample image has multiple sub-regions, the multiple sub-regions correspond one by one to multiple product sub-pictures, the first product sub-picture has first product image information and first product text description information, the first product sub-picture is any product sub-picture of the multiple product sub-pictures, and the first product text description information at least includes the location and brand information of the first product sub-picture; a calling module, used to call a pre-trained multimodal large language model to process the first sample image information and the first sample text description information of the first sample image. Processing is performed to obtain a first sample image feature vector and a first sample text feature vector of the first sample image, respectively, and the first product image information and the first product text description information of the first product sub-image are processed to obtain a first product image feature vector and a first product text feature vector of the first product sub-image, respectively; based on the first sample image feature vector and the instruction information of the first sample image, the generated text description information of the first sample image is obtained; a training module is used to train the pre-trained multimodal large language model based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, and the total loss function to obtain the multimodal large language model.
[0007] In a fourth aspect, the present application provides an application device for a commodity processing model, comprising: a second acquisition module, used to acquire a first image to be identified and instruction information of the first image; a processing module, used to input the first image to be identified and the instruction information of the first image into a pre-trained commodity processing model for identification, and obtain a target output; wherein the commodity processing model is a multimodal large language model trained by the training device as described in the third aspect.
[0008] In a fifth aspect, the present application provides an electronic device, comprising: a memory for storing code; and a processor connected to the memory, for executing the code stored in the memory, so that the electronic device executes the training method as described in the first aspect or the application method as described in the second aspect.
[0009] In a sixth aspect, the present application provides a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the training method as described in the first aspect or the application method as described in the second aspect.
[0010] In a seventh aspect, the present application provides a computer program product, which, when executed on an electronic device, enables the electronic device to implement the training method described in the first aspect or the application method described in the second aspect.
[0011] The embodiment of the present application is a training method for a multimodal large model for retail scenarios, in which a data set of retail scenarios is added to the pre-trained multimodal large model, wherein the first sample image has not only the first sample image information of the overall image, but also the image information of multiple commodity sub-pictures corresponding to multiple sub-regions, and the training sample data contains image-level text descriptions and commodity-level text descriptions, which can accurately capture the details of the commodity level, enhance the model's ability to understand the commodity level, and help improve the model's ability to understand and execute fine-grained tasks (commodity level). The model trained in the embodiment of the present application can be applied to the vertical field of smart retail, and is suitable for natural language instruction generation and task execution in retail-related tasks such as commodity identification, shelf layout analysis, and brand identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in describing the embodiments of the present application are briefly introduced below.
[0013] Figure 1 It is a flowchart of the training method of the commodity processing model provided in the embodiment of the present application.
[0014] Figure 2 It is a working diagram of an image encoder and a text encoder.
[0015] Figure 3 It is a flowchart of the application method of the commodity processing model provided in the embodiment of the present application.
[0016] Figure 4 It is a schematic diagram of the composition of the training device for the commodity processing model provided in an embodiment of the present application.
[0017] Figure 5It is a schematic diagram of the composition of the application device of the commodity processing model provided in the embodiment of the present application.
[0018] Figure 6 It is a schematic diagram of a component unit / partial component unit of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0019] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The same or similar reference numerals are used in the drawings to represent the same or similar modules. It should be understood that the drawings are only schematic, and the protection scope of the present application is not limited thereto.
[0020] First, the application scenarios and related technical terms involved in the embodiments of the present application are introduced.
[0021] Large language models (LLMs) are models trained using deep learning techniques that can process large amounts of text data. These models can predict and generate natural language text by learning from a large corpus, have strong language understanding and generation capabilities, and can handle a variety of natural language tasks, such as text classification, intelligent question-answering, and dialogue.
[0022] Multimodal large language models (MLLMs) are large models that can simultaneously process and understand multiple different types of data (such as images, text, speech, etc.). Traditional artificial intelligence models can usually only process one type of data, while multimodal large language models can use multiple types of data for joint modeling and prediction, thereby more comprehensively understanding and processing input information.
[0023] With the rapid development of artificial intelligence and deep learning technologies, multimodal large models have been widely used in various fields. Multimodal large models can be used for natural language instruction generation and task execution in retail-related tasks such as shelf layout analysis. However, in retail scenarios, models such as contrastive language-image pre-training (CLIP) usually focus on the overall relationship between text and images, but lack accurate capture of details. For example, it is not possible to identify whether the location of the product is correct, whether the product is out of stock, etc. It can be seen that the current multimodal large model has insufficient performance for fine-grained level (product level) tasks in retail scenarios and cannot meet the business usage requirements.
[0024] Therefore, it is necessary to design a technical solution for a multimodal large model that can improve the understanding ability of fine-grained tasks.
[0025] Based on this, the embodiment of the present application proposes a method for training a commodity processing model, and the trained multimodal large language model is oriented to retail scenarios, and the retail scenarios can be physical stores such as clothing stores, shoe stores, and comprehensive shopping malls. The model trained in the embodiment of the present application can be applied to the vertical field of smart retail, and is suitable for natural language instruction generation and task execution in retail-related tasks such as commodity identification, shelf layout analysis, and brand identification, which helps to improve the model's understanding ability and execution effect of fine-grained tasks (commodity level).
[0026] In some implementations, the architecture of the multimodal large language model can adopt the architecture design of a large language and vision assistant (LLAVA), and the model can include three components: a visual encoder, a visual-language connector, and a large language model (LLM). The training is divided into three stages: the first stage is the pre-training stage, training the visual encoder, the second stage is training the visual-language connector, and the third stage is training the connector and LLM.
[0027] Combine the following Figure 1 The training method of the commodity processing model of the embodiment of the present application is described in detail. Figure 1 As shown, the training method of the commodity processing model of the embodiment of the present application may mainly include steps S110 to S140, and these steps are described in detail below.
[0028] It should be pointed out that the size of the serial numbers of the steps in the embodiments of the present application does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0029] In step S110, a first sample image is obtained, and the first sample image is any sample image in the labeled training sample set. The first sample image has first sample image information and first sample text description information, and the first sample image has multiple sub-regions, and the multiple sub-regions correspond to multiple product sub-images one by one. The first product sub-image has first product image information and first product text description information, and the first product sub-image is any product sub-image among the multiple product sub-images. The first product text description information at least includes the location information and brand information of the first product sub-image.
[0030] Compared with focusing only on the overall relationship between text and images, the training sample data contains text descriptions at the image level and product level, which can accurately capture the details at the product level and help improve the model's ability to understand fine-grained tasks (product level).
[0031] Conventional training sample images only include the overall relationship between text and image, and do not mark the commodity area in the image. Before obtaining the first sample image, the training method of the embodiment of the present application may also include: processing the initial first sample image, extracting multiple sub-regions corresponding to multiple commodity sub-images; obtaining the first commodity image information of the first commodity sub-image, marking the first commodity sub-image, and determining the text description information corresponding to the first commodity sub-image; obtaining the marked first sample image. That is, before training, the training sample set is first subjected to commodity-level image area extraction and text annotation. It can be understood that the multiple commodity sub-images corresponding to the multiple sub-regions may be the same or different. For example, the commodity sub-image corresponding to the first sub-region is a picture of a bottle of pure water, and the commodity sub-image corresponding to the second sub-region may be a picture of a can of milk powder.
[0032] In the pre-training stage, the training sample set can be a set of annotated retail product images, and any sample image in the retail product image set is annotated with the location and brand information of each product. Based on any sample image, a text description can be generated for each sub-region corresponding to each product. For example, the first product sub-image can be a bottled water image, and its text description information can be "a bottle of 500ml Wahaha pure water".
[0033] The training sample set may include a public dataset and a dedicated dataset. The public dataset may be, for example, Conceptual 12M (CC12M), which is a large and diverse image-text pair dataset. The dedicated dataset may be a retail commodity dataset suitable for retail scenarios.
[0034] In some specific implementations, for the CC12M image-text pair dataset, each image has a text description corresponding to the entire image. Since multiple sub-regions (targets) in the image are not labeled, the recognize anything model (RAM) can be used to generate the label of the image first, and then GroundingDINO can be used to generate a bounding box for each label. For example, use RAM to generate image labels, such as "grass, puppy", and use GroundingDINO to generate a bounding box for each label. Grounding DINO only needs to input the label name of the item (commodity) to be detected, and it can output the corresponding bounding box. Bounding boxes with confidence levels lower than 0.3 can be filtered out. For each bounding box, after cropping, a general large language model such as Tongyi Qianwen (Qwen2-vl) can be used to generate the corresponding text title. Among them, Grounding DINO adopts the Transformer architecture of the target detector DINO, and draws on the pre-training method of the multimodal GLIP. After deeply integrating language and visual information, any target can be detected based on the text description.
[0035] In the dedicated retail product dataset, each product sub-image can be described by annotating the location coordinates, brand, category and name of each product. For example, Qwen2-vl can be used to generate image-level titles (text description information). For each bounding box, Qwen2-vl is used to generate the corresponding text title after cropping. The annotated category information can be used as a prompt during the generation process.
[0036] In step S120, the pre-trained multimodal large language model is called to process the first sample image information and the first sample text description information of the first sample image to obtain the first sample image feature vector and the first sample text feature vector of the first sample image, respectively. The first product image information and the first product text description information of the first product sub-image are processed to obtain the first product image feature vector and the first product text feature vector of the first product sub-image, respectively.
[0037] In some implementations, the first product image feature vector of the first product sub-image is extracted from the first sample image feature vector using region of interest align (ROIAlign). For example, the initial first sample image may be an image with a resolution of 224*224, the first sample image feature vector may be 196*768, and the region corresponding to the first product on the initial first sample image is assumed to be "100, 100, 150, 150", then the ROIAlign method may be used to extract the features corresponding to this region as the first product image feature vector of the first product sub-image.
[0038] In some implementations, the architecture of the pre-trained multimodal large model can adopt CLIP, which includes a visual encoder and a text encoder.
[0039] A visual encoder is a neural network model that abstracts and extracts the information of each pixel of an input image and converts it into a fixed-length vector representation. A visual encoder is usually composed of alternating convolutional layers and pooling layers. The input of a visual encoder is usually the original image pixel matrix, and the output can be a fixed-length vector representation or a fixed-size visual feature map.
[0040] For example, calling the visual encoder can process the first sample image information of the first sample image to obtain the first sample image feature vector of the first sample image. Calling the visual encoder to process the first product image information of the first product sub-image to obtain the first product image feature vector of the first product sub-image.
[0041] The input of the visual encoder is usually a fixed resolution, for example, it can be 224*224. After the image is compressed, the resolution becomes smaller, and a lot of information will be lost, making it difficult for the model to focus on the details. In some implementations, high resolution can be added on the basis of the image encoder. Specifically, for the first sample image with high resolution, the high-resolution image can be intelligently divided into multiple small blocks (patches) based on the LLaVA-1.5-HD method, and then embedded and merged with the low-resolution image. This helps to improve the resolution of the extracted image feature vector and enables the visual encoder to learn more fine-grained features.
[0042] Multimodal large models usually also include text encoders. Text encoders, also known as text embedding or text representation, can use neural network models to convert input text sequences into fixed-length vector representations with certain semantic features. By converting text information into vector representations, text encoders can directly convert text information into numerical form for processing, and discover the inherent laws of text data through deep learning models, thereby better representing text information.
[0043] Figure 2 This is a working diagram of an image encoder and a text encoder. Figure 2 As shown in the CLIP algorithm framework, the image-text pair (image, text) is an image of a puppy and "Feed the puppy Paipai Pepper". Although "Paipai" in the text is the name of the puppy and cannot be completely aligned with the image, the whole sentence can still be semantically aligned with the image. For the image-text pair, the image data is passed through the image encoder to obtain its representation I i , pass the text data through the text encoder to get its representation T i ,Then, the gradients of the image encoder and text encoder are optimized by contrastive learning loss ,function.
[0044] The text encoder of the pre-trained multimodal large language model can be called to process the first sample text description information of the first sample image to obtain the first sample text feature vector. The first product text description information of the first product sub-image is processed to obtain the first product text feature vector of the first product sub-image. The text encoder is used to generate image-level (dimension) text label embedding and product-level (dimension) text label embedding, and then calculate the image-level contrast loss and product-level contrast loss.
[0045] For example, calculating the contrast loss of the overall image information after the text encoder and the image encoder can be considered as an image-level loss. For example, for an image of a shelf, the corresponding text description information can be "This is a supermarket shelf, and there are many beverages on the shelf." If you want the model to focus on smaller sub-areas, that is, commodity sub-areas, you need to add corresponding text description information to these small sub-areas during training. Therefore, during training, add commodity-level embedding to calculate the commodity-level contrast loss using commodity-level text feature vectors (embeddings) and image feature vectors.
[0046] After pre-training, the obtained multimodal large model includes a text encoder and a visual encoder. Then the visual encoder is connected to the multimodal large language model. The multimodal large language model itself cannot directly understand the representation space of the image feature vector of the CLIP model. Therefore, a fusion module is needed to map the space and map the input features to the representation space that the LLM can understand.
[0047] The vision-language connector is also called the projection layer or fusion module. Its main function is to map the visual feature vector output by the visual encoder to the input space of the large language model, and to project the visual mark into the space of the large language model, so that the visual feature vector can be effectively fused and interacted with the text feature vector, thereby better realizing the integration and understanding of multimodal information. A multilayer perceptron (MLP) can be used as a vision-language connector.
[0048] During the visual-language connector training phase, all visual encoders and large language models can be frozen. For example, only the visual-language connector can be updated based on the CC3M dataset, which is a large dataset containing about 3 million image-text pairs.
[0049] In step S130, based on the first sample image feature vector and the instruction information of the first sample image, the generated text description information of the first sample image is obtained.
[0050] The instruction information of the first sample image may be, for example, information about the brand, location, and type of all commodities in the first sample image. For example, a pre-trained multimodal large language model may be used to obtain the generated text description information of the first sample image based on the first sample image feature vector and the instruction information of the first sample image.
[0051] The visual-language connector transforms and maps the image feature vector to match the semantic space of the language model and concatenates or fuses it with the text feature vector. The fused multimodal features are input into the large language model, which generates corresponding text output based on these features and previous text input as a comprehensive response to the image and text input.
[0052] In step S140, based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, the objective function, and the total loss function, the pre-trained multimodal large language model is trained to obtain a multimodal large language model.
[0053] In the pre-training stage, the objective function of the product sub-image dimension is introduced for supervision, and the contrast loss of the text feature vector and image feature vector of the product sub-image dimension is calculated to optimize the model parameters.
[0054] For the contrast loss of the overall image dimension, the text feature vector (embedded) and image feature vector of the image dimension can be used for calculation. For the contrast loss of the product sub-image dimension, the text feature vector and image feature vector of the product sub-image dimension can be used for calculation. The calculation formula of the total loss function is:
[0055] L=αL image +βL object
[0056] Among them, L image is the overall image loss, L object is the loss of the product sub-image, and α and β are weight coefficients.
[0057] For example, the back propagation algorithm and gradient descent can be used to update the parameters of the text feature vector. According to the loss function value, the gradient of the image feature vector is calculated, and the parameter value of the image feature vector is updated according to the gradient to train and optimize the parameters of the pre-trained multimodal large model.
[0058] The visual-language connector and the large language model (LLM) can be fine-tuned. Large model fine-tuning has become a key technology to improve the performance of models on specific tasks. After the visual-language alignment pre-training, visual instruction fine-tuning is required to help improve the ability of the multimodal large language model to follow instructions and solve tasks. Generally speaking, the input of visual instruction fine-tuning includes an image and a task description text, and the output is the corresponding text response.
[0059] In some implementations, the fine-tuning dataset can use LLaVA's LLaVA-Instruct-150K specified fine-tuning dataset. Questions are answered in a general format, but instructions in retail scenarios cannot be accurately output. In other implementations, the fine-tuning dataset can use LLaVA's LLaVA-Instruct-150K specified fine-tuning dataset, and add a dedicated fine-tuning dataset built for retail scenarios. This helps the multimodal large model to accurately output instructions in retail scenarios, accurately answering in the expected format, rather than outputting general unnecessary content.
[0060] In some implementations, the method of the embodiment of the present application may also include: constructing a dedicated fine-tuning dataset for retail scenarios. Based on the training sample set and image dataset of the already labeled commodity recognition, the constructed dedicated fine-tuning dataset includes at least some or all of the following tasks: counting, commodity positioning, and brand recognition. Adding dedicated retail scenario data and dedicated fine-tuning datasets to the fine-tuned multimodal large model helps to enhance the model's ability to understand and express sub-region and commodity-level information during the pre-training stage.
[0061] The counting tasks include at least: 1) how many shelves are there in the image; 2) how many items are there on each shelf; 3) how many items are there on the X-th shelf; 4) how many empty areas (no items are placed) are there in the image, where X is an integer not less than 1.
[0062] The task of product positioning includes at least: 1) what are the coordinates of the first brand product; 2) what is the product to the right of the first brand product; 3) provide the location information of all products; 4) what is the product in the Xth layer and the Yth column, where Y is an integer not less than 1, and the first brand is any brand among multiple brands in the retail scene.
[0063] The task of brand recognition includes at least: 1) how many brands of goods are there in the image; 2) how many brands of goods are there on the X-th shelf; 3) how many different goods are there in the image; 4) whether there are goods of a specified brand on the k-th shelf, where k is an integer not less than 1.
[0064] The step S140 of training the pre-trained multimodal large language model based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, the objective function, and the total loss function may include: training the pre-trained multimodal large language model based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, the objective function, the total loss function, and a dedicated fine-tuning dataset.
[0065] The main construction process of the dedicated fine-tuning dataset is described below.
[0066] Select a data set (image data set) that is labeled with shelf location, product brand, product name, and product location. For each image, the shelf number and the number of shelf layers can be identified and counted based on the shelf location; the number of products on each shelf layer can be counted based on the product location information; the number of brands and product categories in the entire image can be counted based on the product category and brand, as well as the number of brands and product categories on each shelf layer (row); each product in the image is numbered, the position coordinates are converted to [0-1], and normalized to count the number of empty positions on each shelf in each image; and the above statistical data are formed into a statistical database.
[0067] An instruction fine-tuning dataset is constructed for each task. For example, the task of product positioning is used as an example for explanation.
[0068] Regarding question 1, what are the coordinates of the first brand product? You can query the statistical database, obtain the answer, and construct the corresponding question pair.
[0069] For question 2, what is the product on the right of the first brand product: obtain the product information database of the selected image, select a product to replace the first brand product, select a position pronoun in [left, right, above, below], obtain the answer based on the position pronoun and the product's position number, and construct the corresponding question pair.
[0070] For question 3, given the location information of all products: query the database, obtain the answer, and construct the corresponding question pair.
[0071] For question 4, what is the product in the Xth layer and Yth column? Get the product information database of the selected image, randomly select and replace X and Y according to the number of shelf layers and columns, get the answer, and construct the corresponding question pair.
[0072] The embodiment of the present application is a training method for a multimodal large model for retail scenarios, in which a data set of retail scenarios is added to the pre-trained multimodal large model, wherein the first sample image has not only the first sample image information of the overall image, but also the image information of multiple commodity sub-pictures corresponding to multiple sub-regions, and the training sample data contains image-level text descriptions and commodity-level text descriptions, which can accurately capture the details of the commodity level, enhance the model's ability to understand the commodity level, and help improve the model's ability to understand and execute fine-grained tasks (commodity level). The model trained in the embodiment of the present application can be applied to the vertical field of smart retail, and is suitable for natural language instruction generation and task execution in retail-related tasks such as commodity identification, shelf layout analysis, and brand identification.
[0073] The embodiment of the present application provides an application method of a commodity processing model. Figure 3 Schematic diagram of the application method of the commodity processing model provided in the embodiment of the present application. Figure 3 As shown, the application method of the commodity processing model in the embodiment of the present application may mainly include steps S310 to S320, and these steps are described in detail below.
[0074] In step S310, a first image to be recognized and instruction information of the first image are obtained. The first image is any one of the multiple images to be recognized.
[0075] In step S320, the first image to be recognized and the instruction information of the first image are input into a pre-trained commodity processing model for recognition to obtain a target output quantity. The commodity processing model is a multimodal large language model trained by any of the training methods described above.
[0076] In the embodiment of the present application, the first image not only has image information of the overall image, but also has image information of multiple product sub-images corresponding to multiple sub-regions. The first image contains image-level text descriptions and product-level text descriptions, which can accurately capture the details of the product level, enhance the model's understanding of the product level, and help improve the model's understanding ability and execution effect of fine-grained tasks (product level).
[0077] Combination of the above Figure 1-Figure 3 The method embodiment of the present application is described in detail. Figures 4 to 5 The device embodiment of the present application is described in detail. It should be understood that the description of the device embodiment corresponds to the description of the method embodiment, so the parts not described in detail can refer to the previous method embodiment.
[0078] The embodiment of the present application also provides a training device for a commodity processing model. Figure 4 Schematic diagram of the composition of the training device for the commodity processing model provided in the embodiment of the present application. Figure 4 As shown, the training device 400 for the commodity processing model includes: a first acquisition module 410 , a calling module 420 , and a training module 430 .
[0079] The first acquisition module 410 is used to acquire a first sample image, which is any sample image in the labeled training sample set. The first sample image has first sample image information and first sample text description information, and the first sample image has multiple sub-regions, and the multiple sub-regions correspond to multiple product sub-images one by one, and the first product sub-image has first product image information and first product text description information. The first product sub-image is any product sub-image among the multiple product sub-images, and the first product text description information at least includes the location and brand information of the first product sub-image.
[0080] The calling module 420 is used to call the pre-trained multi-model large language model, process the first sample image information and the first sample text description information of the first sample image, and obtain the first sample image feature vector and the first sample text feature vector of the first sample image; process the first product image information and the first product text description information of the first product sub-image, and obtain the first product image feature vector and the first product text feature vector of the first product sub-image. Based on the first sample image feature vector and the instruction information of the first sample image, the generated text description information of the first sample image is obtained.
[0081] The training module 430 is used to train the pre-trained multimodal large language model based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, and the total loss function to obtain the multimodal large language model.
[0082] Optionally, the training device 400 further includes a preprocessing module. The preprocessing module is used to process the initial first sample image, extract multiple product sub-images in the first sample image, annotate the first product sub-image, determine the text description information of the first product sub-image, and obtain the annotated first sample image. In other words, the preprocessing module is used to process the initial first sample image, extract multiple sub-regions corresponding to multiple product sub-images, annotate the multiple sub-regions, obtain the text description information corresponding to the multiple product sub-images, obtain the image information corresponding to the multiple product sub-images, and obtain the annotated first sample image.
[0083] Optionally, the training device 400 may further include a construction module. The construction module is used to construct a dedicated fine-tuning dataset for retail scenarios. The training module 430 is used to train the pre-trained multimodal large language model based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, the total loss function, and the dedicated fine-tuning dataset. The dedicated fine-tuning dataset constructed by the construction module includes at least some or all of the following tasks: counting, product positioning, and brand recognition.
[0084] The counting task includes at least: how many layers of shelves are there in the image, how many products are there on each layer, how many products are there on the X-th shelf, and how many empty areas (no products are placed) are there in the image, where X is an integer not less than 1. The product positioning task includes at least: what are the coordinates of the first brand product, what is the product to the right of the first brand product, giving the location information of all products, what is the product in the X-th layer and the Y-th column, where Y is an integer not less than 1, and the first brand is any one of the multiple brands in the retail scene. The brand recognition task includes at least: how many brands of products are there in the image, how many brands of products are there on the X-th shelf, how many different products are there in the image, and whether there are products of a specified brand on the i-th shelf, where i is an integer not less than 1.
[0085] The embodiment of the present application also provides an application device of the commodity processing model. Figure 5 Schematic diagram of the composition of the application device of the commodity processing model provided in the embodiment of the present application. Figure 5 As shown, the application device 500 of the commodity processing model includes: a second acquisition module 510 and a processing module 520.
[0086] The second acquisition module 510 is used to acquire the first image to be recognized and instruction information of the first image.
[0087] The processing module 520 is used to input the first image to be identified and the instruction information of the first image into the pre-trained commodity processing model for identification to obtain the target output. The commodity processing model is a multimodal large language model trained by any of the training devices described above.
[0088] Figure 6 Schematic diagram of a component unit / partial component unit of an electronic device provided in an embodiment of the present application. Figure 6 As shown, the electronic device 600 may include: a memory 610 and at least one processor 620 .
[0089] The memory 610 is used to store codes or computer programs.
[0090] The processor 620 is connected to the memory 610 and is used to execute the code or computer program stored in the memory 610 to control the electronic device 600 to execute any of the processing methods described above.
[0091] Exemplarily, the computer program may be divided into one or more modules / units, one or more modules / units are stored in the memory 610 and executed by the processor 620 to complete the present application.
[0092] Those skilled in the art will understand that Figure 6The electronic device 600 is merely an example and does not constitute a limitation on the electronic device. The electronic device 600 may include more or fewer components than shown in the figure, or combine some components, or different components.
[0093] The processor 620 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc.
[0094] The electronic device 600 provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effect are similar, which will not be repeated here.
[0095] An embodiment of the present application further provides a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0096] An embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0097] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent commodity, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above method embodiments, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of each of the above method embodiments can be implemented. Among them, the computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the camera / electronic device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), read-only compact disc (CD-ROM), magnetic tape, floppy disk and optical data storage device. The computer-readable storage medium mentioned in the present application can be a non-volatile storage medium, in other words, it can be a non-transient storage medium.
[0098] It should be noted that the information collection process (such as face image collection process, fingerprint collection process, etc.) / feature extraction process involved in this application is performed with the user's knowledge and permission, that is, the information collection process / feature extraction process complies with the requirements of laws and regulations and does not constitute an act that harms the public interest.
[0099] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0100] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0101] In the embodiments provided in the present application, it should be understood that the disclosed devices / equipment and methods can be implemented in other ways. For example, the device / equipment embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0102] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0103] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0104] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0105] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for training a commodity processing model, characterized in that: include: Acquire a first sample image, the first sample image is any sample image in a labeled training sample set, the first sample image has first sample image information and first sample text description information, and the first sample image has multiple sub-regions, the multiple sub-regions correspond one-to-one to multiple product sub-images, the first product sub-image has first product image information and first product text description information, the first product sub-image is any product sub-image among the multiple product sub-images, and the first product text description information at least includes the location and brand information of the first product sub-image; Calling a pre-trained multimodal large language model, processing the first sample image information and the first sample text description information of the first sample image, respectively obtaining a first sample image feature vector and a first sample text feature vector of the first sample image, and processing the first product image information and the first product text description information of the first product sub-image, respectively obtaining a first product image feature vector and a first product text feature vector of the first product sub-image; Based on the first sample image feature vector and the instruction information of the first sample image, obtaining generated text description information of the first sample image; Based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, and the total loss function, the pre-trained multimodal large language model is trained to obtain the multimodal large language model.
2. The training method according to claim 1, characterized in that: Before acquiring the first sample image, the training method further includes: Processing the initial first sample image to extract multiple product sub-images in the first sample image; Annotate the first product sub-image to determine text description information of the first product sub-image; Get the first labeled sample image.
3. The training method according to claim 1, characterized in that: Also includes: Constructing a dedicated fine-tuning dataset for retail scenarios, wherein the instruction information of the first sample image is one of the instructions in the dedicated fine-tuning dataset; The pre-trained multimodal large language model is trained based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, and the total loss function, including: Training the pre-trained multimodal large language model based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, the total loss function, and the dedicated fine-tuning dataset; The dedicated fine-tuning dataset includes at least part or all of the following tasks: counting, product positioning, and brand recognition; The counting task at least includes: how many layers of shelves are there in the image, how many items are there on each layer, how many items are there on the X-th layer shelf, and how many empty areas are there in the image, where X is an integer not less than 1; The task of positioning the product at least includes: what are the coordinates of the first brand product, what is the product on the right of the first brand product, providing the location information of all products, what is the product in the Xth layer and the Yth column, where Y is an integer not less than 1, and the first brand is any one of the multiple brands in the retail scene; The brand recognition task at least includes: how many brands of goods are there in the image, how many brands of goods are there on the X-th shelf, how many different goods are there in the image, and whether there are goods of a specified brand on the k-th shelf, where k is an integer not less than 1.
4. A method for applying a commodity processing model, characterized in that: include: Acquire a first image to be identified and instruction information of the first image; Inputting the first image to be recognized and the instruction information of the first image into the pre-trained commodity processing model for recognition to obtain a target output amount; The commodity processing model is a multimodal large language model trained by the training method described in any one of claims 1 to 3.
5. A training device for a commodity processing model, characterized in that: include: A first acquisition module is used to acquire a first sample image, wherein the first sample image is any sample image in a labeled training sample set, the first sample image has first sample image information and first sample text description information, and the first sample image has multiple sub-regions, the multiple sub-regions correspond one-to-one to multiple product sub-images, the first product sub-image has first product image information and first product text description information, the first product sub-image is any product sub-image among the multiple product sub-images, and the first product text description information at least includes the location and brand information of the first product sub-image; a calling module, configured to call a pre-trained multimodal large language model, process the first sample image information and the first sample text description information of the first sample image, and obtain a first sample image feature vector and a first sample text feature vector of the first sample image, respectively, and process the first product image information and the first product text description information of the first product sub-image, and obtain a first product image feature vector and a first product text feature vector of the first product sub-image, respectively; Based on the first sample image feature vector and the instruction information of the first sample image, obtaining generated text description information of the first sample image; A training module is used to train the pre-trained multimodal large language model based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, and the total loss function to obtain the multimodal large language model.
6. The training device according to claim 5, characterized in that The training device also includes: The preprocessing module is used to process the initial first sample image, extract multiple product sub-images in the first sample image; mark the first product sub-image, determine the text description information of the first product sub-image; and obtain the marked first sample image.
7. The training device according to claim 5, characterized in that Also includes: A construction module, used to construct a dedicated fine-tuning dataset for retail scenarios, wherein the instruction information of the first sample image is one of the instructions in the dedicated fine-tuning dataset; The training module is used to train the pre-trained multimodal large language model based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, the total loss function, and the dedicated fine-tuning dataset; The dedicated fine-tuning dataset includes at least part or all of the following tasks: counting, product positioning, and brand recognition; The counting task at least includes: how many layers of shelves are there in the image, how many items are there on each layer, how many items are there on the X-th layer shelf, and how many empty areas are there in the image, where X is an integer not less than 1; The task of positioning the product at least includes: what are the coordinates of the first brand product, what is the product on the right of the first brand product, providing the location information of all products, what is the product in the Xth layer and the Yth column, where Y is an integer not less than 1, and the first brand is any one of the multiple brands in the retail scene; The brand recognition task at least includes: how many brands of goods are there in the image, how many brands of goods are there on the X-th shelf, how many different goods are there in the image, and whether there are goods of a specified brand on the k-th shelf, where k is an integer not less than 1.
8. An application device of a commodity processing model, characterized in that: include: A second acquisition module, used to acquire a first image to be identified and instruction information of the first image; A processing module, used for inputting the first image to be identified and the instruction information of the first image into the pre-trained commodity processing model for identification, and obtaining a target output amount; Wherein, the commodity processing model is a multimodal large language model trained by using the training device described in any one of claims 5-7.
9. An electronic device, characterized in that: include: A memory for storing codes; A processor, connected to the memory, is used to execute the code stored in the memory so that the electronic device executes the training method as described in any one of claims 1 to 3 or the application method as described in claim 4.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed, the computer program is used to implement the training method according to any one of claims 1 to 3 or the application method according to claim 4.
Citation Information
Patent Citations
Commodity identification, detection and counting method and system in retail scene
CN114494823A
Model training method, commodity image management method and device
CN114494890A
Cross-modal retrieval method and system based on multi-granularity feature fusion
CN115391625A
Image reconstruction model training method, commodity identification method, device and equipment
CN116468816A
Inventory checking method and device
CN116596450A