Training methods, application methods, and training devices for commodity processing models
Patent Information
- Application Number
- CN202411946355.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2044-12-26
AI Technical Summary
[0003]本申请提供了一种商品处理模型的训练方法、应用方法和训练装置,有助于解决目前的多模态大模型在处理细颗粒度区域任务时性能不足的问题
[0011] This application's embodiment provides a training method for a multimodal large-scale model for retail scenarios. It adds a retail scenario dataset to the pre-trained multimodal large-scale model. The first sample image not only contains the overall image information but also image information for multiple product sub-images corresponding to multiple sub-regions. The training sample data includes image-level and product-level text descriptions, enabling precise capture of product-level details. This enhances the model's understanding of product-level aspects and improves its ability to understand and perform fine-grained tasks (product-level). The model trained according to this embodiment can be applied to the intelligent retail vertical field, suitable for natural language instruction generation and task execution in retail-related tasks such as product recognition, shelf layout analysis, and brand recognition.
Smart Images

Figure CN120014381B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a training method, application method and training device for a commodity processing model. Background Technology
[0002] With the rapid development of artificial intelligence and deep learning technologies, multimodal large models have been widely used in various fields. In retail scenarios, such as comparative language image pre-trained models, the focus is usually on the overall relationship between text and image, lacking precise capture of details. Currently, multimodal large models are insufficient in performance for fine-grained regions or pixel-level tasks, failing to meet business needs. Summary of the Invention
[0003] This application provides a training method, application method, and training device for a commodity processing model, which helps to solve the performance problem of current multimodal large models when processing fine-grained regional tasks. The various aspects involved in this application are described below.
[0004] Firstly, this application provides a training method for a product processing model, comprising: acquiring a first sample image, wherein the first sample image is any sample image in a labeled training sample set, the first sample image has first sample image information and first sample text description information, and the first sample image has multiple sub-regions, each of the multiple sub-regions corresponding to multiple product sub-images, the first product sub-image having first product image information and first product text description information, the first product sub-image being any one of the multiple product sub-images, and the first product text description information including at least the location and brand information of the first product sub-image; and invoking a pre-trained multimodal large language model to process the first sample image information and the first sample text description information of the first sample image. The first sample image feature vector and the first sample text feature vector of the first sample image are obtained. The first product image information and the first product text description information of the first product sub-image are processed to obtain the first product image feature vector and the first product text feature vector of the first product sub-image. Based on the first sample image feature vector and the instruction information of the first sample image, the generated text description information of the first sample image is obtained. Based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, and the total loss function, the pre-trained multimodal large language model is trained to obtain the multimodal large language model.
[0005] Secondly, this application provides an application method for a commodity processing model, comprising: acquiring a first image to be identified and instruction information of the first image; inputting the first image to be identified and instruction information of the first image into the pre-trained commodity processing model for identification, and obtaining a target output quantity; wherein, the commodity processing model is a multimodal large language model trained using the training method described in the first aspect.
[0006] Thirdly, this application provides a training apparatus for a commodity processing model, comprising: a first acquisition module, configured to acquire a first sample image, wherein the first sample image is any sample image in a labeled training sample set, the first sample image has first sample image information and first sample text description information, and the first sample image has multiple sub-regions, each of the multiple sub-regions corresponding to multiple commodity sub-images, the first commodity sub-image having first commodity image information and first commodity text description information, the first commodity sub-image being any one of the multiple commodity sub-images, and the first commodity text description information including at least the location and brand information of the first commodity sub-image; and a calling module, configured to call a pre-trained multimodal large language model to process the first sample image information and the first sample text description information of the first sample image. The first sample image and the first sample text feature vector of the first sample image are obtained respectively. The first product image information and the first product text description information of the first product sub-image are processed to obtain the first product image feature vector and the first product text feature vector of the first product sub-image respectively. Based on the first sample image feature vector and the instruction information of the first sample image, the generated text description information of the first sample image is obtained. The training module is used to train the pre-trained multimodal large language model based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, and the total loss function to obtain the multimodal large language model.
[0007] Fourthly, this application provides an application device for a commodity processing model, comprising: a second acquisition module for acquiring a first image to be recognized and instruction information of the first image; and a processing module for inputting the first image to be recognized and the instruction information of the first image into a pre-trained commodity processing model for recognition, and acquiring a target output; wherein the commodity processing model is a multimodal large language model trained using the training device described in the third aspect.
[0008] Fifthly, this application provides an electronic device, comprising: a memory for storing code; and a processor connected to the memory for executing the code stored in the memory, so that the electronic device performs the training method as described in the first aspect or the application method as described in the second aspect.
[0009] In a sixth aspect, this application provides a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, can implement the training method as described in the first aspect or the application method as described in the second aspect.
[0010] In a seventh aspect, this application provides a computer program product that, when run on an electronic device, enables the electronic device to implement the training method as described in the first aspect or the application method as described in the second aspect.
[0011] This application's embodiment provides a training method for a multimodal large-scale model for retail scenarios. It adds a retail scenario dataset to the pre-trained multimodal large-scale model. The first sample image not only contains the overall image information but also image information for multiple product sub-images corresponding to multiple sub-regions. The training sample data includes image-level and product-level text descriptions, enabling precise capture of product-level details. This enhances the model's understanding of product-level aspects and improves its ability to understand and perform fine-grained tasks (product-level). The model trained according to this embodiment can be applied to the intelligent retail vertical field, suitable for natural language instruction generation and task execution in retail-related tasks such as product recognition, shelf layout analysis, and brand recognition. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.
[0013] Figure 1 This is a flowchart illustrating the training method for the commodity processing model provided in this application embodiment.
[0014] Figure 2 This is a schematic diagram illustrating the operation of an image encoder and a text encoder.
[0015] Figure 3 This is a flowchart illustrating the application method of the commodity processing model provided in the embodiments of this application.
[0016] Figure 4 This is a schematic diagram of the composition of the training device for the commodity processing model provided in the embodiments of this application.
[0017] Figure 5This is a schematic diagram of the composition of the application device for the commodity processing model provided in the embodiments of this application.
[0018] Figure 6 This is a schematic diagram of the constituent units / partial group units of an electronic device provided in an embodiment of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The same or similar reference numerals are used in the drawings to represent the same or similar modules. It should be understood that the drawings are merely illustrative, and the scope of protection of this application is not limited thereto.
[0020] First, the application scenarios and related technical terms involved in the embodiments of this application will be introduced.
[0021] Large language models (LLMs) are models trained using deep learning techniques that can process large-scale text data. These models learn from large corpora and can predict and generate natural language text, possessing strong language understanding and generation capabilities. They can handle various natural language tasks, such as text classification, intelligent question answering, and dialogue.
[0022] Multimodal large language models (MLLMs) are massive models capable of simultaneously processing and understanding multiple different types of data (such as images, text, and speech). Traditional artificial intelligence models typically can only process one type of data, while MLLMs can utilize multiple types of data for joint modeling and prediction, thereby gaining a more comprehensive understanding and processing of input information.
[0023] With the rapid development of artificial intelligence and deep learning technologies, multimodal large models have been widely applied in various fields. Multimodal large models can be used for natural language instruction generation and task execution in retail-related tasks such as shelf layout analysis. However, in retail scenarios, models like contrastive language-image pre-training (CLIP) typically focus on the overall relationship between text and images, lacking precise capture of details. For example, they cannot identify whether products are correctly positioned or out of stock. Therefore, current multimodal large models are insufficient for fine-grained (product-level) tasks in retail scenarios, failing to meet business requirements.
[0024] Therefore, it is necessary to design a technical solution for multimodal large models to improve the understanding of fine-grained tasks.
[0025] Based on this, this application proposes a training method for a product processing model. The trained multimodal large language model is geared towards retail scenarios, such as physical stores like clothing stores, shoe stores, and shopping malls. The model trained according to this application can be applied to the vertical field of smart retail, and is suitable for natural language instruction generation and task execution in retail-related tasks such as product recognition, shelf layout analysis, and brand recognition. This helps improve the model's understanding and execution performance of fine-grained tasks (product level).
[0026] In some implementations, the architecture of a multimodal large language model can adopt the architecture of a large language and vision assistant (LLAVA). The model can include three components: a visual encoder, a visual-language connector, and a large language model (LLM). Training is divided into three phases: the first phase is the pre-training phase, which trains the visual encoder; the second phase trains the visual-language connector; and the third phase trains the connector and the LLM.
[0027] The following is combined with Figure 1 The training method of the commodity processing model in the embodiments of this application is described in detail. For example... Figure 1 As shown, the training method of the commodity processing model in this application embodiment mainly includes steps S110 to S140, which are described in detail below.
[0028] It should be noted that the sequence number of each step in the embodiments of this application does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0029] In step S110, a first sample image is acquired. The first sample image is any sample image from the labeled training sample set. The first sample image has first sample image information and first sample text description information, and the first sample image has multiple sub-regions, each sub-region corresponding to multiple product sub-images. The first product sub-image has first product image information and first product text description information, and the first product sub-image is any one of the multiple product sub-images. The first product text description information includes at least the location information and brand information of the first product sub-image.
[0030] Compared to focusing only on the overall relationship between text and images, the training sample data includes image-level text descriptions and product-level text descriptions, which can accurately capture product-level details and help improve the model's ability to understand fine-grained tasks (product level).
[0031] Conventional training sample images only include the overall relationship between text and images, without labeling the product regions within the images. Before acquiring the first sample image, the training method in this embodiment may further include: processing the initial first sample image to extract multiple sub-regions corresponding to multiple product sub-images; acquiring the first product image information of the first product sub-image, labeling the first product sub-image, and determining the text description information corresponding to the first product sub-image; and acquiring the labeled first sample image. That is, before training, product-level image region extraction and text labeling are first performed on the training sample set. It is understood that the multiple product sub-images corresponding to multiple sub-regions can be the same or different. For example, the product sub-image corresponding to the first sub-region is an image of a bottle of purified water, and the product sub-image corresponding to the second sub-region can be an image of a can of milk powder.
[0032] During the pre-training phase, the training sample set can be a pre-labeled set of retail product images. Each sample image in this set is labeled with the location and brand information of each product. Based on any given sample image, a text description can be generated for each sub-region corresponding to a product. For example, the first product sub-image could be a bottled water image, and its text description could be "A 500ml bottle of Wahaha purified water".
[0033] The training sample set can include a public dataset and a dedicated dataset. A public dataset could be, for example, Conceptual 12M (CC12M), which is a large and diverse image-text pair dataset. A dedicated dataset could be a retail product dataset suitable for retail scenarios.
[0034] In some specific implementations, for the CC12M image-text dataset, each image has a corresponding text description. Since multiple sub-regions (targets) within the image are not labeled, a recognize anything model (RAM) can be used to generate image labels, and then GroundingDINO can be used to generate bounding boxes for each label. For example, RAM can be used to generate image labels such as "grass, puppy," and GroundingDINO can be used to generate bounding boxes for each label. GroundingDINO only needs the label name of the item (product) to be detected to output the corresponding bounding box. Bounding boxes with a confidence level below 0.3 can be filtered out. For each bounding box, after cropping, a general large language model such as Qwen2-vl can be used to generate the corresponding text title. GroundingDINO adopts the Transformer architecture of the object detector DINO and borrows the pre-training method of multimodal GLIP. After deeply fusing language and visual information, it can detect any target based on the text description.
[0035] In a dedicated retail product dataset, each product sub-image can be described using its location coordinates, brand, category, and name. For example, Qwen2-vl can be used to generate image-level titles (text descriptions). For each bounding box, after cropping, Qwen2-vl is used to generate the corresponding text title, and the labeled category information can be used as a prompt during the generation process.
[0036] In step S120, a pre-trained multimodal large language model is invoked to process the first sample image information and the first sample text description information of the first sample image, obtaining the first sample image feature vector and the first sample text feature vector of the first sample image, respectively. Similarly, the first product image information and the first product text description information of the first product sub-image are processed to obtain the first product image feature vector and the first product text feature vector of the first product sub-image, respectively.
[0037] In some implementations, the first product image feature vector of the first product sub-image is extracted from the feature vector of the first sample image using region of interest alignment (ROIAlign). For example, the initial first sample image can be a 224*224 resolution image, and the feature vector of the first sample image can be 196*768. Assuming the region corresponding to the first product in the initial first sample image is "100, 100, 150, 150", then the ROIAlign method can be used to extract the features corresponding to this region as the first product image feature vector of the first product sub-image.
[0038] In some implementations, the architecture of a pre-trained multimodal large model can adopt CLIP, which includes a visual encoder and a text encoder.
[0039] A visual encoder is a neural network model used to abstract and extract information from each pixel of an input image, transforming it into a fixed-length vector representation. Visual encoders typically consist of alternating convolutional and pooling layers. The input to a visual encoder is usually the original image pixel matrix, and the output can be a fixed-length vector representation or a fixed-size visual feature map.
[0040] For example, calling the visual encoder can process the first sample image information of the first sample image to obtain the first sample image feature vector. Similarly, calling the visual encoder can process the first product image information of the first product sub-image to obtain the first product image feature vector of the first product sub-image.
[0041] Visual encoders typically take inputs at a fixed resolution, such as 224x224. Image compression reduces resolution, leading to significant information loss and making it difficult for the model to focus on details. Some implementations add higher resolution to the image encoder. Specifically, for the first high-resolution sample image, the LLaVA-1.5-HD method can be used to intelligently segment the high-resolution image into multiple patches, which are then embedded and merged with the low-resolution image. This helps improve the resolution of the extracted image feature vectors, allowing the visual encoder to learn more fine-grained features.
[0042] Multimodal large models often include text encoders. Text encoders, also known as text embeddings or text representations, use neural network models to transform input text sequences into fixed-length vector representations with specific semantic features. By converting text information into vector representations, text encoders can directly transform text information into numerical form for processing and, through deep learning models, uncover the inherent patterns in text data, thereby better representing the text information.
[0043] Figure 2 This is a schematic diagram illustrating the operation of an image encoder and a text encoder. (Example) Figure 2 As shown, within the CLIP algorithm framework, the image-text pair (image, text) consists of an image of a puppy and the phrase "feed the puppy PiePie peppers," respectively. Although "PiePie" in the text is the puppy's name and cannot be perfectly aligned with the image, the entire sentence can still be semantically aligned with the image. For the image-text pair, the image data is processed by an image encoder to obtain its representation I. i The text data is processed by a text encoder to obtain its representation T. i Then, the gradients of the image encoder and the text encoder are calculated by comparing the learning loss function for optimization.
[0044] The text encoder of a pre-trained multimodal large language model can be invoked to process the first sample text description information of the first sample image, obtaining the first sample text feature vector. The first product text description information of the first product sub-image is also processed to obtain the first product text feature vector of the first product sub-image. The text encoder is then used to generate image-level (dimension-wise) and product-level (dimension-wise) text label embeddings, and subsequently, image-level contrast loss and product-level contrast loss are calculated.
[0045] For example, calculating the contrast loss of the overall image information after text and image encoders can be considered an image-level loss. For instance, for an image of a shelf, the corresponding text description could be "This is a supermarket shelf, and there are many drinks on it." If we want the model to focus on smaller sub-regions, i.e., product sub-regions, we need to add corresponding text descriptions to these smaller sub-regions during training. Therefore, product-level embeddings are added during training to calculate the product-level contrast loss using product-level text feature vectors (embeddings) and image feature vectors.
[0046] After pre-training, the resulting multimodal large model includes a text encoder and a visual encoder. The visual encoder is then fed into the multimodal large language model. However, the multimodal large language model itself cannot directly understand the image feature vector representation space of the CLIP model. Therefore, a fusion module is needed to perform spatial mapping, mapping the input features to a representation space that the LLM can understand.
[0047] The vision-language connector, also known as the projection layer or fusion module, primarily maps the visual feature vectors output by the visual encoder onto the input space of a large language model. This projects visual markers into the language model's space, enabling effective fusion and interaction between visual and textual feature vectors, thus facilitating better integration and understanding of multimodal information. A multilayer perceptron (MLP) can be used as the vision-language connector.
[0048] During the training phase of the visual-language connector, all visual encoders and the large language model can be frozen. For example, based on the CC3M dataset, only the visual-language connector can be updated. The CC3M dataset is a large dataset containing approximately 3 million image-text pairs.
[0049] In step S130, the generated text description information of the first sample image is obtained based on the feature vector of the first sample image and the instruction information of the first sample image.
[0050] The instruction information for the first sample image could be, for example, information such as the brand, location, and type of all goods in the first sample image. For instance, a pre-trained multimodal large language model could be used to obtain the generated text description information for the first sample image based on its feature vector and the instruction information.
[0051] The visual-language connector transforms and maps image feature vectors to match the semantic space of a language model, and then concatenates or fuses them with text feature vectors. The fused multimodal features are input into a large language model, which generates corresponding text outputs based on these features and the previous text inputs, serving as a comprehensive response to the image and text inputs.
[0052] In step S140, the pre-trained multimodal large language model is trained based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, the objective function, and the total loss function to obtain the multimodal large language model.
[0053] During the pre-training phase, a target function for the product sub-image dimension is introduced for supervision, and the contrast loss between the text feature vector and the image feature vector of the product sub-image dimension is calculated to optimize the model parameters.
[0054] The contrast loss for the overall image dimension can be calculated using the text feature vector (embedding) and image feature vector of the image dimension. The contrast loss for the product sub-image dimension can be calculated using the text feature vector and image feature vector of the product sub-image dimension. Therefore, the formula for the total loss function is:
[0055] L=αL image +βL object
[0056] Among them, L image For overall image loss, L object The loss is denoted as α, where α and β are weighting coefficients.
[0057] For example, the parameters of the text feature vector can be updated using backpropagation and gradient descent. Based on the loss function value, the gradient of the image feature vector is calculated, and the parameter values of the image feature vector are updated according to the gradient, thus training and optimizing the parameters of a pre-trained multimodal large model.
[0058] Fine-tuning of the vision-language connector and the Large Language Model (LLM) is possible. Large model fine-tuning has become a key technique for improving model performance on specific tasks. After vision-language alignment pre-training, visual instruction fine-tuning is required, which helps improve the ability of the multimodal large language model to follow instructions and solve tasks. Generally, the input to visual instruction fine-tuning includes an image and a task description text, and the output is a corresponding text response.
[0059] In some implementations, the fine-tuning dataset can be specified using LLaVA-Instruct-150K. This answers questions in a general format but cannot precisely output instructions specific to retail scenarios. In other implementations, the fine-tuning dataset can be specified using LLaVA-Instruct-150K, along with additional dedicated fine-tuning datasets built for retail scenarios. This helps the large multimodal model accurately output instructions for retail scenarios, responding precisely in the desired format, rather than outputting generic, unnecessary content.
[0060] In some implementations, the method of this application embodiment may further include: constructing a dedicated fine-tuning dataset for retail scenarios. Based on the pre-labeled training sample set and image dataset for product recognition, the constructed dedicated fine-tuning dataset includes at least some or all of the following tasks: counting, product localization, and brand recognition. Adding dedicated retail scenario data and a dedicated fine-tuning dataset to the fine-tuning multimodal large model helps enhance the model's understanding and representation capabilities of sub-region and product-level information during the pre-training stage.
[0061] The counting task includes at least the following: 1) how many shelves there are in the image; 2) how many items are on each shelf; 3) how many items are on the Xth shelf; 4) how many empty areas (without items) are in the image, where X is an integer not less than 1.
[0062] The task of product positioning includes at least: 1) What are the coordinates of the first brand product; 2) What is the product to the right of the first brand product; 3) Give the location information of all products; 4) What is the product in the Y-th column of the X-th layer, where Y is an integer not less than 1, and the first brand is any one of the multiple brands in the retail scenario.
[0063] The task of brand recognition includes at least the following: 1) how many brands of goods are in the image; 2) how many brands of goods are on the Xth shelf; 3) how many different kinds of goods are in the image; and 4) whether there are any goods of a specified brand on the kth shelf, where k is an integer not less than 1.
[0064] Step S140, which trains the pre-trained multimodal large language model based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, the objective function, and the total loss function, may include: training the pre-trained multimodal large language model based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, the objective function, the total loss function, and a dedicated fine-tuning dataset.
[0065] The main construction process of the dedicated fine-tuning dataset is described below.
[0066] Select a dataset (image dataset) labeled with shelf location, product brand, product name, and product location. For each image, identify and count the shelf number and shelf layer based on shelf location; count the number of products on each shelf layer based on product location information; count the number of brands and product categories in the entire image based on product category and brand, and count the number of brands and product categories on each shelf layer (row); assign a location number to each product in the image, convert the location coordinates to [0-1], perform normalization processing, and count the number of empty spaces on each shelf in each image; and form a statistical database from the above statistical data.
[0067] For each task, an instruction fine-tuning dataset is constructed. For example, the task of product positioning is used as an example for illustration.
[0068] Regarding question 1, what are the coordinates of the top-selling brand's products? We can query the statistical database to obtain the answer and construct the corresponding question pair.
[0069] Regarding question 2, what is the product to the right of the first brand's product?: Obtain the product information database of the selected image, select a product to replace the first brand's product, select a location referent from [left, right, top, bottom], obtain the answer based on the location referent and the product's location number, and construct the corresponding question pair.
[0070] For question 3, given the location information of all products: query the database, obtain the answer, and construct the corresponding question pair.
[0071] For question 4, what is the product in column Y of shelf X?: Obtain the product information database of the selected image, randomly select and replace X and Y according to the shelf layer and column number, obtain the answer, and construct the corresponding question pair.
[0072] This application's embodiment provides a training method for a multimodal large-scale model for retail scenarios. It adds a retail scenario dataset to the pre-trained multimodal large-scale model. The first sample image not only contains the overall image information but also image information for multiple product sub-images corresponding to multiple sub-regions. The training sample data includes image-level and product-level text descriptions, enabling precise capture of product-level details. This enhances the model's understanding of product-level aspects and improves its ability to understand and perform fine-grained tasks (product-level). The model trained according to this embodiment can be applied to the intelligent retail vertical field, suitable for natural language instruction generation and task execution in retail-related tasks such as product recognition, shelf layout analysis, and brand recognition.
[0073] This application provides a method for applying a commodity processing model. Figure 3 This is a flowchart illustrating the application method of the commodity processing model provided in the embodiments of this application. For example... Figure 3 As shown, the application method of the commodity processing model in this application embodiment can mainly include steps S310 to S320, which are described in detail below.
[0074] In step S310, a first image to be identified and instruction information for the first image are obtained. The first image is any one of the multiple images to be identified.
[0075] In step S320, the first image to be recognized and the instruction information of the first image are input into the pre-trained commodity processing model for recognition to obtain the target output. The commodity processing model is a multimodal large language model trained using any of the training methods described above.
[0076] In this embodiment, the first image not only has image information of the overall image, but also image information of multiple product sub-images corresponding to multiple sub-regions. The first image contains image-level text descriptions and product-level text descriptions, which can accurately capture the details at the product level, enhance the model's understanding of the product level, and help improve the model's understanding and execution of fine-grained tasks (product level).
[0077] The above text combined Figures 1-3 The method embodiments of this application have been described in detail below, in conjunction with... Figures 4 to 5 The apparatus embodiments of this application are described in detail below. It should be understood that the descriptions of the apparatus embodiments correspond to the descriptions of the method embodiments; therefore, any parts not described in detail can be referred to the foregoing method embodiments.
[0078] This application also provides a training device for a commodity processing model. Figure 4 This is a schematic diagram of the composition of the training device for the commodity processing model provided in the embodiments of this application. Figure 4 As shown, the training device 400 for the commodity processing model includes: a first acquisition module 410, a calling module 420, and a training module 430.
[0079] The first acquisition module 410 is used to acquire a first sample image, which is any sample image in the labeled training sample set. The first sample image has first sample image information and first sample text description information, and the first sample image has multiple sub-regions, each corresponding to a multiple product sub-image. Each first product sub-image has first product image information and first product text description information. The first product sub-image is any one of the multiple product sub-images, and the first product text description information includes at least the location and brand information of the first product sub-image.
[0080] The calling module 420 is used to call the pre-trained multi-model large language model to process the first sample image information and the first sample text description information of the first sample image, obtaining the first sample image feature vector and the first sample text feature vector of the first sample image; it also processes the first product image information and the first product text description information of the first product sub-image to obtain the first product image feature vector and the first product text feature vector of the first product sub-image. Based on the first sample image feature vector and the instruction information of the first sample image, the generated text description information of the first sample image is obtained.
[0081] The training module 430 is used to train the pre-trained multimodal large language model based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, and the total loss function, so as to obtain the multimodal large language model.
[0082] Optionally, the training device 400 also includes a preprocessing module. The preprocessing module processes the initial first sample image to extract multiple product sub-images from the first sample image; labels the first product sub-images to determine their textual descriptions; and obtains the labeled first sample image. Alternatively, the preprocessing module processes the initial first sample image to extract multiple sub-regions corresponding to the multiple product sub-images; labels the multiple sub-regions; obtains the textual descriptions corresponding to the multiple product sub-images; obtains the image information corresponding to the multiple product sub-images; and obtains the labeled first sample image.
[0083] Optionally, the training device 400 may further include a construction module. The construction module is used to construct a dedicated fine-tuning dataset for retail scenarios. The training module 430 is used to train the pre-trained multimodal large language model based on first sample image information, first sample text description information, generated text description information of the first sample image, first product image feature vector, first product text feature vector, instruction information of the first sample image, total loss function, and the dedicated fine-tuning dataset. The dedicated fine-tuning dataset constructed by the construction module includes at least some or all of the following tasks: counting, product location, and brand recognition.
[0084] The counting task includes at least the following: how many shelves are there in the image, how many products are on each shelf, how many products are on the Xth shelf, and how many empty areas (no products) are in the image, where X is an integer not less than 1. The product location task includes at least the following: what are the coordinates of the first brand product, what product is to the right of the first brand product, the location information of all products, and what product is in the Yth column of the Xth shelf, where Y is an integer not less than 1, and the first brand is any of the multiple brands in the retail scenario. The brand recognition task includes at least the following: how many brands of products are in the image, how many brands of products are on the Xth shelf, how many different products are in the image, and whether there is a product of a specified brand on the i-th shelf, where i is an integer not less than 1.
[0085] This application also provides an application device for a commodity processing model. Figure 5 This is a schematic diagram of the composition of the application device for the commodity processing model provided in the embodiments of this application. For example... Figure 5 As shown, the application device 500 of the commodity processing model includes: a second acquisition module 510 and a processing module 520.
[0086] The second acquisition module 510 is used to acquire the first image to be identified and the instruction information of the first image.
[0087] The processing module 520 is used to input the first image to be recognized and the instruction information of the first image into a pre-trained product processing model for recognition, and to obtain the target output. The product processing model is a multimodal large language model trained using any of the training devices described above.
[0088] Figure 6 This is a schematic diagram of the constituent units / partial constituent units of an electronic device provided in an embodiment of this application. For example... Figure 6 As shown, the electronic device 600 may include: a memory 610 and at least one processor 620.
[0089] The memory 610 is used to store code or computer programs.
[0090] The processor 620 is connected to the memory 610 and is used to execute the code or computer program stored in the memory 610 to control the electronic device 600 to perform the processing method as described above.
[0091] For example, a computer program may be divided into one or more modules / units, one or more of which are stored in memory 610 and executed by processor 620 to complete this application.
[0092] Those skilled in the art will understand that Figure 6This is merely an example of electronic device 600 and does not constitute a limitation on the electronic device. Electronic device 600 may include more or fewer components than shown, or combine certain components, or use different components.
[0093] The processor 620 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or it can be any conventional processor.
[0094] The electronic device 600 provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effect are similar, so it will not be described again here.
[0095] This application also provides a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps in the various method embodiments described above.
[0096] This application provides a computer program product that, when run on an electronic device, enables the electronic device to perform the steps described in the various method embodiments above.
[0097] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above method embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate form. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a photographic device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, and optical data storage devices. The computer-readable storage medium mentioned in this application can be a non-volatile storage medium; in other words, it can be a non-transient storage medium.
[0098] It should be noted that the information collection process (such as the facial image collection process, fingerprint collection process, etc.) / feature extraction process involved in this application is carried out with the user's knowledge and permission, that is, the information collection process / feature extraction process complies with the requirements of laws and regulations and does not constitute an act that harms the public interest.
[0099] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0100] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0101] In the embodiments provided in this application, it should be understood that the disclosed apparatus / device and method can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0102] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0103] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0104] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0105] Furthermore, in the description of this application and the appended claims, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A training method for a commodity processing model, characterized in that, include: A first sample image is obtained, which is any sample image in the labeled training sample set. The first sample image has first sample image information and first sample text description information. The first sample image has multiple sub-regions, which correspond one-to-one with multiple product sub-images. The first product sub-image has first product image information and first product text description information. The first product sub-image is any one of the multiple product sub-images. The first product text description information includes the location and brand information of the first product sub-image. The pre-trained multimodal large language model is invoked to process the first sample image information and the first sample text description information of the first sample image to obtain the first sample image feature vector and the first sample text feature vector of the first sample image, respectively. The first product image information and the first product text description information of the first product sub-image are also processed to obtain the first product image feature vector and the first product text feature vector of the first product sub-image, respectively. Based on the feature vector of the first sample image and the instruction information of the first sample image, obtain the generated text description information of the first sample image; Based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, and the total loss function, the pre-trained multimodal large language model is trained to obtain the multimodal large language model.
2. The training method according to claim 1, characterized in that, Before acquiring the first sample image, the training method further includes: The initial first sample image is processed to extract multiple product sub-images from the first sample image; The first product sub-image is labeled to determine the text description information of the first product sub-image; Obtain the first labeled sample image.
3. The training method according to claim 1, characterized in that, Also includes: Construct a dedicated fine-tuning dataset for retail scenarios, wherein the instruction information of the first sample image is one of the instructions in the dedicated fine-tuning dataset; The step of training the pre-trained multimodal large language model based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, and the total loss function includes: The pre-trained multimodal large language model is trained based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, the total loss function, and the dedicated fine-tuning dataset. The dedicated fine-tuning dataset includes some or all of the following tasks: counting, product location, and brand recognition; The counting tasks include: how many shelves there are in the image, how many items are on each shelf, how many items are on the Xth shelf, and how many empty areas are in the image, where X is an integer not less than 1; The task of product positioning includes: what are the coordinates of the first brand product, what is the product to the right of the first brand product, providing the location information of all products, and what is the product in the Xth layer and Yth column, where Y is an integer not less than 1, and the first brand is any one of the multiple brands in the retail scenario. The brand recognition task includes: how many brands of goods are in the image, how many brands of goods are on the Xth shelf, how many different kinds of goods are in the image, and whether there are any goods of a specified brand on the kth shelf, where k is an integer not less than 1.
4. A method for applying a commodity processing model, characterized in that, include: Obtain the first image to be identified and the instruction information of the first image; The first image to be identified and the instruction information of the first image are input into the pre-trained commodity processing model for identification, and the target output value is obtained. The commodity processing model is a multimodal large language model trained using the training method described in any one of claims 1-3.
5. A training device for a commodity processing model, characterized in that, include: The first acquisition module is used to acquire a first sample image, which is any sample image in the labeled training sample set. The first sample image has first sample image information and first sample text description information, and the first sample image has multiple sub-regions. Each of the multiple sub-regions corresponds to a multiple product sub-images. The first product sub-image has first product image information and first product text description information. The first product sub-image is any one of the multiple product sub-images. The first product text description information includes the location and brand information of the first product sub-image. The calling module is used to call the pre-trained multimodal large language model to process the first sample image information and the first sample text description information of the first sample image to obtain the first sample image feature vector and the first sample text feature vector of the first sample image, respectively. It also processes the first product image information and the first product text description information of the first product sub-image to obtain the first product image feature vector and the first product text feature vector of the first product sub-image, respectively. Based on the feature vector of the first sample image and the instruction information of the first sample image, obtain the generated text description information of the first sample image; The training module is used to train the pre-trained multimodal large language model based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, and the total loss function, so as to obtain the multimodal large language model.
6. The training device according to claim 5, characterized in that, The training device also includes: The preprocessing module is used to process the initial first sample image, extract multiple product sub-images from the first sample image; annotate the first product sub-images to determine the text description information of the first product sub-images; and obtain the annotated first sample image.
7. The training device according to claim 5, characterized in that, Also includes: A construction module is used to construct a dedicated fine-tuning dataset for retail scenarios, wherein the instruction information of the first sample image is one of the instructions in the dedicated fine-tuning dataset; The training module is used to train the pre-trained multimodal large language model based on the first sample image information, the first sample text description information, the generated text description information of the first sample image, the first product image feature vector, the first product text feature vector, the instruction information of the first sample image, the total loss function, and the dedicated fine-tuning dataset. The dedicated fine-tuning dataset includes some or all of the following tasks: counting, product location, and brand recognition; The counting tasks include: how many shelves there are in the image, how many items are on each shelf, how many items are on the Xth shelf, and how many empty areas are in the image, where X is an integer not less than 1; The task of product positioning includes: what are the coordinates of the first brand product, what is the product to the right of the first brand product, providing the location information of all products, and what is the product in the Xth layer and Yth column, where Y is an integer not less than 1, and the first brand is any one of the multiple brands in the retail scenario. The brand recognition task includes: how many brands of goods are in the image, how many brands of goods are on the Xth shelf, how many different kinds of goods are in the image, and whether there are any goods of a specified brand on the kth shelf, where k is an integer not less than 1.
8. An application device for a commodity processing model, characterized in that, include: The second acquisition module is used to acquire the first image to be identified and the instruction information of the first image; The processing module is used to input the first image to be identified and the instruction information of the first image into the pre-trained commodity processing model for identification and to obtain the target output quantity; The commodity processing model is a multimodal large language model trained using the training device described in any one of claims 5-7.
9. An electronic device, characterized in that, include: Memory, used to store code; A processor, connected to the memory, is configured to execute code stored in the memory to cause the electronic device to perform the training method as described in any one of claims 1-3 or the application method as described in claim 4.
10. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed, is used to implement the training method as described in any one of claims 1-3 or the application method as described in claim 4.
Citation Information
Patent Citations
Text generation method and apparatus
WO2024046189A1
Visual question answering method and apparatus, electronic device and storage medium
WO2024164616A1