Image data augmentation method, electronic device, medium and product
By automatically generating augmented images, the problem of relying on manual shooting and annotation for training data of smart cabinet item recognition models has been solved, achieving low-cost and efficient model training and updating.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING GENKI FOREST BEVERAGE CO LTD
- Filing Date
- 2022-04-27
- Publication Date
- 2026-06-02
AI Technical Summary
The training data requirements of existing smart locker item recognition models rely on manual multi-angle photography and annotation, resulting in high labor costs and slow model update speed, making it impossible to quickly adapt to changes in the types of items in different smart lockers.
The client requests multiple images of the display case and the items to be displayed from different angles from the server. These images are then pasted into the bounding box to automatically generate augmented images, reducing manual intervention and quickly providing training data.
It reduces labor costs, improves model training efficiency and update speed, and can quickly adapt to changes in the types of items inside the smart cabinet.
Smart Images

Figure CN117036839B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and specifically to an image data augmentation method, electronic device, medium, and product. Background Technology
[0002] Given rising labor costs, an increasing number of smart vending machines are appearing on the market to enable automated vending. Unlike traditional mechanical vending machines, existing smart vending machines are lightweight, simple in structure, inexpensive, and user-friendly. Currently, these smart vending machines primarily use item recognition models to identify the items being taken out, thus achieving automated vending. A good item recognition model requires sufficient training data as a learning foundation; this training data mainly includes images of the smart vending machine's displays. Since different smart vending machines display different types of items, the item recognition models in vending machines displaying different types of items need to learn from different display images. The production of these display images and the labeling of items within them require manual multi-angle photography and manual labeling. Summary of the Invention
[0003] This disclosure provides an image data augmentation method, electronic device, medium, and product.
[0004] In a first aspect, this disclosure provides an image data augmentation method.
[0005] Specifically, the image data augmentation method includes:
[0006] Receive an image request message sent by a client, the image request message carrying a target display cabinet identifier and parameters of the item to be displayed, the parameters of the item to be displayed including the identifier of the item to be displayed;
[0007] In response to receiving the image request message, multiple images of the target display cabinet from different angles are obtained based on the target display cabinet identifier, and multiple images of the item to be displayed from different angles are obtained based on the item to be displayed identifier; wherein, the display cabinet images display the bounding boxes of the items to be displayed at each position in the target display cabinet;
[0008] Paste the image of the item to be displayed within the corresponding bounding box of the display case image to generate multiple augmented images;
[0009] The augmented images are returned to the client.
[0010] Secondly, this disclosure provides an image data augmentation method.
[0011] Specifically, the image data augmentation method includes:
[0012] Acquire images of the material display cabinets identified by each display cabinet label from multiple different angles; the display cabinet images display the bounding boxes of the items to be displayed at each position in the target display cabinet;
[0013] Acquire images of the material items identified by each item icon from multiple different angles;
[0014] Receive a material request message sent by the server, the material request message carrying the target display cabinet identifier and the item identifier to be displayed;
[0015] From the images of the material display cases identified by each display case label, find multiple display case images of the target display case identified by the target display case label from different angles;
[0016] From the material items identified by each item identifier, find the images of the item to be displayed identified by the item identifier from multiple different angles among the item images from multiple different angles;
[0017] The server returns multiple images of the target display case from different angles and multiple images of the items to be displayed from different angles, so that the server can generate an augmented image based on the multiple images of the target display case from different angles and the multiple images of the items to be displayed from different angles.
[0018] Thirdly, this disclosure provides an image data augmentation device.
[0019] Specifically, the image data augmentation device includes:
[0020] The first receiving module is configured to receive an image request message sent by a client. The image request message carries a target display cabinet identifier and parameters of the item to be displayed, including the identifier of the item to be displayed.
[0021] The first acquisition module is configured to, in response to receiving the image request message, acquire multiple images of the target display cabinet from different angles based on the target display cabinet identifier, and acquire multiple images of the item to be displayed from different angles based on the item to be displayed identifier; wherein the display cabinet images display the bounding boxes of the items to be displayed at each position in the target display cabinet.
[0022] The generation module is configured to paste the image of the item to be displayed into the corresponding bounding box of the display case image to generate multiple augmented images;
[0023] The first sending module is configured to return the plurality of augmented images to the client.
[0024] Fourthly, this disclosure provides an image data augmentation device.
[0025] Specifically, the image data augmentation device includes:
[0026] The fourth acquisition module is configured to acquire images of the material display cabinets identified by each display cabinet identifier from multiple different angles; the display cabinet images display the bounding boxes of the items to be displayed at each position in the target display cabinet;
[0027] The fifth acquisition module is configured to acquire images of the material items identified by each item identifier from multiple different angles;
[0028] The second receiving module is configured to receive a material request message sent by the server, the material request message carrying the target display cabinet identifier and the item identifier to be displayed;
[0029] The first search module is configured to search for multiple different angle images of the target display cabinet identified by the target display cabinet from the display cabinet images of the material display cabinets identified by each display cabinet identifier.
[0030] The second search module is configured to search for images of the item to be displayed, identified by the item identifier, from multiple different angles of the material items identified by each item identifier.
[0031] The second sending module is configured to return multiple images of the target display cabinet from different angles and multiple images of the items to be displayed from different angles to the server, so that the server can generate an augmented image based on the multiple images of the target display cabinet from different angles and the multiple images of the items to be displayed from different angles.
[0032] Fifthly, this disclosure provides an electronic device including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement any of the methods described above.
[0033] Sixthly, this disclosure provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement any of the methods described above.
[0034] In a seventh aspect, this disclosure provides a computer program product including computer instructions that, when executed by a processor, implement any of the methods described above.
[0035] The above technical solution allows the client to send an image request message to the server when training the model needed for the display case. The client requests the server to return the images required for training the item recognition model. The image request message carries the target display case identifier and parameters of the items to be displayed, including the item identifier. Upon receiving the image request message, the server can obtain multiple images of the target display case from different angles based on the target display case identifier. These images display the bounding boxes of the items to be displayed at various locations within the target display case. The server then obtains multiple images of the items to be displayed from different angles based on the item identifier and pastes these images into the corresponding bounding boxes of the display case images, generating multiple augmented images. These augmented images are then returned to the client for model training. This approach automatically generates a large number of augmented images needed for model training by initially acquiring only multiple images of the display case from different angles and images of various items to be displayed. This reduces labor costs and allows for rapid generation of the large number of images required for model training based on the client's training needs, resulting in faster model updates.
[0036] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0037] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments, taken in conjunction with the accompanying drawings. In the drawings:
[0038] Figure 1A A flowchart illustrating an image data augmentation method applied to a server according to an embodiment of the present disclosure is shown.
[0039] Figure 1B A schematic diagram showing an augmented image according to an embodiment of the present disclosure;
[0040] Figure 1C A schematic diagram showing an augmented image according to an embodiment of the present disclosure;
[0041] Figure 2 A flowchart illustrating an image data augmentation method applied to a material storage terminal according to an embodiment of the present disclosure is shown.
[0042] Figure 3 A schematic diagram of the overall flow of an image data augmentation method according to an embodiment of the present disclosure is shown.
[0043] Figure 4 A schematic diagram of the overall process of another image data augmentation method according to an embodiment of the present disclosure is shown;
[0044] Figure 5 This is a structural block diagram of an image data augmentation device applied to a server according to an embodiment of the present disclosure;
[0045] Figure 6 This is a structural block diagram of an image data augmentation device applied to a material storage terminal according to an embodiment of the present disclosure;
[0046] Figure 7 A schematic structural block diagram of an electronic device according to an embodiment of the present disclosure is shown;
[0047] Figure 8 This is a schematic diagram of the structure of a computer system suitable for implementing an image data augmentation method according to an embodiment of the present disclosure. Detailed Implementation
[0048] In the following, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings to enable those skilled in the art to readily implement them. Furthermore, for clarity, portions unrelated to the description of the exemplary embodiments have been omitted from the drawings.
[0049] In this disclosure, it should be understood that terms such as “comprising” or “having” are intended to indicate the presence of features, figures, steps, behaviors, components, parts or combinations thereof disclosed in this specification, and do not preclude the possibility of the presence or addition of one or more other features, figures, steps, behaviors, components, parts or combinations thereof.
[0050] In this disclosure, it should be understood that the terms "upper", "lower", "vertical", "horizontal", "inner", "outer", "top", "bottom", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0051] In this disclosure, it should be understood that "multiple" means two or more, "greater than," "less than," and "exceeding" are understood to exclude the stated number, while "above," "below," and "within" are understood to include the stated number. If "first" or "second" is used, it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance, or implicitly indicating the number of indicated technical features, or implicitly indicating the order of the indicated technical features.
[0052] It should also be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0053] As mentioned above, due to rising labor costs, more and more smart vending machines are appearing on the market to achieve automated vending functions. Unlike traditional mechanical vending machines, existing smart vending machines are lightweight, simple in structure, inexpensive, and user-friendly. Currently, these smart vending machines primarily use item recognition models to identify the items being taken out, thus enabling automated vending. A good item recognition model requires sufficient training data as a learning foundation; this training data mainly includes images of the smart vending machine's displays. Since different smart vending machines display different types of items, the item recognition models in smart vending machines displaying different types of items need to learn from different display images. The production of these display images and the labeling of items within them require manual multi-angle photography and manual labeling.
[0054] Typically, training a good item recognition model requires hundreds of thousands of display images. Manually taken images are insufficient to meet the basic requirements for model training. Furthermore, manually collected images are too incomplete, requiring significant time and manpower to change the positions and angles of the displayed items to cover all possible outcomes. If the types of items displayed in the smart cabinet change, substantial manual time must be spent updating the training images before the item recognition model can be updated, resulting in slow model updates. Therefore, how to acquire training images quickly and cost-effectively has become a pressing issue.
[0055] To address the aforementioned shortcomings, this disclosure proposes an image data augmentation method. When the client trains the model required for a display case, it can send an image request message to the server, requesting the server to return the images needed to train the item recognition model. The image request message carries a target display case identifier and parameters of the items to be displayed, including the item identifier. Upon receiving the image request message, the server can acquire multiple images of the target display case from different angles based on the target display case identifier. These display case images display the bounding boxes of the items to be displayed at various locations within the target display case. Based on the item identifier, the server can acquire multiple images of the items to be displayed from different angles and paste these images into the corresponding bounding boxes of the display case images, generating multiple augmented images. These augmented images are then returned to the client for model training. This approach automatically generates a large number of augmented images needed for model training by initially acquiring only multiple images of the display case from different angles and images of various items to be displayed. This reduces labor costs and allows for rapid generation of the large number of images required for model training based on the client's training needs, resulting in faster model updates.
[0056] The details of the embodiments of this disclosure are described in detail below through specific examples.
[0057] This disclosure provides an image data augmentation method. Figure 1A This diagram illustrates a flowchart of an image data augmentation method applied to a server according to an embodiment of the present disclosure. The method is applied to the server side, such as... Figure 1A As shown, the image data augmentation method may include the following steps:
[0058] In step S101, an image request message sent by the client is received. The image request message carries the target display cabinet identifier and the parameters of the item to be displayed, including the identifier of the item to be displayed.
[0059] In one possible implementation, the client refers to the client that trains the item recognition model for the display cabinet. When a user acquires a new display cabinet to display items of a certain category, such as various bottled beverages, the client can train an item recognition model for that display cabinet. Alternatively, when a user changes the items displayed in an existing display cabinet, such as previously displaying bottled beverages of brand A and now displaying bottled beverages of brand B, the client needs to update the item recognition model trained for that display cabinet so that it can recognize the placement and removal of bottled beverages of brand B.
[0060] In one possible implementation, the target display case refers to a display case that requires the application of an item recognition model trained by the client, and the item to be displayed refers to the item that the user intends to display in the target display case. The display image required by the client to train the item recognition model is an image of the corresponding item to be displayed within the target display case. Therefore, the client can obtain the target display case identifier and the identifier of the item to be displayed within the display case, and send these two identifiers along with an image request message to the server, so that the server can return a display image of the corresponding item to be displayed within the target display case. The target display case identifier can identify the brand, model, color, etc. of the target display case; alternatively, it can also identify the number of item rows and / or the number of shelves on each shelf within the target display case. Each display case identifier can identify a display case template, which includes a set of multiple display case images from different angles. The item to be displayed identifier can identify the type of item to be displayed.
[0061] It should be noted that the client can obtain the display case identifier and the identifier of the item to be displayed through user input.
[0062] In step S102, in response to receiving the image request message, multiple images of the target display cabinet from different angles are obtained based on the target display cabinet identifier, and multiple images of the item to be displayed from different angles are obtained based on the item to be displayed identifier; wherein, the display cabinet images display the bounding boxes of the items to be displayed at each position in the target display cabinet.
[0063] In one possible implementation, the server can store multiple images of each material display case from different angles, as well as multiple images of various material items from different angles. Thus, after receiving an image request message from the client, the server can, based on the target display case identifier in the image request message, search for multiple images of the target display case from different angles; and based on the item identifier in the image request message, search for multiple images of the item to be displayed from different angles.
[0064] In another possible implementation, to alleviate the storage pressure on the server, a storage device can be set up to store multiple images of the display cases from different angles and multiple images of the various material items from different angles. When the server receives the image request message, it can send a material request message carrying the identifier of the target display case and the identifier of the item to be displayed to the storage device. The storage device can, as described above, search for and obtain multiple images of the target display case from different angles and multiple images of the item to be displayed from different angles based on the identifier of the target display case and the identifier of the item to be displayed, and send the multiple images of the target display case from different angles and the multiple images of the item to be displayed from different angles to the server.
[0065] Here, the multiple images of the display case from different angles can be obtained based on images taken manually from multiple angles. Similarly, the multiple images of the items to be displayed from different angles can be obtained based on images taken by the user of the display case from multiple angles. For example, the user of the display case can take images of the item to be displayed from at least 10 angles, including 8 side images (e.g., front, back, left, right, front left, back left, front right, and back right), 1 top-view image, and 1 bottom image. It should be noted that the background of the images of the items to be displayed is transparent. This can be achieved by automatically making the background transparent after manual photography, and then manually or automatically uploading the images to a server or storage device.
[0066] In step S103, the image of the item to be displayed is pasted into the corresponding bounding box of the display case image to generate multiple augmented images.
[0067] In one possible implementation, for a display case image at a certain angle, the server can determine the angle of the items in each bounding box in the display case image. For a bounding box, the images of each displayed item at the corresponding angle can be randomly pasted into that bounding box. In this way, based on the display case images at multiple angles, the multiple bounding boxes in the display case images, the various types of items to be displayed, and the images of items to be displayed at multiple angles, a large number of augmented images, such as hundreds of thousands, can be randomly pasted and combined to generate, as long as the target display case in the augmented image contains the corresponding items to be displayed.
[0068] In one possible implementation, the images of the items to be displayed in the augmented image can be automatically labeled and recorded when pasted, without the need for manual labeling.
[0069] Example, Figure 1B This diagram illustrates an augmented image according to an embodiment of the present disclosure. The user has previously displayed product A and some other products in display case 11. At this point, the server can generate an image based on an image request message sent by the client. Figure 1B The image shown in the middle left shows an augmented image of product A and some other products displayed in display case 11. Later, when product B is added to display case 11, the server can generate an augmented image using the same set of display case images according to the new image request message sent by the client. Figure 1B The right-middle image shows an augmented view of product B, product A, and some other products displayed in display case 11. Alternatively, Figure 1C The diagram illustrates an augmented image according to an embodiment of the present disclosure. If the image request message sent by the client indicates that the display cabinet 12, as identified by the display cabinet identifier, has 5 rows of items on each shelf, the server can then generate an image based on the image request message sent by the client. Figure 1C The image shown in the middle left shows an augmented image of display case 12 with 5 columns of items on each shelf. If the client sends an image request message indicating that display case 12 has 8 columns of items on each shelf, the server can generate the image according to the client's image request message. Figure 1C The right-middle image shows an augmented view of eight columns of items displayed on each shelf of display case 12.
[0070] In step S104, the plurality of augmented images are returned to the client.
[0071] In one possible implementation, the server can return the generated augmented images to the client after acquiring them, so that the client can train the corresponding model.
[0072] In one possible implementation, the server can rapidly generate augmented images based on the client's model training needs. For example, it can generate a massive number of augmented images to support the client in training large models; it can quickly generate the necessary augmented images for training when users add or change items in the display cases, thus rapidly updating and training new models; it can quickly generate adversarial image sets to support adversarial training; and so on. After the client has trained the corresponding model using the augmented image, it can assign the trained model to the target display case.
[0073] In this embodiment, the client can send an image request message to the server according to its training needs, requesting the server to return the images required for training. The image request message carries the target display case identifier and parameters of the items to be displayed. The parameters of the items to be displayed include the identifier of the items to be displayed. Upon receiving the image request message, the server can obtain multiple images of the target display case from different angles based on the target display case identifier. These display case images show the bounding boxes of the items to be displayed at various locations within the target display case. The server then obtains multiple images of the items to be displayed from different angles based on the item identifiers and pastes these images into the corresponding bounding boxes of the display case images, generating multiple augmented images. These augmented images are then returned to the client for model training. In this way, initially, only multiple images of the display case from different angles and images of various items to be displayed are needed to automatically generate a large number of augmented images required for training the model, resulting in low manual costs. When new items are added to or the display items are changed in the target display case, the client can request the server to generate new augmented images for training. The server can quickly generate a large number of images required for training the model according to the client's training needs, resulting in faster model updates.
[0074] In one embodiment of this disclosure, the parameters of the items to be displayed further include quantity information and / or location information of various items to be displayed. Step S103 in the above-mentioned image data augmentation method pastes the image of the items to be displayed within the corresponding bounding box of the display cabinet image, generating multiple augmented images, including:
[0075] Based on the quantity and / or location information of the various items to be displayed, the images of the items to be displayed are pasted into the corresponding bounding boxes of the display cabinet image to generate multiple augmented images.
[0076] In this implementation, the quantity of various items to be displayed in some display cases is predetermined. For example, the quantity of item A must not be less than a first preset value, or the proportion of item A to the total number of items must exceed a second preset value. Alternatively, the positions of various items to be displayed in the display case are predetermined. For example, item A must be displayed on the top shelf, item B on the middle shelf, and item C on the bottom shelf. Alternatively, both the quantity and position information of the items to be displayed in the display case are predetermined. In this case, in order to ensure that the trained model is more accurate, the quantity and / or position information of the items in the target display case in the generated augmented image must conform to the corresponding regulations.
[0077] In one embodiment of this disclosure, the image request message further carries a predetermined number of augmented images. Step S103 in the above-described image data augmentation method, which involves pasting the image of the item to be displayed within the corresponding bounding box of the display case image to generate multiple augmented images, may include the following steps:
[0078] The image of the item to be displayed is pasted into the corresponding bounding box of the display case image to generate the predetermined number of augmented images.
[0079] In this implementation, the number of augmented images generated by the server can be a default number or a predetermined number indicated by the client. For example, for a normal model, the predetermined number could be hundreds of thousands, while for a large model, the predetermined number could be tens of millions or even hundreds of millions.
[0080] In one embodiment of this disclosure, the server can store multiple images of the display cases from different angles and multiple images of various material items from different angles. In this case, the image data augmentation method may further include the following steps:
[0081] Obtain images of the display cases from multiple different angles, representing the material display cases identified by their respective labels.
[0082] Acquire images of the material items identified by each item icon from multiple different angles;
[0083] The step S103, which involves obtaining multiple images of the target display cabinet from different angles based on the target display cabinet identifier, and the part involving obtaining multiple images of the item to be displayed from different angles based on the item to be displayed identifier, may further include the following steps:
[0084] From the material display cases identified by each display case label, find the target display case identified by the target display case label from multiple different angles of display case images.
[0085] From the material items identified by each item identifier, find the images of the item to be displayed identified by the item identifier from multiple different angles among the item images from multiple different angles.
[0086] In this implementation, images of each material display cabinet and each material item can be obtained by manually taking photos from multiple angles. For example, when a new type of display cabinet is manufactured, the manufacturer can provide the server with images of the display cabinet taken from multiple angles. When a user obtains a new display cabinet and wants to display certain items, or wants to place new items or replace an item in the display cabinet, the user of the display cabinet can take images of the items to be displayed or the new items from multiple angles.
[0087] In this implementation, when the server stores multiple images of display cases from different angles for each material display case and multiple images of various material items from different angles, it can store the multiple images of display cases from different angles for each material display case in a corresponding manner with the display case identifiers for each material display case, and store the multiple images of various material items from different angles in a corresponding manner with the item identifiers for each material item. Thus, after receiving an image request message from the client, the server can search for the target display case identifier from the display case identifiers of each material display case based on the target display case identifier in the image request message, thereby obtaining multiple images of the target display case identified by the target display case identifier from multiple angles; and based on the item identifier to be displayed in the image request message, it can search for the item identifier to be displayed from the item identifiers of each material item, thereby obtaining multiple images of the item to be displayed from multiple angles of the item to be displayed identified by the item identifier to be displayed.
[0088] In one possible implementation, acquiring images of the material items identified by each item identifier from multiple different angles may include the following steps:
[0089] Obtain 3D images of the material items identified by each item identifier;
[0090] Extract images of the material items identified by each item label from the three-dimensional image from multiple different angles.
[0091] In this embodiment, a three-dimensional image of the material item identified by each item identifier can be obtained, and then the item images of each item at multiple different angles can be captured from different angles. In this way, it is only necessary to obtain the three-dimensional image of the corresponding item to automatically capture the item images at multiple different angles, without the need for manual multi-angle shooting, thus reducing labor costs.
[0092] In one embodiment of this disclosure, acquiring images of the material display cases identified by each display case label from multiple different angles may include the following steps:
[0093] Obtain original images of the material display cabinets identified by each display cabinet label from multiple different angles. The original display cabinet images include images of the current items displayed in the material display cabinets.
[0094] Replace each area containing the current item image in the material display cabinet with a bounding box to obtain the display cabinet image of the material display cabinet.
[0095] In this embodiment, original display cabinet images of each material display cabinet can be manually captured from multiple different angles. The original display cabinet images include images of the current items displayed in the material display cabinet. The area where the current item image is located can be identified, the current item image can be subtracted, and each area where the current item image is located can be replaced with each bounding box. In this way, the display cabinet image of the material display cabinet can be obtained.
[0096] It should be noted that a 3D image of the material display case can be obtained, and then images of the display case from multiple different angles can be captured, and then the bounding box can be manually added.
[0097] In one embodiment of this disclosure, to alleviate the storage pressure on the server, a material storage terminal can be set up to store multiple display cabinet images from different angles of each material display cabinet and multiple item images from different angles of various material items. In the above image augmentation method, step S102, namely, responding to receiving the image request message, obtaining multiple display cabinet images from different angles of the target display cabinet based on the target display cabinet identifier, and obtaining multiple item images from different angles of the item to be displayed based on the item identifier, may include the following steps:
[0098] In response to receiving the image request message, a material request message is sent to the material storage terminal, the material request message carrying the target display cabinet identifier and the item identifier to be displayed;
[0099] The system receives multiple images of the target display cabinet from different angles and multiple images of the items to be displayed from different angles, returned by the material storage terminal after a query.
[0100] In this embodiment, after receiving an image request message from the client, the server can send a material request message to the material storage terminal. The material request message carries a target display case identifier and an item identifier indicated by the image request message. The material storage terminal stores multiple display case images from different angles of each material display case and multiple item images from different angles of various material items. The material storage terminal can store the multiple display case images from different angles of each material display case corresponding to the display case identifiers of each material display case, and store the multiple item images from different angles of each material item corresponding to the item identifiers of each material item. Thus, after receiving the material request message from the server, the material storage terminal can, based on the target display case identifier in the material request message, search for the target display case identifier from the display case identifiers of each material display case, thereby obtaining multiple display case images from different angles of the target display case identified by the target display case identifier; and based on the item identifiers to be displayed in the image request message, search for the item identifiers to be displayed from the item identifiers of each material item, thereby obtaining multiple item images from different angles of the item to be displayed identified by the item identifiers to be displayed. Then, the material display cabinet can return multiple images of the target display cabinet from different angles and multiple images of the items to be displayed from different angles to the server, so that the server can generate augmented images accordingly.
[0101] This disclosure also provides an image data augmentation method. Figure 2 This diagram illustrates a flowchart of an image data augmentation method applied to a media storage terminal according to an embodiment of the present disclosure. The method is applied to the media storage terminal, such as... Figure 2 As shown, the image data augmentation method may include the following steps:
[0102] In step S201, images of the material display cabinets identified by each display cabinet identifier are obtained from multiple different angles; the display cabinet images show the bounding boxes of the items to be displayed at each position in the target display cabinet.
[0103] In one possible implementation, the images of each material display case from multiple different angles can be obtained by manually taking photos of the display cases from multiple angles and then manually marking the bounding boxes.
[0104] In another possible implementation, acquiring images of the material display cases identified by each display case label from multiple different angles may include the following steps:
[0105] Obtain original images of the material display cabinets identified by each display cabinet label from multiple different angles. The original display cabinet images include images of the current items displayed in the material display cabinets.
[0106] Replace each area containing the current item image in the material display cabinet with a bounding box to obtain the display cabinet image of the material display cabinet.
[0107] In this embodiment, original display cabinet images of each material display cabinet can be manually captured from multiple different angles. The original display cabinet images include images of the current items displayed in the material display cabinet. The area where the current item image is located can be identified, the current item image can be subtracted, and each area where the current item image is located can be replaced with each bounding box. In this way, the display cabinet image of the material display cabinet can be obtained.
[0108] In another possible implementation, a three-dimensional image of the material display case can be obtained, and then images of the material display case from multiple different angles can be captured from different angles, and then the bounding box can be manually added.
[0109] In one possible implementation, when storing multiple images of each material display cabinet from different angles, the material storage terminal can store the images of each material display cabinet from different angles in correspondence with the display cabinet identifier of each material display cabinet.
[0110] In step S202, the material items identified by each item identifier are obtained from multiple different angles.
[0111] In one possible implementation, images of the materials and items can be taken manually from multiple different angles. For example, when a user obtains a new display case to display certain items or to place new items or replace an item in the display case, the user of the display case can take images of the items to be displayed or the new items from multiple different angles.
[0112] In one possible implementation, acquiring images of the material items identified by each item identifier from multiple different angles includes:
[0113] Obtain 3D images of the material items identified by each item identifier;
[0114] Extract images of the material items identified by each item label from the three-dimensional image from multiple different angles.
[0115] In this embodiment, a three-dimensional image of the material item identified by each item identifier can be obtained, and then the item images of each item at multiple different angles can be captured from different angles. In this way, it is only necessary to obtain the three-dimensional image of the corresponding item to automatically capture the item images at multiple different angles, without the need for manual multi-angle shooting, thus reducing labor costs.
[0116] In one possible implementation, when storing multiple images of each material item from different angles, the material storage terminal can store the multiple images of each material item from different angles in correspondence with the item identifier of each material item.
[0117] In step S203, a material request message sent by the server is received, the material request message carrying the target display cabinet identifier and the identifier of the item to be displayed.
[0118] In this embodiment, after receiving an image request message from the client, the server can send a material request message to the material storage terminal. The material request message carries the target display case identifier and the item identifier to be displayed as indicated by the image request message.
[0119] In step S204, multiple images of the target display cabinet identified by the target display cabinet are searched from the display cabinet images of the material display cabinets identified by each display cabinet identifier.
[0120] In one possible implementation, after receiving a material request message from the server, the material storage terminal can search for the target display cabinet identifier from the display cabinet identifiers of each material display cabinet based on the target display cabinet identifier in the material request message, and then obtain multiple display cabinet images of the target display cabinet identified by the target display cabinet identifier from multiple different angles.
[0121] In step S205, from the material items identified by each item identifier in multiple item images from multiple different angles, the multiple different angle images of the item to be displayed identified by the item to be displayed are searched.
[0122] In one possible implementation, after receiving a material request message from the server, the material storage terminal can search for the item identifier to be displayed from the item identifiers of each material item based on the item identifier in the material request message, and then obtain multiple images of the item to be displayed from different angles identified by the item identifier.
[0123] In step S206, multiple images of the target display cabinet from different angles and multiple images of the items to be displayed from different angles are returned to the server so that the server can generate an augmented image based on the multiple images of the target display cabinet from different angles and the multiple images of the items to be displayed from different angles.
[0124] In one possible implementation, the material storage terminal can send multiple images of the target display cabinet from different angles and multiple images of the items to be displayed from different angles to the server. The server can then paste the images of the items to be displayed into the corresponding bounding boxes of the display cabinet images to generate multiple augmented images. The server then returns the multiple augmented images to the client so that the client can train a model based on the augmented images.
[0125] The implementation process of the above-mentioned image data augmentation method is described in detail below through the following embodiments.
[0126] Example 1:
[0127] This disclosure also provides an image data augmentation method. Figure 3 The diagram illustrates an overall flow chart of an image data augmentation method according to an embodiment of the present disclosure. This method is applied to a system including a server and a client, such as... Figure 3 As shown, the image data augmentation method may include the following steps:
[0128] In step S301, the server acquires images of the material display cabinets identified by each display cabinet identifier from multiple different angles;
[0129] In step S302, the server acquires images of the material items identified by each item identifier from multiple different angles;
[0130] In step S303, the server receives an image request message sent by the client;
[0131] The image request message carries a target display case identifier, parameters of the items to be displayed, and a predetermined number of augmented images. The parameters of the items to be displayed include the identifier of the items to be displayed, as well as quantity information and / or location information of various items to be displayed.
[0132] In step S304, the server searches for multiple different angle images of the target display cabinet identified by the target display cabinet from multiple display cabinet images of the material display cabinet identified by each display cabinet identifier.
[0133] In step S305, the server searches for multiple images of the item to be displayed from multiple different angles of the material item identified by each item identifier in the item images from multiple different angles;
[0134] In step S306, the server pastes the images of the items to be displayed into the corresponding bounding boxes of the display cabinet images according to the quantity information and / or location information of the various types of items to be displayed, thereby generating the predetermined number of augmented images;
[0135] In step S307, the server returns the predetermined number of augmented images to the client.
[0136] Example 2:
[0137] This disclosure also provides an image data augmentation method. Figure 4 The diagram illustrates an overall flow of another image data augmentation method according to an embodiment of the present disclosure. This method is applied to a system including a material storage terminal, a server, and a client, such as... Figure 4 As shown, the image data augmentation method may include the following steps:
[0138] In step S401, the material storage terminal acquires images of the material display cabinets identified by each display cabinet identifier from multiple different angles;
[0139] The image of the display case shows the bounding boxes of the items to be displayed at each position in the target display case;
[0140] In step S402, the material storage terminal acquires material images of the material items identified by each item identifier from multiple different angles;
[0141] In step S403, the server receives an image request message sent by the client;
[0142] The image request message carries a target display case identifier, parameters of the items to be displayed, and a predetermined number of augmented images. The parameters of the items to be displayed include the identifier of the items to be displayed, as well as quantity information and / or location information of various items to be displayed.
[0143] In step S404, in response to receiving the image request message, the server sends a material request message to the material storage terminal;
[0144] The material request message contains the target display case identifier and the identifier of the item to be displayed.
[0145] In step S405, the material storage terminal searches for multiple display cabinet images of the target display cabinet identified by the target display cabinet from the display cabinet images of the material display cabinets identified by each display cabinet identifier;
[0146] In step S406, the material storage terminal searches for multiple images of the item to be displayed from multiple different angles of the material items identified by each item identifier in the item images from multiple different angles;
[0147] In step S407, the material storage terminal returns to the server multiple images of the target display cabinet from different angles and multiple images of the items to be displayed from different angles.
[0148] In step S408, the server pastes the images of the items to be displayed into the corresponding bounding boxes of the display cabinet images according to the quantity information and / or location information of the various types of items to be displayed, thereby generating the predetermined number of augmented images;
[0149] In step S409, the server returns the predetermined number of augmented images to the client.
[0150] The following are embodiments of the apparatus disclosed herein, which can be used to execute embodiments of the method disclosed herein.
[0151] Figure 5 This is a structural block diagram of an image data augmentation device applied to a server according to an embodiment of the present disclosure. This device can be implemented as part or all of an electronic device through software, hardware, or a combination of both. For example... Figure 5 As shown, the image data augmentation device includes:
[0152] The first receiving module 501 is configured to receive an image request message sent by a client. The image request message carries a target display cabinet identifier and parameters of the item to be displayed. The parameters of the item to be displayed include the identifier of the item to be displayed.
[0153] The first acquisition module 502 is configured to, in response to receiving the image request message, acquire multiple images of the target display cabinet from different angles based on the target display cabinet identifier, and acquire multiple images of the item to be displayed from different angles based on the item to be displayed identifier; wherein the display cabinet images display the bounding boxes of the items to be displayed at each position in the target display cabinet.
[0154] The generation module 503 is configured to paste the image of the item to be displayed into the corresponding bounding box of the display case image to generate multiple augmented images;
[0155] The first sending module 504 is configured to return the plurality of augmented images to the client.
[0156] In one possible implementation, the parameters of the items to be displayed also include quantity information and / or location information of various items to be displayed, and the generation module 503 is configured to:
[0157] Based on the quantity and / or location information of the various items to be displayed, the images of the items to be displayed are pasted into the corresponding bounding boxes of the display cabinet image to generate multiple augmented images.
[0158] In one possible implementation, the image request message also carries a predetermined number of augmented images, and the generation module 503 is configured to:
[0159] The image of the item to be displayed is pasted into the corresponding bounding box of the display case image to generate the predetermined number of augmented images.
[0160] In one possible implementation, the device further includes:
[0161] The second acquisition module is configured to acquire images of the material display cabinets identified by each display cabinet logo from multiple different angles.
[0162] The third acquisition module is configured to acquire images of the material items identified by each item identifier from multiple different angles;
[0163] The portion of the first acquisition module 502 that acquires multiple images of the target display cabinet from different angles based on the target display cabinet identifier, and acquires multiple images of the item to be displayed from different angles based on the item to be displayed identifier, is configured as follows:
[0164] From the material display cases identified by each display case label, find the target display case identified by the target display case label from multiple different angles of display case images.
[0165] From the material items identified by each item identifier, find the images of the item to be displayed identified by the item identifier from multiple different angles among the item images from multiple different angles.
[0166] In one possible implementation, the third acquisition module is configured as follows:
[0167] Obtain 3D images of the material items identified by each item identifier;
[0168] Extract images of the material items identified by each item label from the three-dimensional image from multiple different angles.
[0169] In one possible implementation, the second acquisition module is configured as follows:
[0170] Obtain original images of the material display cabinets identified by each display cabinet label from multiple different angles, wherein the original display cabinet images include images of the items in the material display cabinets;
[0171] Replace each area containing the image of each item in the material display case with a bounding box to obtain the display case image of the material display case.
[0172] In one possible implementation, the first acquisition module 502 is configured to:
[0173] In response to receiving the image request message, a material request message is sent to the material storage terminal, the material request message carrying the target display cabinet identifier and the item identifier to be displayed;
[0174] The system receives multiple images of the target display cabinet from different angles and multiple images of the items to be displayed from different angles, returned by the material storage terminal after a query.
[0175] In this embodiment, the image data augmentation device corresponds to the image data augmentation method described above. For specific details, please refer to the description of the image data augmentation method above, which will not be repeated here.
[0176] Figure 6 This is a structural block diagram of an image data augmentation device applied to a material storage terminal according to an embodiment of the present disclosure. The device can be implemented as part or all of an electronic device through software, hardware, or a combination of both.
[0177] like Figure 6 As shown, the image data augmentation device includes:
[0178] The fourth acquisition module 601 is configured to acquire images of the material display cabinets identified by each display cabinet identifier from multiple different angles; the display cabinet images display the bounding boxes of the items to be displayed at each position in the target display cabinet;
[0179] The fifth acquisition module 602 is configured to acquire images of the material items identified by each item identifier from multiple different angles;
[0180] The second receiving module 603 is configured to receive a material request message sent by the server, wherein the material request message carries the target display cabinet identifier and the item identifier to be displayed.
[0181] The first search module 604 is configured to search for multiple different angle images of the target display cabinet identified by the target display cabinet from the display cabinet images of the material display cabinets identified by each display cabinet identifier.
[0182] The second search module 605 is configured to search for multiple images of the item to be displayed from multiple different angles of the material item identified by each item identifier in multiple item images.
[0183] The second sending module 606 is configured to return multiple images of the target display cabinet from different angles and multiple images of the items to be displayed from different angles to the server, so that the server can generate an augmented image based on the multiple images of the target display cabinet from different angles and the multiple images of the items to be displayed from different angles.
[0184] In one possible implementation, the fifth acquisition module 602 is configured as follows:
[0185] Obtain 3D images of the material items identified by each item identifier;
[0186] Extract images of the material items identified by each item label from the three-dimensional image from multiple different angles.
[0187] In one possible implementation, the fourth acquisition module 601 is configured as follows:
[0188] Obtain original images of the material display cabinets identified by each display cabinet label from multiple different angles, wherein the original display cabinet images include images of the items in the material display cabinets;
[0189] Replace each area containing the image of each item in the material display case with a bounding box to obtain the display case image of the material display case.
[0190] This disclosure also discloses an electronic device. Figure 7 A schematic structural block diagram of an electronic device according to an embodiment of the present disclosure is shown, such as... Figure 7 As shown, the electronic device 700 includes a memory 701 and a processor 702; wherein the memory 701 is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor 702 to implement the steps of any of the above-described image augmentation methods.
[0191] Figure 8 This is a schematic diagram of the structure of a computer system suitable for implementing the image augmentation method according to an embodiment of the present disclosure. For example... Figure 8As shown, the computer system 800 includes a processing unit 801, which can execute various processes described above based on a program stored in a read-only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the system 800. The processing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0192] The following components are connected to I / O interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to I / O interface 805 as needed. A removable medium 88, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 810 as needed so that computer programs read from it can be installed into storage section 808 as needed. The processing unit 801 can be implemented as a CPU, GPU, TPU, FPGA, NPU, etc.
[0193] In particular, according to embodiments of this disclosure, the methods described above can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising computer instructions that, when executed by a processor, implement the steps of the methods described above. In such embodiments, the computer program product can be downloaded and installed from a network via communication section 809, and / or installed from removable media 811.
[0194] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0195] The units or modules described in the embodiments of this disclosure can be implemented in software or hardware. The described units or modules can also be located in a processor, and the names of these units or modules do not necessarily constitute a limitation on the unit or module itself.
[0196] In another aspect, this disclosure also provides a computer-readable storage medium, which may be a computer-readable storage medium included in the apparatus described in the above embodiments; or it may be a standalone computer-readable storage medium not assembled into a device. The computer-readable storage medium stores one or more programs that are used by one or more processors to perform the methods described in this disclosure.
[0197] In addition, this disclosure also provides a computer program product storing a computer program that, when executed by a processor, enables the processor to at least implement the methods provided in the foregoing embodiments.
[0198] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
Claims
1. An image data augmentation method, characterized in that, Applied to the server side, including: Pre-acquire images of the material display cases identified by each display case label from multiple different angles; acquire images of the material items identified by each item label from multiple different angles; The step of obtaining material items identified by each item identifier from multiple different angles includes: obtaining three-dimensional images of material items identified by each item identifier; and extracting material items identified by each item identifier from multiple different angles from the three-dimensional images. The process of obtaining display cabinet images of the material display cabinets identified by each display cabinet identifier from multiple different angles includes: obtaining original display cabinet images of the material display cabinets identified by each display cabinet identifier from multiple different angles, wherein the original display cabinet images include images of items in the material display cabinets; replacing each area containing the images of each item in the material display cabinet with bounding boxes to obtain the display cabinet image of the material display cabinet; or obtaining a three-dimensional image of the material display cabinet, cropping the display cabinet images of the material display cabinet from different angles from multiple different angles, and adding bounding boxes; Receive an image request message sent by the client. The image request message carries a target display case identifier and parameters of the item to be displayed. The parameters of the item to be displayed include the identifier of the item to be displayed. Each display case identifier identifies a display case template. A display case template includes a set of multiple display case images from different angles. In response to receiving the image request message, multiple images of the target display cabinet from different angles are obtained based on the target display cabinet identifier, and multiple images of the item to be displayed from different angles are obtained based on the item to be displayed identifier; wherein, the display cabinet images display the bounding boxes of the items to be displayed at each position in the target display cabinet; The images of the items to be displayed are pasted into the corresponding bounding boxes of the display cabinet image to generate multiple augmented images. For a display cabinet image at a certain angle, the angles of the items in each bounding box of the display cabinet image are determined. For each bounding box, the images of each item at the corresponding angle are randomly pasted into the bounding box and labeled and recorded. The augmented images are returned to the client.
2. The method according to claim 1, characterized in that, The parameters of the items to be displayed also include quantity information and / or location information of various items to be displayed. Pasting the image of the items to be displayed within the corresponding bounding box of the display case image generates multiple augmented images, including: Based on the quantity and / or location information of the various items to be displayed, the images of the items to be displayed are pasted into the corresponding bounding boxes of the display cabinet image to generate multiple augmented images.
3. The method according to claim 1 or 2, characterized in that, The image request message also carries a predetermined number of augmented images. The step of pasting the image of the item to be displayed within the corresponding bounding box of the display case image generates multiple augmented images, including: The image of the item to be displayed is pasted into the corresponding bounding box of the display case image to generate the predetermined number of augmented images.
4. The method according to claim 1, characterized in that, The step of responding to receiving the image request message and obtaining multiple images of the target display cabinet from different angles based on the target display cabinet identifier, and obtaining multiple images of the item to be displayed from different angles based on the item to be displayed identifier, includes: In response to receiving the image request message, a material request message is sent to the material storage terminal, the material request message carrying the target display cabinet identifier and the item identifier to be displayed; The system receives multiple images of the target display cabinet from different angles and multiple images of the items to be displayed from different angles, returned by the material storage terminal after a query.
5. An image data augmentation method, characterized in that, Applications include: The process involves acquiring images of the material display cabinets identified by each display cabinet label from multiple different angles. Each display cabinet image displays bounding boxes of items to be displayed at various locations within the target display cabinet. Acquiring these images includes: acquiring original images of the material display cabinets identified by each display cabinet label from multiple different angles, the original images including images of the items in the material display cabinets; replacing each area containing the item images in the material display cabinets with bounding boxes to obtain the display cabinet image of the material display cabinet; or acquiring a 3D image of the material display cabinet, cropping images of the material display cabinet from different angles, and adding bounding boxes. The process involves acquiring images of the material items identified by each item identifier from multiple different angles. This acquisition includes: acquiring three-dimensional images of the material items identified by each item identifier; and extracting images of the material items identified by each item identifier from the three-dimensional images from multiple different angles. The system receives a material request message sent by the server. The material request message carries the target display case identifier and the identifier of the item to be displayed. Each display case identifier identifies a display case template, and a display case template includes a set of display case images from multiple different angles. From the images of the material display cases identified by each display case label, find multiple display case images of the target display case identified by the target display case label from different angles; From the material items identified by each item identifier, find the images of the item to be displayed identified by the item identifier from multiple different angles among the item images from multiple different angles; The server returns multiple images of the target display case from different angles and multiple images of the items to be displayed from different angles, so that the server can generate an augmented image based on these images. Specifically, when generating the augmented image, for a display case image at a certain angle, the angles of the items within each bounding box in that display case image are determined. For each bounding box, images of the items at the corresponding angles are randomly pasted into that bounding box, and then labeled and recorded.
6. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the method of any one of claims 1-5.
7. A computer-readable storage medium storing computer instructions thereon, characterized in that, When executed by a processor, the computer instructions implement the method described in any one of claims 1-5.
8. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the method described in any one of claims 1-5.