Method and apparatus for generating composite images of objects
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-28
- Publication Date
- 2026-08-14
AI Technical Summary
[0006]本发明实施例的物品建模数据产生方法以及物品建模数据产生装置,通过影像合成处理所产生的合成影像用于模拟多种不同物品于平台上彼此相邻或部分重叠的实际情况,可减少对物品建模所需花费的时间。再者,机器学习演算法依据合成影像所产生的物品模型,具有较高的辨识率,可应用于无人商店的商品辨识。
Smart Images

Figure CN114693861B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method and apparatus for generating object modeling data. Background Technology
[0002] In recent years, machine learning algorithms have emerged as a prominent technology. One application involves using these algorithms to build product models and then using these models to identify products on shelves, thus enabling unmanned stores. However, to build a product model with high recognition accuracy, the modeling data provided to the machine learning algorithm must contain sufficient product features. Furthermore, there may be multiple combinations of different product arrangements on shelves. Therefore, product modeling methods have become a current development trend. Summary of the Invention
[0003] This invention provides a method and apparatus for generating item modeling data. The resulting synthetic images can be used to simulate the actual situation of different items being adjacent or overlapping on a platform. The machine learning algorithm, based on the item model generated from the synthetic images, will have a high recognition rate.
[0004] According to an embodiment of the present invention, a method for generating item modeling data is provided, comprising: acquiring a platform image associated with a platform; acquiring a plurality of first item images of a first item disposed on the platform, wherein the first item images correspond to a plurality of different viewpoints; acquiring a plurality of second item images of a second item disposed on the platform, wherein the second item images correspond to the different viewpoints; and performing image compositing processing based on at least one of the first item images and the second item images and the platform image to generate a composite image, the composite image comprising a plurality of adjacent or partially overlapping image regions corresponding to at least a plurality of the different viewpoints, the image regions including first image regions and second image regions, the first image regions including one of the first item images or one of the second item images, and the second image regions including one of the first item images or one of the second item images.
[0005] According to an embodiment of the present invention, an object modeling data generation apparatus is provided, including a platform, at least one camera, and a management host. The platform is used to place a first object and a second object. At least one camera is disposed above the platform and electrically connected to the management host. The management host is used to drive the at least one camera to capture images of the platform, the first object, and the second object to obtain platform images, multiple first object images, and multiple second object images, wherein the first object images correspond to multiple different viewpoints, and the second object images correspond to these viewpoints. The management host is further used to perform image compositing processing based on at least one of the first object images and the second object images and the platform image to generate a composite image. The composite image includes multiple adjacent or partially overlapping image regions corresponding to at least multiple of the different viewpoints. These image regions include first image regions and second image regions. The first image region includes one of the first object images or one of the second object images, and the second image region includes one of the first object images or one of the second object images.
[0006] The item modeling data generation method and device of this invention use image synthesis processing to generate synthetic images that simulate the actual situation where various different items are adjacent or partially overlapping on a platform, thereby reducing the time required for item modeling. Furthermore, the item models generated by machine learning algorithms based on the synthetic images have a high recognition rate and can be applied to product recognition in unmanned stores.
[0007] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments, but this is not intended to limit the present invention. Attached Figure Description
[0008] Figure 1 This is a schematic diagram of an item modeling data generation device according to an embodiment of the present invention;
[0009] Figure 2 This is a schematic diagram of an item modeling data generation device according to an embodiment of the present invention;
[0010] Figure 3 This is a flowchart of a method for generating item modeling data according to an embodiment of the present invention;
[0011] Figure 4 This is a schematic diagram of a platform image related to a platform according to an embodiment of the present invention;
[0012] Figure 5 This is a schematic diagram illustrating the placement of a first item on a platform according to an embodiment of the present invention;
[0013] Figure 6 This is a schematic diagram illustrating the placement of a second item on a platform according to an embodiment of the present invention;
[0014] Figure 7A and Figure 7B for Figure 3 A detailed flowchart of one embodiment of image compositing processing;
[0015] Figure 8A A schematic diagram illustrating one embodiment of multiple first article images;
[0016] Figure 8B A schematic diagram illustrating one embodiment of multiple images of second articles;
[0017] Figure 9 Drawing the first execution Figure 7A A schematic diagram of an embodiment of step S172;
[0018] Figure 10 Drawing the second execution Figure 7A A schematic diagram of an embodiment of step S172;
[0019] Figure 11 Drawing the third execution Figure 7A A schematic diagram of an embodiment of step S172;
[0020] Figure 12 Drawing the fourth execution Figure 7A A schematic diagram of an embodiment of step S172;
[0021] Figure 13 A flowchart illustrating a method for generating item modeling data according to another embodiment of the present invention is shown;
[0022] Figure 14 A comparison chart of the frame extraction rates for object modeling based on synthetic images and single-object modeling.
[0023] Among them, the attached figures are labeled
[0024] 1 platform
[0025] 2 racks
[0026] 3 cameras
[0027] 4. Management Host
[0028] T platform images
[0029] Vertices T1, T2, T3, and T4
[0030] A First Item
[0031] B. Second item
[0032] E1 First row direction
[0033] E2 Second row direction
[0034] E3 Third row direction
[0035] Image of the first item in P1~P9
[0036] Q1~Q9 Second Item Images
[0037] Lengths of L1, L2, L3, and L4
[0038] Widths of W1, W2, W3, and W4
[0039] C1 First Center Point
[0040] C2 Second Center Point
[0041] C3 Third Center Point
[0042] C4 Fourth Center Point Detailed Implementation
[0043] The structural and working principles of the present invention will be described in detail below with reference to the accompanying drawings:
[0044] When using machine learning algorithms to build product models, taking unmanned convenience stores as an example, product features can be obtained from images of products from different perspectives. To obtain images of products from multiple different perspectives, one method is to repeatedly photograph products placed in different positions on the shelf using a camera. Furthermore, in reality, there are often multiple products on a shelf, so if the modeling data includes image data of multiple products overlapping or adjacent to each other, a product model with a high recognition rate can be built. To obtain image data of multiple products overlapping or adjacent to each other, one method is to arrange the products on the shelf in various arrangements and then photograph them. However, since there are many combinations of product arrangements on the shelf, it is difficult to implement all possible arrangements. The following will describe embodiments of the item modeling data generation device and method provided by this invention.
[0045] Figure 1 This is a schematic diagram of an item modeling data generation device according to an embodiment of the present invention, such as... Figure 1 As shown, the object modeling data generation device includes a platform 1, multiple racks 2, multiple cameras 3, and a management host 4. The racks 2 are fixed in an array above the platform 1, for example, to the ceiling or a canopy connected to the platform. The cameras 3 are each fixed to one of the racks 2. The cameras 3 are electrically connected to the management host 4 via wireless or wired communication. The management host 4 includes a processor, a memory device, and a communication device. The management host 4 communicates with the cameras 3 via the communication device to receive image data from the cameras 3, and the processor can execute subsequent operations. Figure 3 , Figures 7A-7BAnd / or the methods or steps of the embodiments of this case, while the memory device can store image data. The lenses of these cameras 3 are all oriented in the same direction toward the platform 1, for example, all pointing downwards, specifically, toward a direction substantially perpendicular to the platform; in one embodiment, the camera 3 is a wide-angle camera. In one embodiment, an object can be placed on the platform 1. In one embodiment, the camera 3 is below and the object is above, or the camera is to the left and the object is to the right; this disclosure is not limited to this.
[0046] Figure 2 This is a schematic diagram of an item modeling data generation device according to an embodiment of the present invention. Figure 2 Implementation examples and Figure 1 Similar, and Figure 2 In this embodiment, the rack 2 and the camera 3 located above the platform 1 are both single units.
[0047] exist Figure 1 In the architecture of the object modeling data generation device, the camera 3 and the object can be fixed in place. These cameras 3 simultaneously (but not limited to simultaneously) capture images of the object once, thus obtaining images of the object from multiple different perspectives. In one embodiment, the number of images from multiple perspectives of the object is the same as the number of cameras. For example, if these cameras 3 are arranged in a 5x5 array, then 25 perspective images can be obtained in one capture. Figure 2 In the architecture of the object modeling data generation device, to obtain images of an object from multiple different perspectives, one implementation involves fixing camera 3 and placing the object sequentially at multiple different positions on the platform, then having camera 3 sequentially photograph the object placed at each position on platform 1. Another implementation involves fixing the object on platform 1 while camera 3 can move to different positions, for example, via a sliding rail mechanism, to sequentially photograph the object from multiple different positions. For ease of explanation, the following will all use... Figure 1 Take, for example, the device for generating data for item modeling.
[0048] Figure 3 The flowchart illustrates a method for generating object modeling data according to an embodiment of the present invention. In step S10, a management host 4 drives one of the cameras 3 to capture images of the platform 1, thereby obtaining a platform image associated with the platform 1. The platform image includes an image of the platform 1, and the driven camera is, for example, one of the cameras 3 aligned with the center of the platform 1. In one embodiment, the management host 4 drives the cameras 3 to capture images of the platform 1, thereby obtaining multiple platform images from different perspectives associated with the platform 1. In one embodiment, in step S10, no object for modeling is placed on the platform.
[0049] Figure 4 This is a schematic diagram of a platform image related to a platform according to an embodiment of the present invention. In this embodiment, as... Figure 4 As shown, the platform image T has first to fourth vertices T1, T2, T3, and T4. The coordinates of the first to fourth vertices T1, T2, T3, and T4 are (0, Ye), (0, 0), (Xe, 0), and (Xe, Ye), respectively, where parameters Xe and Ye are both greater than zero. The lower limit of the horizontal axis coordinate of the platform image T is 0, and the lower limit of the vertical axis coordinate is 0. The upper limit of the horizontal axis coordinate of the platform image T is Xe, and the upper limit of the vertical axis coordinate is Ye. The above coordinate values are only illustrative, and this disclosure does not limit the method of defining coordinate values.
[0050] Figure 5 This is a schematic diagram illustrating the placement of a first item on a platform according to an embodiment of the present invention. (See also:) Figure 3 and Figure 5 In step S11, the first item A is positioned at the center of the platform 1. In step S12, the management host 4 drives the cameras 3 to simultaneously capture images of the first item A to obtain multiple first initial images, wherein each first initial image corresponds to a different viewpoint, and each first initial image includes images of the first item A and the platform 1. In step S13, the management host 4 performs background removal processing on the first initial images based on the platform images to generate multiple first item images, each first item image corresponding to a different viewpoint, and each first item image including an image of the first item A. In one embodiment, the differences between the platform images and the first initial images are compared to generate the first item images, so that the first item images only include images of the first item A and do not include images of the platform 1.
[0051] Figure 6 This is a schematic diagram illustrating the placement of a second item on a platform according to an embodiment of the present invention. (See also: [link to related document]) Figure 3 and Figure 6In step S14, the first item A, which is placed on platform 1, is removed, and the second item B is placed at the center of platform 1. In step S15, the management host 4 drives the cameras 3 to simultaneously capture images of the second item B to obtain multiple second initial images, wherein each second initial image corresponds to a different viewpoint and includes images of the second item B and platform 1. In step S16, the second initial images are processed to remove background based on the platform images to generate multiple second item images, each second item image corresponding to a different viewpoint and each second item image including the second item B. In one embodiment, the differences between the platform images and the second initial images are compared to generate the second item images, so that the second item images only include images of the second item B and do not include images of platform 1. In other embodiments, if there are N types of items, where N is a positive integer greater than or equal to three, the above can be applied similarly. For example, the second item B set on platform 1 is removed and the third item is placed at the center of platform 1. Then, steps S15-S16 are executed to generate multiple images of the third item. The third item is then removed to generate multiple images of N items. In other embodiments, if there are M types of items, where M is a positive integer greater than or equal to two, for example, the second item B set on platform 1 is removed and the third item is placed at the center of platform 1. Then, steps S15-S16 are executed to generate multiple images of the third item. The third item is then removed to generate multiple images of M items.
[0052] like Figure 3As shown, after step S16, step S17 is executed. In step S17, the management host 4 performs image compositing processing based on at least one of the first item images and the second item images, as well as the platform image, to generate a composite image. The composite image includes at least a plurality of adjacent or partially overlapping image regions corresponding to the viewpoints. These image regions include a first image region and a second image region. The first image region includes one of the first item images or one of the second item images, and the second image region includes one of the first item images or one of the second item images. Specifically, two adjacent or partially overlapping item images correspond to different viewpoints but may belong to the same type of item or different types of items. In one embodiment, when there are N types of items, the management host 4 performs image compositing processing based on at least one of the first to Nth item images and a platform image to generate a composite image. The composite image includes at least a plurality of adjacent or partially overlapping image regions corresponding to the viewpoints, and each image region includes one of the first to Nth item images. In another embodiment, when there are M types of items, where M is a positive integer greater than or equal to two, the management host 4 performs image compositing processing based on at least one of the first to Mth item images and a platform image to generate a composite image. The composite image includes at least a plurality of adjacent or partially overlapping image regions corresponding to the viewpoints, and each image region includes one of the first to Mth item images.
[0053] In this embodiment, driving the camera, performing background removal processing, and generating composite images are all performed through the same management host. In other embodiments, driving the camera, performing background removal processing, and generating composite images may be performed through different management hosts.
[0054] Regarding the correspondence between camera and object image perspectives, for example, when there are multiple cameras, the object image from the top left perspective is the image captured by the bottom right camera. Conversely, the object image from the top right perspective is the image captured by the bottom left camera. When there is a single camera and the object's placement is moved, the object image from the top left perspective is captured by the camera fixed in the center, with the object placed in the top left corner of the platform. If there is a single camera and its position is moved, placing the object in the center, the object image from the top left perspective can be obtained when the camera moves to the bottom right corner.
[0055] Figure 7A and Figure 7B for Figure 3 A detailed flowchart of one embodiment of image compositing processing.
[0056] Figure 8A A schematic diagram illustrating one embodiment of multiple first article images, and Figure 8B A schematic diagram illustrating one embodiment of multiple images of second articles. (See diagram below.) Figure 8A and Figure 8B As shown, the cameras 3 are arranged in a 3x3 array. A single shot of the first object can capture nine images of the first object from different perspectives, P1~P9. Similarly, a single shot of the second object can capture nine images of the second object from different perspectives, Q1~Q9. The nine first object images P1~P9 consist of three images P1~P3 in the first row direction E1, three images P4~P6 in the second row direction E2, and three images P7~P9 in the third row direction E3. Likewise, the nine second object images Q1~Q9 consist of three images Q1~Q3 in the first row direction E1, three images Q4~Q6 in the second row direction E2, and three images Q7~Q9 in the third row direction.
[0057] The image compositing processing method in the following embodiments uses two different items as examples, but it is not limited to two different items; it can also be used for three or more different items. In step S171, one of the first item images and the second item images corresponding to the k-th position in the j-th row direction is selected as the k-th selected image in the j-th row direction, where the initial values of parameters j and k are 1. In one embodiment, for example, when j=1 and k=1, the first item image P1 corresponding to the first position in the first row direction is selected from the first item images P1~P9 and the second item images Q1~Q9. In one embodiment, this step can be divided into first selecting an item, and then selecting the corresponding image position; for example, selecting a first item, and then selecting the first item image corresponding to the k-th position in the j-th row direction from the first item images as the k-th selected image in the j-th row direction. In other embodiments, if there are N types of items, where N is a positive integer greater than or equal to three, then step S171 can be changed to selecting one of the k-th positions corresponding to the j-th row direction from the first item images to the N-th item images as the k-th selected image in the j-th row direction. In another embodiment, if there are M types of items, where M is a positive integer greater than or equal to two, then step S171 can be changed to selecting one of the k-th positions corresponding to the j-th row direction from the first item images to the M-th item images as the k-th selected image in the j-th row direction.
[0058] In step S172, the k-th selected image in the j-th row direction is superimposed on the platform image. The k-th selected image in the j-th row direction after being superimposed on the platform image has a center point, and the center point of the k-th selected image in the j-th row direction has a vertical coordinate and a horizontal coordinate. In one embodiment, the superposition can be either applied to the platform image or it can replace the corresponding position of the platform image.
[0059] Figure 9 Drawing the first execution Figure 7A A schematic diagram of an implementation of step S172. First item image P1 is selected as the first selected image from first item images P1~P9 and second item images Q1~Q9, and the first item image P1 is superimposed on the platform image T. The first item image P1 has a first center point C1 and has a length L1 and a width W1 in the X-axis direction and the Y-axis direction, respectively. The first center point C1 has a first vertical axis coordinate and a first horizontal axis coordinate. In one embodiment, the coordinates of the first center point C1 are, for example,... Figure 4 of (0, Ye).
[0060] In step S173, one of the first item images and the second item images corresponding to the (k+1)th position in the j-th row direction is selected as the (k+1)th selected image in the j-th row direction. In one embodiment, a first item image or a second item image corresponding to the (k+1)th position in the j-th row direction is selected from the first item images and the second item images. For example, when j=1 and k=1, the second item image Q2 corresponding to the second position in the first row direction is selected from the first item images P1~P9 and the second item images Q1~Q9. In one embodiment, this step can be divided into first selecting an item and then selecting the corresponding image position. For example, selecting a second item and then selecting the second item image corresponding to the (k+1)th position in the j-th row direction from the second item images as the (k+1)th selected image in the j-th row direction. In one embodiment, if there is no image corresponding to the (k+1)th position in the j-th row direction, for example, the (k+1)th position may be outside the platform range, step S173 can be skipped. In other embodiments, if there are N types of items, where N is a positive integer greater than or equal to three, then step S173 can be changed to selecting one of the (k+1)th positions corresponding to the j-th row direction from the first item images to the Nth item images as the (k+1)th selected image in the j-th row direction. In another embodiment, if there are M types of items, where M is a positive integer greater than or equal to two, then step S173 can be changed to selecting one of the (k+1)th positions corresponding to the j-th row direction from the first item images to the Mth item images as the (k+1)th selected image in the j-th row direction.
[0061] In step S174, the center point of the item after the k+1 selected image in the j-th row direction overlaps with the platform image is estimated based on the width of the k-th selected image in the j-th row direction and the width of the (k+1)-th selected image in the j-th row direction. The center point of the item in the (k+1)-th selected image in the j-th row direction has a horizontal axis coordinate and a vertical axis coordinate. The difference between the estimated center point of the item in the (k+1)-th selected image in the j-th row direction and the center point of the item in the k-th selected image in the j-th row direction is essentially equal to the average of the width of the k-th selected image in the j-th row direction and the width of the (k+1)-th selected image in the j-th row direction. In one embodiment, the x-axis coordinate of the center point of the item in the (k+1)th selected image in the j-th row direction is substantially the same as the x-axis coordinate of the center point of the item in the k-th selected image in the j-th row direction. The difference between the y-axis coordinate of the (k+1)th selected image in the j-th row direction and the y-axis coordinate of the center point of the item in the k-th selected image in the j-th row direction is substantially equal to the average of the width of the k-th selected image in the j-th row direction and the width of the (k+1)th selected image in the j-th row direction. In step S175, it is determined whether the estimated center point of the item in the (k+1)th selected image in the j-th row direction is located outside the range of the platform image T. Specifically, when the second vertex T2 of the platform image T is defined as the origin of the rectangular coordinate system, the lower limit of the y-axis coordinate is 0. Therefore, when the estimated vertical coordinate of the center point of the (k+1)th selected image in the j-th row direction is determined to be not less than the lower limit of the vertical coordinate, it indicates that the center point of the item in the (k+1)th selected image in the j-th row direction is not located outside the range of the platform image T, and then step S176 is executed. In step S176, k is set to k+1. After step S176, the process returns to step S172. In one embodiment, if there is no image corresponding to the (k+1)th position in the j-th row direction in step S173, for example, the (k+1)th position may be outside the platform range, steps S173~175 can be skipped, and step S177 is executed directly, setting k=1.
[0062] Figure 10 Drawing the second execution Figure 7AA schematic diagram of an embodiment of step S172. Specifically, the first item image P1 is superimposed on the platform image T1. Then, in step S173, a second item image Q2 is selected from the first item images P1~P9 and the second item images Q1~Q9 as the second selected image. The second item image Q2 has a second center point C2, which has a second vertical axis coordinate and a second horizontal axis coordinate. The second item image Q2 has a length L2 and a width W2 in the X-axis direction and the Y-axis direction, respectively. In step S174, the processor of the management host 4 estimates the position of the second center point C2 of the second item image Q2. In step S175, the processor of the management host 4 determines that the second vertical axis coordinate of the second center point C2 of the second item image Q2 is not less than the lower limit of the vertical axis coordinate (e.g., 0). Then, in step S176, k = k + 1 is set, and the process returns to step S172 to overlay the second item image Q2 onto the platform image T. The difference between the first center point C1 of the first item image P1 and the second center point C2 of the second item image Q2 is essentially equal to the average of the two widths W1 and W2. In one embodiment, the coordinates of the first center point C1 of the first item image P1 are... Figure 4 In the image Q2, the coordinates of the second center point C2 are (0, Ye-(W1+W2) / 2).
[0063] Figure 11 Drawing the third execution Figure 7A A schematic diagram of an embodiment of step S172. Specifically, the first item image P1 and the second item image Q2 are superimposed on the platform image T1. Then, in step S173, the second item image Q3 is selected from the first item images P1~P9 and the second item images Q1~Q9. The second item image Q3 has a third center point C3, and the third center point C3 has a third vertical axis coordinate and a third horizontal axis coordinate. The second item image Q3 has a length L3 and a width W3 in the X-axis direction and the Y-axis direction, respectively. In step S174, the processor of the management host 4 estimates the position of the third center point C3 of the second item image Q3. In step S175, if the processor of the management host 4 determines that the third vertical axis coordinate of the third center point C3 of the second item image Q3 is not less than the lower limit of the vertical axis coordinate (e.g., 0), then in step S176, k=k+1 is set, and the process returns to step S172 to overlay the second item image Q3 onto the platform image T. The difference between the center point C2 of the second item image Q2 and the third center point C3 of the second item image Q3 is essentially equal to the average of the two widths W2 and W3. In one embodiment, the coordinates of the second center point C2 of the second item image Q2 are, for example, (X2, Y2), and the coordinates of the third center point C3 of the second item image Q3 are, for example, (X2, Y2-(W2+W3) / 2).
[0064] If the center point of the item in the estimated (k+1)th selected image in the j-th row direction is outside the range of the platform image T, then proceed to step S177. In step S177, k=1. In step S178, select one of the first item images and the second item images corresponding to the k-th position in the (j+1)-th row direction as the k-th selected image in the (j+1)-th row. In one embodiment, this step can be divided into first selecting an item, and then selecting the corresponding image position. For example, select the first item, and then select the first item image corresponding to the k-th position in the (j+1)-th row direction from the first item images as the k-th selected image in the (j+1)-th row direction. In one embodiment, if there is no image corresponding to the k-th position in the (j+1)-th row direction, for example, the k-th position may be outside the platform range, step S178 can be skipped. In other embodiments, if the number of item types is changed to N, where N is a positive integer greater than or equal to three, then step S178 can be changed to selecting one of the k-th positions corresponding to the (j+1)-th row direction from the first item images to the N-th item images as the k-th selected image in the (j+1)-th row direction. In another embodiment, if the number of item types is changed to M, where M is a positive integer greater than or equal to two, then step S178 can be changed to selecting one of the k-th positions corresponding to the (j+1)-th row direction from the first item images to the M-th item images as the k-th selected image in the (j+1)-th row direction.
[0065] In step S179, the center point of the item after the k-th selected image in the (j+1)-th row direction overlaps with the platform image is estimated based on the lengths of the k-th selected image in the j-th row direction and the (j+1)-th selected image in the (j+1)-th row direction. The center point of the item in the (j+1)-th row direction has a horizontal axis coordinate and a vertical axis coordinate. The difference between the estimated center point of the k-th selected image in the (j+1)-th row direction and the center point of the k-th selected image in the j-th row direction is essentially equal to the average of the lengths of the k-th selected image in the (j+1)-th row direction and the lengths of the k-th selected image in the j-th row direction. In step S180, it is determined whether the estimated center point of the k-th selected image in the (j+1)-th row direction is located outside the range of the platform image T. Specifically, when the second vertex T2 of the platform image T is defined as the origin of the rectangular coordinate system, the upper limit of the horizontal axis coordinate value is Xe. Therefore, when the estimated x-axis coordinate of the k-th selected image in the (j+1)-th row direction is determined to be less than the upper limit of the x-axis coordinate, it indicates that the center point of the k-th selected image in the (j+1)-th row direction is not located outside the range of the platform image T, and then proceed to step S181. In step S181, set j = j+1. After step S181, return to step S172. In one embodiment, if there is no image corresponding to the k-th position in the (j+1)-th row direction in step S178, for example, the (j+1)-th row direction may have exceeded the platform range, steps S178~180 can be skipped, and step S182 can be executed directly to generate the composite image.
[0066] Figure 12 To illustrate the fourth execution Figure 7A A schematic diagram of an embodiment of step S172. Specifically, the first item image P1, the second item image Q2, and the second item image Q3 are superimposed on the platform image T1. Then, the first item image P4 is selected from the first item images P1~P9 and the second item images Q1~Q9. The first item image P4 has a fourth center point C4, and the fourth center point C4 has a fourth vertical axis coordinate and a fourth horizontal axis coordinate. The first item image P4 has a length L4 and a width W4 in the X-axis direction and the Y-axis direction, respectively. The processor of the management host 4 confirms that the fourth horizontal axis coordinate of the fourth center point C4 of the first item image P4 is less than the upper limit of the horizontal axis coordinate (e.g., Xe), and then superimposes the first item image P4 on the platform image T, wherein the difference between the fourth center point C4 of the first item image P4 and the first center point C1 of the first item image P1 is substantially equal to the average of the two widths L1 and L4. In one embodiment, the coordinates of the first center point C1 of the first item image P1 are Figure 4 In the image (0, Ye), the coordinates of the fourth center point C4 of the first object image P4 are ((L1+L4) / 2, Ye).
[0067] When the center point of the k-th selected image in the (j+1)-th row direction is outside the range of the platform image T, step S182 is executed. In step S182, a composite image is generated based on the overlapped platform images.
[0068] Management host 4 executes multiple times Figure 7A and Figure 7B Image compositing can produce multiple different composite images. These composite images can be used as modeling data, and machine learning algorithms (such as Faster R-CNN, regions with convolution neural network) are trained on these composite images to generate a model of the composite images as a whole. In one embodiment, after the model is generated, the objects placed on the platform can be identified using the model.
[0069] Understandably, the aforementioned Figure 7A and Figure 7B One embodiment of the image compositing process involves overlaying a first item image or a second item image onto a platform image, sequentially from the upper left corner toward the lower right corner. However, the image compositing process of this invention is not limited to this; alternatively, the first item image or the second item image can be overlaid onto the platform image first, sequentially from the left edge toward the center, and then sequentially from the right edge toward the center. In one embodiment, the number of item types is N, where N is a positive integer greater than or equal to three. Figure 7A and Figure 7B An embodiment of the image compositing processing can select item images from the first item image to the Nth item image, and sequentially overlay them onto the platform image from the top left corner to the bottom right corner. In one embodiment, the image compositing processing can select item images from the first item image to the Nth item image, and sequentially overlay them onto the platform image in multiple different row directions, multiple different column directions, from top to bottom, from bottom to top, from left to right, from right to left, first from top to bottom and then from bottom to top, first from bottom to top and then from top to bottom, first from left to right and then from right to left, first from right to left and then from left to right, from inside to outside, from outside to inside, or randomly. The placement of the item images and the proportion of items placed on the platform can be determined by system or user parameters. In another embodiment, there are M types of items, where M is a positive integer greater than or equal to two. Figure 7A and Figure 7BAn embodiment of the image compositing processing can select item images from the first to the Mth item images, choosing corresponding positions, and superimpose them onto the platform image from the top left corner downwards towards the bottom right corner. In one embodiment, the image compositing processing can superimpose item images from the first to the Mth item images, choosing corresponding positions or viewpoints, in a spiral pattern (first placing rows then columns, first columns then rows, top to bottom, bottom to top, left to right, right to left, first top to bottom with a row change then bottom to top, first bottom to top with a row change then top to bottom, first left to right with a column change then right to left, first right to left with a column change then left to right, inside to outside with a column change then left to right, outside to inside with a spiral pattern, or randomly onto the platform image. The placement method of the item images and the proportion of items placed on the platform can be determined by system or user parameters.
[0070] In one embodiment, the method further includes rotating, translating, or adjusting the brightness of the k-th selected image in the j-th row direction overlapping the platform image after step S172 and before step S173. Specifically, when multiple different items are placed on platform 1, two adjacent or partially overlapping items may have different brightness, or the center points of two adjacent or partially overlapping items may not be aligned with the same vertical or horizontal axis coordinate. Therefore, by rotating, translating, or adjusting the brightness of the first or second item image overlapping the platform image, the actual situation of multiple items being placed on the platform can be simulated, thereby improving the recognition rate of the item model.
[0071] In one embodiment, when the management host 4 executes step S172, it further includes deciding whether to overlay the k-th selected image in the j-th row direction onto the platform image. Specifically, the occurrence rate of each image is adjusted to simulate the situation where some positions are not occupied by items, such as items that may have been removed, thereby improving the recognition rate of the item model.
[0072] In one embodiment, a percentage or an upper limit value for the placement coordinates can be set, for example... Figure 7A , 7B After placing items sequentially at coordinates (Xl, Yl), step S182 is executed, where Xl is less than or equal to Xe and Yl is less than or equal to Ye, to simulate a situation where the platform is not full of items, thereby improving the recognition rate of the item model.
[0073] Figure 13 A flowchart illustrating a method for generating item modeling data according to another embodiment of the present invention. Figure 13As shown, in step S20, a platform image associated with the platform is obtained. In one embodiment, the image obtained after the camera captures the platform includes background images of the platform and other objects outside the platform. Background removal processing is performed first to obtain the platform image. In step S21, multiple images of a first item disposed on the platform are obtained, wherein these first item images correspond to multiple different viewpoints. In step S22, multiple images of a second item disposed on the platform are obtained, wherein these second item images correspond to the different viewpoints. In step S23, image compositing processing is performed based on at least one of the first item images and the second item images, and the platform image, to generate a composite image. The composite image includes at least multiple adjacent or partially overlapping image regions corresponding to the different viewpoints. These image regions include a first image region and a second image region. The first image region includes one of the first item images or one of the second item images, and the second image region includes one of the first item images or one of the second item images.
[0074] Figure 14 This is a comparison chart showing the frame extraction rate of object modeling based on synthetic images according to an embodiment of the present invention, and the frame extraction rate of single object modeling. For example... Figure 14 As shown, the horizontal axis represents the threshold value for the item selection ratio. In this embodiment, the item selection ratio is the intersection-over-union (IOU) ratio between the item and the selected item image. When selecting items, if the item selection ratio is greater than the threshold value, it is determined that the item has been selected. The vertical axis represents the recall rate. In this embodiment, the recall rate is the ratio of selected items to the number of items. Curve S1 represents the recall rate for multi-item modeling based on the synthetic image of this invention, while curve S2 represents the recall rate for single-item modeling. Figure 14 It can be seen that the frame extraction rate corresponding to curve S1 is higher than that corresponding to curve S2.
[0075] The item modeling data generation method and apparatus of the present invention allow multiple cameras to capture images of items from multiple different perspectives on a platform with only one shot, reducing the time required for image sampling. Furthermore, the composite images generated through image synthesis processing are used to simulate situations where multiple different items are adjacent or partially overlapping on the platform, eliminating the need for manually arranging multiple items in various combinations on the platform and then photographing the platform with the various items, further reducing the time required for modeling data generation. Moreover, the item models generated based on the composite images using machine learning algorithms have a high recognition rate and can be applied to product recognition in unmanned stores.
[0076] Of course, the present invention may have other various embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and modifications according to the present invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.
Claims
1. A method for generating composite images of objects, characterized in that, include: Obtain platform images associated with the platform; Obtain multiple images of the first item set on the platform, wherein each of the first item images corresponds to multiple different perspectives; Obtain multiple images of the second item set on the platform, each image corresponding to a different viewpoint; Image compositing is performed on at least one of the first object images and the second object images, along with the platform image, to generate a composite image. The composite image includes at least a plurality of adjacent or partially overlapping image regions corresponding to the different viewpoints. These image regions include first image regions and second image regions. The first image region includes one of the first object images or one of the second object images, and the second image region includes one of the first object images or one of the second object images. Obtain multiple images of the third item to the Nth item set on the platform, wherein each of these images corresponds to a different viewpoint. The process of performing the image compositing to generate the composite image based on at least one of the first to Nth item images and the platform image includes: Select one of the k-th positions in the j-th row direction from the first item images to the N-th item images as the k-th selected image in the j-th row direction, where the initial value of parameter j is 1 and the initial value of parameter k is 1. as well as The k-th selected image in the j-th row direction is superimposed on the k-th position in the j-th row direction of the platform image, and the k-th selected image in the j-th row direction has a first center point after being superimposed on the platform image.
2. The method for generating composite images of objects according to claim 1, characterized in that, The images of the first item set on the platform include: Set the first item on the platform; Obtain multiple initial images, each initial image including an image of the first item and the platform; and The initial images of the platform are processed to remove the background to generate images of the first objects.
3. The method for generating composite images of objects according to claim 2, characterized in that, These initial images were captured in a single shot by multiple cameras, with multiple lenses of these cameras pointing towards the platform in the same direction.
4. The method for generating composite images of objects according to claim 2, characterized in that, The first item is placed in multiple locations on the platform in sequence, and the first initial images are obtained by capturing the first item at those locations on the platform using a single camera.
5. The method for generating composite images of objects according to claim 1, characterized in that, It further includes acquiring multiple images of third items to Nth items set on the platform, wherein these images correspond to different viewpoints, and N is a positive integer greater than or equal to three. The image compositing process to generate the composite image is performed based on at least one of these images and the platform image. The platform image is superimposed on at least one of the first object images from different viewpoints, or at least one of the Nth object images, in a spiral pattern from the inside out, or a spiral pattern from the outside in, or in a random manner, according to multiple different row directions, multiple different column directions, from top to bottom, from bottom to top, from left to right, from right to left, from right to left, from inside to outside, or in a spiral pattern from outside to inside.
6. The method for generating composite images of articles according to claim 1, characterized in that, Including: Rotate, translate, adjust brightness, or adjust occurrence rate of the k-th selected image in the j-th row direction.
7. The method for generating composite images of articles according to claim 1, characterized in that, The process of generating the composite image based on at least one of the first to Nth item images and the platform image further includes: Select one of the (k+1)th positions in the j-th row direction from the first item images to the N-th item images as the (k+1)th selected image in the j-th row direction; Based on the width of the k-th selected image in the j-th row direction and the width of the (k+1)-th selected image in the j-th row direction, estimate the second center point of the (k+1)-th selected image in the j-th row direction after it overlaps with the platform image; and Determine whether the second center point of the selected image in the j-th row direction (k+1) is outside the range of the platform image.
8. The method for generating composite images of articles according to claim 7, characterized in that, The process of generating the composite image based on at least one of the first to Nth item images and the platform image further includes: When the second center point of the (k+1)th selected image in the j-th row direction is not located outside the range of the platform image, the (k+1)th selected image in the j-th row direction is superimposed on the platform image.
9. The method for generating composite images of articles according to claim 8, characterized in that, The difference between the second center point of the (k+1)th selected image in the j-th row direction and the first center point of the k-th selected image in the j-th row direction is equal to the average of the width of the k-th selected image in the j-th row direction and the width of the (k+1)th selected image in the j-th row direction.
10. The method for generating composite images of articles according to claim 7, characterized in that, Performing the image compositing process based on at least one of the first to Nth item images and the platform image to generate the composite image further includes: When the second center point of the (k+1)th selected image in the j-th row direction is outside the range of the platform image, one of the k-th positions in the (j+1)th row direction is selected from the first item images to the N-th item images as the k-th selected image in the (j+1)th row direction. Based on the lengths of the k-th selected image in the j-th row direction and the k-th selected image in the (j+1)-th row direction, estimate the third center point of the k-th selected image in the (j+1)-th row direction overlapping the platform image; and Determine whether the third center point of the k-th selected image in the (j+1)-th row direction is outside the range of the platform image. When the third center point of the k-th selected image in the (j+1)-th row direction is not located outside the range of the platform image, the k-th selected image in the (j+1)-th row direction is superimposed on the platform image. When the third center point of the k-th selected image in the (j+1)-th row direction is outside the range of the platform image, the composite image is generated based on all selected images that have been superimposed on the platform image and the platform image.
11. The method for generating composite images of articles according to claim 10, characterized in that, The difference between the third center point of the k-th selected image in the (j+1)-th row direction and the first center point of the k-th selected image in the j-th row direction is equal to the average of the length of the k-th selected image in the (j+1)-th row direction and the length of the k-th selected image in the j-th row direction.
12. An apparatus for generating composite images of objects, characterized in that, include: The platform is used to place the first item and the second item; At least one camera; as well as A management host, electrically connected to the at least one camera, is used to drive the at least one camera to capture images of the platform, the first object, and the second object, wherein the first object images correspond to multiple different viewpoints, and the second object images correspond to these different viewpoints. The management host performs image compositing processing based on at least one of the first object images and the second object images, along with the platform image, to generate a composite image. The composite image includes multiple adjacent or partially overlapping image regions corresponding to at least a plurality of the different viewpoints. These image regions include first image regions and second image regions. The first image regions include one of the first object images or one of the second object images, and the second image regions include one of the first object images or one of the second object images. The platform image has an upper limit for the horizontal axis coordinate and a lower limit for the vertical axis coordinate. The vertical axis coordinate of the center point of the first or second item images overlapping the platform image is greater than or equal to the lower limit of the vertical axis coordinate, and the horizontal axis coordinate of the center point of the first or second item images overlapping the platform image is less than or equal to the upper limit of the horizontal axis coordinate.
13. The article composite image generating apparatus according to claim 12, characterized in that, It also includes a frame, wherein the number of at least one camera is single, and the camera assembly is attached to the frame.
14. The article composite image generating apparatus according to claim 12, characterized in that, It also includes a slide rail mechanism, wherein the number of the at least one camera is a single camera assembly connected to the slide rail mechanism.
15. The article composite image generating apparatus according to claim 12, characterized in that, It also includes multiple racks, and the number of the at least one camera is multiple, with each camera fixed to the rack and facing the platform in the same direction.
16. The article composite image generating apparatus according to claim 15, characterized in that, These cameras are configured in an array.
17. The article composite image generating apparatus according to claim 12, characterized in that, The images of the first objects and the images of the second objects each have different brightness.
Citation Information
Patent Citations
Object image shooting and splicing method
CN104168414A
Method and device for stocktaking simulation and recording medium where stocktaking simulation program is recorded
JP2000163469A
Target positioning system and target positioning method
US20200202163A1