A method, device and storage medium for commodity positioning based on dynamic vision
Through three-frame differential algorithm and preprocessing technology, the problem of environmental interference in the product positioning method is solved, and higher precision product positioning is achieved.
Patent Information
- Application Number
- CN202210096069.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-01-26
AI Technical Summary
Existing product positioning methods are susceptible to environmental interference, making it difficult to accurately locate products.
The three-frame difference algorithm is used to obtain the intersecting images in the video, and through expansion convolution processing and contour extraction, the preset conditional areas are blocked, the appropriate square areas are intercepted as the moving product area, and the target detection model is input to obtain the product category and position.
Effectively reduce the impact of environmental interference on commodity positioning and improve the accuracy and accuracy of commodity positioning.
Smart Images

Figure CN114596330B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a commodity positioning method, device and storage medium based on dynamic vision. Background Art
[0002] Moving object detection involves segmenting a dynamic target of interest from the image background. It is primarily used in image analysis and target tracking. Target motion, background changes, changes in illumination of the target or background, and occlusion between the target and other objects can all affect the detection results. Existing product location methods utilize algorithms such as optical flow, background subtraction, and frame subtraction. However, these methods are susceptible to environmental interference, making accurate product location difficult. Summary of the Invention
[0003] The present invention provides a commodity positioning method, device and storage medium based on dynamic vision, so as to solve the technical problem that the existing commodity positioning method has weak recognition ability and is difficult to accurately locate the commodity.
[0004] An embodiment of the present invention provides a product positioning method based on dynamic vision, comprising:
[0005] Acquire three consecutive frames of images in a video, and use a three-frame difference algorithm to acquire an intersecting image of the three consecutive frames of images;
[0006] Performing dilation convolution processing on the intersecting image to obtain a preprocessed image;
[0007] Masking areas in the preprocessed image that meet preset conditions, and performing contour extraction to obtain connected areas;
[0008] If the width or height of the connected domain is greater than a preset value, a square area with a maximum side length around the center of the connected domain is cut off as the mobile commodity area; if the width or height of the connected domain is not greater than the preset value, a square area with a side length of the preset value around the center of the connected domain is cut off as the mobile commodity area.
[0009] Furthermore, after the mobile commodity area is intercepted, the mobile commodity area is input into a preset target detection model, and the category and location information of the commodity is obtained according to the mobile commodity area.
[0010] Furthermore, the three consecutive image frames include the t-th frame image, the t+1-th frame image, and the t+2-th frame image, and a three-frame difference algorithm is used to obtain an intersecting image of the three consecutive image frames, specifically:
[0011] Performing a differential operation on the t-th frame image and the t+1-th frame image to obtain a first differential image, and performing a differential operation on the t+1-th frame image and the t+2-th frame image to obtain a second differential image;
[0012] performing binarization processing on the first differential image and the second differential image respectively to obtain a first binarized image and a second binarized image;
[0013] A logical AND operation is performed on the first binarized image and the second binarized image to obtain an intersection image of the first binarized image and the second binarized image.
[0014] Furthermore, the intersecting image is subjected to dilation convolution processing to obtain a preprocessed image, specifically:
[0015] A pre-processed image is obtained by selecting a kernel with a width of 5 and a height of 1 to perform dilated convolution processing on the intersecting image.
[0016] Furthermore, the area in the pre-processed image that meets the preset conditions is shielded, and contour extraction is performed to obtain a connected domain, specifically:
[0017] The area around the border of the pre-processed image is shielded, and contour extraction is performed to obtain a connected domain whose area is larger than the maximum area of the background connected domain and smaller than the minimum area of the product.
[0018] Furthermore, the area in the pre-processed image that meets the preset conditions is shielded, and contour extraction is performed to obtain a connected domain, specifically:
[0019] The brightness of each area in the preprocessed image is detected, the areas in the preprocessed image whose brightness is not within a preset threshold range are shielded, and contour extraction is performed to obtain a connected domain whose area is greater than the maximum value of the background connected domain area and smaller than the minimum value of the product area.
[0020] Furthermore, the preset value is 128.
[0021] One embodiment of the present invention provides a commodity positioning device based on dynamic vision, comprising:
[0022] An intersection image acquisition module is used to acquire three consecutive frames of images in a video and to acquire an intersection image of the three consecutive frames of images using a three-frame difference algorithm;
[0023] A preprocessing module, configured to perform dilation convolution processing on the intersecting image to obtain a preprocessed image;
[0024] A connected domain extraction module is used to mask the area that meets the preset conditions in the pre-processed image and perform contour extraction to obtain a connected domain;
[0025] The product positioning module is configured to, if the width or height of the connected domain is greater than a preset value, intercept a square area with a maximum side length around the center of the connected domain as a movable product area; if the width or height of the connected domain is not greater than the preset value, intercept a square area with a side length of the preset value around the center of the connected domain as a movable product area.
[0026] One embodiment of the present invention provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the above-mentioned dynamic vision-based product positioning method.
[0027] The embodiment of the present invention adopts a three-frame difference algorithm to obtain the intersection image of the image, and performs preprocessing to obtain a preprocessed image. The area in the preprocessed image that meets the preset conditions is shielded, which can effectively reduce the interference of the remaining areas on the product area. The embodiment of the present invention performs contour extraction to obtain a connected domain whose area is greater than the maximum value of the background connected domain area and smaller than the minimum value of the product area, which can further reduce the impact of the background area or the remaining areas on the product area, thereby effectively improving the accuracy of product positioning.
[0028] Furthermore, the embodiment of the present invention can also determine the brightness of the preprocessed image and compare it with a preset threshold range, and mask the area in the preprocessed image that will affect product positioning based on the brightness, thereby further improving the accuracy of product positioning. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 1 is a flow chart of a method for commodity positioning based on dynamic vision provided by an embodiment of the present invention;
[0030] Figure 2 3 is a structural diagram of a commodity positioning device based on dynamic vision provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0031] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0032] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "plurality" means two or more.
[0033] In the description of this application, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal connections between two components. Those skilled in the art will understand the specific meanings of the above terms in this application based on the specific circumstances.
[0034] See also Figure 1 One embodiment of the present invention provides a commodity positioning method based on dynamic vision, comprising:
[0035] S1, obtaining three consecutive frames of images in the video, and using a three-frame difference algorithm to obtain an intersecting image of the three consecutive frames of images;
[0036] In an embodiment of the present invention, the three-frame difference algorithm is an improved algorithm based on the two-adjacent-frame difference algorithm. By selecting three consecutive frames of images for difference calculation, the embodiment of the present invention can effectively eliminate the influence of the image background on product recognition due to movement during the movement of the product, and obtain the intersection image of the three consecutive frames of images, so as to accurately extract the motion contour information of the product, which is conducive to improving the accuracy of product positioning.
[0037] S2, performing dilation convolution processing on the intersecting image to obtain a preprocessed image;
[0038] It is understood that goods typically move horizontally. This embodiment of the present invention can perform dilated convolution on the intersecting images with a width of 5 and a height of 1, horizontally connecting the connected domains to ensure the integrity of the area where the goods move. In one specific embodiment, vertical connections are not performed, thereby minimizing the area of the moving region and reducing the resolution of the captured thumbnails.
[0039] S3, shielding the area that meets the preset conditions in the preprocessed image, and performing contour extraction to obtain a connected domain;
[0040] In the implementation of the present invention, the area of the preset conditions can be the border area of the preprocessed image. For example, the preprocessed image is a square with a side length of 100px. After shielding the border area of the preprocessed image, a square with a side length of 90px is obtained. The area of the preset conditions can also be an area whose brightness is not within the preset threshold range. By shielding the area that meets the preset conditions in the preprocessed image, the embodiment of the present invention can effectively reduce the interference of irrelevant areas in the preprocessed image on the positioning of the product, such as irrelevant areas of the border, and the influence of excessively high or low brightness on the positioning of the product. When the connected domain is obtained by performing contour extraction, the motion contour information of the product can be accurately extracted, thereby effectively improving the accuracy of product positioning.
[0041] S4. If the width or height of the connected domain is greater than a preset value, a square area with a maximum side length around the center of the connected domain is cut off as the mobile product area; if the width or height of the connected domain is not greater than the preset value, a square area with a side length of the preset value around the center of the connected domain is cut off as the mobile product area.
[0042] In the embodiment of the present invention, after extracting the connected domain, it is necessary to further determine whether the width or height of the connected domain is greater than a preset value. In the embodiment of the present invention, it is known through experiments that the width or height of the moving area is mainly between 100 and 200. The image length and width being an integer multiple of 32 is conducive to deep learning reasoning, and scaling the image to its original size is not conducive to deep learning reasoning. Figure 1 / 2, the recognition effect is poor, and the effect is almost unchanged when it is scaled to 3 / 4 of the original image. Therefore, in a specific embodiment, the preset value is 128. If the width or height of the connected domain is greater than 128, then based on the center of the connected domain, a square area with the maximum side length is selected around it as the mobile commodity area, that is, the positioning of the commodity during dynamic movement in the smart cabinet is realized; if the width or height of the connected domain is not greater than 128, based on the center of the connected domain, an area with a side length of 128 is directly selected around it as the mobile commodity area. In a specific embodiment, if the side length of the square area with the maximum side length is greater than 128, the square can be further scaled to a side length of 128 in proportion, so that when the category and location information of the commodity is subsequently obtained according to the target detection model, the pressure on the system operation can be effectively reduced while ensuring the detection effect.
[0043] In one embodiment, after the mobile commodity area is captured, the mobile commodity area is input into a pre-set object detection model, and the category and location information of the commodity is obtained based on the mobile commodity area.
[0044] In the embodiment of the present invention, a target detection model can be pre-trained and can be used to detect and classify targets according to the area of the moving goods, so as to further obtain the category information and location information of the goods.
[0045] In one embodiment, the three consecutive frames of images include the t-th frame image, the t+1-th frame image, and the t+2-th frame image. A three-frame difference algorithm is used to obtain an intersecting image of the three consecutive frames of images, specifically:
[0046] Performing a differential operation on the t-th frame image and the t+1-th frame image to obtain a first differential image, and performing a differential operation on the t+1-th frame image and the t+2-th frame image to obtain a second differential image;
[0047] Binarizing the first difference image and the second difference image respectively to obtain a first binarized image and a second binarized image;
[0048] A logical AND operation is performed on the first binarized image and the second binarized image to obtain an intersection image of the first binarized image and the second binarized image.
[0049] In one embodiment, the intersecting image is subjected to dilated convolution processing to obtain a preprocessed image, specifically:
[0050] The preprocessed image is obtained by performing dilated convolution on the intersecting image using a kernel with a width of 5 and a height of 1.
[0051] In a specific embodiment, the kernel with a width of 5 and a height of 1 is:
[0052] (0, 0, 0, 0, 0)
[0053] (0, 0, 0, 0, 0)
[0054] (1, 1, 1, 1, 1)
[0055] (0, 0, 0, 0, 0)
[0056] (0, 0, 0, 0, 0).
[0057] In the embodiment of the present invention, a kernel with a width of 5 and a height of 1 is selected to perform dilated convolution processing on the intersecting images, and the integrity of the commodity movement area is ensured by horizontally connecting the connected domains.
[0058] In one embodiment, the areas in the pre-processed image that meet the preset conditions are shielded, and contour extraction is performed to obtain connected areas, specifically:
[0059] The area around the border of the preprocessed image is shielded, and contour extraction is performed to obtain a connected domain whose area is larger than the maximum area of the background connected domain and smaller than the minimum area of the product.
[0060] It can be understood that the maximum value of the background connected domain area and the minimum value of the commodity cotton neps can be determined by an image recognition method.
[0061] In an embodiment of the present invention, the present invention is applicable to the positioning detection of goods in the process of movement in a smart cabinet. The moving image of the goods is captured by a camera set near the smart cabinet. The moving area of the goods is not located on the border of the moving image. The embodiment of the present invention shields the area around the border of the pre-processed image, which is conducive to reducing the interference of irrelevant areas on the recognition of goods. Furthermore, the embodiment of the present invention extracts a connected domain whose area is greater than the maximum value of the background connected domain area and less than the minimum value of the product area, thereby effectively distinguishing between the background area and the product area, and further reducing the interference of the background area on the recognition of the product area, which is conducive to improving the accuracy of product positioning.
[0062] In one embodiment, the areas in the pre-processed image that meet the preset conditions are shielded, and contour extraction is performed to obtain connected areas, specifically:
[0063] The brightness of each area in the preprocessed image is detected, the areas in the preprocessed image whose brightness is not within the preset threshold range are shielded, and contour extraction is performed to obtain the connected domain whose area is greater than the maximum value of the background connected domain area and smaller than the minimum value of the product area.
[0064] In an embodiment of the present invention, areas corresponding to excessively high and low brightness will not be the commodity areas in the smart cabinet. For example, the area corresponding to high brightness may be the light in the smart cabinet, which will interfere with the recognition of the target movement of the commodity when the smart cabinet is turned on. The embodiment of the present invention shields this area to reduce the impact of areas such as lights on commodity positioning, thereby effectively improving the accuracy of commodity positioning.
[0065] The implementation of the present invention has the following beneficial effects:
[0066] The embodiment of the present invention adopts a three-frame difference algorithm to obtain the intersection image of the image, and performs preprocessing to obtain a preprocessed image. The area in the preprocessed image that meets the preset conditions is shielded, which can effectively reduce the interference of the remaining areas on the product area. The embodiment of the present invention performs contour extraction to obtain a connected domain whose area is greater than the maximum value of the background connected domain area and smaller than the minimum value of the product area, which can further reduce the impact of the background area or the remaining areas on the product area, thereby effectively improving the accuracy of product positioning.
[0067] Furthermore, the embodiment of the present invention can also determine the brightness of the preprocessed image and compare it with a preset threshold range, and mask the area in the preprocessed image that will affect product positioning based on the brightness, thereby further improving the accuracy of product positioning.
[0068] Based on the same inventive concept as the above embodiment, please refer to Figure 2 One embodiment of the present invention provides a commodity positioning device based on dynamic vision, comprising:
[0069] The intersection image acquisition module 21 is used to acquire three consecutive frames of images in the video and use a three-frame difference algorithm to acquire an intersection image of the three consecutive frames of images;
[0070] A preprocessing module 22 is used to perform dilation convolution processing on the intersecting image to obtain a preprocessed image;
[0071] A connected domain extraction module 23 is used to mask the areas that meet the preset conditions in the pre-processed image and perform contour extraction to obtain connected domains;
[0072] The product positioning module 24 is configured to, if the width or height of the connected domain is greater than a preset value, intercept a square area with a maximum side length around the center of the connected domain as the mobile product area; if the width or height of the connected domain is not greater than the preset value, intercept a square area with a side length of the preset value around the center of the connected domain as the mobile product area.
[0073] In one embodiment, the device further includes a target detection module for inputting the mobile commodity area into a pre-set target detection model, and obtaining the category and location information of the commodity according to the mobile commodity area.
[0074] In one embodiment, the intersection image acquisition module 21 is used to:
[0075] Performing a differential operation on the t-th frame image and the t+1-th frame image to obtain a first differential image, and performing a differential operation on the t+1-th frame image and the t+2-th frame image to obtain a second differential image;
[0076] Binarizing the first difference image and the second difference image respectively to obtain a first binarized image and a second binarized image;
[0077] A logical AND operation is performed on the first binarized image and the second binarized image to obtain an intersection image of the first binarized image and the second binarized image.
[0078] In one embodiment, the pre-processing module 22 is configured to:
[0079] The preprocessed image is obtained by performing dilated convolution on the intersecting image using a kernel with a width of 5 and a height of 1.
[0080] In one embodiment, the connected domain extraction module 23 is used to:
[0081] The area around the border of the preprocessed image is shielded, and contour extraction is performed to obtain a connected domain whose area is larger than the maximum area of the background connected domain and smaller than the minimum area of the product.
[0082] In one embodiment, the connected domain extraction module 23 is used to:
[0083] The brightness of each area in the preprocessed image is detected, the areas in the preprocessed image whose brightness is not within the preset threshold range are shielded, and contour extraction is performed to obtain the connected domain whose area is greater than the maximum value of the background connected domain area and smaller than the minimum value of the product area.
[0084] In one embodiment, the preset value is 128.
[0085] One embodiment of the present invention provides a computer-readable storage medium, which includes a stored computer program. When the computer program runs, the device where the computer-readable storage medium is located is controlled to perform the above-mentioned dynamic vision-based product positioning.
[0086] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A commodity positioning method based on dynamic vision, characterized in that: include: Acquire three consecutive frames of images in a video, and use a three-frame difference algorithm to acquire an intersecting image of the three consecutive frames of images; Performing dilation convolution processing on the intersecting image to obtain a preprocessed image; Shielding areas in the preprocessed image that meet preset conditions and performing contour extraction to obtain connected domains; shielding areas in the preprocessed image that meet preset conditions and performing contour extraction to obtain connected domains, specifically: detecting the brightness of each area in the preprocessed image, shielding areas in the preprocessed image whose brightness is not within a preset threshold range, and performing contour extraction to obtain connected domains whose area is greater than the maximum area of the background connected domain and smaller than the minimum area of the product; If the width or height of the connected domain is greater than a preset value, a square area with a maximum side length around the center of the connected domain is cut off as the mobile commodity area; if the width and height of the connected domain are not greater than the preset value, a square area with a side length of the preset value around the center of the connected domain is cut off as the mobile commodity area.
2. The method for commodity positioning based on dynamic vision according to claim 1, characterized in that: After the mobile commodity area is intercepted, the mobile commodity area is input into a preset target detection model, and the category and location information of the commodity is obtained according to the mobile commodity area.
3. The method for commodity positioning based on dynamic vision according to claim 1, characterized in that: The three consecutive frames of images include the t-th frame image, the t+1-th frame image, and the t+2-th frame image. A three-frame difference algorithm is used to obtain an intersecting image of the three consecutive frames of images, specifically: Performing a differential operation on the t-th frame image and the t+1-th frame image to obtain a first differential image, and performing a differential operation on the t+1-th frame image and the t+2-th frame image to obtain a second differential image; performing binarization processing on the first differential image and the second differential image respectively to obtain a first binarized image and a second binarized image; A logical AND operation is performed on the first binarized image and the second binarized image to obtain an intersection image of the first binarized image and the second binarized image.
4. The method for commodity positioning based on dynamic vision according to claim 1, characterized in that: The intersecting image is subjected to dilation convolution processing to obtain a preprocessed image, specifically: A pre-processed image is obtained by selecting a kernel with a width of 5 and a height of 1 to perform dilated convolution processing on the intersecting image.
5. The method for commodity positioning based on dynamic vision according to claim 1, characterized in that: The preset value is 128.
6. A commodity positioning device based on dynamic vision, characterized in that: include: An intersection image acquisition module is used to acquire three consecutive frames of images in a video and to acquire an intersection image of the three consecutive frames of images using a three-frame difference algorithm; A preprocessing module, configured to perform dilation convolution processing on the intersecting image to obtain a preprocessed image; A connected domain extraction module is used to mask the area that meets the preset conditions in the pre-processed image and perform contour extraction to obtain a connected domain; Specifically used for: detecting the brightness of each area in the pre-processed image, shielding the area in the pre-processed image whose brightness is not within a preset threshold range, and performing contour extraction to obtain a connected domain whose area is greater than the maximum value of the background connected domain area and smaller than the minimum value of the product area; The product positioning module is configured to, if the width or height of the connected domain is greater than a preset value, intercept a square area with a maximum side length around the center of the connected domain as a movable product area; if the width and height of the connected domain are not greater than the preset value, intercept a square area with a side length of the preset value around the center of the connected domain as a movable product area.
7. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to perform the dynamic vision-based commodity positioning according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and device for detecting specific identification image in predetermined area
CN105303189A
Automatic vender and automatic vending system
CN105654618A