A method and system for shelf planogram auditing based on pre-roll pictures

CN122821531APending Publication Date: 2026-09-25ZHEJIANG SHANGSHI INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611144216.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-30
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]1.环境敏感与误报率高:主流图像比对方法(如图像差分、结构相似性指标SSIM等)对光照变化、拍摄角度、遮挡物等因素极为敏感,容易产生误报或漏报

Benefits of technology

[0062]1.本发明采用“粗对齐+精容错”双重策略,通过SIFT算法消除摄像头视角变化、货架位移等大尺度几何畸变,再通过局部软匹配机制容忍热胀冷缩、商品微移等残余微小位移,两者层层递进、分工明确,大幅提升系统对物理扰动的鲁棒性,显著降低因非实质性陈列变化导致的误报率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821531A_ABST
    Figure CN122821531A_ABST
Patent Text Reader

Abstract

The application discloses a shelf display checking method and system based on a previous picture, and the method comprises the following steps: a camera regularly shoots a shelf image; a feature point matching algorithm is used to perform spatial alignment on a current frame and a previous frame; the aligned image is input into a twin feature pyramid network, and multi-scale feature maps are extracted through a double-branch structure sharing weights; based on a local soft matching mechanism, the local maximum similarity is calculated in a neighborhood window, and a difference confidence map is generated; binary processing is performed according to a distribution adaptive threshold, and a difference mask is generated; and the difference mask is used to crop a difference block, which is uploaded to a platform. The application adopts a double strategy of 'rough alignment + fine error tolerance', eliminates large-scale geometric distortion through a SIFT algorithm, tolerates residual small displacement through a local soft matching mechanism, improves the robustness of the system to physical disturbance, reduces the false positive rate, and effectively saves network bandwidth by uploading only the difference block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of retail visual recognition technology, and more specifically, to a method and system for inventory management of shelf displays based on preceding images. Background Technology

[0002] With the continuous advancement of retail digitalization and smart store construction, real-time monitoring of shelf merchandise, out-of-stock detection, and display optimization have become key aspects of store operation and management. Accurately and efficiently obtaining the status of merchandise on shelves helps improve store operational efficiency, reduce stockouts, and optimize display compliance.

[0003] Traditional shelf inventory systems typically use fixed cameras to capture real-time images of the entire shelf, uploading the images to the cloud for comprehensive image analysis to identify product types, quantities, locations, and display changes. However, this method has significant technical limitations in practical applications:

[0004] 1. High environmental sensitivity and false alarm rate: Mainstream image comparison methods (such as image differencing and structural similarity index SSIM) are extremely sensitive to factors such as changes in lighting, shooting angle, and obstructions, easily leading to false alarms or missed alarms. Traditional point-to-point feature comparison methods are extremely sensitive to minor camera shakes, slight deformations of shelves, and minute deviations in the angle of product placement; even tiny pixel misalignments can cause widespread "false change" alarms.

[0005] 2. Lack of semantic understanding: Traditional methods are mostly based on pixel-level differences, which cannot accurately identify semantic-level changes such as swapping the positions of similar products or making minor adjustments to packaging.

[0006] 3. Limitations of single-scale features: Shelf items vary greatly in size (e.g., large gift boxes versus tiny chewing gum). Feature extraction based on a fixed patch size often suffers from inconsistencies. If the patch is too large, it is easy to overlook subtle changes; if the patch is too small, it lacks global semantic information about large objects, making it difficult to accurately determine whether an item has been "taken away" or "simply moved."

[0007] 4. Significant waste of data and computing power: Uploading complete high-definition images every day generates a large amount of redundant data traffic. Even if most areas of the shelf remain unchanged, the system still needs to perform comprehensive recognition and analysis on the entire image, resulting in a waste of computing power.

[0008] 5. Difficulty in balancing real-time performance and deployment cost: Running traditional image recognition models on embedded front-end devices is costly and inefficient, making it difficult to adapt to the real-time inventory requirements under high-frequency changes.

[0009] 6. Weak adaptability: It is difficult to guarantee detection robustness under different store scenarios (such as changes in lighting and different camera models), which limits large-scale promotion and implementation.

[0010] Therefore, there is an urgent need for an intelligent shelf inventory method that has multi-scale sensing capabilities, fault tolerance to physical disturbances, high robustness, low computing power consumption, small data transmission volume, and is suitable for deployment on edge devices. Summary of the Invention

[0011] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method and system for inventory management of shelf displays based on preceding images.

[0012] To achieve the above objectives, the present invention adopts the following technical solution:

[0013] A shelf display inventory method based on preceding images includes the following steps:

[0014] Step S1: The camera on the shelf takes a timed picture of the entire shelf.

[0015] Step S2: Obtain the current frame image and the previous frame image, perform coarse spatial alignment using a feature point matching algorithm, and output the geometrically aligned image;

[0016] Step S3: Input the aligned image into the pre-constructed Siamese feature pyramid network, and extract multi-scale feature maps of the current frame and the previous frame respectively through a dual-branch structure with shared weights to generate feature descriptors containing contextual semantic information.

[0017] Step S4: Based on the local soft matching mechanism, calculate the local maximum similarity between the features of the current frame and the previous frame within the neighborhood window of the feature map, and generate a difference confidence map to reflect the position change.

[0018] Step S5: Binarize the difference confidence map according to the distribution adaptive threshold, identify the changing regions, and generate a difference mask based on the changing regions.

[0019] Step S6: Cut out the difference patch based on the difference mask, upload the difference patch and its coordinate information to the platform, and the platform will perform splicing analysis based on the change area information to determine the display change type. The shelf end does not perform semantic level judgment.

[0020] Furthermore, in step S2, the coarse spatial alignment uses the SIFT algorithm to extract key points and calculate the homography matrix, transforming the current frame image to a coordinate system consistent with the previous frame image.

[0021] Furthermore, step S3 specifically includes the following steps:

[0022] Step S31: Construct two ResNet backbone network branches with shared weights, and input the aligned current frame image and the previous frame image into the two ResNet backbone network branches respectively.

[0023] Step S32: Draw side connections from multiple convolutional layers of different depths in the ResNet backbone network to obtain feature maps of different scales;

[0024] Step S33: Construct a twin feature pyramid through top-down paths and lateral connections, and fuse high-level strong semantic features with low-level high-resolution texture features to generate a fused feature map.

[0025] Step S34: The fused feature map is processed by two layers of dilated convolution to obtain shelf features at different scales;

[0026] Step S35: The feature map output by dilated convolution is recalibrated using a self-attention mechanism, and the recalibrated feature values ​​are obtained by weighted summation.

[0027] Step S36: Perform multi-scale stitching or adaptive weighted fusion on the channel dimension of the fused feature map set to generate the final feature vector corresponding to the Patch region of the original image.

[0028] Furthermore, in step S34, the feature output by the dilated convolution is expressed as follows:

[0029]

[0030] In the formula, For the first Features of the output of layered dilated convolution, is the layer number of the dilated convolution, and ; For activation functions; For batch normalization operations; Indicates the void ratio of convolution; For the first Features of the output of layered dilated convolution.

[0031] Furthermore, in step S35, the formula for calculating the self-attention coefficient is:

[0032]

[0033] In the formula, This is the sequence number of the current image patch; For the first The sequence number of the patch within the neighborhood centered on the current patch. For the neighborhood The sequence number of the inner patch; For the first The patch is relative to the first The self-attention coefficient of each region patch; For the first A query vector centered on a patch; For the first The key vector of each region patch; For the first The key vector of each region patch; The dimension of the feature vector; It is the transpose operator; It is an exponential function with the natural number e as its base; For the first Centered on a Patch Neighborhood range.

[0034] Furthermore, in step S35, the expression for the recalibrated eigenvalues ​​is:

[0035]

[0036] In the formula, For the first A vector of patch values ​​for each region; For the recalibrated first The feature vector of each patch.

[0037] Furthermore, in step S4, the local soft matching mechanism includes:

[0038] For the position coordinates on the feature map Using the corresponding position of this coordinate in the feature map of the preceding frame as the center, set a size of... Search neighborhood ;

[0039] Calculate the feature vector of the current frame With search neighborhood all preceding frame feature vectors Cosine similarity;

[0040] The maximum similarity value within the neighborhood is selected as the final change confidence level for that location. The calculation formula is:

[0041]

[0042] In the formula, These are the position coordinates in the feature map of the current frame; Neighborhood in the preceding feature map The internal position coordinates; , The preset fault tolerance radius, To search for the number of pixels with the side length of the neighborhood, , This is the lateral offset. This is the vertical offset. Represents the magnitude of a vector.

[0043] Furthermore, in step S5, the distribution adaptive threshold is calculated as follows:

[0044] Calculate the similarity values ​​of all patches in the entire image and form a similarity vector;

[0045] Based on the mean similarity and standard deviation Calculate the adaptive threshold for the distribution. :

[0046]

[0047]

[0048]

[0049] In the formula, This is the preset sensitivity adjustment coefficient; The total number of patches in the entire graph; Patch number index; For the first The similarity confidence value corresponding to each patch;

[0050] when When this occurs, the location is determined to be a region of change.

[0051] Furthermore, it also includes a baseline update strategy for the preceding frame image:

[0052] If a difference in the state of a certain region is detected in a continuous If a feature vector within a given time period remains stable and its variance is less than a preset stability threshold, then the difference is determined to be a persistent array change. The image of that region in the current frame is then updated in the previous frame as a new comparison benchmark. The preset time period quantity threshold, And it is a positive integer.

[0053] This invention also proposes a shelf display inventory system based on preceding images, the system comprising:

[0054] The image acquisition module is used to control the cameras deployed on the shelves to periodically capture complete images of the shelves and obtain the current frame image;

[0055] The image alignment module is used to read the previously stored frame image, use the SIFT algorithm to extract the key points of the current frame image and the previous frame image and calculate the homography matrix, transform the current frame image to the same coordinate system as the previous frame image, and output the geometrically aligned image.

[0056] The feature extraction module has a built-in pre-built Siamese feature pyramid network, which extracts multi-scale feature maps of the current frame and the previous frame through a dual-branch structure with shared weights.

[0057] The local soft matching module is used to calculate the local maximum cosine similarity between the features of the current frame and the previous frame within the neighborhood window of the feature map, and generate a difference confidence map.

[0058] The adaptive threshold decision module is used to calculate the mean and standard deviation of the similarity distribution of the entire image, calculate the adaptive threshold of the distribution, and perform binarization on the difference confidence map to generate a difference mask.

[0059] The data upload module is used to crop out the difference map based on the difference mask and upload the difference map and its coordinate information to the cloud platform;

[0060] The benchmark update module is used to update the status of a region after detecting a difference. When the image of the region remains stable within a certain time period and the feature variance is less than a preset threshold, the image of that region is updated to the previous frame storage as a new comparison benchmark.

[0061] The beneficial effects of this invention are:

[0062] 1. This invention adopts a dual strategy of "coarse alignment + fine fault tolerance". It uses the SIFT algorithm to eliminate large-scale geometric distortions such as changes in camera view and shelf displacement, and then uses a local soft matching mechanism to tolerate residual small displacements such as thermal expansion and contraction and micro-movement of goods. The two are progressive and have a clear division of labor, which greatly improves the robustness of the system to physical disturbances and significantly reduces the false alarm rate caused by non-substantial display changes.

[0063] 2. This invention constructs a twin feature pyramid network with shared weights, drawing side connections from multiple depths of the backbone network to fuse high-level semantic features with low-level texture features at multiple scales. It simultaneously considers the overall discrimination of large items and the detailed capture of small items, effectively solving the problem of "ignoring minute changes when the patch size is too large, and lacking global semantics when it is too small" under the fixed patch partitioning method, thus improving the overall accuracy of change detection. Furthermore, the two branches of the twin feature pyramid network share weights, ensuring that features from consecutive frames are in the same feature space, giving cosine similarity calculation clear physical meaning and comparability, and avoiding misjudgments caused by inconsistencies in feature spaces.

[0064] 3. This invention introduces two layers of dilated convolution to expand the network's receptive field, enhancing its ability to perceive large-scale changes in shelving. It also combines a self-attention mechanism to adaptively recalibrate the feature map, strengthening significant semantic regions while suppressing background noise and illumination fluctuations, thus significantly improving the detection robustness in complex retail scenarios.

[0065] 4. This invention only performs alignment, feature extraction, and comparison judgment at the edge, and then crops and uploads the difference patch and its coordinates, instead of uploading the complete shelf image. This can save bandwidth significantly and is especially suitable for 4G / 5G IoT environments, significantly reducing network operating costs and cloud computing power consumption.

[0066] 5. This invention utilizes the relatively fixed deployment of cameras and shelves. After the difference image blocks, carrying coordinate information, are uploaded to the platform, they can be accurately stitched to the corresponding positions in the complete shelf image based on the coordinates, achieving accurate mapping between the difference areas and the physical layout of the shelves. This method effectively reduces the complexity of backend processing and the cost of cloud restoration, further improving the actual engineering benefits of bandwidth saving, and possesses good implementation convenience and economy. Attached Figure Description

[0067] Figure 1 This is a flowchart of a shelf display inventory method based on preceding images in this embodiment;

[0068] Figure 2 This is a flowchart of a twin feature pyramid network in this embodiment;

[0069] Figure 3 This is a framework diagram of a shelf display and inventory system based on preceding images in this embodiment.

[0070] Figure labels: 1. Image acquisition module; 2. Image alignment module; 3. Feature extraction module; 4. Local soft matching module; 5. Adaptive threshold decision module; 6. Data upload module; 7. Benchmark update module. Detailed Implementation

[0071] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0072] Example: A shelf display inventory method based on preceding images. This method uses the SIFT algorithm to coarsely align the current frame with the preceding frame to eliminate large-scale geometric distortion. Then, it extracts multi-scale semantic features from the two frames using a shared-weighted Siamese feature pyramid network. Next, a local soft matching mechanism searches for the best match within a neighborhood window to tolerate residual small displacements after coarse alignment. Finally, it determines the changed areas based on a distribution-adaptive threshold and generates a difference mask, uploading only the difference patches and coordinate information to the platform. This method employs a dual strategy of "coarse alignment + fine fault tolerance," significantly saving network bandwidth and computing resources while ensuring detection accuracy.

[0073] like Figure 1 As shown, the method includes the following steps:

[0074] Step S1: The camera on the shelf takes a timed picture of the entire shelf.

[0075] Specifically, cameras are deployed at fixed points on the shelves, and the cameras take a complete picture of the shelf at preset time intervals (e.g., every 5 minutes) to obtain the current frame image. The size is ,in, Image width, Image height, It has three color channels: RGB. Previous frame image. (That is, the image taken in the previous time period, with the same size) Read from local storage.

[0076] Step S2: Obtain the current frame image and the previous frame image, perform coarse spatial alignment using a feature point matching algorithm, and output the geometrically aligned image.

[0077] This step aims to eliminate large-scale geometric distortions caused by changes in camera viewpoint, overall shelf displacement, and thermal expansion and contraction of mounting brackets, providing a preliminary spatial alignment basis for subsequent feature comparison.

[0078] Specifically, acquire the current frame image. and reading preceding frame images from local storage Then, the Scale-Invariant Feature Transform (SIFT) algorithm is used to extract key points from these two frames. The homography matrix is ​​then calculated through feature point matching, and the current frame image is then... Transform to the previous frame image Obtain the geometrically aligned current frame image in a consistent coordinate system. .

[0079] The SIFT algorithm detects extreme points in scale space to extract keypoints and generates 128-dimensional feature descriptors. Keypoint matching is then performed by comparing the Euclidean distance between the descriptors. Based on the matched keypoint pairs, the Random Sample Consensus (RANSAC) algorithm is used to estimate the homography matrix, making... .

[0080] SIFT coarse alignment can eliminate most geometric distortions, but residual offsets of 1-3 pixels may still exist due to the following reasons: 1. Uneven distribution of SIFT matching points at the edge of the shelf, resulting in insufficient local alignment accuracy; 2. Slight positional movement of the product itself after being slightly touched by a customer; 3. Subpixel-level thermal expansion and contraction deformation of the camera bracket due to temperature changes. These residual offsets will be tolerated by the local soft matching mechanism in the subsequent step S4.

[0081] Step S3: Input the aligned image into the pre-constructed Siamese feature pyramid network, and extract multi-scale feature maps of the current frame and the previous frame through a dual-branch structure with shared weights to generate feature descriptors containing contextual semantic information.

[0082] This step aims to extract deep semantic features from two consecutive frames of images and simultaneously obtain global category information and local texture details of the product through a multi-scale fusion strategy, providing distinctive and comparable feature descriptors for subsequent semantic-level change discrimination.

[0083] Specifically, such as Figure 2 As shown, this step includes the following steps:

[0084] Step S31, Backbone Network Feature Extraction: Construct a ResNet backbone network branch containing two shared weights, and extract geometrically aligned features from the current frame image. and preceding frame images Input two ResNet backbone branches.

[0085] The ResNet backbone uses ResNet-18, which includes an initial convolutional layer and four residual stages (Layer 1-Layer 4), each consisting of several residual blocks. Each branch outputs a preliminary feature map, with each location on the feature map corresponding to a local region in the original image. The feature vector at each location has a dimension of 512.

[0086] Step S32, Multi-level Feature Output: Instead of using only the output of the last layer of the ResNet backbone, side connections are drawn from multiple convolutional layers of different depths in the ResNet backbone to obtain feature maps of different scales. , , , , These correspond to feature maps output from different stages of the backbone network. As the index increases, the spatial resolution of the feature maps halve sequentially (stride sizes of 8, 16, 32, and 64), while the semantic level gradually increases. It retains rich spatial texture details. This implies high-level semantic category information.

[0087] Step S33, Siamese Feature Pyramid Fusion: A siamese feature pyramid is constructed through top-down paths and lateral connections to fuse high-level semantic features with low-level high-resolution texture features. The fusion process is as follows:

[0088] First, the characteristics of high-level buildings After a The convolutional layer performs channel dimensionality reduction, which serves as the initial fusion feature. Then, the fusion features will be... Upsampling is performed to improve its spatial resolution and features. Consistency, and then with the process Dimensionality reduction after convolution Element-by-element addition yields the fusion feature. Similarly, the fusion features Upsampling and features Fusion, resulting in fusion characteristics .

[0089] The fusion formula is expressed as:

[0090]

[0091] In the formula, This refers to the hierarchical number of the twin feature pyramid. ; Indicates the use of channel dimensionality reduction Convolution operation, which is about to The number of channels is uniformly set to 256; This indicates the ResNet backbone network's... The characteristic diagram of the stage; This indicates an upsampling operation, implemented using bilinear interpolation.

[0092] Through the above fusion, the fusion characteristics , , High-level semantic information and low-level texture information are fused at different resolutions, providing a rich feature base for subsequent multi-scale change detection.

[0093] Step S34, Dilated Convolution Enhancement: Based on the fused feature map set, two layers of dilated convolution are used to increase the network's receptive field, helping the model better capture global information about objects, thereby improving its ability to perceive large-scale changes in shelving. Dilated convolution expands the receptive field by inserting zero values ​​between elements of the convolution kernel, obtaining a larger context range without increasing the number of parameters and computational cost. Its calculation formula is as follows:

[0094]

[0095] In the formula, For the first Features of the output of layered dilated convolution, The layer number is the one representing the dilated convolution layer. In this embodiment, That is, the dilation rate of the first dilated convolution layer is 2, and the dilation rate of the second dilated convolution layer is 3; For activation functions, such as ReLU; This is a batch normalization operation used to accelerate convergence and prevent overfitting; Indicates the void ratio of convolution; For the first Features of the output of layered dilated convolution.

[0096] After two layers of dilated convolution, the receptive field of the network is significantly increased, enabling it to aggregate a wider range of contextual information without sacrificing feature map resolution, thereby improving the comprehensive perception of changes in both large and small products.

[0097] Step S35, Self-Attention Mechanism Recalibration: The feature map output by dilated convolution is recalibrated using a self-attention mechanism. Recalibration involves reweighting the original features using attention weights, allowing the model to focus on regions that are more discriminative for change detection. The self-attention mechanism models the correlation between each location on the feature map and all other locations (or locations within their neighborhood), adaptively adjusting the weights of each location. This enables the model to better focus on semantic regions important for change detection while suppressing background noise interference, exhibiting excellent robustness in complex retail shelf scenarios.

[0098] The formula for calculating the self-attention coefficient is:

[0099]

[0100] In the formula, This is the index of the current image patch, that is, the index of the current position on the feature map; For the first The sequence number of the patch within the neighborhood centered on the current patch. For the neighborhood The sequence number of the inner patch; For the first The patch is relative to the first The self-attention coefficient of each region patch; For the first The query vector centered on the patch is composed of the first patch. The features of each patch are obtained through linear transformation; For the first The key vector of each region patch is obtained by linear transformation of the features at the corresponding positions; For the first The key vector of each region patch is obtained by linear transformation of the features at the corresponding positions; In this embodiment, the dimension of the feature vector is... ; It is the transpose operator; It is an exponential function with the natural number e as its base; For the first Centered on a Patch Neighborhood range.

[0101] The recalibrated feature values ​​are obtained by weighted summing of the attention coefficients over the value vectors of each neighborhood patch:

[0102]

[0103] In the formula, For the first Each region's Patch value vector is obtained by linear transformation of the features at the corresponding locations; For the recalibrated first The feature vector of each patch retains the content information of the original features while incorporating the semantic association weights of the neighborhood context.

[0104] Step S36, Multi-scale Feature Fusion: The obtained fused feature map is spliced ​​across multiple scales along the channel dimension, i.e., the feature vectors at the same spatial location at different scales are concatenated end-to-end. To ensure that the spatial dimensions of the features at each scale are consistent after splicing, the fused features are first... and fusion features Upsampling is performed to adjust its resolution and fused features. The same features are then concatenated along the channel dimension to generate a unified feature vector corresponding to each patch region of the original image. and , The feature vector of the current frame (after SIFT alignment). This is the feature vector of the preceding frame.

[0105] Because the two ResNet backbone branches share network weights, the feature vectors and For elements at the same position in the same feature space, point-by-point cosine similarity can be calculated directly.

[0106] Step S3 involves constructing a Siamese feature pyramid network with shared weights to extract semantic-level features and ensure the comparability of features between consecutive frames. This Siamese feature pyramid network comprises two branches with identical structures and shared network parameters, each processing the current frame image. and preceding frame images The fact that the two branches share weights means that no matter which frame the image comes from, the same image content will be mapped to the same location in the feature space after passing through the network, thus ensuring the rationality of subsequent cosine similarity calculations.

[0107] Step S4: Based on the local soft matching mechanism, calculate the local maximum similarity between the features of the current frame and the previous frame within the neighborhood window of the feature map, and generate a difference confidence map to reflect the positional change.

[0108] This step aims to tolerate the 1-3 pixel residual micro-displacement that may still exist after the coarse alignment in step S2. By searching for the best match within the neighborhood window for each position, it avoids false alarms caused by non-substantial display changes such as slight camera vibrations, bracket deformation due to thermal expansion and contraction, and minor adjustments in the position of goods due to slight touches by customers. This significantly improves the system's robustness to physical disturbances.

[0109] The local soft matching mechanism includes:

[0110] For the position coordinates on the feature map Using the corresponding position of this coordinate in the feature map of the preceding frame as the center, set a size of... Search neighborhood , , In this embodiment, the preset fault tolerance radius is used. (Value can be 1 or 2) To search for the number of pixels with the side length of the neighborhood, , This is the lateral offset. This represents the vertical offset.

[0111] Calculate the feature vector of the current frame With search neighborhood all preceding frame feature vectors The cosine similarity.

[0112] The maximum similarity value within the neighborhood is selected as the final change confidence level for that location. The calculation formula is:

[0113]

[0114] In the formula, These are the position coordinates in the feature map of the current frame; Neighborhood in the preceding feature map The internal position coordinates; Represents the magnitude of a vector.

[0115] The above formula means that if a feature vector highly similar to that of the same location in the previous frame is found within the neighborhood (i.e., the "nearby" location) of a certain position in the current frame, we consider that the product has not undergone substantial changes (no shortage of stock, nor replacement with a new product), thus filtering out spurious changes caused by residual offsets. Only when the maximum similarity value within the neighborhood is found... Only when the value is low is it determined that the location has undergone a substantial change (such as the goods being taken away or replaced with a different category).

[0116] Step S5: Binarize the difference confidence map according to the distribution adaptive threshold, identify the change region, and generate a difference mask based on the change region.

[0117] This step aims to adaptively calculate the discrimination threshold based on the statistical characteristics of the overall image similarity distribution, eliminating the overall similarity shift caused by global style differences such as weather, lighting, and time of day, so that the determination of changing regions remains stable under different scene conditions, avoiding missed detections or false detections caused by fixed thresholds.

[0118] Considering that in real-world retail scenarios, factors such as weather changes, indoor lighting adjustments, and variations in natural light at different times of day can cause global differences in color, brightness, and contrast between two images, resulting in an overall shift (either an increase or decrease) in the similarity values ​​of all patches across the entire image. Using a fixed threshold cannot adapt to these variations, easily leading to a large number of false positives or false negatives. Therefore, a distributed adaptive threshold is adopted. Dynamic adjustments are made.

[0119] Specifically, the similarity values ​​of all patches in the entire image are counted to form a similarity vector, which is calculated in step S4. Then, the mean of the similarity vector is calculated. and standard deviation The calculation formula is as follows:

[0120]

[0121]

[0122] In the formula, The total number of patches in the entire graph; Patch number index; For the first The similarity confidence score for each patch. Mean Reflects the overall level of similarity across the entire image; standard deviation It reflects the degree of dispersion of similarity across the entire map, that is, the magnitude of the difference in the degree of change at each location.

[0123] Based on the calculated mean of the similarity vector and standard deviation Calculate the adaptive threshold for the distribution. :

[0124]

[0125] In the formula, This is a preset sensitivity adjustment coefficient used to control the judgment sensitivity. The larger the value, the more adaptive the distribution threshold. The lower the value, the more lenient the judgment of the change area, and more areas are judged as changes; The smaller the value, the more adaptive the distribution threshold. The higher the value, the more stringent the determination of the change region. In this embodiment, The value range is 1.0-1.5, and the specific value can be adjusted according to the actual scenario.

[0126] The judgment rule is: when the similarity confidence value at a certain location... If the position is considered a changed region, the corresponding position in the difference mask is set to 1; otherwise, it is considered unchanged and the corresponding position is set to 0.

[0127] Step S6: Cut out the difference patch based on the difference mask, upload the difference patch and its coordinate information to the platform, and the platform will perform splicing analysis based on the change area information to determine the display change type. The shelf end does not perform semantic level judgment.

[0128] Specifically, based on the difference mask generated in step S5, the device performs a matting operation, using the minimum bounding rectangle algorithm to crop out all image patches identified as change regions. Only these difference patches and their normalized coordinates in the original image are then processed. The packaged and compressed file was uploaded to the cloud server. , The normalized coordinates of the top-left corner of the difference patch in the original image. , These represent the width and height of the difference image block, respectively. Subsequent processing tasks such as product category identification, out-of-stock judgment, and display compliance analysis are completed by the cloud platform. For example, it may identify that a product in a difference image block has changed from "Cola" to "Sprite", or from "product available" to "empty space".

[0129] Furthermore, to avoid the cumulative errors caused by the lack of updates to preceding images over a long period (such as persistent changes like overall store display adjustments or seasonal product replacements that cause the initial preceding frames to gradually lose their value as effective comparison benchmarks), this embodiment also provides a benchmark update strategy:

[0130] If a difference in the state of a certain region is detected in a continuous If the feature vector within a given time period remains stable (i.e., the variance of the feature vector within that region is less than a preset stability threshold), then the difference is determined to be a persistent display change (i.e., not a momentary disturbance caused by a customer briefly taking an item and then putting it back, or by a stock clerk temporarily replenishing stock). The image of that region in the current frame is then updated to the corresponding region in the previous frame image as a new comparison benchmark.

[0131] in, The preset time period quantity threshold, And the result must be a positive integer. This strategy can effectively filter out momentary interference such as the stock clerk's arm obstructing the view or customers briefly touching the goods, ensuring that the inventory results reflect the steady-state display status.

[0132] This embodiment also proposes a shelf display inventory system based on preceding images, such as... Figure 3 As shown, the system includes an image acquisition module 1, an image alignment module 2, a feature extraction module 3, a local soft matching module 4, an adaptive threshold decision module 5, a data upload module 6, and a benchmark update module 7.

[0133] Among them, the image acquisition module 1 is responsible for controlling the camera deployed on the shelf to take full pictures of the shelf at regular intervals and obtain the current frame image.

[0134] Image alignment module 2 is responsible for reading the previous frame image stored locally, using the SIFT algorithm to extract the key points of the current frame image and the previous frame image and calculate the homography matrix, transforming the current frame image to the same coordinate system as the previous frame image, and outputting the geometrically aligned image.

[0135] Feature extraction module 3 incorporates a pre-built Siamese feature pyramid network, which extracts multi-scale feature maps of the current frame and previous frames through a dual-branch structure with shared weights. The Siamese feature pyramid network includes a ResNet backbone, feature pyramid fusion layers, dilated convolutional layers, and self-attention layers.

[0136] The local soft matching module 4 is responsible for calculating the local maximum cosine similarity between the features of the current frame and the previous frame within the neighborhood window of the feature map, and generating a difference confidence map.

[0137] The adaptive threshold decision module 5 is responsible for calculating the mean and standard deviation of the similarity distribution of the entire image, calculating the adaptive threshold of the distribution, and binarizing the difference confidence map to generate a difference mask.

[0138] The data upload module 6 is responsible for cropping the difference map based on the difference mask and uploading the difference map and its coordinate information to the cloud platform.

[0139] The baseline update module 7 is responsible for updating the baseline status of a region after detecting a difference in the continuous state. When the image of the region remains stable within a certain time period and the feature variance is less than a preset threshold, the image of that region is updated to the previous frame storage as a new comparison benchmark.

[0140] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A method for inventorying shelf displays based on preceding images, characterized in that, Includes the following steps: Step S1: The camera on the shelf takes a timed picture of the entire shelf. Step S2: Obtain the current frame image and the previous frame image, perform coarse spatial alignment using a feature point matching algorithm, and output the geometrically aligned image; Step S3: Input the aligned image into the pre-constructed Siamese feature pyramid network, and extract multi-scale feature maps of the current frame and the previous frame respectively through a dual-branch structure with shared weights to generate feature descriptors containing contextual semantic information. Step S4: Based on the local soft matching mechanism, calculate the local maximum similarity between the features of the current frame and the previous frame within the neighborhood window of the feature map, and generate a difference confidence map to reflect the position change. Step S5: Binarize the difference confidence map according to the distribution adaptive threshold, identify the changing regions, and generate a difference mask based on the changing regions. Step S6: Cut out the difference patch based on the difference mask, upload the difference patch and its coordinate information to the platform, and the platform will perform splicing analysis based on the change area information to determine the display change type. The shelf end does not perform semantic level judgment.

2. The shelf display inventory method based on preceding images according to claim 1, characterized in that, In step S2, coarse spatial alignment uses the SIFT algorithm to extract key points and calculate the homography matrix, transforming the current frame image to a coordinate system consistent with the previous frame image.

3. The shelf display inventory method based on preceding images according to claim 1, characterized in that, Step S3 specifically includes the following steps: Step S31: Construct two ResNet backbone network branches with shared weights, and input the aligned current frame image and the previous frame image into the two ResNet backbone network branches respectively. Step S32: Draw side connections from multiple convolutional layers of different depths in the ResNet backbone network to obtain feature maps of different scales; Step S33: Construct a twin feature pyramid through top-down paths and lateral connections, and fuse high-level strong semantic features with low-level high-resolution texture features to generate a fused feature map. Step S34: The fused feature map is processed by two layers of dilated convolution to obtain shelf features at different scales; Step S35: The feature map output by dilated convolution is recalibrated using a self-attention mechanism, and the recalibrated feature values ​​are obtained by weighted summation. Step S36: Perform multi-scale stitching or adaptive weighted fusion on the channel dimension of the fused feature map set to generate the final feature vector corresponding to the Patch region of the original image.

4. The shelf display inventory method based on preceding images according to claim 3, characterized in that, In step S34, the feature output by the dilated convolution is expressed as follows: In the formula, For the first Features of the output of layered dilated convolution, is the layer number of the dilated convolution, and ; For activation functions; For batch normalization operations; Indicates the void ratio of convolution; For the first Features of the output of layered dilated convolution.

5. The shelf display inventory method based on preceding images according to claim 3, characterized in that, In step S35, the formula for calculating the self-attention coefficient is: In the formula, This is the sequence number of the current image patch; For the first The sequence number of the patch within the neighborhood centered on the current patch. For the neighborhood The sequence number of the inner patch; For the first The patch is relative to the first The self-attention coefficient of each region patch; For the first A query vector centered on a patch; For the first The key vector of each region patch; For the first The key vector of each region patch; The dimension of the feature vector; It is the transpose operator; It is an exponential function with the natural number e as its base; For the first Centered on a Patch Neighborhood range.

6. The shelf display inventory method based on preceding images according to claim 5, characterized in that, In step S35, the expression for the recalibrated eigenvalues ​​is: In the formula, For the first A vector of patch values ​​for each region; For the recalibrated first The feature vector of each patch.

7. The shelf display inventory method based on preceding images according to claim 1, characterized in that, In step S4, the local soft matching mechanism includes: For the position coordinates on the feature map Using the corresponding position of this coordinate in the feature map of the preceding frame as the center, set a size of... Search neighborhood ; Calculate the feature vector of the current frame With search neighborhood all preceding frame feature vectors Cosine similarity; The maximum similarity value within the neighborhood is selected as the final change confidence level for that location. The calculation formula is: In the formula, These are the position coordinates in the feature map of the current frame; Neighborhood in the preceding feature map The internal position coordinates; , The preset fault tolerance radius, To search for the number of pixels with a side length in the neighborhood, , This is the lateral offset. This is the vertical offset; Represents the magnitude of a vector.

8. The shelf display inventory method based on preceding images according to claim 1, characterized in that, In step S5, the distribution adaptive threshold is calculated as follows: Calculate the similarity values ​​of all patches in the entire image and form a similarity vector; Based on the mean similarity and standard deviation Calculate the adaptive threshold for the distribution. : In the formula, This is the preset sensitivity adjustment coefficient; The total number of patches in the entire graph; Patch number index; For the first The similarity confidence value corresponding to each patch; when When this occurs, the location is determined to be a region of change.

9. The shelf display inventory method based on preceding images according to claim 1, characterized in that, It also includes a baseline update strategy for the preceding frame image: If a difference in the state of a certain region is detected in a continuous If a feature vector within a given time period remains stable and its variance is less than a preset stability threshold, then the difference is determined to be a persistent array change. The image of that region in the current frame is then updated in the previous frame as the new comparison benchmark. The preset time period quantity threshold, And it is a positive integer.

10. A shelf display and inventory system based on preceding images for implementing the method of claim 1, characterized in that, The system includes: Image acquisition module (1) is used to control the camera deployed on the shelf to take full shelf images at regular intervals and obtain the current frame image; Image alignment module (2) is used to read the previous frame image stored locally, extract the key points of the current frame image and the previous frame image using the SIFT algorithm and calculate the homography matrix, transform the current frame image to the same coordinate system as the previous frame image, and output the geometrically aligned image. The feature extraction module (3) has a built-in pre-built twin feature pyramid network, which extracts multi-scale feature maps of the current frame and the previous frame respectively through a dual-branch structure with shared weights. The local soft matching module (4) is used to calculate the local maximum cosine similarity between the features of the current frame and the previous frame within the neighborhood window of the feature map and generate a difference confidence map. The adaptive threshold decision module (5) is used to calculate the mean and standard deviation of the similarity distribution of the whole image, calculate the adaptive threshold of the distribution, and perform binarization on the difference confidence map to generate a difference mask. The data upload module (6) is used to cut out the difference map according to the difference mask and upload the difference map and its coordinate information to the cloud platform; The benchmark update module (7) is used to update the benchmark status continuously when a difference in a certain region is detected. When the image of the region remains stable within a certain time period and the feature variance is less than a preset threshold, the image of that region is updated to the previous frame storage as a new comparison benchmark.