Customer behavior identification method and device based on multi-view camera
Through the customer behavior recognition method based on multi-eye camera, a three-dimensional cost volume and depth map is constructed, the three-dimensional bounding box is detected and moving targets is tracked, and the problem of product loss in the existing self-service shopping system is solved, achieving high accuracy and high efficiency behavior recognition.
Patent Information
- Application Number
- CN202411803644.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-05-13
AI Technical Summary
In the absence of monitoring, existing self-service shopping systems are prone to missed and missed products, resulting in product losses. The existing image analysis technology is not accurate in product identification and customer behavior judgment, and the real-time performance is insufficient, so it cannot intervene in time.
The customer behavior recognition method based on multi-eye cameras is adopted to obtain images through multi-eye cameras, extract image features, construct three-dimensional cost volume and depth maps, extract pseudo-point cloud data, divide grid cells, extract grid cell features, detect three-dimensional bounding boxes, track motion targets, and analyze behavior.
Real-time identification of customer behavior in self-service shopping scenarios is achieved, the effect of preventing product losses is improved, and the accuracy and efficiency of behavior recognition is improved.
Smart Images

Figure CN119992640A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of behavior recognition technology, and in particular to a customer behavior recognition method and device based on a multi-camera. Background Art
[0002] This section is intended to provide a background or context to the embodiments of the invention recited in the claims. No admission is made that the description herein is prior art by inclusion in this section.
[0003] With the rapid development of the retail industry, self-service shopping systems have been widely used. Existing self-service shopping systems usually rely on customers to actively scan the barcodes of goods, and then the system automatically calculates the total price and completes the payment. However, there are some significant problems in the use of such systems. Especially in the absence of monitoring, it is easy to miss or scan the wrong goods, resulting in the loss of goods.
[0004] In order to deal with these problems, some retailers have introduced surveillance cameras and simple image analysis technology to assist manual inspection, but these methods still have significant shortcomings. First, the existing image analysis technology has low accuracy in identifying goods and judging customer behavior, especially when there are many types of goods and customers operate in a variety of ways. Misjudgments and missed judgments often occur. Second, traditional methods rely more on manual monitoring and post-event tracing, which lacks real-time performance and cannot intervene in time before losses occur. In addition, the existing system lacks unified data analysis and processing capabilities and cannot efficiently process massive video data, thus limiting the large-scale application of the system.
[0005] With the rapid development of artificial intelligence and computer vision technology, the use of AI technology to monitor customer behavior in self-service checkout scenarios in real time and identify abnormal behavior has become an industry development trend. Through the application of deep learning algorithms, accurate identification of customer behavior and automatic prompts of abnormal events can be achieved, thereby effectively reducing the risk of commodity loss. However, there is currently a lack of a solution on the market that can fully cover a variety of abnormal behavior scenarios and has high accuracy and real-time response capabilities. Summary of the invention
[0006] The embodiment of the present invention provides a customer behavior recognition method based on a multi-camera, which is applied to a shopping cart and can recognize customer behavior in real time in a self-service shopping scenario to effectively prevent commodity loss. The method includes:
[0007] Obtain the image of the moving target captured by each camera of the multi-view camera;
[0008] Extract image features from the images of the moving target captured by each camera;
[0009] Constructing a three-dimensional cost volume based on the extracted image features;
[0010] Construct a depth map based on the three-dimensional cost volume;
[0011] Extracting pseudo point cloud data according to the depth map;
[0012] Divide the pseudo point cloud data into multiple grid cells of fixed size, each grid cell covers a predetermined three-dimensional space area;
[0013] Extracting a grid cell feature representation of each grid cell, and merging the grid cell feature representations of all grid cells into a grid cell feature map;
[0014] Detecting a three-dimensional bounding box of a moving target according to the grid unit feature map;
[0015] Track the three-dimensional bounding box of the moving target to obtain the three-dimensional tracking trajectory of the moving target;
[0016] The three-dimensional tracking trajectory is analyzed to analyze the behavior of the moving target.
[0017] The embodiment of the present invention provides a customer behavior recognition device based on a multi-camera, which is applied to a shopping cart and can recognize customer behavior in real time in a self-service shopping scenario to effectively prevent commodity loss. The device includes:
[0018] An image acquisition module is used to obtain an image of a moving target captured by each camera of a multi-camera;
[0019] An image feature extraction module is used to extract image features from the image of the moving target taken by each camera;
[0020] A three-dimensional cost volume construction module is used to construct a three-dimensional cost volume according to the extracted image features;
[0021] A depth map construction module, used to construct a depth map based on a three-dimensional cost volume;
[0022] A pseudo point cloud data extraction module, used to extract pseudo point cloud data according to the depth map;
[0023] A grid unit division module is used to divide the pseudo point cloud data into a plurality of grid units of a fixed size, each grid point covering a predetermined three-dimensional space area;
[0024] A grid feature extraction module, used to extract a grid cell feature representation of each grid cell, and merge the grid cell feature representations of all grid cells into a grid cell feature map;
[0025] A bounding box detection module, used to detect a three-dimensional bounding box of a moving target based on the grid unit feature map;
[0026] A tracking module is used to track the three-dimensional bounding box of the moving target to obtain the three-dimensional tracking trajectory of the moving target;
[0027] The behavior analysis module is used to analyze the three-dimensional tracking trajectory and analyze the behavior of the moving target.
[0028] An embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned customer behavior recognition method based on multi-camera when executing the computer program.
[0029] An embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned customer behavior recognition method based on multi-camera is implemented.
[0030] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the method for identifying customer behavior based on a multi-camera is implemented.
[0031] In an embodiment of the present invention, an image of a moving target captured by each camera of a multi-eye camera is obtained; image features are extracted from the image of the moving target captured by each camera; a three-dimensional cost volume is constructed according to the extracted image features; a depth map is constructed according to the three-dimensional cost volume; pseudo point cloud data is extracted according to the depth map; the pseudo point cloud data is divided into a plurality of grid units of fixed size, each grid point covers a predetermined three-dimensional space area; the grid unit feature representation of each grid unit is extracted, and the grid unit feature representations of all grid units are merged into a grid unit feature map; according to the grid unit feature map, a three-dimensional bounding box of the moving target is detected; the three-dimensional bounding box of the moving target is tracked to obtain a three-dimensional tracking trajectory of the moving target; the three-dimensional tracking trajectory is analyzed to analyze the behavior of the moving target. Compared with the prior art, the present invention uses a plurality of cameras of a multi-eye camera for shooting, so that the captured image can obtain more spatial information; a three-dimensional cost volume is constructed according to the extracted image features, and after obtaining the depth information, the three-dimensional bounding box of the moving target can be detected, providing rich information for subsequent behavior analysis, thereby improving the accuracy and efficiency of behavior recognition, so as to fully identify abnormal behavior and effectively prevent commodity loss. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0033] Figure 1 is a flow chart of a customer behavior recognition method based on a multi-camera in an embodiment of the present invention;
[0034] Figure 2 A flowchart of detecting a three-dimensional bounding box of a moving target in an embodiment of the present invention;
[0035] Figure 3 Schematic diagram of the structure of a customer behavior recognition device based on a multi-camera in an embodiment of the present invention;
[0036] Figure 4 2 is another structural schematic diagram of a customer behavior recognition device based on a multi-camera in an embodiment of the present invention;
[0037] Figure 5 Schematic diagram of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION
[0038] To make the purpose, technical solution and advantages of the embodiments of the present invention more clear, the embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings. Here, the exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0039] The inventors found that when a monocular camera identifies and locates multiple objects, it is easily affected by factors such as viewing angle, lighting, and object overlap, making it difficult to accurately determine the three-dimensional structure and spatial distribution of objects. This limits its ability to accurately identify objects and understand space in complex scenes. When the lighting conditions of a monocular camera vary greatly (such as strong light, backlight, shadows, etc.), the imaging effect is easily affected, resulting in a decrease in image quality, which in turn affects the accuracy of the detection results. Monocular cameras only provide two-dimensional image data and lack three-dimensional information. When processing complex scenes (such as 3D reconstruction, environmental modeling, etc.), the data is not rich enough, which can easily lead to inaccurate or incomplete results.
[0040] Therefore, the embodiment of the present invention relates to a behavior recognition method based on a multi-camera, which aims to obtain the three-dimensional pseudo point cloud information of image pixels from two images collected from different perspectives through a multi-camera stereo matching algorithm; perform three-dimensional moving target detection on the three-dimensional pseudo point cloud to obtain the three-dimensional bounding box information of the moving target (bounding box category, three-dimensional position of the center point of the bounding box, length, width, height of the bounding box, yaw angle, etc.); track and analyze the three-dimensional bounding box to obtain the tracking trajectory; analyze the tracking trajectory of the goods to determine the specific behavior of the target, "put in the goods behavior" and "take out the goods behavior". This method is often used in self-service shopping scenarios and autonomous driving scenarios.
[0041] Figure 1 The flowchart of the customer behavior recognition method based on multi-camera in an embodiment of the present invention is applied to a shopping cart, including:
[0042] Step 101, obtaining an image of a moving target captured by each camera of a multi-camera;
[0043] Step 102, extracting image features from the image of the moving target captured by each camera;
[0044] Step 103, constructing a three-dimensional cost volume according to the extracted image features;
[0045] Step 104, constructing a depth map according to the three-dimensional cost volume;
[0046] Step 105, extracting pseudo point cloud data according to the depth map;
[0047] Step 106, dividing the pseudo point cloud data into a plurality of grid cells of a fixed size, each grid cell covering a predetermined three-dimensional space area;
[0048] Step 107, extracting a grid cell feature representation of each grid cell, and merging the grid cell feature representations of all grid cells into a grid cell feature map;
[0049] Step 108, detecting a three-dimensional bounding box of the moving target according to the grid unit feature map;
[0050] Step 109, tracking the three-dimensional bounding box of the moving target to obtain a three-dimensional tracking trajectory of the moving target;
[0051] Step 110: Analyze the three-dimensional tracking trajectory to analyze the behavior of the moving target.
[0052] Compared with the prior art, the present invention uses a multi-camera to shoot, so that the captured images can obtain more spatial information; after constructing the three-dimensional cost volume based on the extracted image features and obtaining the depth information, the three-dimensional bounding box of the moving target can be detected, providing rich information for subsequent behavior analysis, thereby improving the accuracy and efficiency of behavior recognition, fully identifying abnormal behaviors, and effectively preventing commodity losses. Each step is introduced in detail below.
[0053] In step 101, an image of a moving target captured by each camera of a multi-camera is obtained;
[0054] In an embodiment of the present invention, the multi-eye camera includes at least two cameras. The more cameras there are, the more accurate the spatial information of the image obtained. The binocular camera is installed at the center of the handlebar of the shopping cart. However, in other embodiments, the binocular camera can also be installed at other positions or different angles of the shopping cart to adapt to different shopping cart designs or detection requirements. The multi-eye camera can further enhance the accuracy of 3D information acquisition and behavior recognition. By increasing the number of cameras, the area around the shopping cart can be better covered, thereby improving the density of the pseudo point cloud and the accuracy of 3D target detection. For example, four cameras can be installed at the four corners of the shopping cart, and a more accurate 3D depth map can be obtained through a multi-view stereo matching algorithm.
[0055] In one embodiment, before obtaining the image of the moving target captured by each camera of the multi-eye camera, the method further includes:
[0056] Calibrate the multi-camera and calibrate the origin of the multi-camera coordinate system to the stereo center of the shopping cart;
[0057] After obtaining the image of the moving target captured by each camera of the multi-camera, it also includes:
[0058] Time synchronization of images of moving targets;
[0059] Extracting image features from the image of the moving target captured by each camera, including: extracting image features from the time-synchronized image.
[0060] Specifically, the multi-camera is first calibrated. The purpose of calibration is to unify the coordinate system of the camera and map it to the stereo center position of the shopping cart. Through this step, it can be ensured that the image data acquired by the multi-camera has a consistent coordinate reference in subsequent processing. In order to ensure the accuracy of multi-camera stereo matching, images from multiple cameras need to be time synchronized. In this embodiment, synchronization is performed by timestamp alignment to ensure that all images can be correctly matched at the same time point.
[0061] In step 102, image features are extracted from the image of the moving object captured by each camera;
[0062] In one embodiment, extracting image features from the image of the moving target captured by each camera includes:
[0063] A shared convolutional neural network (CNN) is used to extract multi-scale image features from images of moving targets.
[0064] Specifically, commonly used convolutional neural network architectures include ResNet, VGG, ViT, etc. The image of each camera is extracted through a shared CNN network to extract high-dimensional image features. This image feature can be regarded as an encoding of the high-level semantic information of the original image. Two-dimensional feature extraction and matching are performed here. In other embodiments, other types of neural networks or feature extraction techniques can also be used, such as Vision Transformer (ViT), Graph Neural Network (GNN) or feature extraction algorithms based on optical flow to improve the efficiency and accuracy of feature extraction. In addition, the feature matching algorithm can also choose an end-to-end method based on deep learning, or combine traditional feature point matching methods such as SIFT, ORB, etc.
[0065] In step 103, a three-dimensional cost volume is constructed according to the extracted image features;
[0066] In one embodiment, constructing a three-dimensional cost volume according to the extracted image features includes:
[0067] The image features of the images taken by multiple cameras are matched to construct a three-dimensional cost volume, where each point of the three-dimensional cost volume is the matching cost of the corresponding pixels of the multiple images under a given disparity.
[0068] Specifically, given the parallax, that is, the assumption within the parallax range, taking two cameras as an example, the features of the left image can be matched with the image features of the right image to construct a three-dimensional cost volume. Then each point of the three-dimensional cost volume represents the matching cost of a pixel of the left image with the corresponding pixel of the right image under a given parallax, where the matching cost can be understood as point-to-point feature subtraction between the image features of the left image and the image features of the right image.
[0069] In one embodiment, after constructing the three-dimensional cost volume according to the extracted image features, the method further includes:
[0070] The 3D cost volume is processed using a 3D convolutional neural network to fuse global feature information;
[0071] Performing a pooling operation on the processed three-dimensional cost volume to fuse local feature information and obtain a refined three-dimensional cost volume;
[0072] Constructing a depth map according to the three-dimensional cost volume, including: constructing a depth map according to the refined three-dimensional cost volume.
[0073] Specifically, the convolution operation is performed in three-dimensional space to aggregate more contextual information and integrate global feature information. The residual module of the three-dimensional convolutional neural network helps the network learn deep feature representations through jump connections to avoid gradient disappearance. The pooling operation can integrate local feature information, which is equivalent to integrating global feature information with local feature information, refining the disparity estimate, and obtaining a refined three-dimensional cost volume.
[0074] In step 104, a depth map is constructed based on the three-dimensional cost volume;
[0075] In one embodiment, constructing a depth map according to a three-dimensional cost volume includes:
[0076] Use the softmax operation to convert the 3D cost volume into a probability distribution;
[0077] Calculate the disparity value according to the probability distribution to obtain a disparity map;
[0078] At least one of the following methods is used to calculate a depth map according to the disparity map:
[0079] Camera calibration parameters, supervised learning of disparity maps based on lidar, and deep learning of disparity maps using convolutional neural networks.
[0080] Specifically, the disparity value is calculated according to the probability distribution, that is, the expected calculation is performed on the cost volume to obtain the final disparity map. In addition, the acquisition of disparity maps and depth maps can be completed by adaptive algorithms, which dynamically adjust the calculation accuracy according to the complexity of the scene and the availability of computing resources. For example, a hierarchical disparity estimation method can be used in the detection process: a fast but low-precision disparity calculation is used in the initial estimation stage, and after the potential target is detected, a high-precision disparity estimation method is used for refinement, thereby improving the overall computational efficiency.
[0081] In addition to the above-mentioned scheme for constructing a depth map, the present invention also proposes the following two schemes for constructing a depth map:
[0082] (1) Using lidar supervision information, the depth of image features is estimated to construct a depth map;
[0083] (2) Through the classification algorithm, the depth information in the image is classified and predicted to obtain the probabilities of different depths, and then the depth map corresponding to the depths of different probabilities is obtained through information fusion.
[0084] In step 105, pseudo point cloud data is extracted according to the depth map. Specifically, the pseudo point cloud data of the image can be obtained by using the camera internal parameters and the depth map. The pseudo point cloud represents the position of each pixel in the three-dimensional space.
[0085] Subsequently, the pseudo point cloud data can be processed by the grid unit method to obtain three-dimensional features; the pseudo point cloud data can also be processed by the voxel method; or by processing the point data of the 3D point cloud. The following are the specific steps for processing pseudo point cloud data by the grid unit method.
[0086] In one embodiment, after extracting the pseudo point cloud data, the foreground may be extracted from the pseudo point cloud data using a foreground-background difference algorithm to obtain valid point cloud data of the moving target, thereby reducing the amount of calculation for the subsequent three-dimensional bounding box.
[0087] In step 106, the pseudo point cloud data is divided into a plurality of fixed-size grid cells, each grid cell covering a predetermined three-dimensional space region; for example, the three-dimensional space region can be represented as 0.2mx0.2mx 0.2m; this division decomposes the large-scale pseudo point cloud data into smaller blocks that are easier to process.
[0088] In one embodiment, after dividing the pseudo point cloud data into a plurality of grid units of fixed size, the method further includes:
[0089] According to preset requirements, the points in each grid unit are standardized, and the standardization processing includes integrity processing and uniformity processing.
[0090] Among them, the integrity processing includes using techniques such as nearest neighbor interpolation to fill in the missing data of the point cloud, and the uniformity processing includes restricting the point cloud to ensure that the number of points in each grid cell does not exceed a predetermined value.
[0091] In step 107, a grid cell feature representation of each grid cell is extracted, and the grid cell feature representations of all grid cells are merged into a grid cell feature map;
[0092] In one embodiment, extracting a grid cell feature representation of each grid cell includes:
[0093] Use a multi-layer perceptron encoder to convert the features of the points in each grid cell into a high-dimensional feature vector;
[0094] The high-dimensional feature vectors of the points in each grid cell are aggregated through the pooling operation to generate a fixed-dimensional grid cell feature representation.
[0095] Specifically, the features of the points in each grid cell are converted into a high-dimensional feature vector, which realizes the conversion of the original point cloud data into a more meaningful feature representation. Pooling operations such as maximum pooling or weighted average help to reduce the data dimension and retain important features. The grid cell feature representations of all grid cells are merged into a grid cell feature map, which provides a comprehensive feature representation of the entire scene.
[0096] In step 108, a three-dimensional bounding box of the moving object is detected according to the grid unit feature map;
[0097] See also Figure 2 The flowchart of detecting a three-dimensional bounding box of a moving target in an embodiment of the present invention is as follows. In one embodiment, detecting a three-dimensional bounding box of a moving target according to the grid unit feature map includes:
[0098] Step 201, using a convolutional neural network to process the grid unit feature map to obtain a convolution-processed grid unit feature map; this step includes convolution operations, activation function application, and pooling techniques to enhance the information representation capability in the feature map; specifically, there are the following processing methods:
[0099] (1) Convolution can be performed directly in the 3D space of the grid cell feature map to capture spatial and contextual information;
[0100] (2) The grid unit feature graph can be processed through graph convolution operations to extract the relationship information between nodes;
[0101] (3) It can optimize convolution for sparse data in grid cell feature maps, improve computational efficiency, and is suitable for high-dimensional feature processing;
[0102] (4) Self-supervised learning methods can be used to automatically extract effective information from grid cell feature maps and enhance feature expression;
[0103] (5) Feature selection methods (such as PCA, t-SNE, etc.) can be used to reduce the dimension of the grid cell feature map to improve the efficiency and effect of subsequent processing.
[0104] Step 202, fusing the convolution-processed grid unit feature map with the context feature of the grid unit to obtain a fused feature map;
[0105] Among them, contextual features can be other types of sensor data (such as ultrasonic, radar, infrared sensors, etc.), such as bird's-eye view feature maps, and multimodal data fusion can be performed to improve the accuracy of target behavior recognition. For example, the temperature information obtained by the infrared sensor can be combined with the 3D point cloud data obtained by the binocular camera to help identify different types of objects, or distinguish the temperature characteristics of human hands from other objects.
[0106] Fusion technology helps to comprehensively utilize feature information from different sources and improve detection performance. When fused with context features, features can be extracted from different scales and fused to capture richer context information. After fusion, the following processing can be performed:
[0107] (1) Perform attention weighting on each channel of the fused feature map, so that the fused feature map can adaptively focus on the features of different channels;
[0108] (2) Weighting different regions of the grid cell feature map in the spatial dimension to highlight important regions;
[0109] (3) The features of the fused feature map are weighted;
[0110] (4) Use the attention mechanism to dynamically weight different contextual features, highlighting important features while suppressing unimportant features;
[0111] (5) Concatenating feature maps in a specific dimension can retain all information and is often used for multi-scale feature fusion in convolutional neural networks.
[0112] Step 203, inputting the fused feature map into a three-dimensional convolutional neural network to extract spatial features;
[0113] Specifically, 3D convolutional neural networks can effectively capture complex relationships in space;
[0114] Step 204: Use the target detection head to predict the three-dimensional bounding box of the moving target from the spatial features and obtain three-dimensional bounding box information, wherein the three-dimensional bounding box information includes at least one of the following information:
[0115] The category, size information, 3D coordinates of the center point, and yaw angle of the bounding box.
[0116] The object detection head focuses on predicting detailed information about the object from the extracted features. The classification of the bounding box is the process of classifying the object, such as hand, commodity, etc. Classification assigns specific category labels to the detected objects.
[0117] In one embodiment, after detecting the three-dimensional bounding box of the moving target, the method further includes:
[0118] Apply the non-maximum suppression algorithm to remove duplicates from the detected 3D bounding box to obtain the filtered 3D bounding box;
[0119] Calibrate the filtered 3D bounding box;
[0120] Tracking the three-dimensional bounding box of the moving target includes: tracking the filtered three-dimensional bounding box.
[0121] Specifically, the non-maximum suppression algorithm can retain the best detection results, help reduce duplicate frames in the detection results, and improve the reliability of the results. In addition, when using the non-maximum suppression algorithm to remove duplicates, the score can also be dynamically adjusted when processing each 3D bounding box based on the overlap (IoU) with the 3D bounding box with the highest score, rather than directly suppressing 3D bounding boxes with scores below a certain threshold. This method can retain some boxes that may be misjudged. The overlap (IoU) can also be adjusted dynamically, and the threshold can be adaptively selected based on the density of the 3D bounding boxes and the characteristics of the scene to more effectively remove redundant frames.
[0122] In addition, other deduplication methods can be used. For example, clustering algorithms (such as K-means, DBSCAN, etc.) can be used to cluster 3D bounding boxes and merge overlapping or close boxes to reduce redundancy. 3D bounding boxes can also be regarded as nodes in a graph, and a graph cut algorithm can be used to optimize the selection of boxes and remove redundancy.
[0123] Calibrate the selected 3D bounding box to improve the accuracy and reliability of detection. The calibration process includes adjusting the position, size and other parameters of the detection box to match the target in the real scene.
[0124] When making a specific correction, the first 40 frames of the filtered 3D bounding box can be saved, or according to the actual situation, the detection results of the first n frames can be saved, and a smoothing operation or a filtering operation can be performed on the detected target box to make the detection box continuous.
[0125] In step 109, the three-dimensional bounding box of the moving target is tracked to obtain the three-dimensional tracking trajectory of the moving target; after obtaining the three-dimensional tracking trajectory, the three-dimensional tracking trajectory can also be corrected. A tracking algorithm can be used to save the tracking trajectory of the previous n frames, smooth the tracking trajectory, or perform a filtering operation to correct the trajectory.
[0126] In step 110, the three-dimensional tracking trajectory is analyzed to analyze the behavior of the moving target.
[0127] Specifically, the analysis of the behavior of the moving target can be performed according to preset rules. For example, if the point on the three-dimensional tracking trajectory of the product is getting closer and closer to the center point of the shopping cart, it is judged as "putting the product in", otherwise it is "taking the product out". Alternatively, the three-dimensional tracking trajectory of the product can be sent to a binary classification model to determine whether the trajectory is "putting the product in" or "taking the product out".
[0128] In the embodiments of the present invention, multiple classification models can also be used to identify more types of behaviors. For example, the classification model can be expanded to identify operations such as rotation, dropping, and flipping of goods, or to identify simultaneous operations of multiple goods. These enhanced classification models can be based on more complex neural network architectures, or combined with sequence models such as LSTM, Transformer, etc. to process time series information.
[0129] In the embodiment of the present invention, the coordinate system calibration of the shopping cart is no longer static, but a dynamic calibration method is used to dynamically adjust the calibration parameters according to the movement of the shopping cart in the real-time scene. This method is suitable for automatically adjusting the parallax calculation and the generation of the 3D pseudo point cloud during the movement of the shopping cart, thereby improving the robustness and adaptability of the system.
[0130] It should be noted that although the embodiments of the present invention are mainly applied to self-service shopping scenarios, the method can also be extended to other scenarios such as autonomous driving, industrial robots, smart homes, etc. For example, in autonomous driving, the method proposed in the embodiments of the present invention can be used to identify specific behaviors of pedestrians or vehicles, such as pedestrians crossing the road or vehicles turning; in smart homes, the method proposed in the embodiments of the present invention can be used to identify the behaviors of different items in the home being moved, placed, or taken away, thereby realizing intelligent management.
[0131] The embodiment of the present invention further proposes a customer behavior recognition device based on a multi-camera, the principle of which is similar to the customer behavior recognition method based on a multi-camera, and will not be repeated here.
[0132] Figure 3 The schematic diagram of a customer behavior recognition device based on a multi-camera in an embodiment of the present invention is applied to a shopping cart, including:
[0133] An image acquisition module 301 is used to obtain an image of a moving target captured by each camera of a multi-camera;
[0134] An image feature extraction module 302 is used to extract image features from the image of the moving target captured by each camera;
[0135] A three-dimensional cost volume construction module 303 is used to construct a three-dimensional cost volume according to the extracted image features;
[0136] A depth map construction module 304, configured to construct a depth map according to the three-dimensional cost volume;
[0137] A pseudo point cloud data extraction module 305 is used to extract pseudo point cloud data according to the depth map;
[0138] A grid unit division module 306 is used to divide the pseudo point cloud data into a plurality of grid units of a fixed size, each grid point covering a predetermined three-dimensional space area;
[0139] A grid feature extraction module 307 is used to extract a grid cell feature representation of each grid cell and merge the grid cell feature representations of all grid cells into a grid cell feature map;
[0140] A bounding box detection module 308, configured to detect a three-dimensional bounding box of a moving target based on the grid unit feature map;
[0141] A tracking module 309 is used to track the three-dimensional bounding box of the moving target to obtain a three-dimensional tracking trajectory of the moving target;
[0142] The behavior analysis module 310 is used to analyze the three-dimensional tracking trajectory and analyze the behavior of the moving target.
[0143] See also Figure 4 FIG. 4 is another structural diagram of a customer behavior recognition device based on a multi-camera in an embodiment of the present invention. In one embodiment, the device further includes a calibration module 401 for:
[0144] Before obtaining the image of the moving target captured by each camera of the multi-camera, the multi-camera is calibrated, and the origin of the multi-camera coordinate system is calibrated to the stereo center position of the shopping cart;
[0145] The device further comprises a time synchronization module 402, configured to:
[0146] After obtaining the image of the moving target captured by each camera of the multi-view camera, the image of the moving target is time synchronized;
[0147] The image feature extraction module is used to extract image features from time-synchronized images.
[0148] In one embodiment, the image feature extraction module is used to:
[0149] A shared convolutional neural network is used to extract multi-scale image features from images of moving targets.
[0150] In one embodiment, the three-dimensional cost volume construction module is used to:
[0151] The image features of the images taken by multiple cameras are matched to construct a three-dimensional cost volume, where each point of the three-dimensional cost volume is the matching cost of the corresponding pixels of the multiple images under a given disparity.
[0152] In one embodiment, the apparatus further includes a fusion module 403, configured to:
[0153] The 3D cost volume is processed using a 3D convolutional neural network to fuse global feature information;
[0154] Performing a pooling operation on the processed three-dimensional cost volume to fuse local feature information and obtain a refined three-dimensional cost volume;
[0155] The depth map construction module is used to construct a depth map based on the refined three-dimensional cost volume.
[0156] In one embodiment, the depth map construction module is used to:
[0157] Use the softmax operation to convert the 3D cost volume into a probability distribution;
[0158] Calculate the disparity value according to the probability distribution to obtain a disparity map;
[0159] At least one of the following methods is used to calculate a depth map according to the disparity map:
[0160] Camera calibration parameters, supervised learning of disparity maps based on lidar, and deep learning of disparity maps using convolutional neural networks.
[0161] In one embodiment, the grid unit division module is further used for:
[0162] After the pseudo point cloud data is divided into a plurality of grid units of fixed size, the points in each grid unit are standardized according to preset requirements, and the standardized processing includes integrity processing and uniformity processing.
[0163] In one embodiment, the mesh feature extraction module is used to:
[0164] Use a multi-layer perceptron encoder to convert the features of the points in each grid cell into a high-dimensional feature vector;
[0165] The high-dimensional feature vectors of the points in each grid cell are aggregated through the pooling operation to generate a fixed-dimensional grid cell feature representation.
[0166] In one embodiment, the bounding box detection module is used to:
[0167] Using a convolutional neural network to process the grid unit feature map, to obtain a convolution-processed grid unit feature map;
[0168] The convolution-processed grid unit feature map is fused with the context feature of the grid unit to obtain a fused feature map;
[0169] Input the fused feature map into the 3D convolutional neural network to extract spatial features;
[0170] Using the target detection head, predicting the three-dimensional bounding box of the moving target from the spatial features, and obtaining three-dimensional bounding box information, wherein the three-dimensional bounding box information includes at least one of the following information:
[0171] The category, size information, 3D coordinates of the center point, and yaw angle of the bounding box.
[0172] In one embodiment, the apparatus further comprises a calibration module 404, configured to:
[0173] After detecting the 3D bounding box of the moving target, a non-maximum suppression algorithm is applied to remove duplicates from the detected 3D bounding box to obtain a filtered 3D bounding box;
[0174] Calibrate the filtered 3D bounding box;
[0175] The tracking module is used to track the filtered three-dimensional bounding box.
[0176] In summary, the method and device proposed in the embodiments of the present invention have the following beneficial effects:
[0177] The present invention uses a multi-camera to shoot, so that the captured images can obtain more spatial information; after constructing a three-dimensional cost volume based on the extracted image features and obtaining depth information, the three-dimensional bounding box of the moving target can be detected, providing rich information for subsequent behavior analysis, thereby improving the accuracy and efficiency of behavior recognition, fully identifying abnormal behaviors, and effectively preventing commodity losses.
[0178] An embodiment of the present invention further provides a computer device, Figure 5 It is a schematic diagram of a computer device in an embodiment of the present invention. The computer device 500 includes a memory 510, a processor 520, and a computer program 530 stored in the memory 510 and executable on the processor 520. When the processor 520 executes the computer program 530, the above-mentioned customer behavior recognition method based on multi-camera is implemented.
[0179] An embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned customer behavior recognition method based on multi-camera is implemented.
[0180] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the method for identifying customer behavior based on a multi-camera is implemented.
[0181] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0182] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0183] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0184] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0185] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A customer behavior recognition method based on multi-camera, characterized in that: Applied to shopping cart, including: Obtain the image of the moving target captured by each camera of the multi-view camera; Extract image features from the images of the moving target captured by each camera; Constructing a three-dimensional cost volume based on the extracted image features; Construct a depth map based on the three-dimensional cost volume; Extracting pseudo point cloud data according to the depth map; Divide the pseudo point cloud data into multiple grid cells of fixed size, each grid cell covers a predetermined three-dimensional space area; Extracting a grid cell feature representation of each grid cell, and merging the grid cell feature representations of all grid cells into a grid cell feature map; Detecting a three-dimensional bounding box of a moving target according to the grid unit feature map; Track the three-dimensional bounding box of the moving target to obtain the three-dimensional tracking trajectory of the moving target; The three-dimensional tracking trajectory is analyzed to analyze the behavior of the moving target.
2. The method according to claim 1, characterized in that Before obtaining the image of the moving target captured by each camera of the multi-eye camera, it also includes: Calibrate the multi-camera and calibrate the origin of the multi-camera coordinate system to the stereo center of the shopping cart; After obtaining the image of the moving target captured by each camera of the multi-camera, it also includes: Time synchronization of images of moving targets; Extracting image features from the image of the moving target captured by each camera, including: extracting image features from the time-synchronized image.
3. The method according to claim 1, characterized in that Image features are extracted from the images of the moving target taken by each camera, including: A shared convolutional neural network is used to extract multi-scale image features from images of moving targets.
4. The method according to claim 1, characterized in that The three-dimensional cost volume is constructed based on the extracted image features, including: The image features of the images taken by multiple cameras are matched to construct a three-dimensional cost volume, where each point of the three-dimensional cost volume is the matching cost of the corresponding pixels of the multiple images under a given disparity.
5. The method according to claim 1, characterized in that After constructing the 3D cost volume based on the extracted image features, it also includes: The 3D cost volume is processed using a 3D convolutional neural network to fuse global feature information; Performing a pooling operation on the processed three-dimensional cost volume to fuse local feature information and obtain a refined three-dimensional cost volume; Constructing a depth map according to the three-dimensional cost volume, including: constructing a depth map according to the refined three-dimensional cost volume.
6. The method according to claim 1, characterized in that According to the 3D cost volume, a depth map is constructed, including: Use the softmax operation to convert the 3D cost volume into a probability distribution; Calculate the disparity value according to the probability distribution to obtain a disparity map; At least one of the following methods is used to calculate a depth map according to the disparity map: Camera calibration parameters, supervised learning of disparity maps based on lidar, and deep learning of disparity maps using convolutional neural networks.
7. The method according to claim 1, characterized in that After dividing the pseudo point cloud data into multiple fixed-size grid cells, it also includes: According to preset requirements, the points in each grid unit are standardized, and the standardization processing includes integrity processing and uniformity processing.
8. The method according to claim 1, characterized in that Extract the grid cell feature representation of each grid cell, including: Use a multi-layer perceptron encoder to convert the features of the points in each grid cell into a high-dimensional feature vector; The high-dimensional feature vectors of the points in each grid cell are aggregated through the pooling operation to generate a fixed-dimensional grid cell feature representation.
9. The method according to claim 1, characterized in that Detecting a three-dimensional bounding box of a moving target according to the grid unit feature map includes: Using a convolutional neural network to process the grid unit feature map, to obtain a convolution-processed grid unit feature map; The convolution-processed grid unit feature map is fused with the context feature of the grid unit to obtain a fused feature map; Input the fused feature map into the 3D convolutional neural network to extract spatial features; Using the target detection head, predicting the three-dimensional bounding box of the moving target from the spatial features, and obtaining three-dimensional bounding box information, wherein the three-dimensional bounding box information includes at least one of the following information: The category, size information, 3D coordinates of the center point, and yaw angle of the bounding box.
10. The method according to claim 1, characterized in that After detecting the 3D bounding box of the moving object, it also includes: Apply the non-maximum suppression algorithm to remove duplicates from the detected 3D bounding box to obtain the filtered 3D bounding box; Calibrate the filtered 3D bounding box; Tracking the three-dimensional bounding box of the moving target includes: tracking the filtered three-dimensional bounding box.
11. A customer behavior recognition device based on multi-camera, characterized in that: Applied to shopping cart, including: An image acquisition module is used to obtain an image of a moving target captured by each camera of a multi-camera; An image feature extraction module is used to extract image features from the image of the moving target taken by each camera; A three-dimensional cost volume construction module is used to construct a three-dimensional cost volume according to the extracted image features; A depth map construction module, used to construct a depth map based on a three-dimensional cost volume; A pseudo point cloud data extraction module, used to extract pseudo point cloud data according to the depth map; A grid unit division module is used to divide the pseudo point cloud data into a plurality of grid units of a fixed size, each grid point covering a predetermined three-dimensional space area; A grid feature extraction module, used to extract a grid cell feature representation of each grid cell, and merge the grid cell feature representations of all grid cells into a grid cell feature map; A bounding box detection module, used to detect a three-dimensional bounding box of a moving target based on the grid unit feature map; A tracking module is used to track the three-dimensional bounding box of the moving target to obtain the three-dimensional tracking trajectory of the moving target; The behavior analysis module is used to analyze the three-dimensional tracking trajectory and analyze the behavior of the moving target.
12. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 10 is implemented.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
14. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.