Intelligent video identification method and device, equipment and storage medium
By combining product gravity timing data and video recognition technology, the intelligent video recognition method of vending machines can accurately identify product pick-up and placement actions and user abnormal behaviors, solving the identification accuracy problem in complex scenarios in traditional technology, and improving the security and management efficiency of vending machines.
Patent Information
- Application Number
- CN202510584189.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The video recognition technology of existing vending machines has low accuracy in product abnormal status and user abnormal behavior recognition in complex scenarios. Traditional gravity sensors cannot handle the sensitivity of multiple products to take and place and product form changes. Camera recognition is limited by the blind spots of viewing angle coverage and target segmentation and trajectory tracking accuracy defects during multi-user operation.
Combining the product gravity timing data and target video data, the product pick-and-place behavior is determined through the weight recognition module. The video recognition module constructs user behavior portraits, and the abnormal behavior processing module generates abnormal behavior information. It uses the product gravity timing data to accurately identify the product pick-and-place action, automatically selects the target recognition area image, builds user behavior portraits and recognizes abnormal behaviors.
It improves the accuracy of identification of commodity pick-up and release behaviors, reduces computing resources and time costs, promptly detects abnormal behaviors, ensures the safe operation of vending machines, reduces commodity losses and economic losses, and provides data support system optimization.
Smart Images

Figure CN120472369A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of video recognition technology, and more specifically, relates to an intelligent video recognition method and apparatus, device, and storage medium. Background Art
[0002] As a core platform for new retail, vending machines utilize intelligent video recognition technology, which is crucial for automating commodity transactions. Existing solutions primarily rely on single gravity sensors or fixed-angle cameras. While gravity sensor solutions can initially detect weight fluctuations by determining whether items are being removed, they cannot handle complex interactions such as the simultaneous placement and retrieval of multiple items, or the return of some items. Furthermore, they are sensitive to weight fluctuations caused by factors such as stacked beverage bottles and packaging variations, making it difficult to accurately correlate user behavior with product status.
[0003] Traditional video recognition solutions rely solely on cameras to capture video for visual recognition. However, they are limited by the blind spots covered by fixed viewing angles and the accuracy defects of target segmentation and trajectory tracking during multi-user parallel operations, resulting in low recognition accuracy.
[0004] In summary, traditional video recognition technology still has the problem of low accuracy in identifying abnormal product status and abnormal user behavior in complex scenarios. Summary of the Invention
[0005] The purpose of this application is to provide an intelligent video recognition method and device, equipment, and storage medium to improve the accuracy of identifying products and user behavior in complex scenarios.
[0006] A first aspect of the embodiments of the present application provides an intelligent video recognition method, comprising: Determining the product placement behavior of the vending machine based on the product gravity time series data; if the product placement behavior is abnormal, selecting a target recognition area image from the target video data based on the product placement behavior; Determine a user behavior profile based on the target recognition area image, and identify abnormal user behavior based on the user behavior profile; Generate abnormal behavior information based on the identified abnormal user behavior.
[0007] A second aspect of the embodiments of the present application provides an intelligent video recognition device, comprising: A weight recognition module is configured to determine the product placement behavior of the vending machine based on the product gravity time series data; if the product placement behavior is abnormal, a target recognition area image is selected from the target video data based on the product placement behavior; A video recognition module is used to determine a user behavior profile based on the target recognition area image, and identify abnormal user behavior based on the user behavior profile; The exception handling module is used to generate abnormal behavior information based on the identified abnormal user behavior.
[0008] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of the above-mentioned intelligent video recognition method when executing the computer program.
[0009] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned intelligent video recognition method are implemented.
[0010] The beneficial effects of the intelligent video recognition method, device, equipment, and storage medium provided by the embodiments of the present application are as follows: in determining the behavior of picking and placing goods, the embodiments of the present application utilize the time-series data of the gravity of the goods to more accurately identify the picking and placing actions of the goods, effectively overcoming the shortcomings of traditional gravity sensors that cannot handle the simultaneous picking and placing of multiple goods, partial return, and are sensitive to changes in the shape of the goods, thereby improving transaction accuracy and inventory management reliability. When abnormal product behavior is detected, the embodiments of the present application can automatically select the target recognition area image from the target video data, focusing on the problem area in a targeted manner, avoiding ineffective analysis of the entire video, saving computing resources and time costs, and improving processing efficiency.
[0011] This embodiment of the application constructs a user behavior profile using target recognition area images, and based on this, identifies abnormal behavior, enabling in-depth analysis of user operating habits. This not only enables timely detection of abnormal behaviors such as aggressive pickup, prolonged delays, and multiple pick-ups and drops without settlement, but also prevents potential risks, ensuring the safe operation of vending machines and reducing product loss and economic losses. Finally, the generated abnormal behavior information provides strong data support for subsequent system optimization and strengthened management, promoting the intelligent and efficient development of the vending industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0013] Figure 1 A flowchart of an intelligent video recognition method provided in one embodiment of the present application; Figure 2 A structural block diagram of an intelligent video recognition device provided in one embodiment of the present application; Figure 3 A schematic block diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0014] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0015] In order to make the purpose, technical solutions and advantages of this application clearer, specific embodiments will be described below with reference to the accompanying drawings.
[0016] Please refer to Figure 1 , Figure 1 This is a flow chart of an intelligent video recognition method provided in one embodiment of the present application. The method may include S101 to S103.
[0017] S101: Determine the product placement behavior of the vending machine based on the product gravity time series data. If the product placement behavior is abnormal, select a target recognition area image from the target video data based on the product placement behavior.
[0018] In this embodiment, the commodity gravity time series data refers to a data sequence of commodity weight changes continuously collected by gravity sensors in each shelf area of the vending machine at fixed time intervals.
[0019] For example, each shelf layer is divided into 4-8 detection zones. A thin-film gravity sensor is attached to the center of each zone. The gravity sensor is connected to the edge computing module via a cable. A hardware clock synchronization module is also included to ensure timestamp accuracy. The edge computing module reads the sensor weight value at a fixed frequency, associates the timestamp with the zone number, and stores it.
[0020] In this embodiment, the product picking and placing behavior can include product picking behavior, product returning behavior, and abnormal product behavior. This embodiment can determine whether the user is picking up or returning the product, and whether abnormal behavior such as product displacement occurs by analyzing the dynamic changes in gravity time series data.
[0021] For example, this embodiment sets a weight change threshold and a minimum duration threshold based on the average weight of the goods. Using a time series analysis algorithm, the gravity data is analyzed frame by frame. When the weight change of goods in a monitored area meets the threshold, the direction of the change is considered to determine the pick-and-place behavior. If weight changes occur simultaneously in adjacent areas, this embodiment can consider the product dimensions in the product library to eliminate misjudgments caused by cross-area movement of goods.
[0022] In this embodiment, abnormal product behavior refers to abnormal actions that do not conform to the normal product picking and placing pattern, which may include violent picking, returning foreign objects, frequent tentative picking and placing, etc.
[0023] For example, this embodiment can establish a normal product behavior model based on rules such as the weight variation range and operation time interval of the product when it is picked up and placed. In this embodiment, real-time product gravity time series data is compared with the normal product behavior model. When the weight change amplitude, rate, or duration exceeds the normal range, it is determined to be abnormal behavior.
[0024] For example, this embodiment can calculate the mean and standard deviation of weight changes for each area based on historical data on normal product placement and handling, thereby determining a normal weight fluctuation range. When weight changes in a particular area exceed the normal range, or the rate of change exceeds a preset threshold, this can be flagged as abnormal product behavior. Specifically, the magnitude of change, rate, and historical operation records can be combined to distinguish different types of abnormalities.
[0025] In this embodiment, the target video data refers to the video stream containing the product shelf area captured in real time by the vending machine's internal camera, as well as the video stream of the area surrounding the vending machine captured in real time by an external camera, used to record user operations. The target identification area image refers to a partial image of a specific shelf area captured from the target video data based on the location and time of abnormal product behavior.
[0026] For example, this embodiment uses camera calibration to establish a mapping between the physical coordinates of the shelf and the pixel coordinates of the video. When abnormal behavior occurs, the corresponding pixel area in the video is located and an image is captured based on the gravity sensor's area number. The camera calibration process can use Zhang's calibration method to obtain the camera's intrinsic parameter matrix and determine the pixel range of each gravity detection area in the video. This embodiment extracts the target area image of the corresponding frame from the video stream by receiving the area number and timestamp of the abnormal behavior.
[0027] When abnormal behavior is detected in a gravity zone, this embodiment captures an image of that area from the current video frame based on a pre-stored pixel range. In addition to the current frame, the 10 frames before and after the abnormality are captured, forming an image sequence that encompasses the entire process before and after the abnormality, which is used for subsequent user behavior analysis and abnormality tracing.
[0028] S102: Determine a user behavior profile based on the target recognition area image, and identify abnormal user behavior based on the user behavior profile.
[0029] In this embodiment, abnormal user behavior includes time sequence anomalies and risky behavior. Identifying abnormal user behavior based on user behavior profiles includes: Utilize statistical models of normal user behavior to identify time series anomalies and risky behaviors in user behavior profiles. Time series anomalies include abnormal operation time points and operation intervals, while risky behaviors include abnormal product placement and handling, and abnormal action risks.
[0030] This embodiment uses images of the target recognition area to extract user behavior characteristics and construct a behavioral profile that includes information such as operation time, interval, product placement, and movement. This embodiment builds a statistical model of normal user behavior based on historical normal operation data, including parameters such as the normal distribution range, mean, and standard deviation of each behavior characteristic.
[0031] During the anomaly identification phase, for timing anomalies, this embodiment can compare the time points and intervals of the user's real-time operations with the normal time distribution and interval range in the model. If the operation time deviates significantly from the high-frequency period, or the operation interval exceeds the normal fluctuation range, it is determined to be a timing anomaly. For risky behaviors, this embodiment can compare the user's real-time number of goods taken and put, weight changes, and movement speed, strength and other characteristics with the preset standards and thresholds in the model. If the number and weight of goods taken and put are abnormal, or the action has a violent tendency, etc., which exceeds the safety threshold, it is determined to be a risky behavior. Finally, a comprehensive judgment is made to identify the user's abnormal behavior.
[0032] In this embodiment, the user behavior profile is a digital description of the user's operating habits and behavior patterns constructed through target recognition area image analysis, which is used for comparison and identification of abnormal behaviors.
[0033] User behavior profiles can include temporal features, spatial features, and action features. This embodiment can be based on historical behavior data to collect statistics on the user's normal operation mode in time, space, and action dimensions to form a multi-dimensional feature vector as a baseline model for anomaly detection.
[0034] For example, this embodiment can establish a user behavior database containing characteristics such as time, region, product, and action. This embodiment uses statistical analysis methods to calculate the mean, standard deviation, and quantile of each characteristic in the user behavior database, and constructs a range of normal behavior, such as an hourly distribution histogram of operation time, the proportion of operation frequency in each region, and the average pickup speed and fluctuation range.
[0035] For example, timing anomaly identification can include abnormal operation time points and abnormal operation intervals. An abnormal operation time point refers to a user's operation time significantly deviating from their normal high-frequency period. An abnormal operation interval refers to the time interval between two consecutive pickups exceeding the normal fluctuation range.
[0036] This embodiment sets reasonable thresholds based on the time distribution and interval statistics of historical user operations. When real-time operation data exceeds the threshold, it is identified as an anomaly. Time point anomalies must also meet the requirements of interval anomalies and action characteristics, such as nighttime operations accompanied by quick pickup, to avoid occasional misjudgments.
[0037] Risky behavior identification can include abnormal product placement and handling, as well as unusual action risks. Abnormal product placement refers to situations where the quantity or weight of items being picked up or placed does not conform to normal patterns, such as when the number of items picked up in a single transaction exceeds the historical maximum or when the weight exceeds the standard value by more than 2 times. Unusual action risks refer to situations where user actions exhibit violent tendencies or violations, such as hand movement speed exceeding a preset threshold or repeated bumping of the shelf.
[0038] This embodiment can identify unsafe picking and placing behaviors and dangerous actions by combining preset standard weights in a product library with historical user pick-and-place data and real-time motion characteristics. For example, a standard weight library for products and action risk thresholds can be established, such as hand speeds exceeding 1.5 m / s, which indicates rapid picking. This embodiment can use a posture estimation algorithm to calculate hand motion parameters and compare them with risk thresholds to identify abnormalities.
[0039] Abnormal product picking and placing can be detected by comparing the real-time weight change with the standard value of the product library. If the real-time weight change is less than 0.7 times the standard value or greater than 1.3 times the standard value, and the number of goods picked up is greater than the historical maximum value, it is determined to be a picking and placing abnormality.
[0040] Abnormal action risk detection can be achieved by calculating the average speed of the hand from the time it touches the product to the time it is removed from the shelf. If it is greater than 2m / s and accompanied by abnormal joint angles, such as an elbow bending angle of less than 30 degrees (indicating violent pulling), it is determined to be an abnormal action risk.
[0041] When an anomaly is triggered, this embodiment synchronously saves the target recognition area image, gravity time series data, and operation log to form a traceable risk evidence package.
[0042] S103: Generate abnormal behavior information based on the identified abnormal behavior of the user.
[0043] In this embodiment, an intelligent video recognition method further includes: The commodity weight data is detected. If a change in the commodity weight is detected, the first monitoring frequency is updated to obtain a third monitoring frequency, and the second monitoring frequency is updated to obtain a fourth monitoring frequency.
[0044] The third monitoring frequency refers to the frequency of collecting target video data of the product and the user when the weight of the product changes.
[0045] The fourth monitoring frequency refers to the frequency of collecting commodity gravity time series data for commodities whose weight has changed when the weight of the commodities has changed.
[0046] The first monitoring frequency refers to the frequency of collecting video data of the product and the user when the weight of the product has not changed, and the second monitoring frequency refers to the frequency of collecting weight data of the product when the weight of the product has not changed.
[0047] The third monitoring frequency is greater than the first monitoring frequency, and the fourth monitoring frequency is greater than the second monitoring frequency.
[0048] In this embodiment, the first monitoring frequency refers to the frequency at which the vending machine collects video data of products and users under normal conditions. The second monitoring frequency refers to the frequency at which the machine collects weight data under normal conditions. The first and second monitoring frequencies are used for routine monitoring, balancing data collection accuracy and device resource consumption.
[0049] For example, this embodiment can control the camera to shoot at a first monitoring frequency by setting a timer in the video acquisition module. A timed acquisition program can be set in the gravity sensor's data acquisition circuit to read weight data at a second monitoring frequency. After the vending machine is started, the video acquisition module and weight acquisition module begin periodically collecting data at the first monitoring frequency and the second monitoring frequency, respectively, and store the data in a local storage device.
[0050] A change in the weight of an item occurs when the item is removed or returned, causing the weight value detected by the gravity sensor to differ from the previous value. This embodiment uses a gravity sensor to monitor the weight of the item in real time. When the weight value exceeds a set fluctuation range, it is determined to be a weight change. This embodiment can set a weight change threshold in the data processing program. When the weight data changes beyond this threshold, the corresponding processing flow is triggered.
[0051] When the weight of an item changes, this embodiment increases the frequency of video and weight data collection. The third monitoring frequency is used to more frequently collect video data of the item and user, capturing detailed behavioral information. The fourth monitoring frequency is used to more frequently collect weight data for items with weight changes, accurately recording the weight change process.
[0052] Exemplarily, when a weight change signal is detected, the control programs of the video acquisition module and the weight acquisition module shorten the acquisition interval to the time corresponding to the third monitoring frequency and the fourth monitoring frequency.
[0053] The target video data is video data collected at the third monitoring frequency after the weight of the product changes, including the product and user behavior. The product gravity time series data is a data sequence collected at the fourth monitoring frequency, showing the weight of the product changing over time. The target video data may include video frames and capture timestamps. The product gravity time series data may include weight values and capture timestamps.
[0054] In the high-frequency acquisition mode, the video acquisition device and gravity sensor of this embodiment work continuously, and the acquired data are stored as target video data and product gravity time series data respectively, for subsequent use in determining product picking and placing behavior and user abnormal behavior.
[0055] As can be seen from the above, this embodiment utilizes product gravity time-series data to more accurately identify product placement actions, effectively overcoming the shortcomings of traditional gravity sensors, such as their inability to handle simultaneous placement and retrieval of multiple products, partial returns, and sensitivity to changes in product shape. This improves transaction accuracy and inventory management reliability. When abnormal product behavior is detected, this embodiment can automatically select the target identification area image from the target video data, focusing on the problem area in a targeted manner, avoiding ineffective analysis of the entire video, saving computing resources and time costs, and improving processing efficiency.
[0056] This embodiment constructs user behavior profiles from target recognition area images, identifies abnormal behavior based on these profiles, and enables in-depth analysis of user operating habits. This not only enables timely detection of abnormal behaviors such as aggressive pickup, prolonged delays, and multiple pick-ups and put-aways without settlement, but also prevents potential risks, ensuring the safe operation of vending machines and reducing product loss and economic losses. Finally, the generated abnormal behavior information provides strong data support for subsequent system optimization and strengthened management, promoting the intelligent and efficient development of the vending industry.
[0057] In one embodiment of the present application, determining the product placement behavior of a vending machine based on product gravity time series data includes: Based on the commodity gravity time series data, the commodity weight change characteristics, change trend characteristics and pick-up and placement time characteristics are extracted.
[0058] Based on the change trend characteristics, the target time node where the change trend changes is selected.
[0059] Determine multiple time periods based on the target time node.
[0060] According to multiple time periods, the weight change characteristics, change trend characteristics and pick-up and placement time characteristics of the goods are divided into weight feature subsets corresponding to different time periods.
[0061] The picking and placing behavior of the target product is obtained based on the weight feature subsets corresponding to different time periods.
[0062] The target product is the product that is taken out and placed in the vending machine.
[0063] In this embodiment, the product picking and placing behavior of the target product is obtained based on the weight feature subsets corresponding to different time periods, including: The weight feature subset corresponding to each time period is input into the commodity behavior recognition model to obtain the commodity picking and placing behavior corresponding to each time period.
[0064] The product picking and placing behaviors corresponding to different time periods are spliced together according to the chronological order to obtain the product picking and placing behaviors of the target product.
[0065] In this example, by analyzing time-series data on product weight, key features reflecting product placement behavior are extracted. Trend features refer to the increase or decrease in product weight over time. By selecting the time point at which the trend changes as the target time node, different product placement phases can be identified.
[0066] Based on the target time node, the entire data time range is divided into multiple consecutive time periods. Product pick-up and placement behaviors are consistent within each time period. The extracted features are grouped according to the divided time periods to obtain the feature subset corresponding to each time period. The feature subset for each time period is input into the product behavior recognition model to obtain the product pick-up and placement behaviors for that time period, and then the complete product pick-up and placement behaviors are spliced together.
[0067] For example, this embodiment obtains product gravity time series data to ensure data accuracy and completeness. This embodiment calculates the weight difference between adjacent time points using the product gravity time series data to obtain product weight change characteristics. The trend of change characteristics is determined based on the positive or negative weight difference, and the time of each data point is recorded as the pick-and-place time characteristic.
[0068] This embodiment traverses the trend feature array. When the trend changes, such as from rising to falling and exceeding a set threshold, the time point is recorded as the target time node. This embodiment divides the entire time range into multiple time periods based on the target time nodes. For example, the first time period is from the start time to the first target time node, the second time period is from the first target time node to the second target time node, and so on.
[0069] This embodiment groups product weight change characteristics, change trend characteristics, and access time characteristics based on time periods, obtaining a weight feature subset corresponding to each time period. This embodiment inputs the feature subset for each time period into a product behavior recognition model. The product behavior recognition model can be a classification model trained based on machine learning or deep learning, and outputs product access behaviors, such as pickup and placement, for that time period. After obtaining the product access behaviors corresponding to multiple time periods, this embodiment concatenates these behaviors in chronological order to obtain the complete access behavior for the target product.
[0070] This embodiment uses multi-dimensional feature extraction and precise time period division to accurately capture changes in product placement behavior and improve recognition accuracy. Based on segmented processing of target time nodes, this embodiment can effectively cope with complex and changing product placement scenarios. This embodiment uses a trained product behavior recognition model to achieve automated and intelligent behavior judgment. This embodiment splices behaviors from different time periods to fully restore the product placement process, providing strong support for abnormal behavior monitoring and inventory management of vending machines.
[0071] In one embodiment of the present application, the target video data includes target product video data and target user video data.
[0072] Select target recognition area images from target video data based on product pick-up and placement behavior, including: Determine the target storage area corresponding to the target product where the product pick-up and placement behavior occurs.
[0073] The product image sequence data corresponding to the target storage area is selected from the target product video data, and the user image sequence data is selected from the target user video data.
[0074] The product image sequence data and user image sequence data are used as target recognition area images.
[0075] In this embodiment, the target storage area refers to the specific location of the target product where the product is taken and placed in the vending machine. This embodiment can locate the storage area where the target product is located based on the previously obtained product taking and placing behavior combined with the fixed product layout of the vending machine.
[0076] Selecting product image sequence data refers to extracting a series of image data corresponding to the target storage area from the target product video data. Selecting user image sequence data refers to obtaining an image sequence that records user behavior from the target user video data.
[0077] For example, this embodiment acquires target product video data, target user video data, and a product layout diagram of a vending machine. Combining product placement behavior information with the product layout diagram, the coordinates of the target product's storage area are determined. Based on the target storage area coordinates, image frames corresponding to the target product video data are captured to form product image sequence data. This embodiment can directly extract image frames from the target user video data to form user image sequence data. The product image sequence data and the user image sequence data are then combined to obtain the final target recognition area image.
[0078] This embodiment can accurately locate the target product storage area, focusing on key image information, avoiding interference from invalid video data, and improving the efficiency of subsequent analysis. The product image sequences and user image sequences obtained in this embodiment provide rich and targeted data for building user behavior profiles, helping to more accurately identify abnormal user behavior and enhance the security and management efficiency of vending machines.
[0079] In one embodiment of the present application, the target recognition area image includes commodity image sequence data and user image sequence data.
[0080] Determine user behavior profiles based on target recognition area images, including: Determine the product pick-up and placement results based on the product image sequence data.
[0081] A user behavior sequence is determined based on the user image sequence data.
[0082] The product picking and placing results and user behavior sequence are fused to obtain the user behavior profile.
[0083] This embodiment can use image processing and analysis technology to analyze the product image sequence data to determine whether the target product has been taken away or put back. For example, target detection, image comparison, etc. can be used to identify the state changes of the product in the image, thereby determining the product removal result.
[0084] This embodiment can identify a series of user actions during the operation of a vending machine from user image sequence data. For example, this embodiment uses human gesture recognition technology to determine the user's gestures, and then uses a behavior classification model to convert these gestures into specific actions to form a behavior sequence.
[0085] This embodiment integrates two different types of information: product placement results and user behavior sequences, to form a comprehensive profile describing user behavior. For example, feature fusion algorithms such as feature concatenation and weighted summation are used to fuse the two types of information, resulting in a profile that comprehensively reflects the user's interaction with the product.
[0086] For example, this embodiment can perform preprocessing operations such as grayscale conversion, noise reduction, and contrast enhancement on the product image sequence data to improve the accuracy of subsequent analysis. This embodiment uses a target detection algorithm to detect and track target products in the product image sequence. This embodiment determines whether the product has been removed or returned by comparing its position and status in different frames. Based on the results of target detection and tracking, this embodiment determines the product removal and placement result, such as "removed," "returned," or "not operated."
[0087] This embodiment uses human pose estimation algorithms such as OpenPose to estimate human poses in user image sequences, extracting key point information. Based on this extracted key point information, a deep learning model is used to classify user behaviors and identify various user actions, such as reaching, taking, and putting down. This embodiment groups the identified user actions into a user action sequence in chronological order.
[0088] This embodiment extracts features from both the product placement results and the user behavior sequence, such as the time and number of product placements, and the duration and frequency of user behavior. This embodiment fuses the extracted product placement results features with the user behavior sequence features, perhaps using simple concatenation or weighted summation. This embodiment inputs the fused features into a classification or clustering model to model the user's behavior and generate a user behavior profile.
[0089] This embodiment uses image processing to determine product placement results, accurately capturing product dynamics. This embodiment utilizes human posture recognition and behavior classification to construct behavioral sequences, accurately capturing user actions. This embodiment fuses these two features to form a comprehensive and integrated portrait that reflects the user's interaction with the product. This helps vending machine operators gain insight into user behavior habits and identify abnormal behavior, thereby optimizing product layout, improving service quality, and enhancing operational safety and efficiency.
[0090] Corresponding to an intelligent video recognition method of the above embodiment, Figure 2 This is a block diagram of the structure of an intelligent video recognition device provided by an embodiment of the present application. For ease of explanation, only the parts related to the embodiment of the present application are shown. Figure 2 The intelligent video recognition device 20 includes: a weight recognition module 21, a video recognition module 22 and an exception handling module 23.
[0091] The weight recognition module 21 is used to determine the product placement behavior of the vending machine based on the product gravity time series data. If the product placement behavior is abnormal, a target recognition area image is selected from the target video data based on the product placement behavior.
[0092] The video recognition module 22 is used to determine the user behavior profile based on the target recognition area image, and identify abnormal user behavior based on the user behavior profile.
[0093] The exception handling module 23 is configured to generate abnormal behavior information based on the identified abnormal behavior of the user.
[0094] In one embodiment of the present application, the weight identification module 21 is specifically configured to extract commodity weight change characteristics, change trend characteristics, and pick-up and placement time characteristics based on commodity gravity time series data.
[0095] Based on the change trend characteristics, the target time node where the change trend changes is selected.
[0096] Determine multiple time periods based on the target time node.
[0097] According to multiple time periods, the weight change characteristics, change trend characteristics and pick-up and placement time characteristics of the goods are divided into weight feature subsets corresponding to different time periods.
[0098] The picking and placing behavior of the target product is obtained based on the weight feature subsets corresponding to different time periods.
[0099] The target product is the product that is taken out and placed in the vending machine.
[0100] In one embodiment of the present application, the weight recognition module 21 is further configured to input the weight feature subset corresponding to each time period into the commodity behavior recognition model to obtain the commodity picking and placing behavior corresponding to each time period.
[0101] The product picking and placing behaviors corresponding to different time periods are spliced together according to the chronological order to obtain the product picking and placing behaviors of the target product.
[0102] In one embodiment of the present application, the target video data includes target product video data and target user video data. The weight recognition module 21 is further configured to determine a target storage area corresponding to the target product where the product picking and placing behavior occurs.
[0103] The product image sequence data corresponding to the target storage area is selected from the target product video data, and the user image sequence data is selected from the target user video data.
[0104] The product image sequence data and user image sequence data are used as target recognition area images.
[0105] In one embodiment of the present application, the target recognition area image includes product image sequence data and user image sequence data. The video recognition module 22 is specifically configured to determine a product pick-up and placement result based on the product image sequence data.
[0106] A user behavior sequence is determined based on the user image sequence data.
[0107] The product picking and placing results and user behavior sequence are fused to obtain the user behavior profile.
[0108] In one embodiment of the present application, an intelligent video recognition device 20 also includes: a data acquisition module for detecting product weight data. If a change in product weight is detected, the first monitoring frequency is updated to obtain a third monitoring frequency, and the second monitoring frequency is updated to obtain a fourth monitoring frequency.
[0109] The third monitoring frequency refers to the frequency of collecting target video data of the product and the user when the weight of the product changes.
[0110] The fourth monitoring frequency refers to the frequency of collecting commodity gravity time series data for commodities whose weight has changed when the weight of the commodities has changed.
[0111] The first monitoring frequency refers to the frequency of collecting video data of the product and the user when the weight of the product has not changed, and the second monitoring frequency refers to the frequency of collecting weight data of the product when the weight of the product has not changed.
[0112] The third monitoring frequency is greater than the first monitoring frequency, and the fourth monitoring frequency is greater than the second monitoring frequency.
[0113] In one embodiment of the present application, abnormal user behavior includes timing anomalies and risky behavior. The video recognition module 22 is further configured to utilize a statistical model of normal user behavior to identify timing anomalies and risky behavior within a user's behavior profile. Timing anomalies include abnormal operation timing and operation intervals, while risky behavior includes abnormal product placement and handling, and the risk of abnormal actions.
[0114] See also Figure 3 , Figure 3 This is a schematic block diagram of an electronic device provided in one embodiment of the present application. Figure 3 The electronic device 300 in the embodiment shown may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memory 304 is used to store computer programs, which include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. The processor 301 is configured to call the program instructions to execute the functions of the modules in the above-mentioned device embodiments, such as Figure 2 The functions of the weight recognition module 21, the video recognition module 22 and the exception handling module 23 are shown.
[0115] It should be understood that in the embodiment of the present application, the processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0116] The input device 302 may include a touchpad, a fingerprint collection sensor (for collecting user fingerprint information and fingerprint direction information), a microphone, etc. The output device 303 may include a display (LCD, etc.), a speaker, etc.
[0117] The memory 304 may include a read-only memory and a random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.
[0118] In a specific implementation, the processor 301, input device 302, and output device 303 described in the embodiments of the present application can execute the implementation methods described in the first and second embodiments of an intelligent video recognition method provided in the embodiments of the present application, and can also execute the implementation methods of the electronic device 300 described in the embodiments of the present application, which will not be repeated here.
[0119] In another embodiment of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, all or part of the process of the method in the above embodiment is implemented. The computer program can also be used to instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of each of the above method embodiments are implemented. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium.
[0120] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the aforementioned embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, the computer-readable storage medium can include both an internal storage unit of the electronic device and an external storage device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or is about to be output.
[0121] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0122] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the electronic devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0123] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces or units, or can be an electrical, mechanical or other form of connection.
[0124] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present application.
[0125] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0126] The above are only specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. An intelligent video recognition method, characterized in that: include: Determine the product placement behavior of the vending machine based on the product gravity time series data; If the product picking and placing behavior is abnormal, selecting a target recognition area image from the target video data based on the product picking and placing behavior; Determine a user behavior profile based on the target recognition area image, and identify abnormal user behavior based on the user behavior profile; Generate abnormal behavior information based on the identified abnormal user behavior.
2. The intelligent video recognition method according to claim 1, wherein: The method of determining the product taking and placing behavior of the vending machine based on the product gravity time series data includes: Extract commodity weight change characteristics, change trend characteristics, and pick-up and placement time characteristics based on commodity gravity time series data; Selecting a target time point at which a change trend occurs based on the change trend characteristics; Determining multiple time periods based on the target time node; Dividing the commodity weight change characteristics, the change trend characteristics, and the pick-up and put-down time characteristics into weight feature subsets corresponding to different time periods according to the multiple time periods; The commodity picking and placing behavior of the target commodity is obtained based on the weight feature subsets corresponding to different time periods; the target commodity is the commodity for which the commodity picking and placing behavior occurs in the vending machine.
3. The intelligent video recognition method according to claim 2, wherein: The obtaining of the picking and placing behavior of the target commodity based on the weight feature subsets corresponding to different time periods includes: The weight feature subset corresponding to each time period is input into the product behavior recognition model to obtain the product picking and placing behavior corresponding to each time period; The product picking and placing behaviors corresponding to different time periods are spliced together according to the chronological order to obtain the product picking and placing behaviors of the target product.
4. The intelligent video recognition method according to claim 1, wherein: The target video data includes target product video data and target user video data; The selecting a target recognition area image from the target video data based on the product picking and placing behavior includes: Determine the target storage area corresponding to the target product where the product pick-up and placement behavior occurs; Selecting product image sequence data corresponding to the target storage area from the target product video data, and selecting user image sequence data from the target user video data; The product image sequence data and the user image sequence data are used as target recognition area images.
5. The intelligent video recognition method according to claim 1, wherein: The target recognition area image includes product image sequence data and user image sequence data; The determining of the user behavior profile based on the target recognition area image includes: Determining a product pick-up and placement result based on the product image sequence data; determining a user behavior sequence based on the user image sequence data; The product picking and placing results and the user behavior sequence are subjected to feature fusion to obtain a user behavior profile.
6. The intelligent video recognition method according to claim 1, wherein: Also includes: The commodity weight data is detected. If a change in the commodity weight is detected, the first monitoring frequency is updated to obtain a third monitoring frequency, and the second monitoring frequency is updated to obtain a fourth monitoring frequency; The third monitoring frequency refers to the frequency of collecting target video data of the product and the user when the weight of the product changes; The fourth monitoring frequency refers to the frequency of collecting commodity gravity time series data for commodities whose weight changes when the commodity weight changes; The first monitoring frequency refers to the frequency of collecting video data of the product and the user when the weight of the product has not changed, and the second monitoring frequency refers to the frequency of collecting weight data of the product when the weight of the product has not changed; The third monitoring frequency is greater than the first monitoring frequency, and the fourth monitoring frequency is greater than the second monitoring frequency.
7. The intelligent video recognition method according to claim 1, wherein: The abnormal user behavior includes timing anomalies and risky behaviors; The identifying of abnormal user behavior based on the user behavior profile includes: Utilize the user's normal behavior statistical model to identify time series anomalies and risky behaviors of the user's behavior profile; The timing anomalies include abnormal operation time points and abnormal operation intervals, and the risky behaviors include abnormal product placement and abnormal action risks.
8. An intelligent video recognition device, characterized in that: include: A weight recognition module is used to determine the product placement behavior of the vending machine based on the product gravity time series data; If the product picking and placing behavior is abnormal, selecting a target recognition area image from the target video data based on the product picking and placing behavior; A video recognition module is used to determine a user behavior profile based on the target recognition area image, and identify abnormal user behavior based on the user behavior profile; The exception handling module is used to generate abnormal behavior information based on the identified abnormal user behavior.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Mirror image vision identification self-service vending machine system
CN108230553A
Vending machine
CN108765702A
Method and system for processing abnormal behaviors of customers in unmanned store
CN110147723A
Article terminal user behavior monitoring method and device and computer storage medium
CN111460855A
Image and gravity dual-mode automatic commodity identification system for door-opening self-taking type vending cabinet
CN111815852A
Cited By
Pick-and-place identification method, storage medium and unmanned retail equipment
CN120977046A