An intelligent video recognition method and device, equipment, and storage medium
By combining product gravity time-series data and video recognition technology, the system identifies products and abnormal user behavior in vending machines, solving the accuracy problem of traditional video recognition technology in complex scenarios and improving transaction accuracy and security.
Patent Information
- Application Number
- CN202510584189.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-05-07
AI Technical Summary
Traditional video recognition technology has low accuracy in recognizing abnormal product states and abnormal user behaviors in complex scenarios, and cannot effectively handle weight fluctuations caused by simultaneous picking and placing of multiple products, partial return of products, and changes in product shape.
By combining product gravity time-series data and target video data, the weight recognition module and video recognition module determine product handling behavior and user behavior profiles, identify abnormal behavior, and generate abnormal information.
It improves the accuracy of identifying product retrieval and placement behavior, reduces computing resources and time costs, promptly detects abnormal behavior, and ensures the safe operation and inventory management of vending machines.
Smart Images

Figure CN120472369B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of video recognition, and more particularly relates to an intelligent video recognition method and device, equipment and a storage medium. BACKGROUND
[0002] As a core carrier of new retail, the intelligent video recognition technology of the vending machine is a key to realize the automation of commodity transactions. The existing scheme mainly relies on a single gravity sensor or a fixed angle camera. The gravity sensor scheme determines the taking of goods by the change of the weight of the shelf. Although it can initially detect the change of the weight, it cannot handle complex interaction scenarios such as taking and placing multiple goods at the same time, and placing some back, and is sensitive to the non-take-and-place fluctuations of the weight caused by the stacking of beverage bottles and the change of the packaging form, and it is difficult to accurately associate the user behavior and the state of the goods.
[0003] The traditional video recognition scheme simply relies on the camera to collect video for visual recognition, and is limited by the fixed angle coverage blind area and the defects of target segmentation and trajectory tracking accuracy when multiple users operate in parallel, and the recognition accuracy is low.
[0004] In summary, the traditional video recognition technology still has the problem of low recognition accuracy of abnormal state of goods and abnormal behavior of users in complex scenarios. SUMMARY
[0005] The application aims to provide an intelligent video recognition method and device, equipment and a storage medium to improve the recognition accuracy of goods and user behavior in complex scenarios.
[0006] The first aspect of the embodiment of the application provides an intelligent video recognition method, comprising:
[0007] determining the taking and placing behavior of the goods of the vending machine based on the time sequence data of the gravity of the goods; if the taking and placing behavior of the goods is an abnormal behavior of the goods, selecting a target recognition area image from target video data based on the taking and placing behavior of the goods;
[0008] determining a user behavior portrait based on the target recognition area image, and performing abnormal behavior recognition of the user based on the user behavior portrait;
[0009] generating abnormal behavior information based on the recognized abnormal behavior of the user.
[0010] The second aspect of the embodiment of the application provides an intelligent video recognition device, comprising:
[0011] a weight recognition module configured to determine the taking and placing behavior of the goods of the vending machine based on the time sequence data of the gravity of the goods; if the taking and placing behavior of the goods is an abnormal behavior of the goods, the weight recognition module is configured to select a target recognition area image from target video data based on the taking and placing behavior of the goods;
[0012] a video recognition module, configured to determine a user behavior portrait based on the target recognition region image, and perform user abnormal behavior recognition based on the user behavior portrait;
[0013] an abnormality processing module, configured to generate abnormal behavior information based on the identified user abnormal behavior.
[0014] In a third aspect, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor implements the steps of the intelligent video recognition method when running the computer program.
[0015] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program implements the steps of the intelligent video recognition method when executed by a processor.
[0016] The intelligent video recognition method and device, the electronic device, and the computer readable storage medium provided by the embodiments of the present application have the following beneficial effects: in terms of commodity taking and placing behavior judgment, the embodiments of the present application use commodity gravity time sequence data to more accurately identify commodity taking and placing actions, effectively overcome the drawbacks of traditional gravity sensors that cannot handle multiple commodities being taken and placed at the same time, part of commodities being placed back, and sensitivity to commodity shape changes, and improve transaction accuracy and reliability of inventory management. When detecting abnormal commodity behavior, the embodiments of the present application can automatically select a target recognition region image from target video data, focus on the problem area, avoid invalid analysis of the entire video, save computing resources and time cost, and improve processing efficiency.
[0017] The embodiments of the present application construct a user behavior portrait through a target recognition region image, and identify abnormal behavior based on the user behavior portrait, which can deeply analyze user operation habits. Not only can the embodiments of the present application timely discover abnormal behaviors such as violent taking of goods, long-time retention, and multiple taking and placing without settlement, but also can prevent potential risks, ensure safe operation of vending machines, and reduce commodity loss and economic loss. Finally, the generated abnormal behavior information provides strong data support for subsequent system optimization and management strengthening, and promotes the intelligent and efficient development of the vending industry. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0019] Figure 1A flowchart of an intelligent video recognition method provided by an embodiment of the present application is shown in the figure.
[0020] Figure 2 A structural block diagram of an intelligent video recognition device provided by an embodiment of the present application is shown in the figure.
[0021] Figure 3 A schematic block diagram of an electronic device provided by an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0022] In the following description, specific details are set forth, such as particular system configurations, techniques, etc., in order to provide a thorough understanding of the embodiments of the present application. However, persons skilled in the art will understand that the present application can be practiced without these specific details. In other instances, well-known systems, structures, circuits, and techniques have not been shown in detail in order not to obscure the understanding of this description.
[0023] In order to make the objectives, technical solutions and advantages of the present application clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0024] Reference is made to Figure 1 , Figure 1 A flowchart of an intelligent video recognition method provided by an embodiment of the present application is shown in the figure. The method can include S101-S103.
[0025] S101: Determine the commodity taking and placing behavior of the vending machine based on commodity gravity time series data. If the commodity taking and placing behavior belongs to abnormal commodity behavior, select a target recognition region image from the target video data based on the commodity taking and placing behavior.
[0026] In the present embodiment, the commodity gravity time series data refers to the commodity weight change data sequence continuously collected by the gravity sensor of each shelf region in the vending machine at fixed time intervals.
[0027] For example, 4-8 detection zones are divided on each shelf layer, a thin film gravity sensor is pasted at the center of each region, the gravity sensor is connected to the edge computing module through a wire harness, and a hardware clock synchronization module is provided to ensure the timestamp accuracy. The edge computing module reads the sensor weight value at a fixed frequency, and stores it after associating the timestamp and region number.
[0028] In the present embodiment, the commodity taking and placing behavior can include commodity taking behavior, commodity placing behavior, and commodity abnormal behavior. By analyzing the dynamic change of the gravity time series data, the present embodiment can determine whether the user performs commodity taking or placing operation, and whether abnormal behavior such as commodity displacement occurs.
[0029] Exemplarily, the embodiment sets the weight change threshold and the minimum duration threshold according to the average weight of the goods, analyzes the gravity data frame by frame by using the time series analysis algorithm, and determines the taking and placing behavior in combination with the change direction when the weight change of the goods in a certain monitoring area meets the threshold condition. If the weight changes in adjacent areas at the same time, the embodiment can exclude the misjudgment caused by the movement of the goods across the area in combination with the size of the goods in the goods library.
[0030] In the embodiment, the abnormal goods behavior refers to an abnormal action that does not conform to the normal goods taking and placing mode, which can include violent taking of goods, return of foreign objects, and frequent exploratory taking and placing.
[0031] Exemplarily, the embodiment can establish a normal goods behavior model by rules such as the weight change range of goods taking and placing and the operation time interval. The embodiment compares the real-time goods gravity time series data with the normal goods behavior model, and determines that the behavior is abnormal when the weight change amplitude, rate or duration exceeds the normal range.
[0032] Exemplarily, the embodiment can calculate the mean and standard deviation of the weight change of the goods in each area based on the historical normal goods taking and placing data, and determine the normal weight fluctuation range. When the weight change in a certain area exceeds the normal range or the change rate exceeds the preset threshold, it can be marked as an abnormal goods behavior. Specifically, different abnormal types can be distinguished in combination with the change amplitude, rate and historical operation record.
[0033] In the embodiment, the target video data refers to the video stream containing the goods area of the shelf taken by the built-in camera of the vending machine in real time and the video stream of the surrounding area of the vending machine taken by the external camera in real time, which is used to record the user operation picture. The target recognition area image refers to the local image of a specific shelf area extracted from the target video data according to the occurrence position and time of the abnormal goods behavior.
[0034] Exemplarily, the embodiment establishes the mapping relationship between the physical coordinates of the shelf and the video pixel coordinates by camera calibration, locates the corresponding pixel area in the video according to the region number of the gravity sensor when the abnormal behavior occurs, and extracts the image. The camera calibration process can use Zhang's calibration method to obtain the camera intrinsic matrix and determine the pixel range of each gravity detection area in the video. The embodiment extracts the target area image of the corresponding frame from the video stream by receiving the region number and timestamp of the abnormal behavior.
[0035] When an abnormal behavior is detected in a certain gravity area, the embodiment extracts the image of the area from the current video frame according to the pre-stored pixel range. In addition to the current frame, 10 frames of images before and after the occurrence of the abnormal behavior are additionally extracted to form an image sequence containing the process before and after the behavior, which is used for subsequent user behavior analysis and abnormal behavior tracing.
[0036] S102: Determine a user behavior portrait based on the target recognition area image, and perform user abnormal behavior recognition based on the user behavior portrait.
[0037] In this embodiment, the user abnormal behavior includes timing abnormality and risk behavior. The user abnormal behavior recognition based on the user behavior portrait includes:
[0038] The user behavior portrait is subjected to timing abnormality recognition and risk behavior recognition by using the user normal behavior statistical model. The timing abnormality includes operation time point abnormality and operation interval abnormality, and the risk behavior includes abnormality of taking and placing goods and abnormal motion risk.
[0039] This embodiment extracts user behavior features by using the target recognition area image, and constructs a behavior portrait containing operation time, interval, taking and placing goods, and motion information. This embodiment establishes a user normal behavior statistical model according to historical normal operation data, and covers normal distribution range, mean value, and standard deviation of each behavior feature.
[0040] In the abnormality recognition stage, for timing abnormality, this embodiment can compare the time point and interval of real-time operation of the user with the normal time distribution and interval range in the model. If the operation time significantly deviates from the high-frequency period, or the operation interval exceeds the normal fluctuation range, it is determined as timing abnormality. For risk behavior, this embodiment can compare the real-time taking and placing quantity, weight change, and motion speed, force, and other features of the user with the preset standards and thresholds in the model. If the taking and placing quantity and weight are abnormal, or the motion has a violent tendency, etc., which exceeds the safety threshold, it is determined as risk behavior. Finally, the abnormal behavior of the user is recognized through comprehensive judgment.
[0041] In this embodiment, the user behavior portrait is a digital description of user operation habits and behavior patterns constructed by target recognition area image analysis, which is used for comparison and identification of abnormal behavior.
[0042] The user behavior portrait can include timing features, spatial features, and motion features. This embodiment can statistically analyze the normal operation mode of the user in the time, space, and motion dimensions based on historical behavior data, form a multi-dimensional feature vector, and use it as a benchmark model for abnormality detection.
[0043] For example, this embodiment can establish a user behavior database containing time, area, goods, and motion features. This embodiment uses statistical analysis methods to calculate the mean value, standard deviation, and quantile of each feature of the user behavior database, and constructs the interval range of normal behavior, such as the hour distribution histogram of operation time, the proportion of operation frequency in each area, the average taking speed and fluctuation range, etc.
[0044] For example, timing anomaly identification can include anomalies in operation time points and anomalies in operation intervals. Anomalies in operation time points refer to user operation times that deviate significantly from their normal high-frequency periods. Anomalies in operation intervals refer to time intervals between two consecutive pickups that exceed the normal fluctuation range.
[0045] This embodiment sets a reasonable threshold based on the time distribution and interval statistics of users' historical operations. When real-time operation data exceeds the threshold, it is judged as abnormal. Anomalies at a specific time point must simultaneously meet the criteria of interval abnormalities and action characteristic abnormalities, such as nighttime operations accompanied by rapid order pickup, to avoid occasional misjudgments.
[0046] Risk behavior identification can include abnormal product handling and abnormal action risks. Abnormal product handling refers to changes in the quantity or weight of products handled that do not conform to normal patterns, such as a single quantity exceeding the historical maximum or a weight change exceeding twice the standard value for the product. Abnormal action risks refer to user actions exhibiting violent tendencies or illegal characteristics, such as hand movement speed exceeding a preset normal threshold or multiple impacts with shelves by the user.
[0047] This embodiment can identify unsafe picking and placing behaviors and dangerous actions by combining preset standard weights in the product database with users' historical picking and placing data and real-time motion characteristics. For example, a standard weight database for products and motion risk thresholds can be established, such as defining rapid picking as a hand speed greater than 1.5 m / s. This embodiment can use a posture estimation algorithm to calculate hand motion parameters and compare them with the risk threshold to determine anomalies.
[0048] Anomalies in product retrieval and placement can be detected by comparing real-time weight changes with the standard value in the product inventory. If the real-time weight change is less than 0.7 times or greater than 1.3 times the standard value, and the number of items retrieved is greater than the historical maximum value, then anomalies in retrieval and placement are identified.
[0049] Abnormal movement risk detection can be achieved by calculating the average speed of the hand from touching the product to removing it from the shelf. If it is greater than 2m / s and accompanied by abnormal joint angles, such as an elbow bending angle of less than 30 degrees (indicating violent pulling), it is judged as an abnormal movement risk.
[0050] When an anomaly is triggered, this embodiment simultaneously saves the target identification area image, gravity time series data, and operation log to form a traceable risk evidence package.
[0051] S103: Generate abnormal behavior information based on the identified abnormal user behavior.
[0052] In this embodiment, an intelligent video recognition method further includes:
[0053] The product weight data is monitored. If a change in product weight is detected, the first monitoring frequency is updated to obtain the third monitoring frequency, and the second monitoring frequency is updated to obtain the fourth monitoring frequency.
[0054] The third monitoring frequency refers to the frequency of collecting target video data of the commodity and the user when the weight of the commodity changes.
[0055] The fourth monitoring frequency refers to the frequency of collecting commodity gravity time-series data of the commodity whose weight changes when the weight of the commodity changes.
[0056] The first monitoring frequency refers to the frequency of collecting video data of the commodity and the user when the weight of the commodity does not change, and the second monitoring frequency refers to the frequency of collecting weight data of the commodity when the weight of the commodity does not change.
[0057] The third monitoring frequency is greater than the first monitoring frequency, and the fourth monitoring frequency is greater than the second monitoring frequency.
[0058] In this embodiment, the first monitoring frequency refers to the frequency of collecting video data of the commodity and the user by the vending machine in the normal state. The second monitoring frequency is the frequency of collecting weight data of the commodity in the normal state. The first monitoring frequency and the second monitoring frequency are used for daily monitoring, balancing data collection accuracy and device resource consumption.
[0059] For example, in this embodiment, a timer can be set in the video collection module to control the camera to shoot at the first monitoring frequency. A timing collection program is set in the data collection circuit of the gravity sensor to read the weight data at the second monitoring frequency. After the vending machine device is started, the video collection module and the weight collection module start periodic data collection at the first monitoring frequency and the second monitoring frequency, respectively, and store the data in the local storage device.
[0060] The change of the weight of the commodity refers to the removal or return of the commodity, resulting in a difference in the weight value detected by the gravity sensor compared with the previous one. In this embodiment, the gravity sensor is used to monitor the weight of the commodity in real time, and when the weight value detected exceeds the set fluctuation range, it is determined that the weight of the commodity has changed. In this embodiment, a weight change threshold can be set in the data processing program, and when the change of the weight data exceeds the threshold, the corresponding processing flow is triggered.
[0061] When the weight of the commodity changes, this embodiment increases the frequency of video data collection and weight data collection. The third monitoring frequency is used to collect video data of the commodity and the user more frequently to obtain detailed behavior information. The fourth monitoring frequency is used to collect weight data of the commodity whose weight changes more intensively to accurately record the weight change process.
[0062] For example, after detecting the weight change signal, the control program of the video collection module and the weight collection module shortens the collection interval time to the time corresponding to the third monitoring frequency and the fourth monitoring frequency.
[0063] The target video data is video data containing the commodity and user behavior collected at a third monitoring frequency after the weight of the commodity changes. The commodity gravity time series data is a data sequence of the weight of the commodity changing over time collected at a fourth monitoring frequency. The target video data can include video frame images and shooting time stamps. The commodity gravity time series data can include weight values and collection time stamps.
[0064] In the high-frequency collection mode, the video collection device and the gravity sensor of the embodiment continue to work, and the collected data are stored as target video data and commodity gravity time series data, respectively, for subsequent determination of commodity taking and placing behavior and user abnormal behavior.
[0065] From the above, in terms of commodity taking and placing behavior judgment, the embodiment can more accurately identify commodity taking and placing actions by using commodity gravity time series data, effectively overcoming the drawbacks of traditional gravity sensors that cannot handle multiple commodities being taken and placed at the same time, partial return, and sensitivity to commodity shape changes, improving transaction accuracy and reliability of inventory management. When abnormal commodity behavior is detected, the embodiment can automatically select target recognition area images from the target video data, focus on the problem area, avoid invalid analysis of the entire video, save computing resources and time cost, and improve processing efficiency.
[0066] The embodiment constructs a user behavior portrait through target recognition area images and identifies abnormal behavior accordingly, which can deeply analyze user operation habits. Not only can it timely discover abnormal behaviors such as violent taking of goods, long-term retention, and multiple taking and placing without settlement, but also can prevent potential risks, ensure the safe operation of the vending machine, and reduce commodity loss and economic loss. Finally, the generated abnormal behavior information provides strong data support for subsequent system optimization and management strengthening, promoting the intelligent and efficient development of the vending industry.
[0067] In an embodiment of the present application, the commodity taking and placing behavior of the vending machine is determined based on commodity gravity time series data, comprising:
[0068] The commodity weight change feature, change trend feature, and taking and placing time feature are extracted based on the commodity gravity time series data.
[0069] The target time node at which the change trend changes is selected based on the change trend feature.
[0070] A plurality of time periods are determined based on the target time node.
[0071] The commodity weight change feature, change trend feature, and taking and placing time feature are divided into weight feature subsets corresponding to different time periods according to the plurality of time periods.
[0072] The commodity taking and placing behavior of the target commodity is obtained based on the weight feature subsets corresponding to different time periods.
[0073] The target commodity is a commodity in which a commodity taking and placing behavior occurs in a vending machine.
[0074] In this embodiment, the commodity taking and placing behavior of the target commodity is obtained based on the weight feature subsets corresponding to different time periods.
[0075] The weight feature subset corresponding to each time period is respectively input into a commodity behavior recognition model to obtain the commodity taking and placing behavior corresponding to each time period.
[0076] The commodity taking and placing behaviors corresponding to different time periods are spliced according to the time sequence to obtain the commodity taking and placing behavior of the target commodity.
[0077] In this embodiment, the key features reflecting the commodity taking and placing behavior are extracted by analyzing the commodity gravity time series data. The change trend feature refers to the increasing or decreasing trend of the commodity weight over time. The time point at which the trend changes is selected as the target time node, and different commodity taking and placing stages can be divided.
[0078] According to the target time node, the entire data time range is divided into multiple continuous time periods, and the commodity taking and placing behavior in each time period is consistent. The extracted features are grouped according to the divided time periods to obtain the feature subset corresponding to each time period. The feature subset of each time period is input into the commodity behavior recognition model to obtain the commodity taking and placing behavior of the time period, and then spliced into the complete commodity taking and placing behavior.
[0079] For example, the commodity gravity time series data is obtained in this embodiment to ensure the accuracy and integrity of the data. In this embodiment, the weight difference between adjacent time points is calculated based on the commodity gravity time series data to obtain the commodity weight change feature. The change trend feature is determined according to the positive and negative of the weight difference, and the time of each data point is recorded as the taking and placing time feature.
[0080] This embodiment traverses the change trend feature array, and when the trend changes, such as from rising to falling and exceeds the set threshold, the time point is recorded as the target time node. In this embodiment, the entire time range is divided into multiple time periods according to the target time node, for example, the first time period is from the starting time to the first target time node, the second time period is from the first target time node to the second target time node, and so on.
[0081] The embodiment divides the commodity weight change feature, the change trend feature and the taking and placing time feature according to the time period division, and obtains the weight feature subset corresponding to each time period. The embodiment inputs the feature subset of each time period into the commodity behavior recognition model. The commodity behavior recognition model can be a classification model trained based on machine learning or deep learning, and outputs the commodity taking and placing behavior of the time period, such as taking goods and placing goods. After obtaining the commodity taking and placing behaviors corresponding to multiple time periods, the embodiment splices these behaviors according to the time sequence to obtain the complete commodity taking and placing behavior of the target commodity.
[0082] The embodiment can accurately capture the commodity taking and placing behavior change and improve the recognition accuracy through multi-dimensional feature extraction and accurate time period division. The embodiment can effectively cope with complex and variable commodity taking and placing scenes based on the segmentation processing of the target time node. The embodiment realizes automatic and intelligent behavior judgment with the help of the trained commodity behavior recognition model. The embodiment splices the behaviors of each time period to completely restore the commodity taking and placing process, and provides strong support for automatic vending machine abnormal behavior monitoring and inventory management.
[0083] In an embodiment of the present application, the target video data includes target commodity video data and target user video data.
[0084] Selecting the target recognition region image from the target video data based on the commodity taking and placing behavior includes:
[0085] Determining the target storage area corresponding to the target commodity taking and placing behavior.
[0086] Selecting the commodity image sequence data corresponding to the target storage area from the target commodity video data, and selecting the user image sequence data from the target user video data.
[0087] Taking the commodity image sequence data and the user image sequence data as the target recognition region image.
[0088] In the embodiment, the target storage area refers to the specific storage position of the target commodity taking and placing behavior in the automatic vending machine. The embodiment can locate the storage area of the target commodity according to the previously obtained commodity taking and placing behavior and in combination with the fixed commodity layout of the vending machine.
[0089] Selecting the commodity image sequence data refers to extracting a series of image data corresponding to the target storage area from the target commodity video data. Selecting the user image sequence data refers to obtaining the image sequence recording the user behavior from the target user video data.
[0090] Exemplarily, the embodiment obtains target commodity video data, target user video data, and a commodity layout diagram of the vending machine. In combination with commodity taking and placing behavior information and the commodity layout diagram, a storage area coordinate of the target commodity is determined. According to the target storage area coordinate, an image frame at a corresponding position in the target commodity video data is intercepted to form commodity image sequence data. The embodiment can directly extract an image frame in the target user video data to form user image sequence data, and the commodity image sequence data and the user image sequence data are combined to obtain a final target recognition area image.
[0091] The embodiment can accurately locate the target commodity storage area, focus on key image information, avoid invalid video data interference, and improve subsequent analysis efficiency. The commodity image sequence and the user image sequence obtained by the embodiment provide rich and targeted data for constructing a user behavior portrait, which helps to more accurately identify user abnormal behavior and enhance the security and management efficiency of the vending machine.
[0092] In an embodiment of the present application, the target recognition area image includes commodity image sequence data and user image sequence data.
[0093] Determining a user behavior portrait based on the target recognition area image includes:
[0094] Determining a commodity taking and placing result based on the commodity image sequence data.
[0095] Determining a user behavior sequence based on the user image sequence data.
[0096] Feature fusion of the commodity taking and placing result and the user behavior sequence to obtain a user behavior portrait.
[0097] The embodiment can analyze the commodity image sequence data using image processing and analysis technology to determine whether the target commodity is taken away or placed back. For example, target detection, image comparison, and other means can be used to identify the state change of the commodity in the image, thereby determining the commodity taking and placing result.
[0098] The embodiment can identify a series of behavior actions of the user in the process of operating the vending machine from the user image sequence data. For example, the embodiment determines the action posture of the user through human posture recognition technology, and then converts these postures into specific behavior actions through a behavior classification model to form a behavior sequence.
[0099] The embodiment integrates the two different types of information of the commodity taking and placing result and the user behavior sequence together to form a portrait that comprehensively describes the user behavior. For example, feature splicing, weighted summation, and other feature fusion algorithms are used to fuse the two types of information, so that the portrait can comprehensively reflect the interaction behavior of the user and the commodity.
[0100] Exemplarily, the embodiment can perform preprocessing operations such as graying, noise reduction, and contrast enhancement on the commodity image sequence data to improve the accuracy of subsequent analysis. The embodiment uses a target detection algorithm to detect target commodities in the commodity image sequence and tracks them. The embodiment determines whether the commodities are taken out or put back by comparing the positions and states of the commodities in different frames. The embodiment determines the commodity taking and putting results, such as “taken out”, “put back”, or “no operation”, based on the results of target detection and tracking.
[0101] The embodiment estimates the human poses in the user image sequence using a human pose estimation algorithm such as OpenPose, extracts human key point information, and classifies the user's behavior using a deep learning model based on the extracted human key point information to identify various behavior actions of the user such as reaching out, taking, and putting back. The embodiment combines the identified behavior actions of the user into a user behavior sequence in chronological order.
[0102] The embodiment extracts features such as the time and frequency of commodity taking and putting, and the duration and frequency of user behavior from the commodity taking and putting results and the user behavior sequence. The embodiment fuses the extracted commodity taking and putting result features and user behavior sequence features, which can be performed by simple concatenation, weighted summation, and the like. The embodiment inputs the fused features into a classification or clustering model to model the user's behavior and obtain a user behavior portrait.
[0103] The embodiment determines the commodity taking and putting results through image processing and can accurately grasp the commodity dynamics. The embodiment uses human pose recognition and behavior classification to construct a behavior sequence and can accurately capture the user's operation actions. The embodiment fuses the features of the two to form a portrait, which comprehensively and comprehensively reflects the interaction between the user and the commodity. This helps the operator of the vending machine to understand user behavior habits, identify abnormal behavior, and thus optimize commodity layout, improve service quality, and enhance the safety and efficiency of operation.
[0104] A smart video recognition method corresponding to the above embodiment, Figure 2 A structural block diagram of a smart video recognition device provided by an embodiment of the present application. For ease of illustration, only parts related to the embodiments of the present application are shown. For reference Figure 2 The smart video recognition device 20 includes a weight recognition module 21, a video recognition module 22, and an abnormality processing module 23.
[0105] The weight recognition module 21 is configured to determine the commodity taking and putting behavior of the vending machine based on the commodity gravity time series data. If the commodity taking and putting behavior is an abnormal commodity behavior, the target recognition region image is selected from the target video data based on the commodity taking and putting behavior.
[0106] The video recognition module 22 is configured to determine a user behavior portrait based on the target recognition region image, and perform user abnormal behavior recognition based on the user behavior portrait.
[0107] The abnormal processing module 23 is configured to generate abnormal behavior information based on the recognized user abnormal behavior.
[0108] In an embodiment of the present application, the weight recognition module 21 is specifically configured to extract a commodity weight change feature, a change trend feature and a taking and placing time feature based on commodity gravity time sequence data.
[0109] The target time node at which the change trend changes is selected based on the change trend feature.
[0110] A plurality of time periods are determined based on the target time node.
[0111] The commodity weight change feature, the change trend feature and the taking and placing time feature are divided into weight feature subsets corresponding to different time periods according to the plurality of time periods.
[0112] The commodity taking and placing behavior of the target commodity is obtained based on the weight feature subsets corresponding to different time periods.
[0113] The target commodity is a commodity that performs a commodity taking and placing behavior in the vending machine.
[0114] In an embodiment of the present application, the weight recognition module 21 is specifically further configured to input the weight feature subset corresponding to each time period into a commodity behavior recognition model respectively, to obtain a commodity taking and placing behavior corresponding to each time period.
[0115] The commodity taking and placing behaviors corresponding to different time periods are spliced according to time sequence to obtain the commodity taking and placing behavior of the target commodity.
[0116] In an embodiment of the present application, the target video data includes target commodity video data and target user video data. The weight recognition module 21 is specifically further configured to determine a target storage region corresponding to a target commodity that performs a commodity taking and placing behavior.
[0117] The commodity image sequence data corresponding to the target storage region is selected from the target commodity video data, and user image sequence data is selected from the target user video data.
[0118] The commodity image sequence data and the user image sequence data are taken as target recognition region images.
[0119] In an embodiment of the present application, the target recognition region image includes the commodity image sequence data and the user image sequence data. The video recognition module 22 is specifically configured to determine a commodity taking and placing result based on the commodity image sequence data.
[0120] Determine user behavior sequences based on user image sequence data.
[0121] The user behavior profile is obtained by fusing the product retrieval and placement results with the user behavior sequence.
[0122] In one embodiment of this application, an intelligent video recognition device 20 further includes: a data acquisition module, used to detect commodity weight data; if a change in commodity weight is detected, the first monitoring frequency is updated to obtain a third monitoring frequency, and the second monitoring frequency is updated to obtain a fourth monitoring frequency.
[0123] The third monitoring frequency refers to the frequency at which target video data is collected from the product and the user when the product weight changes.
[0124] The fourth monitoring frequency refers to the frequency at which time-series gravity data of goods whose weight changes when their weight changes.
[0125] The first monitoring frequency refers to the frequency at which video data of the product and the user is collected when the product weight has not changed, and the second monitoring frequency refers to the frequency at which weight data of the product is collected when the product weight has not changed.
[0126] The third monitoring frequency is greater than the first monitoring frequency, and the fourth monitoring frequency is greater than the second monitoring frequency.
[0127] In one embodiment of this application, abnormal user behavior includes temporal anomalies and risky behaviors. The video recognition module 22 is further configured to utilize a statistical model of normal user behavior to identify temporal anomalies and risky behaviors in the user behavior profile. Temporal anomalies include abnormal operation timing and abnormal operation intervals, while risky behaviors include abnormal product handling and abnormal action risks.
[0128] See Figure 3 , Figure 3 This is a schematic block diagram of an electronic device provided according to an embodiment of this application. Figure 3 The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of the modules in the aforementioned device embodiments, for example... Figure 2 The functions of the weight recognition module 21, video recognition module 22, and anomaly handling module 23 are shown.
[0129] It should be understood that, in the embodiments of the present application, the processor 301 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0130] The input device 302 can include a touchpad, a fingerprint collection sensor (for collecting fingerprint information and direction information of a fingerprint of a user), a microphone, etc., and the output device 303 can include a display (LCD, etc.), a speaker, etc.
[0131] The memory 304 can include a read-only memory and a random access memory, and provide instructions and data for the processor 301. A part of the memory 304 can also include a non-volatile random access memory. For example, the memory 304 can also store device type information.
[0132] In specific implementations, the processor 301, the input device 302 and the output device 303 described in the embodiments of the present application can execute the implementation manners described in the first and second embodiments of the intelligent video recognition method provided by the embodiments of the present application, and can also execute the implementation manners of the electronic device 300 described in the embodiments of the present application, which will not be described here.
[0133] In another embodiment of the present application, a computer readable storage medium is provided, which stores a computer program. The computer program includes program instructions, which, when executed by a processor, implement all or part of the processes of the above-mentioned embodiment methods. The computer program can also instruct related hardware to complete the implementation. The computer program can be stored in a computer readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate form. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0134] The computer readable storage medium can be an internal storage unit of the electronic device of any of the preceding embodiments, such as a hard disk or a memory of the electronic device. The computer readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the computer readable storage medium can include both the internal storage unit and the external storage device of the electronic device. The computer readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer readable storage medium can also be used to temporarily store data that has been output or will be output.
[0135] Those skilled in the art can appreciate that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0136] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the electronic device and the units described above can refer to the corresponding processes in the above-mentioned method embodiments, which will not be described here.
[0137] In several embodiments provided in the present application, it should be understood that the disclosed electronic device and method can be implemented in other manners. For example, the embodiments of the apparatus described above are merely schematic, and the division of units is merely logical function division, and there can be other division manners in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection between units can be indirect coupling or communication connection through some interfaces, or can be electrical, mechanical or other forms of connection.
[0138] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments of the present application.
[0139] In addition, each functional unit in the various embodiments of the present application can be integrated in one processing unit, or each unit can exist physically as a separate unit, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware, or in the form of a software functional unit.
[0140] The above is merely specific embodiments of the present application, and the protection scope of the present application is not limited thereto, and any modification or replacement within the technical scope disclosed in the present application can be easily thought by those skilled in the art, and these modifications or replacements should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method of intelligent video recognition, the method comprising: The method comprises the following steps: determining the product taking and placing behavior of the vending machine based on the product gravity time series data; if the product taking and placing behavior is an abnormal product behavior, selecting a target recognition area image from the target video data based on the product taking and placing behavior; determining a user behavior portrait based on the target recognition area image, and identifying the user abnormal behavior based on the user behavior portrait; generating abnormal behavior information based on the identified user abnormal behavior; the method for determining the product taking and placing behavior of the vending machine based on the product gravity time series data comprises the following steps: extracting product weight change features, change trend features and taking and placing time features based on the product gravity time series data; selecting a target time node where the change trend changes based on the change trend features; determining a plurality of time periods based on the target time node; dividing the product weight change features, the change trend features and the taking and placing time features into weight feature subsets corresponding to different time periods according to the plurality of time periods; inputting the weight feature subset corresponding to each time period into a product behavior recognition model respectively to obtain the product taking and placing behavior corresponding to each time period; splicing the product taking and placing behaviors corresponding to different time periods according to the time sequence to obtain the product taking and placing behavior of the target product; the target product is the product that has the product taking and placing behavior in the vending machine; the method further comprises the following steps: detecting the product weight data, if the product weight changes, updating the first monitoring frequency to obtain the third monitoring frequency, and updating the second monitoring frequency to obtain the fourth monitoring frequency; the third monitoring frequency refers to the frequency of collecting target video data of the product and the user when the product weight changes; the fourth monitoring frequency refers to the frequency of collecting product gravity time series data of the product whose weight changes when the product weight changes; the first monitoring frequency refers to the frequency of collecting video data of the product and the user when the product weight does not change, and the second monitoring frequency refers to the frequency of collecting weight data of the product when the product weight does not change; the third monitoring frequency is greater than the first monitoring frequency, and the fourth monitoring frequency is greater than the second monitoring frequency.
2. The intelligent video recognition method of claim 1, wherein, the target video data comprises target product video data and target user video data; the method for selecting a target recognition area image from the target video data based on the product taking and placing behavior comprises the following steps: determining the target storage area corresponding to the target product that has the product taking and placing behavior; selecting product image sequence data corresponding to the target storage area from the target product video data, and selecting user image sequence data from the target user video data; taking the product image sequence data and the user image sequence data as the target recognition area image.
3. The intelligent video recognition method of claim 1, wherein, the target recognition area image comprises product image sequence data and user image sequence data; the method for determining a user behavior portrait based on the target recognition area image comprises the following steps: determining the product taking and placing result based on the product image sequence data; determining the user behavior sequence based on the user image sequence data; performing feature fusion on the product taking and placing result and the user behavior sequence to obtain the user behavior portrait.
4. The intelligent video recognition method of claim 1, wherein, The user abnormal behavior includes time sequence abnormality and risk behavior; The user abnormal behavior identification based on the user behavior portrait includes: The time sequence abnormality identification and the risk behavior identification of the user behavior portrait are performed by using a user normal behavior statistical model; The time sequence abnormality includes operation time point abnormality and operation interval abnormality, and the risk behavior includes abnormal commodity taking and placing and abnormal motion risk.
5. An intelligent video recognition apparatus, characterized by comprising: The method comprises the steps of: The data acquisition module is configured to detect commodity weight data, and if it is detected that the commodity weight has changed, update the first monitoring frequency to obtain a third monitoring frequency and update the second monitoring frequency to obtain a fourth monitoring frequency; The third monitoring frequency refers to the frequency of target video data acquisition of the commodity and the user when the commodity weight changes, and the fourth monitoring frequency refers to the frequency of commodity weight time sequence data acquisition of the commodity whose weight has changed when the commodity weight changes; The first monitoring frequency refers to the frequency of video data acquisition of the commodity and the user when the commodity weight does not change, and the second monitoring frequency refers to the frequency of weight data acquisition of the commodity when the commodity weight does not change; the third monitoring frequency is greater than the first monitoring frequency, and the fourth monitoring frequency is greater than the second monitoring frequency; The weight identification module is configured to determine a commodity taking and placing behavior of the vending machine based on the commodity weight time sequence data, and if the commodity taking and placing behavior is an abnormal commodity behavior, select a target recognition area image from the target video data based on the commodity taking and placing behavior; The weight identification module is specifically configured to extract a commodity weight change feature, a change trend feature, and a taking and placing time feature based on the commodity weight time sequence data; Select a target time node at which the change trend changes based on the change trend feature; Determine a plurality of time periods based on the target time node; Divide the commodity weight change feature, the change trend feature, and the taking and placing time feature into weight feature subsets corresponding to different time periods according to the plurality of time periods; Input the weight feature subset corresponding to each time period into a commodity behavior identification model respectively to obtain a commodity taking and placing behavior corresponding to each time period; According to the time sequence, splice the commodity taking and placing behaviors corresponding to different time periods to obtain a commodity taking and placing behavior of a target commodity; The target commodity is a commodity in the vending machine that has a commodity taking and placing behavior; The video identification module is configured to determine a user behavior portrait based on the target recognition area image and identify a user abnormal behavior based on the user behavior portrait; The abnormality processing module is configured to generate abnormal behavior information based on the identified user abnormal behavior.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, The processor executes the computer program to implement the steps of the method of any one of claims 1 to 4.
7. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 6. The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 4.
Citation Information
Patent Citations
Method and system for processing abnormal behaviors of customers in unmanned store
CN110147723A
Interaction behavior identification method and system for intelligent unmanned selling cabinet
CN119672856A