Intelligent turnover cabinet safety management method and system based on image recognition technology
By using image recognition technology and deep learning methods in intelligent turnover cabinets, separating user behavior and environmental factors and building a turnover cabinet usage pattern map, the problem of low recognition accuracy of traditional turnover cabinet management systems in safety management and complex environments is solved, accurate identification and early warning of abnormal behavior is achieved, and the system's safety management capabilities are improved.
Patent Information
- Application Number
- CN202510551723.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional turnover cabinet management systems are difficult to meet the high requirements for material safety management in modern production environments, and there are safety hazards such as material loss, mistaking, and misuse. In addition, the recognition accuracy of the existing smart turnover cabinet systems has decreased in complex lighting and multi-person collaboration scenarios, and lacks deep learning and prediction capabilities.
The intelligent turnover cabinet safety management method based on image recognition technology is adopted, and the user interaction image sequence is collected through the intelligent turnover cabinet, and the lighting standardization and target area cropping is performed. The user behavior and environmental factors are separated by a two-level Transformer encoder, and multi-scale time characteristics are extracted in combination with the timing analysis network, and the turnover cabinet usage pattern map is constructed, abnormal mode detection results are generated and the management system is reported in real time through the RESTful interface.
It realizes accurate identification and early warning of user abnormal behavior, provides safety prompts through storage indicators, realizes a multi-level security early warning mechanism, improves the system's understanding of complex operating sequences, and enhances the spatial and temporal sequence prediction capability of cabinet usage patterns.
Smart Images

Figure CN120088900A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition technology, and particularly to an intelligent turnover cabinet security management method and system based on image recognition technology. Background Art
[0002] Traditional turnover cabinet management mainly relies on manual records and simple electronic lock control, which is difficult to meet the high requirements of modern production environments for material security management. Especially in the management process of high-value components, tooling, and important production materials, the lack of real-time monitoring and behavior analysis capabilities during use leads to frequent safety hazards such as material loss, misappropriation, and misuse, causing unnecessary economic losses and decreased production efficiency for enterprises.
[0003] Currently, the security management systems of intelligent turnover cabinets mainly use single technical means such as RFID tag recognition, barcode scanning, and biometric recognition for identity verification and switch cabinet control. However, these methods have obvious limitations in practical applications. For example, these technologies cannot effectively capture the complete process of user interaction with the cabinet body, making it difficult to identify unconventional operation behaviors; the lack of separate analysis of environmental factors and user behaviors leads to a decrease in recognition accuracy in scenarios such as complex lighting and multi-person collaboration; at the same time, existing systems lack the ability to deeply learn and predict the usage patterns of the cabinet body and cannot detect potential security risks in advance. Summary of the Invention
[0004] This application provides an intelligent turnover cabinet security management method and system based on image recognition technology. This application enables the intelligent turnover cabinet to accurately identify abnormal behaviors deviating from the normal mode, and can provide intuitive safety prompts through different colors and flashing modes of the storage location indicator lights according to the detection results of the abnormal mode. At the same time, it reports to the management system in real time through the RESTful interface to implement a multi-level security warning mechanism.
[0005] In the first aspect, this application provides an intelligent turnover cabinet security management method based on image recognition technology. The intelligent turnover cabinet security management method based on image recognition technology includes: Collecting a sequence of user interaction images through the intelligent turnover cabinet, and performing illumination normalization and target area cropping processing on the sequence of user interaction images to obtain a normalized image sequence; Inputting the normalized image sequence into a two-stage Transformer encoder for separating user behavior and environmental factors to obtain user behavior feature data; Inputting the user behavior feature data into a time series analysis network for multi-scale time feature extraction to obtain user operation behavior features; Construct a turnover cabinet usage pattern map based on the user operation behavior characteristics, and generate cabinet usage prediction results and abnormal pattern detection results based on the turnover cabinet usage pattern map. The abnormal pattern detection results are used to trigger the storage location indicator to give safety prompts in different colors and flashing modes, and report to the management system in real time through the RESTful interface.
[0006] In a second aspect, the present application provides an intelligent turnover cabinet safety management system based on image recognition technology. The intelligent turnover cabinet safety management system based on image recognition technology includes: An acquisition module, configured to acquire a user interaction image sequence through an intelligent turnover cabinet, and perform illumination normalization and target area cropping processing on the user interaction image sequence to obtain a normalized image sequence; A processing module, configured to input the normalized image sequence into a two-stage Transformer encoder for separating user behavior and environmental factors to obtain user behavior feature data; A feature extraction module, configured to input the user behavior feature data into a time series analysis network for multi-scale time feature extraction to obtain user operation behavior characteristics; A generation module, configured to construct a turnover cabinet usage pattern map according to the user operation behavior characteristics, and generate cabinet usage prediction results and abnormal pattern detection results based on the turnover cabinet usage pattern map. The abnormal pattern detection results are used to trigger the storage location indicator to give safety prompts in different colors and flashing modes, and report to the management system in real time through the RESTful interface.
[0007] In the technical solution provided by this application, the effective separation of user behavior and environmental factors is achieved through the dual-level Transformer encoder technology, solving the problem of low accuracy of behavior recognition in complex environments. The light normalization processing and target area cropping can reduce the influence of environmental light changes and interference from irrelevant areas, making the behavior feature extraction more accurate. The multi-scale temporal feature extraction mechanism can capture the short-term, medium-term, and long-term user operation behavior features simultaneously, and calculate the temporal correlation features through the cross-attention mechanism, significantly improving the system's understanding ability of complex operation sequences. The three-layer heterogeneous graph structure (spatiotemporal behavior layer, user feature layer, and cabinet association layer) based on the turnover cabinet usage pattern atlas can comprehensively represent the user interaction pattern. Combining with the GraphSAGE graph neural network for representation learning enables the system to accurately identify abnormal behaviors deviating from the normal pattern. The system can provide intuitive safety prompts according to the abnormal pattern detection results through different colors and flashing modes of the storage location indicator lights, and at the same time report to the management system in real time through the RESTful interface, realizing a multi-level safety warning mechanism. Through the spatiotemporal sequence prediction of the cabinet usage pattern, the system can predict the cabinet usage frequency and possible abnormal situations in advance, providing data support for management decisions and realizing the transformation of the management mode from passive response to active prevention. Based on the Pearson correlation coefficient matrix analysis of the usage correlation between multiple cabinets in the same area, the system can discover the spatial correlation features of cabinet usage, and achieve seamless integration with the upper-level management system through the standardized RESTful interface, supporting real-time data reporting and information sharing, and enhancing the collaborative working ability of the intelligent turnover cabinet and other enterprise information systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0009] Figure 1 FIG. is a schematic diagram of an embodiment of the intelligent turnover cabinet safety management method based on image recognition technology in the embodiments of this application; Figure 2 FIG. is a schematic diagram of an embodiment of the intelligent turnover cabinet safety management system based on image recognition technology in the embodiments of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0010] The embodiments of the present application provide an intelligent turnover cabinet security management method and system based on image recognition technology. Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims, and the above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the term "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0011] For ease of understanding, the specific process of the embodiments of the present application will be described below. Please refer to Figure 1 , an embodiment of the intelligent turnover cabinet security management method based on image recognition technology in the embodiments of the present application includes: Step S101: Collect a user interaction image sequence through the intelligent turnover cabinet, and perform illumination normalization and target area cropping processing on the user interaction image sequence to obtain a normalized image sequence; It can be understood that the execution subject of the present application can be an intelligent turnover cabinet security management system based on image recognition technology, or a terminal or a server. Specifically, it is not limited here. The embodiments of the present application will be described by taking the server as the execution subject as an example.
[0012] Specifically, a 170° wide-angle camera is installed inside the intelligent turnover cabinet system to cover the entire cabinet area with a large viewing angle and capture all interaction behaviors of users when picking up and placing equipment. A 120° viewing angle camera is installed outside the turnover cabinet to collect external behaviors of users when operating the turnover cabinet, including identity authentication, hand operations, and position changes during the interaction process. The dual-view layout forms an internally and externally linked image acquisition network. During the data acquisition process, the wide-angle camera and the viewing angle camera record image data at acquisition frequencies of 10 frames per second and 15 frames per second respectively. Gray value remapping is performed on the image sequences acquired from the internal and external dual views, and the image brightness under different lighting environments is normalized through the adaptive histogram equalization algorithm, so that the images maintain consistent contrast and detail performance under different lighting conditions, improving the accuracy of user operation area recognition. The YOLOv5 object detection algorithm is used to identify the target area of the image after lighting standardization. This algorithm precisely locates the user operation area (such as hands, assets, storage locations, etc.) as the target area, and crops out the target area image. The cropped area image is uniformly standardized to a fixed size of 224×224 pixels to form a region standardized image. Temporal frame difference processing is performed on the region standardized image, and a dynamic change map is generated by calculating the pixel changes between adjacent image frames frame by frame, highlighting the key information of the regional changes during user operations and effectively reflecting the change characteristics of the operation area at different time points. The dynamic change map is segmented by time window, and the image data is divided into different analysis window sequences according to a fixed duration (such as 1.5 seconds, 3 seconds, or 4.5 seconds). Each analysis window retains the image frame data within a fixed time, and organizes these image frame data into structured metadata, including the time stamp of the image frame, the cabinet ID, the storage location number, and the coordinate information of the user operation area. These metadata are combined with the analysis window sequence to form a standardized image sequence.
[0013] Step S102: Input the standardized image sequence into a two-stage Transformer encoder for separating user behaviors and environmental factors to obtain user behavior feature data; Specifically, the standardized image sequence is input into a two-stage Transformer encoder, which includes an environmental context encoding branch and a user behavior encoding branch. The environmental context encoding branch is used to extract background information and environmental features in the standardized image sequence, while the user behavior encoding branch is used to extract behavior features during the user's operation. The two process the image data in parallel and capture the feature information of the interaction scenario from different dimensions. In the environmental context encoding branch, a 6-layer Vision Transformer network is used to extract environmental features from the standardized image sequence. Each layer of the Vision Transformer contains 8 attention heads, and each attention head calculates information in different dimensions. The spatial features in the image sequence are captured through the multi-head self-attention mechanism, and the spatial position information is retained through position encoding. After the stacking of multiple layers of Transformer networks, the feature vectors extracted by the environmental context encoding branch accurately capture the context information of the image background, including static features such as the spatial distribution of assets in the cabinet, lighting conditions, and storage location status. At the same time, the user behavior encoding branch adopts an 8-layer temporal Vision Transformer structure. By inputting the standardized image sequence into the temporal encoding network, the user's interaction behavior information is captured frame by frame, and the self-attention mechanism is used to capture the dependencies in the temporal dimension between frames, obtaining the user behavior spatial features. The user behavior spatial features are input into the temporal Transformer structure for temporal analysis, and the spatio-temporal behavior features are refined by capturing the temporal change patterns in the image sequence. The temporal Transformer uses multi-scale time windows for feature extraction to capture short-term interaction behaviors (such as instantaneous operations of single pick-and-place equipment) and long-term behavior patterns (such as abnormal stays or illegal pick-and-place after multiple operations). The environmental feature vector and the spatio-temporal behavior features are separated through a feature decoupling layer. The feature decoupling layer separates the environmental features from the user behavior features through a feature disentanglement mechanism, removing the interference of environmental factors on the user behavior features, and obtaining the purified behavior features after separation. The purified behavior features after separation are subjected to feature comparison calculation. By comparing the current behavior features with the reference features of normal operation behaviors, it is determined whether the user behavior is abnormal, and thus the user behavior feature data is output.
[0014] Step S103: Input the user behavior feature data into a temporal analysis network for multi-scale time feature extraction to obtain the user operation behavior features; Specifically, the user behavior feature data is input into a time series analysis network, which adopts a multi-scale modeling strategy to capture short-term, medium-term, and long-term time series information respectively, forming multi-scale behavior representations. By modeling the interaction behaviors at different time scales, complex patterns are identified, including both instantaneous pick-and-place device operations, as well as longer stay behaviors and continuous operation behaviors. During the process of multi-scale feature extraction, the user behavior feature data is segmented by a sliding window, and the continuous time series data is divided into F groups of behavior sequence segments at fixed time intervals, where F equals 3, corresponding to short-term, medium-term, and long-term time periods respectively. The F groups of behavior sequence segments are respectively input into the corresponding time series convolutional modules for feature mapping. The time series convolutional modules adopt multi-layer one-dimensional convolutions, and capture time features at different scales through convolutional kernels of different sizes (such as 3, 5, 7), obtaining multi-scale time series mapping features. Cross-attention mechanism calculations are performed on the multi-scale time series mapping features to identify the dependencies between different time segments. By dynamically adjusting the feature weights, the key features related to the user's true behavior pattern are highlighted, generating time series correlation features. The time series correlation features are subjected to feature transformation through an adaptive pooling layer and a three-layer fully connected network. The adaptive pooling layer dynamically adjusts the size of the pooling window according to the length of the time series features, thereby reducing the feature dimension without losing key information and improving the calculation efficiency. After feature transformation through the three-layer fully connected network, the abstract ability of the features is enhanced through non-linear transformation, obtaining more advanced behavior abstract features. The advanced behavior abstract features are input into a behavior semantic classifier, which is based on a multi-layer perceptron structure. Through classification mapping of the behavior features, the final user operation behavior features are output. The behavior semantic classifier has classification ability through pre-training with a large amount of operation behavior data and can identify various behavior patterns such as normal operations, unauthorized pick-and-place, and long stays.
[0015] Preprocess the F-group behavior sequence segments, including zero-padding and normalization. During zero-padding, ensure that sequences with insufficient length are made to have the same length by padding zeros at both ends, thereby eliminating the impact of sequence length differences on the temporal convolutional module. Normalization, on the other hand, normalizes the amplitude of the behavior feature data, mapping all feature values to the same scale range, so that the features have a unified scale specification when input into the temporal convolutional module. Input the standardized behavior sequence segments into the first one-dimensional convolutional layer of the corresponding temporal convolutional module for feature processing. This layer uses convolutional kernels of different sizes (such as 3, 5, 7) to capture short-term, medium-term, and long-term behavior features, and performs feature mapping on the sequence segments along the time dimension through one-dimensional convolution operations to generate the first-layer feature map. Use the first-layer feature map as input and continue to pass it to the second one-dimensional convolutional layer of the corresponding temporal convolutional module to capture higher-dimensional temporal features. Convolutional kernels of different scales can more comprehensively cover the temporal variation characteristics of user interaction behaviors, extract the behavior patterns at the medium-term time scale, and generate the second-layer feature map. Input the second-layer feature map into the third one-dimensional convolutional layer for feature refinement. The convolutional kernel of this layer has a larger receptive field, can capture long-term behavior patterns, identify abnormal behavior trends that occur during long-term interactions of users, and output the third-layer feature map. Input the third-layer feature map into the fourth one-dimensional convolutional layer. Through further feature mapping, make the feature representation more temporally consistent, and generate the final F time-scale output features. Each feature corresponds to a behavior representation at a time scale, covering short-term, medium-term, and long-term behavior patterns during user interactions. Align and pad the F time-scale output features in the time dimension, and dynamically adjust the weights of different-length sequences through the time mask mechanism to ensure that multi-scale features can be accurately aligned in the time dimension and avoid feature misalignment caused by length differences. Concatenate in the feature channel dimension, and form a multi-scale temporal feature representation by merging the feature mapping results of different time scales.
[0016] Step S104: Construct a turnover cabinet usage pattern map based on the user operation behavior characteristics, and generate cabinet usage prediction results and abnormal pattern detection results based on the turnover cabinet usage pattern map. The abnormal pattern detection results are used to trigger the storage location indicator lights for safety prompts with different colors and flashing patterns, and are reported to the management system in real time through the RESTful interface.
[0017] Specifically, the user operation behavior characteristics are deeply analyzed and classified in the spatio-temporal dimension. By analyzing the behaviors of picking up and placing devices, interaction patterns, and device usage frequencies of users in different time periods, a turnover cabinet usage pattern map containing a spatio-temporal behavior layer, a user feature layer, and a cabinet association layer is constructed. The spatio-temporal behavior layer records the changes in behavior sequences within each time window during the user interaction process. The user feature layer captures information such as user identity, operation proficiency, and interaction habits. The cabinet association layer establishes the spatial association patterns between different storage locations. These three-layer structures together form the turnover cabinet usage pattern map. The turnover cabinet usage pattern map is input into a three-layer GraphSAGE graph neural network for graph-structured representation learning. Through the aggregation mechanism of the graph neural network, the association information between different nodes in the map is effectively captured, the feature representations of nodes at different levels are learned, and a graph-structured representation containing spatio-temporal behavior features, user features, and cabinet association features is generated. At the same time, the GraphSAGE network deeply analyzes the spatial dependence relationship between cabinets. Through the convergence and update mechanism of cross-node information, spatial association features are generated, reflecting the interaction dependence patterns between different storage locations. When performing temporal modeling on the spatial association features and spatio-temporal behavior features, a temporal convolutional network is used to extract features and model the spatio-temporal sequence data, capture the operation patterns within a long time scale, and dynamically model the temporal features of the user behavior sequence to generate spatio-temporal sequence prediction data. Through the feature mapping and time perception processing of the multi-layer convolutional network, an accurate prediction result of the cabinet usage frequency is obtained. This prediction result predicts the potential patterns such as usage peaks and abnormal pick-up and placement in the future time period of the cabinet 24 hours in advance. Based on the turnover cabinet usage pattern map and the cabinet usage prediction result, the deviation metric value between the current usage pattern and the historical normal pattern is calculated. By comparing the similarity between the current user operation behavior and the historical operation pattern, it is identified whether there is an abnormal usage pattern. To detect the abnormal pattern, the deviation metric threshold is used to judge the degree of abnormality of the user behavior. When the deviation metric value exceeds the set threshold, an abnormal usage pattern detection result is generated. The detection result includes abnormal situations such as unauthorized access, device misplacement, long stay, and frequent pick-up and placement. According to the detected abnormal pattern, the storage location indicator lights are triggered to give safety prompts in different colors and flashing modes. Among them, the green light indicates normal, the yellow light flashes at a frequency of 2 Hz to prompt unskilled operation or minor abnormality, the orange light flashes at a frequency of 4 Hz to warn of obvious abnormal behavior, and the red light indicates a serious safety threat and triggers the cabinet door to lock and start video recording and storage. At the same time, the abnormal detection result is reported to the management system in real time through the RESTful interface, and the abnormal information is pushed to the upper-layer business management platform through the URL so that the management personnel can intervene in time and take measures.
[0018] Integrate timestamps and map spatial coordinates for user operation behavior characteristics. By matching the data of user interaction behaviors with corresponding timestamps and combining the storage location codes inside the cabinet and the geographical location information of the cabinet, map the interaction data into a fixed spatial coordinate system to generate structured spatio-temporal behavior data with time and space attributes. Perform recursive clustering on the structured spatio-temporal behavior data. Through density-based spatio-temporal behavior similarity measurement algorithms (such as DBSCAN or HDBSCAN), automatically aggregate similar user operation behavior segments and generate a set of spatio-temporal behavior prototypes. These prototype sets represent the most common interaction patterns of different users during the use of the turnover cabinet, such as normal equipment picking and placing, long stays, frequent picking and placing, or abnormal access and other behavior patterns. Based on the set of spatio-temporal behavior prototypes, create a spatio-temporal behavior layer. Map each behavior pattern to a node in the graph and model the temporal transition relationship between different behavior patterns as an edge between nodes to form the spatio-temporal behavior layer. Analyze the interaction frequency, operation proficiency, and usage habit patterns of users based on the set of spatio-temporal behavior prototypes. By statistically analyzing the interaction behaviors of users in different time periods, the temporal sequence rules of equipment picking and placing, and the probability of abnormal behaviors occurring, generate user behavior portrait data, which reflects the personalized behavior characteristics of users during the operation of the turnover cabinet and reveals the usage habits and potential risk preferences of users in different environments. According to the user behavior portrait data, create a user feature layer. The nodes in this layer represent different user portraits, and the edges represent the similarity of operation habits between users, forming an association network of user features. To construct the cabinet association layer, combine the user behavior portrait data with the geographical location information of the cabinet. Generate a cabinet association matrix by calculating the usage correlation between multiple cabinets in the same area. This process is calculated using the Pearson correlation coefficient matrix. Analyze the switching frequency of users between different cabinets, the usage patterns of adjacent cabinets in the same time period, and the correlation of changes in cabinet storage locations to obtain the cabinet association matrix, and create the cabinet association layer accordingly. The nodes of this layer represent different cabinets, and the edges represent the usage association strength between cabinets. Perform inter-layer relationship mapping on the spatio-temporal behavior layer, user feature layer, and cabinet association layer. By defining the inter-layer connection weights and the association function between nodes, achieve the fusion of features at different levels. Calculate the importance of inter-layer nodes through the multi-layer attention mechanism of the heterogeneous graph and dynamically adjust the contribution of features at different levels to the overall pattern. By calculating the node similarity weight and the path importance score, measure the association strength of different nodes in the heterogeneous graph spectrum and the importance of the information propagation path. The node similarity weight measures the similarity degree between behavior patterns, user features, and cabinet usage patterns, while the path importance score is used to judge the important paths of information flow propagation in the graph. Based on the node similarity weight and the path importance score, fuse the three-layer information to form a turnover cabinet usage pattern spectrum.
[0019] In the embodiments of the present application, the effective separation of user behavior and environmental factors is achieved through the dual-stage Transformer encoder technology, solving the problem of low accuracy of behavior recognition in complex environments. The light normalization processing and target area cropping can reduce the influence of ambient light changes and interference from irrelevant areas, making the behavior feature extraction more accurate. The multi-scale time feature extraction mechanism can capture the short-term, medium-term, and long-term user operation behavior features simultaneously, and calculate the temporal correlation features through the cross-attention mechanism, significantly improving the system's understanding ability of complex operation sequences. The three-layer heterogeneous graph structure (spatiotemporal behavior layer, user feature layer, and cabinet association layer) based on the turnover cabinet usage pattern atlas can comprehensively represent the user interaction pattern. Combining with the GraphSAGE graph neural network for representation learning enables the system to accurately identify abnormal behaviors deviating from the normal pattern. The system can provide intuitive safety prompts through different colors and blinking modes of the storage location indicator according to the abnormal pattern detection results, and at the same time report to the management system in real time through the RESTful interface, realizing a multi-level safety warning mechanism. Through the spatiotemporal sequence prediction of the cabinet usage pattern, the system can predict the cabinet usage frequency and possible abnormal situations in advance, providing data support for management decisions and realizing the transformation of the management mode from passive response to proactive prevention. Based on the Pearson correlation coefficient matrix analysis of the usage correlation between multiple cabinets in the same area, the system can discover the spatial correlation features of cabinet usage, and achieve seamless integration with the upper-layer management system through the standardized RESTful interface, supporting real-time data reporting and information sharing, and enhancing the collaborative working ability of the intelligent turnover cabinet and other enterprise information systems.
[0020] In a specific embodiment, the process of executing step S101 may specifically include the following steps: Install a wide-angle camera inside the intelligent turnover cabinet and a perspective camera outside the turnover cabinet, and collect the internal and external dual-perspective image sequences through the wide-angle camera and the perspective camera; Perform gray value remapping on the internal and external dual-perspective image sequences to obtain the light-normalized images, and perform user operation area recognition on the light-normalized images to obtain the area-normalized images; Perform temporal frame difference processing on the area-normalized images to obtain the dynamic change diagrams, and perform time window segmentation on the dynamic change diagrams to obtain the analysis window sequences; Record the time stamp, cabinet ID, and operation area coordinates for each image frame in the analysis window sequence to form structured metadata, and combine the structured metadata with the analysis window sequence to obtain the standardized image sequence.
[0021] Specifically, a 170° wide-angle camera is installed inside the turnover cabinet to monitor the entire storage space comprehensively, and at the same time capture the characteristic information of the position change, storage position occupancy status, and abnormal placement behavior of the device during the picking and placing process. A 120° perspective camera is installed outside the turnover cabinet to capture the user's operation behavior, including hand picking and placing actions, identity recognition process, interaction duration, and other behavior characteristics. The dual-view image acquisition scheme inside and outside ensures complete coverage of the user interaction behavior and the cabinet device status, forming multi-dimensional and multi-perspective interaction behavior data. Gray value remapping is performed on the dual-view image sequences inside and outside to achieve illumination normalization processing. Since the turnover cabinet is in different lighting environments, such as natural light, artificial light source, or low-light environment, the original image sequences collected have problems such as uneven brightness and poor contrast. The adaptive histogram equalization algorithm is used to remap the image gray values. This algorithm makes the image gray histogram distribution uniform, so that the brightness of the image remains consistent under different lighting conditions, thereby enhancing image details and improving feature recognition accuracy, and obtaining illumination-normalized images. User operation area recognition is performed on the illumination-normalized images. Through the YOLOv5 object detection algorithm, the user's hand, storage position, and device boundary are accurately located in the image, and corresponding bounding boxes are generated. The illumination-normalized image is input into the YOLOv5 model for target area recognition. Image features are obtained through the feature extraction network, and the user operation area is located in combination with the boundary regression module to obtain area-normalized images. Temporal frame difference processing is performed on the area-normalized images to capture the dynamic change characteristics during the user interaction process. Temporal frame difference identifies the moving areas in the image sequence by calculating the pixel differences between adjacent frames. This process highlights the operation changes of the user picking and placing the device and the changes in the storage position status, obtaining a dynamic change map, which reflects the operation behavior trajectory of the user at different time periods, including the hand movement when picking and placing the device, the change in the device position, and the temporal characteristics of abnormal operation behaviors. Through temporal frame difference processing, the static background information is removed, and only the dynamic change areas related to the user interaction are retained, thereby improving the accuracy of temporal behavior analysis. The dynamic change map is segmented by time windows to obtain an analysis window sequence. Time window segmentation divides the dynamic change map into different time segments through a time sliding window of a fixed length. Each time window covers a fixed time length, such as 1.5 seconds, 3 seconds, or 4.5 seconds, so as to capture the change characteristics of the user interaction behavior at different time scales. This time window segmentation mechanism helps to retain the temporal continuity of the user operation behavior and provides multi-time scale data support for subsequent behavior pattern analysis. To ensure the integrity of the data in each time window, frame alignment is performed on each analysis window sequence, and necessary zero-padding is added to make them of the same length, improving the accuracy of temporal behavior modeling.Record the timestamp, cabinet ID, and operation area coordinate information for each image frame in the analysis window sequence. These information form structured metadata. The timestamp is used to record the shooting time of each frame of the image, providing a time reference for the sequential analysis of user operation behaviors. The cabinet ID is used to identify the position of the turnover cabinet corresponding to the current image, thereby achieving spatial differentiation of interaction behaviors in a multi-cabinet environment. The operation area coordinates are used to record the position changes of the user's hand operations and the equipment picking and placing areas, ensuring accurate tracking of the user's interaction trajectory. By combining the structured metadata with the analysis window sequence, a standardized image sequence is formed.
[0022] In a specific embodiment, the process of executing step S102 may specifically include the following steps: Input the standardized image sequence into a two-stage Transformer encoder respectively. The two-stage Transformer encoder includes an environmental context encoding branch and a user behavior encoding branch; Apply a 6-layer Vision Transformer network in the environmental context encoding branch to extract environmental features from the standardized image sequence, obtaining an environmental feature vector; Apply an 8-layer sequential Vision Transformer structure in the user behavior encoding branch to extract user behavior features from the standardized image sequence, obtaining user behavior spatial features; Input the user behavior spatial features into a sequential Transformer structure for sequential analysis, obtaining spatio-temporal behavior features; Separate the environmental feature vector and the spatio-temporal behavior features through a feature decoupling layer, obtaining the purified behavior features after separation, and perform feature comparison calculation on the purified behavior features after separation, outputting user behavior feature data.
[0023] Specifically, the standardized image sequence is input into a two-stage Transformer encoder, which consists of two parallel branches. One is the environmental context encoding branch that captures the static background and environmental information in the image, and the other is the user behavior encoding branch that analyzes the dynamic behavior characteristics during the user's interaction with the turnover cabinet. The two-branch structure can effectively distinguish the background information and user interaction information in the image, ensuring that the process of extracting behavior characteristics is not interfered by the environment, thereby improving the accuracy of behavior recognition. In the environmental context encoding branch, a 6-layer Vision Transformer network is used to extract environmental features from the standardized image sequence. Each layer of the Vision Transformer network includes a self-attention mechanism, a multi-head attention module, a feed-forward network, and a residual connection mechanism. Through the multi-head self-attention mechanism, the system captures the context correlation information in different regions of the image, and the attention weights can be dynamically adjusted to focus on the background features related to the environment, thereby extracting a high-dimensional environmental feature vector. During the process of feature extraction, the standardized image sequence is position-encoded so that the image sequence can retain spatial position information when passing through the Transformer encoder. Subsequently, through the feature stacking of multiple layers of Transformer blocks, the changing characteristics of environmental information in the spatial dimension are captured, and a fixed-length environmental feature vector is generated at the output layer. These environmental feature vectors reflect the static information such as the storage location distribution, lighting conditions, and equipment status inside the turnover cabinet. At the same time, the user behavior encoding branch adopts an 8-layer temporal Vision Transformer (ViT-T) structure to extract user behavior characteristics from the standardized image sequence. The ViT-T structure is a Transformer model based on the time dimension, which can perform temporal modeling on the user interaction behavior in the image sequence and extract the operation trajectory, action posture, and equipment picking and placing mode of the user at different time points. The user behavior encoding branch divides the standardized image sequence into image blocks of a fixed size and flattens these image blocks into a sequence form. By position-encoding, the temporal relationship of the image sequence is ensured to be maintained. Then the sequence is input into an 8-layer temporal Vision Transformer network. Each layer of the Transformer module includes a temporal self-attention mechanism, a feed-forward network, and a residual connection. Through the multi-head temporal self-attention mechanism, the system identifies the interaction characteristics of the user in different time windows and generates user behavior spatial features containing the time dimension through feature correlation across time steps. The user behavior spatial features are input into the temporal Transformer structure for temporal analysis. The temporal Transformer can model the behavior characteristics in the time series and capture the long-term dependence relationship of the user interaction behavior. Through multiple layers of temporal self-attention mechanisms, the change patterns of the user's behavior at different time scales are identified, and spatio-temporal behavior features are generated.The temporal Transformer structure can dynamically adjust the feature weights of different time windows and amplify the features of abnormal operation behaviors, ensuring that during long-term interactions, abnormal behavior signals can be accurately captured, and the sensitivity and accuracy of anomaly detection are improved. The environmental feature vector and spatio-temporal behavior features are separated through a feature decoupling layer. Based on the self-attention mechanism and the feature decoupling loss function, the environmental features and user behavior features are decoupled by means of feature decomposition and reconstruction, thereby eliminating the interference of environmental information on user behavior features and obtaining the purified behavior features after separation. The feature decoupling layer first calculates the feature distance between the environmental features and the spatio-temporal behavior features, and maps the feature vectors to different feature subspaces through a feature orthogonality mechanism to achieve complete decoupling of the features. The purified behavior features after feature decoupling more truly reflect the user's operation behavior pattern, eliminate the interference of environmental background noise, and make the behavior features more representative. After the feature decoupling is completed, the purified behavior features after separation are subjected to feature comparison calculation. Through algorithms such as feature cosine similarity or Euclidean distance measurement, the similarity between the current behavior features and the historical behavior features is compared, so as to judge whether there are abnormal situations in the user's behavior. Feature comparison calculation can quickly identify the pattern changes of the user's behavior, judge whether the current operation conforms to the normal interaction rules, and output the final user behavior feature data.
[0024] In a specific embodiment, the process of executing step S103 may specifically include the following steps: Input the user behavior feature data into the temporal analysis network to capture short-term, medium-term, and long-term temporal information respectively, and obtain multi-scale behavior representations; Segment the user behavior feature data by sliding windows to obtain F groups of behavior sequence segments, where F is 3; Input the F groups of behavior sequence segments into the corresponding temporal convolutional modules for feature mapping respectively to obtain multi-scale temporal mapping features; Perform cross-attention mechanism calculation on the multi-scale temporal mapping features to obtain temporal correlation features; Perform feature transformation on the temporal correlation features through an adaptive pooling layer and a three-layer fully connected network to obtain high-level behavior abstraction features; Apply a behavior semantic classifier to label the high-level behavior abstraction features and output the user operation behavior features.
[0025] Specifically, the user behavior feature data is input into the time series analysis network. Based on the multi-scale time series feature extraction framework, this network captures features and recognizes patterns of the user behavior feature data within short-term, medium-term, and long-term time windows. The short-term time series information is used to capture the instantaneous operation patterns when the user picks up and places the device, such as the quick pick-up and placement of the device or multiple interactions within a short period. The medium-term time series information is used to identify the interaction changes of the user over a period of time, including the operation frequency, stay duration, and regularity of device placement. The long-term time series information can capture the behavior trends during the user's long-term interaction, such as the operation habits formed by multiple pick-up and placement operations of the device or abnormal retention behaviors. The user behavior feature data is segmented by a sliding window, and the continuous behavior feature data is divided into F groups of behavior sequence segments, where F is 3, corresponding to short-term, medium-term, and long-term time series segments respectively. The sliding window mechanism can ensure the time continuity of the behavior data within each time window, while avoiding the loss of feature information caused by window division. The length and step size of the sliding window are dynamically adjusted according to different application scenarios. For example, the short-term window is set to 1.5 seconds, the medium-term window is 3 seconds, and the long-term window is set to 4.5 seconds, so as to ensure the accurate capture of behavior changes at different time scales. After completing the sliding window segmentation, the F groups of behavior sequence segments are respectively input into the corresponding time series convolutional modules for feature mapping. The time series convolutional modules adopt a multi-layer one-dimensional convolutional structure. Each time series convolutional module extracts features for the behavior sequence segments at different time scales, captures the time series features of short-term, medium-term, and long-term behavior segments through convolutional kernels of different sizes (such as 3, 5, 7), and performs feature mapping on the basis of convolutional operations to obtain multi-scale time series mapping features. The cross-attention mechanism is calculated for the multi-scale time series mapping features to capture the correlation between features at different time scales. The cross-attention mechanism calculates the correlation weights between features in different time segments, weights and fuses the features with high correlation degrees to generate time series correlation features. The cross-attention mechanism models the interaction of short-term, medium-term, and long-term features through the multi-head self-attention mechanism. By calculating the attention weight matrix between features at different time scales, it dynamically adjusts the weights of different feature segments, highlighting the key features closely related to the user behavior changes, thus effectively solving the problems of information redundancy or feature conflict that occur during the fusion process of features at different time scales. The introduction of the cross-attention mechanism enables the system to not only capture the feature dependency relationships within the same time window, but also identify the behavior association patterns across time scales, thereby improving the accuracy of behavior recognition.The time-series correlation features are transformed through an adaptive pooling layer and a three-layer fully connected network. The adaptive pooling layer dynamically adjusts the size of the pooling window according to the length of the time-series features, thereby mapping the time-series correlation features to a fixed length to adapt to the behavior data with different time window lengths. The pooled features are transformed through a three-layer fully connected network. The fully connected network adopts a layer-by-layer feature transformation mechanism, enhances the abstract ability of the features through non-linear mapping, and enhances the non-linear expression ability of the features through activation functions. This process maps the time-series correlation features to high-level behavior abstract features with a fixed dimension and strengthens the time-series dependence relationship between features at different time scales. A behavior semantic classifier is applied to label the high-level behavior abstract features. The behavior semantic classifier classifies the features based on a multi-layer perceptron or an XGBoost model and combines historical behavior pattern data to refine the classification of user operation behaviors. The behavior semantic classifier is pre-trained with a large amount of interaction behavior data and has the ability to accurately identify different user behavior patterns. The classification results can identify normal device picking and placing behaviors and capture abnormal behaviors such as unauthorized picking and placing, long-term staying, and device misalignment, thereby outputting user operation behavior features.
[0026] In a specific embodiment, the process of inputting the F groups of behavior sequence segments into the corresponding time-series convolution modules for feature mapping to obtain multi-scale time-series mapping features may specifically include the following steps: Perform zero-padding and normalization processing on the F groups of behavior sequence segments respectively to obtain standardized behavior sequences; Input the standardized behavior sequences into the first one-dimensional convolutional layer of the corresponding time-series convolution module for feature processing to obtain the first-layer feature maps; Input the first-layer feature maps into the second one-dimensional convolutional layer of the corresponding time-series convolution module for feature processing to obtain the second-layer feature maps; Input the second-layer feature maps into the third one-dimensional convolutional layer of the corresponding time-series convolution module for feature processing to obtain the third-layer feature maps; Input the third-layer feature maps into the fourth one-dimensional convolutional layer of the corresponding time-series convolution module for feature processing to obtain output features at F time scales; Align and pad the output features at F time scales in the time dimension, process the time-series differences between sequences of different lengths through a time mask mechanism, and splice them in the channel dimension to obtain multi-scale time-series mapping features.
[0027] Specifically, length alignment and zero-padding processing are performed on the behavior sequence segments in group F. By calculating the maximum length of each group of behavior segments, the behavior segments with insufficient length are padded by adding zero values at the end of the sequence, so that the lengths of all behavior segments reach the same level. At the same time, in order to eliminate the amplitude differences of different feature dimensions in the behavior sequence segments, the zero-padded sequence segments are normalized. The normalization adopts the mean-variance normalization method, which maps the data of each feature dimension to a standardized distribution with a mean of 0 and a variance of 1, thereby ensuring that the behavior feature data has the same feature scale under different time scales, which is conducive to the convolutional network's learning and modeling of features at different scales. The standardized behavior sequences are respectively input into the first one-dimensional convolutional layer of the corresponding temporal convolutional module for feature processing. The first one-dimensional convolutional layer of the temporal convolutional module uses multiple convolutional kernels of different sizes (such as 3, 5, 7) to capture the local features of short-term, medium-term, and long-term behavior segments. Through the sliding convolutional window, step-by-step feature extraction is performed on the behavior sequence segments, and the ReLU activation function is combined to introduce non-linear feature transformation, generating the first-layer feature map, which reflects the change patterns of the user's interaction behavior under different time scales and has high temporal resolution ability. The first-layer feature map is input into the second one-dimensional convolutional layer of the corresponding temporal convolutional module for feature processing. This layer performs deeper feature extraction in the time dimension and combines batch normalization to adjust the scale of the features to prevent the convolutional network from experiencing gradient vanishing or explosion during training. The second-layer feature map retains the dynamic change information of the temporal features in a longer time dimension and enhances the temporal dependence relationship of the features. The second-layer feature map is input into the third one-dimensional convolutional layer of the corresponding temporal convolutional module for feature processing. The third convolutional layer captures the temporal features of the user's behavior within a longer time window through a convolutional kernel with a larger receptive field. The convolutional operation of this layer can identify potential changes in behavior patterns during long-term interactions and map the local behavior features to a higher-dimensional feature space, generating the third-layer feature map. The third-layer feature map is input into the fourth one-dimensional convolutional layer of the corresponding temporal convolutional module for feature processing. This convolutional layer further performs deep feature extraction on the feature map, and by capturing the feature correlations across time steps, it forms output features at F time scales. The output features at F time scales are aligned and padded in the time dimension, and a time masking mechanism is introduced to handle the temporal differences between sequences of different lengths. The time masking mechanism calculates the length differences between features at different time scales and adds masking marks at the end of the shorter feature sequences. These masking marks will be automatically ignored during subsequent feature processing, thus ensuring that no invalid data is introduced when features of different lengths are concatenated. This mechanism can effectively solve the problem of inconsistent lengths between features at different time scales, enabling multi-scale features to maintain consistent temporal characteristics during alignment.After completing the alignment and padding in the time dimension, the output features of F time scales are concatenated in the channel dimension. By stacking the feature maps of different time scales along the feature dimension, a unified feature representation containing multi-scale temporal features is formed.
[0028] In a specific embodiment, the process of executing step S104 may specifically include the following steps: Perform spatio-temporal dimensional analysis and classification on the user operation behavior characteristics, and construct a turnover cabinet usage pattern map including a spatio-temporal behavior layer, a user feature layer, and a cabinet association layer; Input the turnover cabinet usage pattern map into a three-layer GraphSAGE graph neural network for representation learning to obtain a graph-structured representation, and perform spatial dependence relationship analysis between cabinets on the graph-structured representation to obtain spatial association features; Perform temporal modeling on the spatial association features and the spatio-temporal behavior features to obtain spatio-temporal sequence prediction data, and apply a temporal convolutional network to predict the usage frequency of the spatio-temporal sequence prediction data, and output the cabinet usage prediction result; Based on the turnover cabinet usage pattern map and the cabinet usage prediction result, calculate the deviation metric value between the current usage pattern and the historical normal pattern, generate an abnormal usage pattern detection result, and the abnormal pattern detection result is used to trigger the storage location indicator to give safety prompts in different colors and blinking modes, and report to the management system in real time through the RESTful interface.
[0029] Specifically, the dynamic analysis of user operation behavior characteristics in the time dimension is carried out. By analyzing the behavior of picking up and placing devices, interaction frequency, and operation mode changes of users in different time periods, time series behavior characteristics are formed. Combining with the location information in the space dimension, the interaction operations of users in different storage locations are spatially encoded to form a spatial feature matrix. These time and space characteristics, combined with user identity, storage location status, and asset change information, construct a three-layer graph structure including a spatio-temporal behavior layer, a user feature layer, and a cabinet association layer. The spatio-temporal behavior layer records the dynamic changes of user operation behavior in the time and space dimensions, including key characteristics such as the time interval, interaction sequence, and residence duration of users picking up and placing devices in different storage locations. The user feature layer extracts static characteristics such as user identity, operation proficiency, and behavior habits to depict the user's interaction portrait. The cabinet association layer generates a spatial association matrix by analyzing the usage association relationships between different storage locations, thereby constructing a graph of the turnover cabinet usage pattern. The graph of the turnover cabinet usage pattern is input into a three-layer GraphSAGE graph neural network for representation learning. GraphSAGE is a graph neural network model that can perform node feature aggregation and spatial relationship modeling on heterogeneous graph structures, and capture the feature association information between nodes in different layers through the graph convolution mechanism. The first layer of the GraphSAGE network aggregates the features of adjacent nodes through a feature aggregation function (such as mean, LSTM, or pooling), weights and fuses the features from the spatio-temporal behavior layer, the user feature layer, and the cabinet association layer to generate an initial node representation vector. Then, the second layer captures the cross-layer feature dependence relationship between nodes in different layers through a cross-layer connection mechanism, dynamically adjusts the importance of different node features through an attention mechanism, and strengthens the feature representation of key nodes. The third layer updates the node representation based on feature aggregation to generate a graph-structured representation. The graph-structured representation contains the time series change information of user operation behavior characteristics and captures the spatial association relationship between different storage locations. Analyze the spatial dependence relationship between cabinets for the graph-structured representation. By calculating the feature similarity and operation frequency correlation between different storage locations, a spatial association feature matrix is formed. The spatial association feature matrix captures the operation dependence pattern between adjacent storage locations by calculating features such as the interaction frequency of users in different storage locations, the time interval of picking up and placing devices, and the change of storage location status. This process can identify frequently used storage location combinations and discover potential abnormal association relationships, providing information support in the space dimension for predicting the cabinet usage pattern. Perform time series modeling on the spatial association features and spatio-temporal behavior features to obtain spatio-temporal sequence prediction data. Time series modeling uses gated recurrent units or long short-term memory networks to model the spatio-temporal behavior features in the time dimension. By capturing the change trend of user behavior in different time windows, spatio-temporal sequence features of user interaction behavior are generated. Combining with the spatial association feature matrix, spatial information and time information are fused to generate spatio-temporal sequence prediction data.Apply a temporal convolutional network to predict the usage frequency of spatio-temporal sequence prediction data. The temporal convolutional network extracts features from spatio-temporal sequence data through a one-dimensional convolutional kernel, and combines dilated convolutions to capture behavior changes at different time scales, generating a high-dimensional usage frequency prediction result. Based on the turnover cabinet usage pattern map and the cabinet usage prediction result, calculate the deviation metric value between the current usage pattern and the historical normal pattern. The deviation metric value is calculated by comparing the similarity between the current user interaction behavior characteristics and the historical operation pattern, and uses metric methods such as cosine similarity, dynamic time warping, or Euclidean distance to compare the current usage characteristics with the historical normal pattern. When the deviation metric value exceeds the set threshold, an abnormal usage pattern detection result is generated. The abnormal usage pattern detection result includes various abnormal behaviors such as unauthorized access, long stays, misplacement of equipment, and abnormal changes in storage locations. When an abnormal pattern is detected, the system triggers the storage location indicator to give safety prompts in different colors and flashing modes, where the green light indicates the normal state, the yellow light flashes at a frequency of 2 Hz to indicate a minor abnormality, the orange light flashes at a frequency of 4 Hz to warn of an obvious abnormality, and the red light is constantly on or flashing to indicate a serious abnormality. At the same time, the abnormal pattern detection result is reported to the management system in real time through a RESTful interface, and the abnormal information is pushed to the upper management system through a POST request to ensure that the management personnel receive the abnormal alarm in time and take corresponding measures, thereby realizing real-time safety monitoring during the use of the turnover cabinet.
[0030] In a specific embodiment, the process of performing steps to analyze and classify the user operation behavior characteristics in spatio-temporal dimensions and construct a turnover cabinet usage pattern map including a spatio-temporal behavior layer, a user feature layer, and a cabinet association layer may specifically include the following steps: Integrate the timestamps and map the spatial coordinates of the user operation behavior characteristics to obtain structured spatio-temporal behavior data; Perform recursive clustering on the structured spatio-temporal behavior data to obtain a set of spatio-temporal behavior prototypes, and create a spatio-temporal behavior layer based on the set of spatio-temporal behavior prototypes; Analyze the user interaction frequency, operation proficiency, and habitual patterns based on the set of spatio-temporal behavior prototypes to obtain user behavior portrait data, and create a user feature layer based on the user behavior portrait data; Based on the user behavior portrait data and the cabinet geographical location information, calculate the Pearson correlation coefficient matrix to analyze the usage correlation between multiple cabinets in the same area, obtain the cabinet association matrix, and create a cabinet association layer based on the cabinet association matrix; Perform an inter-layer relationship mapping on the spatio-temporal behavior layer, the user feature layer, and the cabinet association layer, and at the same time define the inter-layer connection weights and the node-to-node association functions to construct a three-layer heterogeneous graph structure; Calculate the node similarity weights and path importance scores according to the three-layer heterogeneous graph structure, and fuse the information of the three layers based on the node similarity weights and path importance scores to obtain the turnover cabinet usage pattern map.
[0031] Specifically, timestamp integration is performed on the user operation behavior characteristics. By marking the time for each frame of image, the precise time information during the user's operations such as picking and placing equipment, opening and closing cabinet doors, and changing storage locations is recorded. At the same time, spatial coordinate mapping is performed on the user operation behavior characteristics. The interaction positions of the user picking and placing equipment, the hand trajectories, and the storage location numbers are spatially encoded to form a spatial coordinate matrix, reflecting the interaction behaviors of the user at different positions in the turnover cabinet. By combining timestamp integration and spatial coordinate mapping, structured spatio-temporal behavior data is formed. Recursive clustering is performed on the structured spatio-temporal behavior data. Through the methods of recursive splitting and aggregation, similar user interaction behavior sequences are clustered to form a set of spatio-temporal behavior prototypes. The recursive clustering process uses a distance metric algorithm based on dynamic time warping or cosine similarity. By calculating the similarity between different behavior sequences, the behavior data is adaptively aggregated. The system dynamically adjusts the number of behavior categories during the clustering process and selects the optimal clustering division scheme according to the maximum likelihood estimation to generate a set of spatio-temporal behavior prototypes representing different user interaction patterns. Each spatio-temporal behavior prototype represents a typical user interaction pattern, such as normal equipment picking and placing, long-term stay, unauthorized access, etc. Based on the set of spatio-temporal behavior prototypes, a spatio-temporal behavior layer is created. The spatio-temporal behavior layer represents the operation patterns of the user at different time points and storage locations in a graph-structured manner by node-based modeling of different behavior prototypes, and provides spatio-temporal feature support for subsequent graph neural network modeling. Based on the set of spatio-temporal behavior prototypes, the interaction frequency, operation proficiency, and behavior habits of the user are analyzed to form user behavior portrait data. By analyzing the interaction records between the user and the turnover cabinet, the operation frequency of the user at different storage locations is calculated, and the operation proficiency of the user is evaluated, such as the accuracy of equipment picking and placing operations, the stability of operation time, and the sensitivity to storage location changes. At the same time, the operation habits of the user are captured, including behavior characteristics such as preferred storage locations, picking and placing orders, and high-frequency usage patterns. These information form the user behavior portrait data. Based on the user behavior portrait data, a user feature layer is created. The user feature layer represents the user identity, behavior pattern, and operation characteristics by nodes and is associated with the spatio-temporal behavior layer to form a user interaction feature network. At the same time, based on the user behavior portrait data and the geographical location information of the cabinet, the usage correlation between multiple cabinets in the same area is calculated. By calculating the Pearson correlation coefficient matrix, the spatial correlation between different storage locations is analyzed. The Pearson correlation coefficient matrix identifies storage location combinations with highly similar usage patterns by comparing the usage frequencies, user interaction times, and equipment picking and placing time intervals in the time dimension, and generates a cabinet association matrix.The cabinet association matrix can capture the storage location interaction association information in the spatial dimension, providing key feature support for the system to identify the usage dependency relationship between cabinets. Based on the cabinet association matrix, a cabinet association layer is created. The cabinet association layer represents the association relationship between storage locations through nodes, models the usage patterns and spatial correlations between different storage locations, and forms a storage location association graph structure. Perform inter-layer relationship mapping on the spatio-temporal behavior layer, user feature layer, and cabinet association layer, define the inter-layer connection weights and node-to-node association functions. Through the cross-layer connection mechanism, dynamically aggregate the node features of the spatio-temporal behavior layer, user feature layer, and cabinet association layer, and calculate the feature similarity between nodes in different layers through the association function, forming a three-layer heterogeneous graph structure. The inter-layer connection weights reflect the association strength between nodes in different layers, while the node-to-node association function is used to calculate the influence of different node features in the process of cross-layer information propagation, thus ensuring the full integration of multi-layer information in the graph neural network modeling process. To achieve the fusion of multi-layer information and the construction of the final turnover cabinet usage pattern map, calculate the node similarity weights and path importance scores according to the three-layer heterogeneous graph structure, calculate the similarity between different nodes through the combination of the graph convolutional network and the attention mechanism, and perform information propagation through the importance scores of the weighted paths. The path importance scores perform priority weight allocation for key paths by calculating the weights of cross-layer paths and the feature aggregation effect, thereby enhancing the feature propagation ability of key nodes. The node similarity weights cluster nodes with similar behavioral characteristics by measuring the feature similarity between different nodes, and enhance the association between similar nodes during the information propagation process. This mechanism ensures that the system can capture the information of key nodes and important paths during the multi-layer information fusion process, generating a turnover cabinet usage pattern map. Based on the three-layer heterogeneous graph structure, fuse the spatio-temporal behavior characteristics, user operation patterns, and storage location association relationships to form a turnover cabinet usage pattern map with spatio-temporal analysis ability, user portrait modeling ability, and storage location association modeling ability. This map accurately reflects the interaction pattern between users and turnover cabinets and identifies potential abnormal usage behaviors.
[0032] The intelligent turnover cabinet security management method based on image recognition technology in the embodiments of the present application has been described above. Next, the intelligent turnover cabinet security management system based on image recognition technology in the embodiments of the present application will be described. Please refer to Figure 2 , an embodiment of the intelligent turnover cabinet security management system based on image recognition technology in the embodiments of the present application includes: An acquisition module 201, configured to acquire a user interaction image sequence through an intelligent turnover cabinet, and perform illumination normalization and target area cropping processing on the user interaction image sequence to obtain a normalized image sequence; A processing module 202, configured to input the normalized image sequence into a two-stage Transformer encoder for separating user behavior and environmental factors to obtain user behavior feature data; The feature extraction module 203 is configured to input user behavior feature data into a time series analysis network for multi-scale time feature extraction to obtain user operation behavior features; The generation module 204 is configured to construct a turnover cabinet usage pattern map based on the user operation behavior features, and generate a cabinet usage prediction result and an abnormal pattern detection result based on the turnover cabinet usage pattern map. The abnormal pattern detection result is used to trigger the storage location indicator to provide safety prompts in different colors and flashing modes, and report to the management system in real time through the RESTful interface.
[0033] Through the collaborative cooperation of the above-mentioned components, the effective separation of user behavior and environmental factors is achieved through the dual-level Transformer encoder technology, solving the problem of low accuracy of behavior recognition in complex environments. The light normalization processing and target area cropping can reduce the influence of environmental light changes and interference from irrelevant areas, making the behavior feature extraction more accurate. The multi-scale time feature extraction mechanism can capture short-term, medium-term, and long-term user operation behavior features simultaneously, and calculate the time series correlation features through the cross-attention mechanism, significantly improving the system's understanding ability of complex operation sequences. The three-layer heterogeneous graph structure (spatio-temporal behavior layer, user feature layer, and cabinet association layer) based on the turnover cabinet usage pattern map can comprehensively represent the user interaction pattern. Combining with the GraphSAGE graph neural network for representation learning enables the system to accurately identify abnormal behaviors deviating from the normal pattern. The system can provide intuitive safety prompts through different colors and flashing modes of the storage location indicator according to the abnormal pattern detection result, and at the same time report to the management system in real time through the RESTful interface, realizing a multi-level safety warning mechanism. Through the spatio-temporal sequence prediction of the cabinet usage pattern, the system can predict the cabinet usage frequency and possible abnormal situations in advance, providing data support for management decisions, and realizing the transformation of the management mode from passive response to active prevention. Based on the Pearson correlation coefficient matrix analysis of the usage correlation between multiple cabinets in the same area, the system can discover the spatial correlation features of cabinet usage, and achieve seamless integration with the upper-level management system through the standardized RESTful interface, supporting real-time data reporting and information sharing, and enhancing the collaborative working ability of the intelligent turnover cabinet and other enterprise information systems.
[0034] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described system, system, and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0035] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable an intelligent turnover cabinet safety management device based on image recognition technology (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0036] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of the present invention.
Claims
1. A method for the safety management of intelligent turnover cabinets based on image recognition technology, characterized in that: include: The user interaction image sequence is collected through the intelligent turnover cabinet, and the user interaction image sequence is subjected to illumination standardization and target area cropping to obtain a standardized image sequence; Inputting the standardized image sequence into a two-stage Transformer encoder to separate user behavior from environmental factors to obtain user behavior feature data; Inputting the user behavior feature data into a time series analysis network to extract multi-scale time features to obtain user operation behavior features; A turnover cabinet usage pattern map is constructed according to the user operation behavior characteristics, and cabinet usage prediction results and abnormal pattern detection results are generated based on the turnover cabinet usage pattern map. The abnormal pattern detection results are used to trigger the storage location indicator lights to provide safety prompts with different colors and flashing modes, and report to the management system in real time through the RESTful interface.
2. The intelligent turnover cabinet safety management method based on image recognition technology according to claim 1 is characterized in that: The user interaction image sequence is collected by the intelligent turnover cabinet, and the user interaction image sequence is subjected to illumination standardization and target area cropping to obtain a standardized image sequence, including: A wide-angle camera is installed inside the intelligent turnover cabinet, and a viewing angle camera is installed outside the turnover cabinet, and internal and external dual-viewing angle image sequences are collected by the wide-angle camera and the viewing angle camera; Remapping the grayscale values of the internal and external dual-view image sequence to obtain a lighting standardized image, and identifying the user operation area of the lighting standardized image to obtain a regional standardized image; Performing time-series frame difference processing on the regional standardized image to obtain a dynamic change map, and performing time window segmentation on the dynamic change map to obtain an analysis window sequence; A timestamp, a cabinet ID and operation area coordinates are recorded for each image frame in the analysis window sequence to form structured metadata, and the structured metadata is combined with the analysis window sequence to obtain a standardized image sequence.
3. The intelligent turnover cabinet safety management method based on image recognition technology according to claim 1 is characterized in that: The step of inputting the standardized image sequence into a two-stage Transformer encoder to separate user behavior from environmental factors to obtain user behavior feature data includes: Inputting the standardized image sequences into a two-stage Transformer encoder respectively, wherein the two-stage Transformer encoder includes an environment context encoding branch and a user behavior encoding branch; Applying the 6-layer Vision Transformer network in the environmental context encoding branch to extract environmental features from the standardized image sequence, thereby obtaining an environmental feature vector; Applying the 8-layer temporal Vision Transformer structure in the user behavior encoding branch to extract user behavior features from the standardized image sequence, thereby obtaining user behavior spatial features; Input the user behavior spatial features into the time series Transformer structure for time series analysis to obtain spatiotemporal behavior features; The environmental feature vector and the spatiotemporal behavior feature are separated through a feature decoupling layer to obtain separated purified behavior features, and feature comparison calculation is performed on the separated purified behavior features to output user behavior feature data.
4. The intelligent turnover cabinet safety management method based on image recognition technology according to claim 1 is characterized in that: The step of inputting the user behavior feature data into a time series analysis network to extract multi-scale time features to obtain user operation behavior features includes: Inputting the user behavior feature data into a time series analysis network to capture short-term, medium-term and long-term time series information respectively to obtain a multi-scale behavior representation; Segment the user behavior feature data using a sliding window to obtain F groups of behavior sequence segments, where F is 3; Inputting the F groups of behavior sequence fragments into corresponding temporal convolution modules for feature mapping respectively to obtain multi-scale temporal mapping features; Performing cross-attention mechanism calculation on the multi-scale temporal mapping features to obtain temporal correlation features; The temporal correlation features are transformed through an adaptive pooling layer and a three-layer fully connected network to obtain high-level behavioral abstract features; The application behavior semantic classifier annotates the high-level behavior abstract features and outputs user operation behavior features.
5. The intelligent turnover cabinet safety management method based on image recognition technology according to claim 4 is characterized in that: The F groups of behavior sequence fragments are respectively input into corresponding time series convolution modules for feature mapping to obtain multi-scale time series mapping features, including: Zero-filling and normalizing the F groups of behavior sequence fragments to obtain standardized behavior sequences; Input the standardized behavior sequences into the first one-dimensional convolution layer of the corresponding temporal convolution module for feature processing to obtain a first-layer feature map; Inputting the first layer feature map into the second layer one-dimensional convolution layer of the corresponding temporal convolution module for feature processing to obtain a second layer feature map; Inputting the second layer feature map into the third layer one-dimensional convolution layer of the corresponding temporal convolution module for feature processing to obtain a third layer feature map; Inputting the third layer feature map into the fourth layer one-dimensional convolution layer of the corresponding temporal convolution module for feature processing to obtain output features of F time scales; The output features of F time scales are aligned and padded in the time dimension, the timing differences between sequences of different lengths are processed through the time mask mechanism, and they are spliced in the channel dimension to obtain multi-scale time series mapping features.
6. The intelligent turnover cabinet safety management method based on image recognition technology according to claim 1 is characterized in that: The method of constructing a turnover cabinet usage pattern map according to the user operation behavior characteristics, and generating a cabinet usage prediction result and an abnormal mode detection result based on the turnover cabinet usage pattern map, wherein the abnormal mode detection result is used to trigger the storage position indicator light to provide safety prompts with different colors and flashing modes, and report to the management system in real time through the RESTful interface, including: Performing spatiotemporal analysis and classification on the user operation behavior characteristics, and constructing a turnover cabinet usage pattern map including a spatiotemporal behavior layer, a user feature layer, and a cabinet association layer; Input the turnover cabinet usage pattern map into the three-layer GraphSAGE graph neural network for representation learning to obtain a graph structured representation, and perform spatial dependency relationship analysis between cabinets on the graph structured representation to obtain spatial correlation features; The spatial correlation features and the spatiotemporal behavior features are subjected to time series modeling to obtain spatiotemporal series prediction data, and a time convolutional network is applied to predict the usage frequency of the spatiotemporal series prediction data, and a cabinet usage prediction result is output; Based on the turnover cabinet usage pattern map and the cabinet usage prediction results, the deviation measurement value between the current usage pattern and the historical normal pattern is calculated, and an abnormal usage pattern detection result is generated. The abnormal pattern detection result is used to trigger the storage location indicator light to provide safety prompts with different colors and flashing patterns, and report to the management system in real time through the RESTful interface.
7. The intelligent turnover cabinet safety management method based on image recognition technology according to claim 6 is characterized in that: The spatiotemporal dimension analysis and classification of the user operation behavior characteristics are performed to construct a turnover cabinet usage pattern map including a spatiotemporal behavior layer, a user feature layer, and a cabinet association layer, including: Performing timestamp integration and spatial coordinate mapping on the user operation behavior features to obtain structured spatiotemporal behavior data; Recursively clustering the structured spatiotemporal behavior data to obtain a spatiotemporal behavior prototype set, and creating a spatiotemporal behavior layer according to the spatiotemporal behavior prototype set; Analyze the user interaction frequency, operation proficiency and habit pattern based on the spatiotemporal behavior prototype set to obtain user behavior portrait data, and create a user feature layer based on the user behavior portrait data; Based on the user behavior portrait data and the cabinet geographical location information, a Pearson correlation coefficient matrix is calculated to analyze the usage correlation between multiple cabinets in the same area, a cabinet association matrix is obtained, and a cabinet association layer is created according to the cabinet association matrix; Performing inter-layer relationship mapping on the spatiotemporal behavior layer, the user feature layer, and the cabinet association layer, while defining inter-layer connection weights and inter-node association functions to construct a three-layer heterogeneous graph structure; The node similarity weights and path importance scores are calculated according to the three-layer heterogeneous graph structure, and the three-layer information is fused based on the node similarity weights and the path importance scores to obtain a turnover cabinet usage pattern map.
8. An intelligent turnover cabinet safety management system based on image recognition technology, characterized in that: Used to implement the intelligent turnover cabinet safety management method based on image recognition technology as described in any one of claims 1 to 7, the intelligent turnover cabinet safety management system based on image recognition technology includes: A collection module, used to collect a user interaction image sequence through an intelligent turnover cabinet, and perform illumination standardization and target area cropping processing on the user interaction image sequence to obtain a standardized image sequence; A processing module, used for inputting the standardized image sequence into a two-stage Transformer encoder to separate user behavior from environmental factors, and obtaining user behavior feature data; A feature extraction module, used for inputting the user behavior feature data into a time series analysis network to perform multi-scale time feature extraction to obtain user operation behavior features; A generation module is used to construct a turnover cabinet usage pattern map according to the user operation behavior characteristics, and generate cabinet usage prediction results and abnormal pattern detection results based on the turnover cabinet usage pattern map. The abnormal pattern detection results are used to trigger the storage location indicator lights to provide safety prompts with different colors and flashing modes, and report to the management system in real time through the RESTful interface.
9. An intelligent turnover cabinet safety management device based on image recognition technology, characterized in that: It includes a memory and a processor, the memory stores a computer program that can be run on the processor, and the processor implements the intelligent turnover cabinet safety management method based on image recognition technology as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the processor is enabled to execute the intelligent turnover cabinet safety management method based on image recognition technology as described in any one of claims 1 to 7.
Citation Information
Cited By
Image recognition method and device, model training method and device, equipment and medium
CN120375098A
Trajectory abnormal route detection method and system based on self-supervised trajectory representation learning
CN120724351A