Behavior recognition method and device, electronic equipment, storage medium and product
By combining instance segmentation models and video behavior recognition models, the problem of low accuracy in hand identification in existing technologies is solved, and more efficient detection of unauthorized proxy operations is achieved.
Patent Information
- Application Number
- CN202511060184.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-11
AI Technical Summary
In existing technologies, the method of determining hand ownership by combining the personnel information detected by the target detection model with the detected hand position information has low accuracy, resulting in low accuracy in detecting valet behavior.
An instance segmentation model is used to identify hand features in video images, GhostNetV2 lightweight image classification network is used to determine identity information, DeepSort target tracking algorithm is used to obtain video images of employees' hands, and finally SlowFast video behavior recognition model is used to determine whether there is any valet operation behavior.
This improves the accuracy of hand identification, which in turn improves the accuracy of behavior recognition, enabling more effective detection of unauthorized proxy operations.
Smart Images

Figure CN120932301A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of computer vision technology, and in particular to a behavior recognition method, device, electronic device, storage medium and product. Background Technology
[0002] With the rapid development of the financial industry, banking operations have become increasingly complex and diverse. Ensuring the safety of customer funds and operational compliance are crucial in the daily operations of banks. However, unauthorized transactions on behalf of customers occur frequently, which not only seriously damages the legitimate rights and interests of customers but also brings significant reputational risks and economic losses to banks.
[0003] Bank transactions involve large amounts of capital flows and sensitive customer information. Unauthorized actions by employees on behalf of customers, such as conducting account transactions, entering passwords, or signing documents, can lead to the misappropriation of customer funds and asset losses. This also undermines the fair order of the financial market and the foundation of trust between banks and customers. However, banks currently lack efficient and accurate detection methods to promptly identify and prevent such unauthorized actions in their actual operations.
[0004] Existing technology proposes a method for recognizing valet service behavior, which includes identifying employees, non-employees, and hands through object detection algorithms; determining whether a hand belongs to an employee based on a first distance and a second distance formed between the hand and the employee or non-employee; and finally, detecting the movement and changes of the employee's hand to determine whether the employee is performing valet service. However, this scheme has low accuracy in determining hand ownership by combining the personnel information detected by the object detection model with the detected hand position information, resulting in low accuracy in valet service behavior detection. Summary of the Invention
[0005] This invention provides a behavior recognition method, device, electronic device, storage medium, and product to solve the problem that the accuracy of determining hand ownership by using target detection models to identify personnel information and the detected hand position information is low, resulting in low accuracy of valet behavior detection.
[0006] According to one aspect of the present invention, a behavior recognition method is provided, comprising:
[0007] Acquire video images, including images of the operator's appearance and hands;
[0008] The hand mask is obtained by identifying hand features in the video image using an instance segmentation model;
[0009] The hand mask and the video image are input into an image classification network, and the identity information of the smart teller machine operator is determined through the image classification network;
[0010] If the operator of the smart teller machine is an employee, then the hand features are tracked to obtain video images of the employee's hands.
[0011] The video image of the employee's hand is input into the video behavior recognition model to determine whether there is any surrogate operation behavior.
[0012] According to another aspect of the present invention, a behavior recognition device is provided, comprising:
[0013] The first acquisition module is used to acquire video images, including images of the operator's appearance and hands.
[0014] The first recognition module is used to identify hand features in the video image using an instance segmentation model to obtain a hand mask;
[0015] The determination module is used to input the hand mask and the video image into an image classification network, and determine the identity information of the smart teller machine operator through the image classification network;
[0016] The second acquisition module is used to perform target tracking on the hand features to obtain video images of the employee's hands if the identity information of the operator of the smart teller machine is that of an employee.
[0017] The second recognition module is used to input the video image of the employee's hand into the video behavior recognition model to determine whether there is any valet operation behavior.
[0018] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0019] At least one processor;
[0020] and a memory communicatively connected to the at least one processor;
[0021] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the behavior recognition method according to any embodiment of the present invention.
[0022] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the behavior recognition method described in any embodiment of the present invention.
[0023] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the behavior recognition method described in any embodiment of the present invention.
[0024] The technical solution of this invention identifies hand features in video images through an instance segmentation model, solving the problem of low accuracy in hand identification leading to low behavior recognition accuracy in existing technologies. This achieves the beneficial effect of improving the accuracy of hand identification, thereby improving the accuracy of behavior recognition.
[0025] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a flowchart illustrating a behavior recognition method provided in Embodiment 1 of the present invention;
[0028] Figure 2 This is a GhostNetV2 network structure diagram provided in Embodiment 1 of the present invention;
[0029] Figure 3 This is a structural diagram of the SlowFast model provided in Embodiment 1 of the present invention;
[0030] Figure 4 This is a structural diagram of an optimized YOLOv8-seg instance segmentation model provided in Embodiment 2 of the present invention;
[0031] Figure 5 This is a structural diagram of the cross-layer feature fusion module provided in Embodiment 2 of the present invention;
[0032] Figure 6 A flowchart of a behavior recognition method provided in a specific embodiment of the present invention;
[0033] Figure 7 This is a schematic diagram of the structure of a behavior recognition device provided in Embodiment 3 of the present invention;
[0034] Figure 8 This is a schematic diagram of the structure of an electronic device for a behavior recognition method according to an embodiment of the present invention. Detailed Implementation
[0035] To enable those skilled in the art to better understand the present invention, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention. It should be understood that the various steps described in the method embodiments of the present invention can be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0036] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0037] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0038] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0039] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0040] Example 1
[0041] Figure 1This is a flowchart illustrating a behavior recognition method provided in Embodiment 1 of the present invention. This method is applicable to identifying situations involving unauthorized valet operations. The method can be executed by a behavior recognition device, which can be implemented by software and / or hardware and is generally integrated into an electronic device. In this embodiment, the electronic device includes, but is not limited to, a computer device.
[0042] like Figure 1 As shown, an embodiment of the present invention provides a behavior recognition method, which includes the following steps:
[0043] S110. Acquire video images, including images of the operator's appearance and hands.
[0044] The video images can be videos of people operating bank smart teller machines, and the videos can include the person's appearance and the person's hands.
[0045] In this embodiment, video images can be obtained from the monitoring camera of the bank's smart teller machine.
[0046] S120. Use an instance segmentation model to identify hand features in the video image to obtain a hand mask.
[0047] Instance segmentation refers to dividing image pixels into different object categories with specific semantic meanings, further separating individual instances belonging to the same category, classifying all pixels in the image, and distinguishing different instances of the same category based on semantic segmentation.
[0048] The instance segmentation model takes video images as input and outputs information such as target category, target mask, and confidence level. It determines whether there is a hand image in the video image based on the output data.
[0049] Preferably, the instance segmentation model can be an optimized YOLOv8-seg model, a model obtained by optimizing the feature splicing layer of the Neck layer network in YOLOv8-seg, or a model obtained by optimizing the feature splicing layer and upsampling operator of the Neck layer network in YOLOv8-seg.
[0050] S130. Input the hand mask and the video image into an image classification network, and determine the identity information of the smart teller machine operator through the image classification network.
[0051] In this process, by inputting the hand mask and video image into the video classification network, the identity of the hand can be determined. If the confidence level of the hand belongs to an employee is greater than a threshold, it is determined that there is an employee's hand image in the video image.
[0052] In this embodiment, the GhostNetV2 lightweight image classification network is used, and its network results are as follows: Figure 2 As shown, Figure 2 This is a GhostNetV2 network structure diagram provided in Embodiment 1 of the present invention. The GhostNetV2 network can extract hand identity features using cropped hand region image information, efficiently determining the identity of the hand in a video image. If the hand belongs to an employee, then the identity information of the smart teller machine operator is determined to be that of an employee.
[0053] S140. If the identity information of the operator of the smart teller machine is that of an employee, then target tracking is performed on the hand features to obtain video images of the employee's hands.
[0054] In this embodiment, after confirming the presence of an employee's hand image in the video image, a target tracking algorithm can be invoked to track the employee's hand, obtaining a video image of the employee's hand. The target tracking algorithm can be DeepSort. DeepSOrt is an algorithm for multi-target tracking; its full name is Deep Learning based SORT (Simple Online and Realtime Tracking). DeepSOrt combines deep learning with the traditional SORT algorithm, enabling real-time tracking of multiple targets in a video.
[0055] S150. Input the video image of the employee's hand into the video behavior recognition model to determine whether there is any valet operation behavior.
[0056] The process involves enhancing the video images of employees' hands, scaling the sub-images, and storing them in a buffer. Once the number of employee hand video images cached in the buffer reaches a threshold, all cached employee hand video images are input into a video behavior recognition model to identify whether there is any behavior of employees acting on behalf of customers in the employee hand video images.
[0057] Preferably, the video behavior recognition model can be a SlowFast network model. The SlowFast network model achieves efficient extraction of video information through a two-stream structure. The structure of the SlowFast network model is as follows: Figure 3 As shown, Figure 3The diagram shows the structure of the SlowFast model provided in Embodiment 1 of this invention. In the SlowFast model, the slow branch receives low frame rate video frames as input, aiming to capture the spatial semantic information of the video, while the fast branch takes high frame rate video frames as input, focusing on extracting dynamic features from the video. During the feature extraction process of the slow and fast branches, they interact with each other in a specific fusion stage, ultimately outputting a prediction result for the video behavior. This algorithm can simultaneously extract both spatial semantic information and motion information from the video. Compared to existing methods for recognizing surrogate actions based on manually defined rules, the algorithm of this invention has higher recognition accuracy.
[0058] In this embodiment, the input data of the video behavior recognition model are low frame rate video frames and high frame rate video frames of the employee's hand video image, and the output data of the video behavior recognition model is the prediction result, that is, the behavior in the image is a valet operation or a normal operation.
[0059] This invention provides a behavior recognition method. First, a video image is acquired, including an image of the operator's appearance and hands. Second, an instance segmentation model is used to identify hand features in the video image to obtain a hand mask. Then, the hand mask and the video image are input into an image classification network to determine the operator's identity information. If the operator is identified as an employee, target tracking is performed on the hand features to obtain a video image of the employee's hands. Finally, the employee's hand video image is input into a video behavior recognition model to determine if surrogacy is involved. This method introduces hand identification based on instance segmentation, which improves the accuracy of hand identification and thus enhances the accuracy of behavior recognition.
[0060] Based on the above embodiments, modified embodiments of the above embodiments are proposed. It should be noted that, in order to keep the description brief, only the differences from the above embodiments are described in the modified embodiments.
[0061] In one embodiment, the training method for the instance segmentation model includes:
[0062] The system collects surveillance video data from smart teller machines, including images of employees' hands and customers' hands. Preprocessing of the video data, such as frame extraction, deduplication, and brightness adjustment, yields image data. Image data is then labeled to obtain training samples. Target masks and identity categories are annotated in the image data to obtain true labels. The model is trained based on the training samples, and the training results are compared with the true labels. Model parameters are adjusted according to the comparison results until the accuracy of the training results reaches a preset value.
[0063] For example, an improved YOLOv8-seg instance segmentation model is used for training. To accelerate model training and inference, smaller parameter settings, such as a network size of scale s, can be used. YOLOv8-seg uses a backbone network for multi-scale feature extraction, performs multi-scale feature fusion at bottlenecks, and finally performs instance segmentation at multiple scales, outputting information such as the target category, mask, and confidence score for each hand object. This embodiment uses the Adam optimizer for training, with a learning rate of 0.001, a batch size of 128, and a total training time of 1000 epochs before convergence.
[0064] In one embodiment, the training method for the image classification network includes:
[0065] The system collects surveillance video data from smart teller machines, including images of employees' hands and customers' hands. Preprocessing of the video data, such as frame extraction, deduplication, and brightness adjustment, yields image data. Image data is then labeled to obtain training samples. Target masks and identity categories are annotated in the image data to obtain true labels. The model is trained based on the training samples, and the training results are compared with the true labels. Network parameters are adjusted based on the comparison results until the accuracy of the training results reaches a preset value.
[0066] In one embodiment, the SlowFast network model is optimized using stochastic gradient descent with a learning rate of 0.005. During training, the slow branch inputs video training samples at a rate of one frame every 15 frames, ensuring that all input frames are keyframes; while the fast branch inputs video training samples at a rate of one frame every two frames, including both keyframes and non-keyframes. The entire training process lasts for 500 rounds to achieve sufficient optimization and convergence of the model parameters.
[0067] Example 2
[0068] Figure 2 This is a flowchart illustrating a behavior recognition method according to Embodiment 2 of the present invention. Embodiment 2 is an optimization based on the above embodiments. For details not covered in this embodiment, please refer to Embodiment 1.
[0069] like Figure 2 As shown, the behavior recognition method provided in Embodiment 2 of the present invention includes the following steps:
[0070] S210. Acquire video images, including images of the operator's appearance and hands.
[0071] S220. Input the video image into the instance segmentation model.
[0072] Preferably, the instance segmentation model is an optimized YOLOv8-seg network model, which includes a Backbone network, a Neck layer network, and a Head layer network; wherein, the feature splicing layer in the Neck layer network is optimized into a cross-layer feature fusion module, which has simple and efficient skip connections.
[0073] Figure 4 This is a structural diagram of an optimized YOLOv8-seg instance segmentation model provided in Embodiment 2 of the present invention. In this model, the cross-layer feature fusion module takes the feature maps of different layers as input and injects high-level semantic features and low-level detail features through term-by-term multiplication, thereby enhancing the detail features of the image at that layer. This data can be transferred to subsequent network layers to reconstruct the resolution and perform the final image segmentation.
[0074] The cross-layer feature fusion module is used to traverse the N-layer feature maps in the M-layer feature maps output by the Backbone network, adjust the size of the j-th layer feature map in the N-layer feature maps to be the same as the size of the output layer feature map, and obtain the output feature map; perform a standard convolution operation on the output feature map to obtain multiple i-th level feature maps; adjust the multiple i-th level feature maps to the same resolution and then perform a product term by term; where 1≤i≤M≤N.
[0075] Figure 5 This is a structural diagram of the cross-layer feature fusion module provided in Embodiment 2 of the present invention, as shown below. Figure 5 As shown, given an input image I, the feature map output by YOLOv8-seg through the neck network has N layers. M of these layers are used as input to the cross-layer feature fusion module, and the feature map of the i-th layer in the input is labeled as f. i 0 1≤i≤M≤N, the features of different layers in the cross-layer feature fusion module can be represented as M takes the value 3.
[0076] First, the cross-layer feature fusion module iterates through the input layers, adjusting the size of the feature map of the j-th layer to be the same as the size of the feature map of the i-th output layer, respectively. D, I, and U represent downsampling, identity mapping, and upsampling to resolution H, respectively. i ×W i Where 1≤i,j≤M, the specific calculation formula is as follows:
[0077]
[0078] Then, the cross-layer feature fusion module uses a 3x3 convolution on the feature map. Smoothing process θ ij The parameters representing this convolution are calculated using the following formula:
[0079]
[0080] Finally, the cross-layer feature fusion module adjusts all i-th level feature maps to the same resolution and then performs a term-by-term product P on all resized feature maps to enhance the i-th level features with more semantic information and finer details. The calculation formula is as follows:
[0081]
[0082] S230. Identify targets in the image using the instance segmentation model, and output the target category, target mask, and confidence level.
[0083] The target can include objects such as faces, hands, and eyes.
[0084] S240. If the target category is a hand and the confidence level corresponding to the target category is greater than the threshold, then the target mask is determined to be a hand mask.
[0085] S250. Input the hand mask and the video image into an image classification network, and determine the identity information of the smart teller machine operator through the image classification network.
[0086] S260. If the identity information of the operator of the smart teller machine is that of an employee, then target tracking is performed on the hand features to obtain video images of the employee's hands.
[0087] S270. Input the video image of the employee's hand into the video behavior recognition model to determine whether there is any valet operation behavior.
[0088] The second embodiment of the present invention provides a behavior recognition method, which specifies the specific network structure and function of the instance segmentation model. The optimized YOLOv8-seg can quickly and accurately extract the image information of the customer's hand region in the intelligent video image. By enhancing the hand features, the accuracy of the subsequent image classification network in confirming the hand identity is improved.
[0089] Furthermore, the method also includes: replacing the standard convolution operation with a GSConv convolution operation; performing the GSConv convolution operation on the output feature map to obtain multiple i-th level feature maps.
[0090] To further accelerate network processing, GSConv convolution can be used to replace the standard convolution operation of the cross-layer feature fusion module.
[0091] Furthermore, the method also includes: adding a content-aware upsampling operator to the Neck layer network, wherein the content-aware upsampling operator samples the content-aware feature recombination method to scale the feature map.
[0092] In the YOLOv8 neck network layer, upsampling uses the nearest neighbor difference method. The principle is to determine the nearest neighbor pixel in the target image based on the position of each pixel in the original image, and replace the value of that pixel in the original image with the value of the nearest neighbor pixel. However, the image quality after scaling using the nearest neighbor difference method is poor and may have obvious jagged edges.
[0093] In this embodiment, a content-aware upsampling operator is introduced into the neck network of YOLOv8-seg. This operator replaces the nearest neighbor interpolation method with a content-aware feature recombination method, thereby enhancing the semantic feature extraction capability of the neural network in the instance segmentation task.
[0094] Based on the technical solutions of the above embodiments, this invention provides a specific implementation method. As one specific implementation method of this invention, Figure 6 A flowchart of a behavior recognition method provided in a specific embodiment of the present invention is shown below. Figure 6 As shown, the process includes the following:
[0095] The ATM surveillance video data is input into an instance segmentation model to perform hand feature recognition and obtain a hand mask. The hand mask and the ATM surveillance video data are then input into an image classification network to determine whether an employee's hand has been identified. If so, the employee's hand is tracked. The tracked image is then enhanced, sub-images are scaled, and stored in a buffer. It is determined whether the buffered images have accumulated to a full X frames. If they have accumulated to a full X frames, the buffered images are input into a video behavior recognition model to identify valet operations. If valet operations are identified, an alarm is sent. If they have not accumulated to a full X frames, target tracking of the employee's hand continues until the buffered images have accumulated to a full X frames.
[0096] Example 3
[0097] Figure 7 This is a schematic diagram of a behavior recognition device provided in Embodiment 3 of the present invention. The device is applicable to identifying unauthorized valet operations. The device can be implemented by software and / or hardware and is generally integrated into an electronic device.
[0098] like Figure 7 As shown, the device includes: an acquisition module 110, an identification module 120, a determination module 130, an acquisition module 140, and an identification module 150.
[0099] The first acquisition module 110 is used to acquire video images, including images of the appearance and hands of the operator of the smart teller machine;
[0100] The first recognition module 120 is used to use an instance segmentation model to recognize hand features in the video image to obtain a hand mask;
[0101] The determination module 130 is used to input the hand mask and the video image into an image classification network, and determine the identity information of the smart teller machine operator through the image classification network;
[0102] The second acquisition module 140 is used to perform target tracking on the hand features to obtain a video image of the employee's hand if the identity information of the operator of the smart teller machine is that of an employee.
[0103] The second recognition module 150 is used to input the video image of the employee's hand into the video behavior recognition model to determine whether there is any valet operation behavior.
[0104] In this embodiment, the device first acquires video images through a first acquisition module 110, the video images including the appearance image and hand image of the smart teller machine operator; secondly, a first recognition module 120 uses an instance segmentation model to identify hand features in the video images to obtain a hand mask; then, a determination module 130 inputs the hand mask and the video image into an image classification network to determine the identity information of the smart teller machine operator; subsequently, a second acquisition module 140 performs target tracking on the hand features if the smart teller machine operator's identity information is that of an employee, to obtain a video image of the employee's hand; finally, a second recognition module 150 inputs the video image of the employee's hand into a video behavior recognition model to determine whether there is any valet operation behavior.
[0105] This embodiment provides a behavior recognition device that can improve the accuracy of hand identification, thereby improving the accuracy of behavior recognition.
[0106] Furthermore, the first identification module 120 includes:
[0107] The input submodule is used to input the video image into the instance segmentation model;
[0108] The output submodule is used to identify targets in an image using the instance segmentation model and output the target category, target mask, and confidence score.
[0109] The determination submodule is used to determine the target mask as a hand mask if the target category is a hand and the confidence level corresponding to the target category is greater than a threshold.
[0110] Based on the above optimizations, the instance segmentation model is an optimized YOLOv8-seg network model, which includes a Backbone network, a Neck layer network, and a Head layer network. The feature splicing layer in the Neck layer network is optimized into a cross-layer feature fusion module, which has simple and efficient skip connections.
[0111] Based on the above technical solution, the cross-layer feature fusion module is used to traverse the N-layer feature maps in the M-layer feature maps output by the Backbone network, adjust the size of the j-th layer feature map in the N-layer feature maps to be the same as the size of the output layer feature map, and obtain the output feature map; perform standard convolution operation on the output feature map to obtain multiple i-th level feature maps; adjust the multiple i-th level feature maps to the same resolution and then perform item-by-item multiplication; where 1≤i≤M≤N.
[0112] Furthermore, the device also includes:
[0113] The replacement module is used to replace the standard convolution operation with the GSConv convolution operation;
[0114] The convolution module is used to perform the GSConv convolution operation on the output feature map to obtain multiple i-th level feature maps.
[0115] Furthermore, the device also includes an adding module, which is used to add a content-aware upsampling operator to the Neck layer network. The content-aware upsampling operator samples content-aware feature recombination methods to scale the feature map.
[0116] The above-described behavior recognition device can execute the behavior recognition method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method.
[0117] Example 4
[0118] Figure 8 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0119] like Figure 8As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0120] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0121] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as behavior recognition methods.
[0122] In some embodiments, the behavior recognition method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the behavior recognition method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the behavior recognition method by any other suitable means (e.g., by means of firmware).
[0123] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0124] In some embodiments, the behavior recognition method may be implemented as a computer program, which is implicitly included in a computer program product. When executed by a processor, the computer program implements the behavior recognition method of the present invention. The computer program product can be understood as a software product that primarily implements its solution through a computer program. The computer program used to implement the method of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer program causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer program may be executed entirely on a machine, partially on a machine, partially on a remote machine as a standalone software package, or entirely on a remote machine or server.
[0125] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0126] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0127] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0128] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0129] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0130] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A behavior recognition method, characterized in that, The method includes: Acquire video images, including images of the operator's appearance and hands; The hand mask is obtained by identifying hand features in the video image using an instance segmentation model; The hand mask and the video image are input into an image classification network, and the identity information of the smart teller machine operator is determined through the image classification network; If the operator of the smart teller machine is an employee, then the hand features are tracked to obtain video images of the employee's hands. The video image of the employee's hand is input into the video behavior recognition model to determine whether there is any surrogate operation behavior.
2. The method according to claim 1, characterized in that, The step of using an instance segmentation model to identify hand features in the video image to obtain a hand mask includes: The video image is input into the instance segmentation model; The instance segmentation model identifies targets in an image and outputs the target category, target mask, and confidence score. If the target category is a hand, and the confidence level corresponding to the target category is greater than the threshold, then the target mask is determined to be a hand mask.
3. The method according to claim 1 or 2, characterized in that, The instance segmentation model is an optimized YOLOv8-seg network model, which includes a Backbone network, a Neck layer network, and a Head layer network. The feature splicing layer in the Neck layer network is optimized into a cross-layer feature fusion module, which has simple and efficient skip connections.
4. The method according to claim 3, characterized in that, The cross-layer feature fusion module is used to traverse the N-layer feature maps in the M-layer feature maps output by the Backbone network, adjust the size of the j-th layer feature map in the N-layer feature maps to be the same as the size of the output layer feature map, and obtain the output feature map; perform a standard convolution operation on the output feature map to obtain multiple i-th level feature maps; adjust the multiple i-th level feature maps to the same resolution and then perform a product term by term; where 1≤i≤M≤N.
5. The method according to claim 4, characterized in that, The method further includes: Replace the standard convolution operation with the GSConv convolution operation; The GSConv convolution operation is performed on the output feature map to obtain multiple i-th level feature maps.
6. The method according to claim 3, characterized in that, The method further includes: A content-aware upsampling operator is added to the Neck layer network. The content-aware upsampling operator samples the content-aware feature reconstruction method to scale the feature map.
7. A behavior recognition device, characterized in that, The device includes: The first acquisition module is used to acquire video images, including images of the operator's appearance and hands. The first recognition module is used to identify hand features in the video image using an instance segmentation model to obtain a hand mask; The determination module is used to input the hand mask and the video image into an image classification network, and determine the identity information of the smart teller machine operator through the image classification network; The second acquisition module is used to perform target tracking on the hand features to obtain video images of the employee's hands if the identity information of the operator of the smart teller machine is that of an employee. The second recognition module is used to input the video image of the employee's hand into the video behavior recognition model to determine whether there is any valet operation behavior.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the behavior recognition method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the behavior recognition method according to any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the behavior recognition method according to any one of claims 1-6.