Method, device and equipment for identifying and evaluating handover scene of cash box of cash truck and medium
By using an improved YOLOv5 model and algorithm, high-precision automatic identification of cash box handover scenarios in armored trucks has been achieved, solving the problems of misjudgment due to manual monitoring and sensor limitations in existing technologies, and improving the security and standardization of cash box handover in bank armored trucks.
Patent Information
- Application Number
- CN202511066411.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies cannot achieve high-precision, automated security monitoring in the cash box handover scenario of bank armored vehicles. They are prone to misjudgment and inaccurate security standard determination due to human fatigue or sensor limitations.
By employing an improved YOLOv5 model based on an attention mechanism, combined with coordinate attention mechanism and Soft-NMS algorithm, accurate identification and security assessment of armored trucks, cash boxes, and security personnel can be achieved.
It improves the real-time performance and reliability of the monitoring system, reduces labor costs, and ensures the safety and standardization of the handover process.
Smart Images

Figure CN120997761A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of target detection technology, and in particular to a method, apparatus, equipment and medium for recognizing and evaluating cash box handover scenes in armored trucks. Background Technology
[0002] Currently, there are several methods for security monitoring of bank armored truck and cash box handover operations. The first method relies primarily on manual monitoring. Staff must constantly monitor the surveillance camera feeds to assess the status of the armored truck, cash boxes, and security personnel, determining whether to proceed with the handover and whether the operation meets security standards. However, this method is highly manpower-dependent, prone to errors due to staff fatigue or negligence, and inefficient. The second method uses cash box sensors to monitor the cash box's status. However, this method only monitors the cash box itself and cannot comprehensively assess surrounding key elements such as the armored truck and security personnel.
[0003] The third approach could be a video analytics-based surveillance system. However, this system is limited by the performance of its target detection algorithms. In scenarios with complex lighting, dynamic occlusion, and dense distribution of multiple targets, it struggles to accurately identify armored trucks, cash boxes, and security personnel, leading to misjudgments of handover process triggering conditions and inaccurate security standard determinations. Therefore, existing technical solutions cannot meet the urgent needs of financial escort scenarios for automated, high-precision monitoring. Summary of the Invention
[0004] This invention provides a method, apparatus, equipment, and medium for identifying and evaluating cash box handover scenarios in armored trucks, in order to ensure the security and standardization of the handover process in such scenarios.
[0005] According to one aspect of the present invention, a method for recognizing and evaluating cash box handover scenes in armored trucks is provided, comprising:
[0006] Receive video stream of the cash box handover scene in the armored truck;
[0007] The video stream is input into the target recognition model, which then identifies the video stream to obtain the corresponding recognition result. The target recognition model is a YOLOv5 model based on an improved attention mechanism, and the recognition result includes information about the armored truck, cash box, and security personnel corresponding to the video stream.
[0008] The security of the cash transfer process in armored vehicles is assessed based on the identification results.
[0009] According to another aspect of the present invention, a device for recognizing and evaluating cash box handover scenes in armored trucks is provided, comprising:
[0010] The video stream receiving module is used to receive video streams of the cash box handover scene in the armored truck;
[0011] The recognition result determination module is used to input the video stream into the target recognition model, and to identify the video stream through the target recognition model to obtain the recognition result corresponding to the video stream; wherein, the target recognition model is a YOLOv5 model based on an improved attention mechanism, and the recognition result includes information on the armored truck, cash box and security personnel corresponding to the video stream;
[0012] The security assessment module is used to assess the security of the cash box handover process in armored vehicles based on the identification results.
[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0014] At least one processor;
[0015] and memory that is communicatively connected to at least one processor;
[0016] The memory stores a computer program that can be executed by at least one processor. The computer program is executed by at least one processor so that at least one processor can execute the cash box handover scene recognition and evaluation method of any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement the cash box handover scene recognition and evaluation method of any embodiment of the present invention.
[0018] The technical solution of this invention involves receiving a video stream of the cash handover scene from an armored truck; inputting the video stream into a target recognition model; using the target recognition model to identify the video stream and obtain the corresponding recognition result; and assessing the security of the cash handover process based on the recognition result. This technical solution, employing a target recognition model, can achieve automatic identification and monitoring, solving the problems of manual monitoring being prone to misjudgments due to staff fatigue or negligence, low efficiency, and the inability of sensors to fully perceive key surrounding elements such as the armored truck and security personnel. Specifically, the target recognition model is a YOLOv5 model based on an improved attention mechanism, and the recognition result includes information about the armored truck, cash box, and security personnel corresponding to the video stream. This technical solution incorporates a coordinate attention mechanism into the YOLOv5 model structure to obtain a target recognition model. This allows the model to focus more on task-related areas within the image, accurately locating targets (such as armored trucks and cash boxes) and perceiving the distribution of surrounding personnel. It solves the technical problem in existing technologies where the performance of target detection algorithms is limited, making it difficult to accurately identify armored trucks, cash boxes, and security personnel in scenarios with complex lighting, dynamic occlusion, and dense distribution of multiple targets. This leads to misjudgments of handover process trigger conditions and inaccurate security standard determinations. It can significantly reduce labor costs, improve the real-time performance and reliability of monitoring systems, provide an intelligent solution for financial security scenarios, and ensure the security and standardization of the handover process.
[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A flowchart illustrating a method for recognizing and evaluating cash box handover scenarios in armored trucks, provided in an embodiment of the present invention;
[0022] Figure 2 A flowchart illustrating another method for recognizing and evaluating cash box handover scenarios in armored vehicles, provided by an embodiment of the present invention;
[0023] Figure 3 A flowchart illustrating another method for recognizing and evaluating cash box handover scenarios in armored vehicles, provided by an embodiment of the present invention;
[0024] Figure 4 A structural diagram of the target recognition model provided in an embodiment of the present invention;
[0025] Figure 5 A network structure diagram of the coordinate attention mechanism provided in an embodiment of the present invention;
[0026] Figure 6 This is a schematic diagram of the structure of a cash transport vehicle cash box handover scene recognition and evaluation device provided in an embodiment of the present invention;
[0027] Figure 7 A schematic diagram of the electronic device used to implement the cash box handover scene recognition and evaluation method of the armored truck according to an embodiment of the present invention. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0030] Figure 1 This is a flowchart illustrating a method for identifying and evaluating cash handover scenes in armored trucks, provided by an embodiment of the present invention. This embodiment is applicable to identifying and evaluating cash handover scenes in armored trucks to ensure security and compliance. The method can be executed by an armored truck cash handover scene identification and evaluation device, which can be implemented in hardware and / or software and can be configured in a computer device. Figure 1 As shown, the method specifically includes the following steps:
[0031] S110: Receives the video stream of the cash box handover scene from the armored truck.
[0032] Among them, the video stream refers to the video stream captured by the image acquisition device corresponding to the handover scene of the cash box of the armored truck. The image acquisition device can be a surveillance camera.
[0033] S120. Input the video stream into the target recognition model, and use the target recognition model to recognize the video stream to obtain the recognition result corresponding to the video stream.
[0034] The target recognition model is a YOLOv5 model based on an improved attention mechanism. The recognition results include information about the armored truck, cash box, and security personnel corresponding to the video stream.
[0035] Understandably, bank branch surveillance equipment typically uses an overhead camera approach. While this method covers a wide area, it often results in cash boxes and security personnel appearing as small figures in the frame, incomplete capture of armored trucks, and susceptibility to obstruction by other objects. This can cause the recognition model to incorrectly focus on background content, thus affecting accuracy. Furthermore, the limited computing power of mobile devices or edge servers deployed in bank branches makes it difficult to deploy complex network models with a large number of parameters.
[0036] To address the aforementioned issues, this invention, based on the YOLOv5 target detection model, incorporates a coordinate attention mechanism into its structure to obtain a target recognition model. This allows the target recognition model to focus more on task-related areas within the image, enabling it to accurately locate targets (such as armored trucks or cash boxes) and perceive the distribution of people in the surrounding area.
[0037] In some alternative embodiments, in order to save computing power on mobile devices or edge servers and improve detection efficiency, the video stream can be processed by cross-sampling technology after receiving the video stream of the cash box handover scene in the armored truck, so as to reduce the video frame rate and / or resolution of the video stream.
[0038] S130. Assess the security of the cash box handover process in armored vehicles based on the identification results.
[0039] Specifically, the handover process of cash boxes in armored vehicles can be assessed based on the identification results to determine whether it is safe and compliant.
[0040] In some embodiments, assessing the security of the cash box handover process in an armored truck based on the identification results includes: for a single video frame in a video stream, determining whether an armored truck exists in the video frame based on the identification results; if an armored truck exists in the video frame, determining whether a cash box exists in the video frame based on the identification results; if a cash box exists in the video frame, determining the number of security personnel in the video frame, and conducting a security assessment based on the number of security personnel.
[0041] Figure 2A flowchart of another method for recognizing and evaluating cash box handover scenarios in armored vehicles provided by an embodiment of the present invention is shown below. Figure 2 As shown:
[0042] (1) The system receives real-time video streams from surveillance cameras as input.
[0043] (2) Preprocess the input video stream by reducing the video frame rate or resolution through cross-sampling techniques to reduce computation and improve processing speed.
[0044] (3) A coordinate attention mechanism was added to the network of the YOLOv5 object detection model, and the modified object detection model was used as the object recognition model. The preprocessed video frames were input into the object recognition model. The object recognition model performed deep feature extraction and classification on the video frames and output the recognition results, including information such as the location and number of armored trucks, cash boxes, and security personnel.
[0045] (4) Based on the output of the target recognition model, determine whether there is an armored truck in the current frame. If it does not exist, the process ends; if it does exist, continue to the next step.
[0046] (5) After confirming the existence of the armored truck, further determine whether the cash box exists. If the cash box does not exist, the process ends; if it exists, proceed to the next step.
[0047] (6) Statistical analysis of the number of security personnel output by the model to assess the security during the handover process.
[0048] (7) Compare the counted number of security personnel with the preset threshold. If the number of security personnel is lower than the threshold, it means that the handover was not configured with enough security personnel as required, the handover operation is risky, and the handover process does not meet the standards. At this time, the system will generate an abnormal alarm. If the number of security personnel is higher than the threshold, it means that the handover process meets the personnel standards in the security regulations, and the system will generate a cash box handover prompt.
[0049] The technical solution of this invention involves receiving a video stream of the cash handover scene from an armored truck; inputting the video stream into a target recognition model; using the target recognition model to identify the video stream and obtain the corresponding recognition result; and assessing the security of the cash handover process based on the recognition result. This technical solution, employing a target recognition model, can achieve automatic identification and monitoring, solving the problems of manual monitoring being prone to misjudgments due to staff fatigue or negligence, low efficiency, and the inability of sensors to fully perceive key surrounding elements such as the armored truck and security personnel. Specifically, the target recognition model is a YOLOv5 model based on an improved attention mechanism, and the recognition result includes information about the armored truck, cash box, and security personnel corresponding to the video stream. This technical solution incorporates a coordinate attention mechanism into the YOLOv5 model structure to obtain a target recognition model. This allows the model to focus more on task-related areas within the image, accurately locating targets (such as armored trucks and cash boxes) and perceiving the distribution of surrounding personnel. It solves the technical problem in existing technologies where the performance of target detection algorithms is limited, making it difficult to accurately identify armored trucks, cash boxes, and security personnel in scenarios with complex lighting, dynamic occlusion, and dense distribution of multiple targets. This leads to misjudgments of handover process trigger conditions and inaccurate security standard determinations. It can significantly reduce labor costs, improve the real-time performance and reliability of monitoring systems, provide an intelligent solution for financial security scenarios, and ensure the security and standardization of the handover process.
[0050] Figure 3 This is a flowchart illustrating another method for recognizing and evaluating cash box handover scenes in armored trucks, provided by an embodiment of the present invention. This embodiment further optimizes the video stream processing method. For example... Figure 3 As shown, the method specifically includes the following steps:
[0051] S210: Receive the video stream of the cash box handover scene from the armored truck.
[0052] S220. Perform feature extraction and target prediction on the video stream through the backbone network, neck network and head network to obtain at least two candidate detection boxes corresponding to the video stream.
[0053] like Figure 4 The diagram shown illustrates the structure of the target recognition model provided in this embodiment of the invention. The target recognition model includes a backbone network, a neck network, and a head network; it also utilizes an improved non-maximum suppression algorithm. Candidate detection boxes are understood as boxes within video frames, used to define regions where targets may exist in the video frames.
[0054] Specifically, feature extraction and target prediction can be performed on the video stream through the backbone network, neck network, and head network of the target recognition model to obtain candidate detection boxes in the video frames.
[0055] In some embodiments, feature extraction and target prediction are performed on the video stream through a backbone network, a neck network, and a head network to obtain at least two candidate detection boxes corresponding to the video stream, including: inputting the video stream into the backbone network to obtain basic features output by the backbone network; inputting the basic features into the neck network to obtain deep features output by the neck network; and inputting the deep features into the head network to obtain at least two candidate detection boxes output by the head network.
[0056] Among them, backbone network refers to Figure 4 The backbone portion; the neck network refers to Figure 4 The neck part; the head network refers to Figure 4 The head section of the algorithm is as follows: Specifically, the video stream first passes through the backbone section to extract preliminary features. The extracted features are then processed by multi-scale pooling through the SPPF layer. The features processed by the backbone section enter the neck section for feature fusion and upsampling. The fused features are then passed to the head section for final prediction. Through multi-scale feature extraction and fusion, the model's ability to detect targets of different sizes is improved.
[0057] In some embodiments, the neck network includes at least one coordinate attention module, which processes the input feature map to obtain a target feature map corresponding to the input feature map.
[0058] It should also be noted that the neck network can include multiple Coordination Attention (CA) modules. The input of each Coordination Attention module is the input feature map, and the output is the target feature map. The processing procedure of each Coordination Attention module can be the same.
[0059] Understandably, coordinate attention (CA) is a lightweight attention mechanism that enhances a model's spatial perception of targets by fusing channel attention with location information. Compared to the loss of location information caused by global pooling in traditional channel attention, CA captures long-range dependencies by decomposing spatial dimensions and using low-cost computation, significantly improving localization accuracy in tasks such as object detection and semantic segmentation. It can also be seamlessly embedded into lightweight networks such as MobileNetV2 and EfficientNet. Figure 4 CAAtt is the embedded coordinate attention module, whose orientation sensitivity enables the model to more accurately identify the location of targets such as armored trucks and cash boxes, as well as the distribution of security personnel.
[0060] In some embodiments, processing the input feature map to obtain the target feature map corresponding to the input feature map includes: globalizing the input feature map in one dimension along the horizontal and vertical directions to obtain feature encoding vectors in two directions; concatenating the encoding vectors in the two directions, generating attention weights through 1×1 convolution and nonlinear activation, and decomposing them into a horizontal weight map and a vertical weight map; and weighting the input feature map based on the horizontal weight map and the vertical weight map to obtain the target feature map.
[0061] Specifically, the coordinate self-attention module's network structure comprises two core stages: coordinate information embedding and coordinate attention generation. Coordinate information embedding involves performing one-dimensional global pooling on the input feature map along both the horizontal (X-axis) and vertical (Y-axis) directions, generating feature encoding vectors for both directions. This operation decomposes the spatial dimension into independent coordinate axes, preserving precise location information while avoiding the positional loss caused by traditional global pooling. For an input of size C×H×W, pooling in the X-direction yields a C×H×1 encoding vector, and pooling in the Y-direction yields a C×1×W encoding vector.
[0062] Coordinate attention generation involves concatenating the encoded vectors of two directions, then generating attention weights through 1×1 convolution and non-linear activation (such as h-sigmoid), decomposing them into a horizontal weight map (C×H×1) and a vertical weight map (C×1×W). Finally, these two weight maps are element-wise multiplied with the original feature map to achieve joint spatial-channel enhancement. This operation captures long-range dependencies through a direction-sensitive mechanism, enabling the model to accurately locate targets (such as armored vehicles and cash boxes) and perceive the distribution of surrounding people. The network structure of the coordinate attention mechanism is as follows: Figure 5 As shown.
[0063] In this embodiment of the invention, a coordinate attention mechanism is embedded into the original backbone network of the YOLOv5 object detection model to improve the model's ability to perceive the location and discriminate features of targets in complex scenes. The core innovation of this solution lies in adding a spatial-channel joint perception structure to the traditional YOLOv5 backbone network. In the backbone network, the input feature map is subjected to one-dimensional global pooling along the horizontal (X-axis) and vertical (Y-axis) directions to generate direction-sensitive feature encoding vectors. Attention weights are generated through lightweight convolution and non-linear activation to guide the model to focus on key areas of the target (such as the edge of an armored truck or the outline of a cash box). Compared to the native YOLOv5, which relies on local convolutional kernels to implicitly learn spatial relationships, coordinate attention explicitly preserves the precise coordinate information of the target, improving the average detection accuracy (mAP) in complex scenes such as occlusion and small targets. At the same time, the module only adds a minimal amount of computation, is compatible with real-time detection requirements, and provides higher-precision target localization capabilities for security monitoring systems.
[0064] S230. Based on the improved nonmaximum suppression algorithm, target detection boxes are selected from at least two candidate detection boxes, and the recognition result corresponding to the video stream is determined based on the target detection boxes.
[0065] Among them, Soft-NMS (Soft Non-Maximum Suppression) is an improved non-maximum suppression (NMS) algorithm. Traditional non-maximum suppression (NMS) algorithms are prone to mistakenly deleting valid detection boxes when dealing with highly overlapping targets, affecting counting accuracy. Existing technical solutions cannot meet the urgent needs of automated and high-precision monitoring in financial escort scenarios.
[0066] Traditional NMS algorithms select the prediction box with the highest confidence score from among many overlapping prediction boxes and compare its confidence score with a set threshold. If the overlap exceeds the threshold, the prediction box score is set to zero. However, in images captured by bank branch cameras, overlapping or occlusion of people and targets frequently occur. Therefore, if the prediction box score is directly set to zero when the score of an adjacent detection box is greater than the threshold, in scenarios involving the handover of cash boxes, partial occlusion of cash boxes / security personnel, and target overlap frequently occur. This can easily lead to missed detections and false detections, resulting in reduced detection accuracy.
[0067] This invention proposes using the Soft-NMS algorithm to replace the traditional NMS algorithm. By introducing a confidence decay mechanism instead of the direct suppression strategy of traditional NMS, it addresses the issue of missed detections in dense target scenes. The core process involves sorting the detection boxes by confidence score, and then penalizing the highly overlapping boxes adjacent to the highest-scoring box with confidence (e.g., linear or Gaussian decay), rather than directly removing them. Compared to the native NMS algorithm used in YOLOv5, this algorithm retains more potentially valid detection boxes, making it particularly suitable for scenes with occlusion and densely distributed small targets. The function of the Soft-NMS algorithm is shown below:
[0068]
[0069] Where b i Let represent the i-th detection box, Si represent the score of the i-th detection box, Nt represent the NMS threshold, and M represent the detection box with the highest score.
[0070] S240. Assess the security of the cash box handover process in armored vehicles based on the identification results.
[0071] This invention integrates a coordinate attention module and the Soft-NMS algorithm into the YOLOv5 framework to construct an intelligent security detection system for financial escort scenarios, which has the following core advantages:
[0072] 1. High-precision spatial awareness: After embedding a coordinate attention module into the YOLOv5 network, the model explicitly models the target coordinate information through one-dimensional global pooling in the X / Y axis directions. This allows the model to pay more attention to the characteristics of key areas such as the edges of the armored truck and the outline of the cash box. Compared to the native YOLOv5, the improved model has improved detection accuracy in occluded scenes (such as security personnel partially obscuring the cash box).
[0073] 2. Dense target detection optimization: The Soft-NMS algorithm is adopted to replace the original NMS. The effective detection boxes in overlapping areas are retained through the Gaussian decay strategy, which improves the recognition recall rate in scenarios with stacked cash boxes and people blocking the view.
[0074] 3. Real-time performance and lightweight compatibility: The improved solution has excellent computational efficiency. The number of parameters added to the coordinate attention module is less than 1%, and the increase in Soft-NMS processing time is less than 3ms / frame. While ensuring model performance, it meets the computing power limitations of bank branch mobile devices and edge hosts.
[0075] 4. Intelligent security decision-making capability: By dynamically detecting the number of armored trucks, cash boxes, and security personnel, the system can automatically trigger the cash box handover operation judgment logic: when ≥1 armored truck, ≥3 cash boxes, and ≥4 security personnel are detected, it is determined that the security standard is met; when the number of security personnel is <4, a violation warning is issued.
[0076] This invention presents an improved YOLOv5-based cash box handover scene recognition model for armored trucks, enabling automated recognition and security compliance assessment of the cash box handover process. By embedding a coordinate attention module in the backbone network, the model's ability to encode target spatial location information is enhanced, improving the detection accuracy of small and dense targets. The Soft-NMS algorithm optimizes the post-processing workflow, effectively preserving valid detection boxes for overlapping targets and ensuring the accuracy of multi-target counting. Combining multi-target detection results with preset logical rules, the cash box handover process is automatically triggered, and compliance is assessed based on a security personnel number threshold. This invention significantly reduces labor costs, improves the real-time performance and reliability of monitoring systems, provides an intelligent solution for financial security scenarios, and ensures the security and standardization of the handover process.
[0077] Figure 6 This is a schematic diagram of a cash transport vehicle cash box handover scene recognition and evaluation device provided in an embodiment of the present invention. Figure 6 As shown, the device includes:
[0078] The video stream receiving module 310 is used to receive the video stream of the cash box handover scene in the armored truck;
[0079] The recognition result determination module 320 is used to input the video stream into the target recognition model, and to recognize the video stream through the target recognition model to obtain the recognition result corresponding to the video stream; wherein, the target recognition model is a YOLOv5 model based on an improved attention mechanism, and the recognition result includes information on the armored truck, cash box and security personnel corresponding to the video stream;
[0080] The security assessment module 330 is used to assess the security of the cash box handover process in armored vehicles based on the identification results.
[0081] The technical solution of this invention involves receiving a video stream of the cash handover scene from an armored truck; inputting the video stream into a target recognition model; using the target recognition model to identify the video stream and obtain the corresponding recognition result; and assessing the security of the cash handover process based on the recognition result. This technical solution, employing a target recognition model, can achieve automatic identification and monitoring, solving the problems of manual monitoring being prone to misjudgments due to staff fatigue or negligence, low efficiency, and the inability of sensors to fully perceive key surrounding elements such as the armored truck and security personnel. Specifically, the target recognition model is a YOLOv5 model based on an improved attention mechanism, and the recognition result includes information about the armored truck, cash box, and security personnel corresponding to the video stream. This technical solution incorporates a coordinate attention mechanism into the YOLOv5 model structure to obtain a target recognition model. This allows the model to focus more on task-related areas within the image, accurately locating targets (such as armored trucks and cash boxes) and perceiving the distribution of surrounding personnel. It solves the technical problem in existing technologies where the performance of target detection algorithms is limited, making it difficult to accurately identify armored trucks, cash boxes, and security personnel in scenarios with complex lighting, dynamic occlusion, and dense distribution of multiple targets. This leads to misjudgments of handover process trigger conditions and inaccurate security standard determinations. It can significantly reduce labor costs, improve the real-time performance and reliability of monitoring systems, provide an intelligent solution for financial security scenarios, and ensure the security and standardization of the handover process.
[0082] In some embodiments, the armored vehicle cash box handover scene recognition and evaluation device further includes a cross-sampling processing module, used for:
[0083] After receiving the video stream of the cash box handover scene in the armored truck, the video stream is processed using cross-sampling technology to reduce the video frame rate and / or resolution.
[0084] In some embodiments, the target recognition model includes a backbone network, a neck network, a head network, and an improved nonmaximum suppression algorithm. The recognition result determination module 320 includes:
[0085] The feature extraction and recognition submodule is used to perform feature extraction and target prediction on the video stream through the backbone network, neck network and head network to obtain at least two candidate detection boxes corresponding to the video stream.
[0086] The recognition result determination submodule is used to filter out target detection boxes from at least two candidate detection boxes based on an improved nonmaximum suppression algorithm, and determine the recognition result corresponding to the video stream based on the target detection boxes.
[0087] In some embodiments, the feature extraction and recognition submodule is specifically used for:
[0088] The video stream is input into the backbone network to obtain the basic features output by the backbone network;
[0089] The basic features are input into the neck network to obtain the deep features output by the neck network;
[0090] Deep features are input into the head network to obtain at least two candidate detection boxes output by the head network.
[0091] In some embodiments, the neck network includes at least one coordinate attention module, which processes the input feature map to obtain a target feature map corresponding to the input feature map.
[0092] In some embodiments, the coordinate attention module is specifically used for:
[0093] The input feature map is processed to obtain the target feature map corresponding to the input feature map, including:
[0094] The input feature map is globalized in one dimension along the horizontal and vertical directions to obtain feature encoding vectors in two directions;
[0095] After concatenating the encoding vectors from the two directions, attention weights are generated through 1×1 convolution and non-linear activation, and then decomposed into a horizontal weight map and a vertical weight map.
[0096] The target feature map is obtained by weighting the input feature map based on the horizontal and vertical weight maps.
[0097] In some embodiments, the security assessment module 330 is specifically used for:
[0098] For a single video frame in a video stream, determine whether an armored truck exists in the video frame based on the recognition results;
[0099] If an armored truck is present in a video frame, determine whether a cash box is present in the video frame based on the recognition results;
[0100] If a cash box is present in a video frame, determine the number of security personnel in the video frame and conduct a security assessment based on the number of security personnel.
[0101] The cash transport vehicle cash box handover scene recognition and evaluation device provided in the embodiments of the present invention can execute the cash transport vehicle cash box handover scene recognition and evaluation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method.
[0102] Figure 7 This is a schematic diagram of the electronic device used to implement the cash box handover scene recognition and evaluation method of the armored truck according to embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0103] like Figure 7 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0104] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0105] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the method for recognizing and evaluating cash box handover scenes in armored trucks.
[0106] In some embodiments, the armored truck cash box handover scene recognition and evaluation method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the armored truck cash box handover scene recognition and evaluation method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the armored truck cash box handover scene recognition and evaluation method by any other suitable means (e.g., by means of firmware).
[0107] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0108] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0109] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0110] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0111] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0112] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0113] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0114] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for recognizing and evaluating cash box handover scenes in armored trucks, characterized in that, include: Receive video stream of the cash box handover scene in the armored truck; The video stream is input into a target recognition model, and the target recognition model identifies the video stream to obtain the recognition result corresponding to the video stream; wherein, the target recognition model is a YOLOv5 model based on an improved attention mechanism, and the recognition result includes information about the armored truck, cash box, and security personnel corresponding to the video stream; The security of the cash handling process in armored trucks is assessed based on the identification results.
2. The method according to claim 1, characterized in that, After receiving the video stream of the cash box handover scene from the armored truck, it also includes: The video stream is processed using cross-sampling techniques to reduce the video frame rate and / or resolution of the video stream.
3. The method according to claim 1, characterized in that, The target recognition model includes a backbone network, a neck network, a head network, and an improved nonmaximum suppression algorithm. The step of using the target recognition model to identify the video stream and obtain the corresponding recognition result includes: Feature extraction and target prediction are performed on the video stream through the backbone network, the neck network and the head network to obtain at least two candidate detection boxes corresponding to the video stream; Based on an improved nonmaximum suppression algorithm, a target detection box is selected from at least two candidate detection boxes, and the recognition result corresponding to the video stream is determined based on the target detection box.
4. The method according to claim 3, characterized in that, The process of performing feature extraction and target prediction on the video stream through the backbone network, the neck network, and the head network to obtain at least two candidate detection boxes corresponding to the video stream includes: The video stream is input into the backbone network to obtain the basic features output by the backbone network; The basic features are input into the neck network to obtain the deep features output by the neck network; The deep features are input into the head network to obtain at least two candidate detection boxes output by the head network.
5. The method according to claim 4, characterized in that, The neck network includes at least one coordinate attention module, which is used to process the input feature map to obtain the target feature map corresponding to the input feature map.
6. The method according to claim 5, characterized in that, The process of processing the input feature map to obtain the target feature map corresponding to the input feature map includes: The input feature map is globalized in one dimension along the horizontal and vertical directions to obtain feature encoding vectors in two directions; After concatenating the encoding vectors from the two directions, attention weights are generated through 1×1 convolution and non-linear activation, and then decomposed into a horizontal weight map and a vertical weight map. The input feature map is weighted based on the horizontal weight map and the vertical weight map to obtain the target feature map.
7. The method according to claim 1, characterized in that, The security assessment of the cash transport vehicle's cash box handover process based on the identification results includes: For a single video frame in the video stream, determine whether an armored truck exists in the video frame based on the recognition result; If the armored truck is present in the video frame, determine whether a cash box is present in the video frame based on the identification result; If the cash box is present in the video frame, determine the number of security personnel in the video frame and conduct a security assessment based on the number of security personnel.
8. A device for recognizing and evaluating the handover scene of cash boxes in armored vehicles, characterized in that, include: The video stream receiving module is used to receive video streams of the cash box handover scene in the armored truck; The recognition result determination module is used to input the video stream into the target recognition model, and to recognize the video stream through the target recognition model to obtain the recognition result corresponding to the video stream; wherein, the target recognition model is a YOLOv5 model based on an improved attention mechanism, and the recognition result includes information about the armored truck, cash box, and security personnel corresponding to the video stream; The security assessment module is used to assess the security of the cash box handover process in armored vehicles based on the identification results.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the cash transport vehicle cash box handover scene recognition and evaluation method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the cash transport vehicle cash box handover scene recognition and evaluation method according to any one of claims 1-7.