Bank counter double-person adjacent cabinet detection method and device and computer program product

By extracting the key areas of video frames and using the Resnet network for analysis, the problem of increasing computing costs in non-critical areas in the prior art is solved, real-time compliance status monitoring and alerting of bank counters is achieved, and detection efficiency and security are improved.

CN119942411APending Publication Date: 2025-05-06中国邮政储蓄银行股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510084093.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art directly uses the image as input when analyzing video frames, resulting in additional computational costs and time in non-critical areas, wasted resources, and lack of comprehensive considerations for the compliance status of tellers and counter doors.

Method used

By obtaining video frames and extracting the location feature set, multiple key areas corresponding to the first video frame and the second video frame are extracted respectively, and feature change analysis is performed. If the feature change value is greater than or equal to the mutation threshold, use the Resnet network for in-depth analysis to determine whether there is a counter violation and issue an alarm.

Benefits of technology

It effectively reduces the waste of computing resources in non-critical areas, improves video image recognition efficiency, can monitor and alert the compliance status of the counter in real time, and improves security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942411A_ABST
    Figure CN119942411A_ABST
Patent Text Reader

Abstract

The invention provides a double-person cabinet approaching detection method and device for a bank counter and a computer program product. The method comprises the following steps: acquiring a first video frame and a second video frame; respectively extracting a plurality of first key areas and a plurality of second key areas corresponding to the first video frame and the second video frame according to the position feature set; performing feature change analysis on all the first key areas and the corresponding second key areas to obtain corresponding feature change values; under the condition that any characteristic change value is greater than or equal to a sudden change threshold value, analyzing the first target key area and the second target key area by adopting a Resnet network to obtain an analysis result; and under the condition that the analysis result is that the counter violation behavior exists, alarm information is sent out. According to the method and the device, the problem of resource waste caused by additional increase of calculation cost and time in a non-key area by directly taking the whole image as input and directly analyzing each image of continuous video frames when the video frames are analyzed in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent compliance detection for double-person visit to a bank counter, and in particular to a double-person visit to a bank counter detection method, a double-person visit to a bank counter detection device, a computer-readable storage medium and a computer program product. Background Art

[0002] Faced with a large number of monitoring screens, relevant personnel are unable to quickly and promptly identify abnormal operating behaviors with the naked eye, which provides opportunities for certain risky behaviors and increases the possibility of counter operation risks.

[0003] The counter area is one of the areas that the monitoring center personnel focus on. Real-time dynamic inspections and remote video monitoring ensure the safety and compliance of the counter during working hours. The main violations of the counter include: only one or no teller on duty for a continuous period of time; the teller is in an invalid duty area; the counter door is not closed when no one is on duty or during non-working hours. The traditional approach of manual investigation cannot achieve accurate and real-time alerts for risky behaviors. At present, deep learning is becoming more and more mature. With the help of video image analysis technology, risky behaviors can be monitored in real time and intelligently identified. The existing counter detection method based on video monitoring mainly captures video images at a certain frequency, and then uses algorithms to analyze the captured images to determine the position and behavior of the human body. Human position detection is currently divided into two categories: one is the detection method based on traditional machine learning. This method achieves accurate pedestrian position detection by designing representative human features, such as texture, color, histogram of oriented gradients (HOG), Haar features, scale-invariant features, etc., and efficient feature classifiers (such as support vector machines, random forests, etc.); the other is the pedestrian detection method based on deep learning. This method uses the learning ability of neural networks to extract pedestrian features and compare the extracted image features to identify whether there is a human body in the area. The focus of these two methods is only on human feature recognition for each frame of the image. The shortcomings of the existing technology: (1) When analyzing the video frame, the entire image is directly used as input, which requires additional computing cost and time for non-critical areas, which is not conducive to the recognition of video images in continuous time; (2) Directly performing AI analysis on each image of the continuous video frame will consume a lot of computing resources for scenes where most of the time is normal behavior, resulting in resource waste; (3) There is a lack of comprehensive consideration of the compliance status of the teller and the counter door, and most of the focus is on human recognition. Summary of the invention

[0004] The main purpose of the present application is to provide a method for detecting two people at a bank counter, a device for detecting two people at a bank counter, a computer-readable storage medium and a computer program product, so as to at least solve the problem in the prior art that when analyzing video frames, the entire image is directly taken as input and each image of continuous video frames is directly analyzed, thereby increasing additional computing costs and time in non-critical areas and causing waste of resources.

[0005] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a method for detecting two people at a bank counter is provided, including: an acquisition step, acquiring a first video frame and a second video frame, the first video frame being the current video frame, and the second video frame being the next video frame of the first video frame; an extraction step, respectively extracting a plurality of first key areas and a plurality of second key areas corresponding to the first video frame and the second video frame according to a position feature set, the position feature set representing a set of features of the effective positions of the counter and the teller on the image, the effective position being a position that meets the set requirements, the first key area being an area corresponding to when the effective positions of the counter and the teller in the first video frame meet the position feature set, the first key area being an area corresponding to when the effective positions of the counter and the teller in the second video frame meet the position feature set, and there is a one-to-one correspondence between the first key area and the second key area; a first analysis step, performing feature change analysis on all the first key areas and the corresponding second key areas to obtain corresponding a feature change value, one of the first key areas and one of the corresponding second key areas corresponds to one feature change value; a second analysis step, when any one of the feature change values ​​is greater than or equal to the mutation threshold, using the Resnet network to analyze the first target key area and the second target key area to obtain an analysis result, the first target key area is the first key area in the first video frame where the feature change value is greater than or equal to the mutation threshold, the second target key area is the second key area in the second video frame where the feature change value is greater than or equal to the mutation threshold, the analysis result is one of the following: there is a counter violation or there is no counter violation, the counter violation is at least that there is only one teller or no teller on duty for a continuous time, the teller is in an invalid on-duty area, and the door is in a non-closed state when no one is on duty or during non-working hours; an alarm step, when the analysis result is that the counter violation exists, an alarm message is issued, and the alarm message is used to prompt that there is a safety hazard in the current counter.

[0006] Optionally, after acquiring the first video frame and the second video frame, the method further includes: in the case where the first video frame is the first video frame among all the video frames, using an edge detection operator method to extract all valid position features, the valid position features being features representing the valid positions of the counter and the teller on the image; performing area division based on all the valid position features with the counter as a unit to obtain the position feature set, wherein a position feature in the position feature set corresponds to a key area in the video frame, and the number of the position features is the same as the number of the counters.

[0007] Optionally, an edge detection operator method is used to extract all valid position features, including: performing a planar convolution operation on the first video frame with a horizontal matrix and a vertical matrix image, respectively, to obtain a first gradient approximation and a second gradient approximation, wherein the first gradient approximation is a gradient approximation of each pixel point in the first video frame in a horizontal direction, and the first gradient approximation is a gradient approximation of each pixel point in the first video frame in a vertical direction; calculating the gradient size of each pixel point according to a first formula based on the first gradient approximation and the second gradient approximation of each pixel point, wherein the expression of the first formula is: G x represents the first gradient approximation, G y represents the second gradient approximation, G represents the gradient magnitude; the gradient direction of each pixel is calculated according to the second formula, and the second formula represents θ represents the gradient direction; all edges in the first video frame are identified according to the gradient magnitude and the gradient direction of each pixel point; and the position features formed by all the identified edges are determined as valid position features.

[0008] Optionally, a feature change analysis is performed on all the first key areas and the corresponding second key areas to obtain corresponding feature change values, including: using Euclidean distance to calculate the similarity between the color features and texture features of all the first key areas and the corresponding second key areas to obtain the corresponding feature change values.

[0009] Optionally, after performing feature change analysis on all the first key areas and the corresponding second key areas to obtain corresponding feature change values, the method further includes: when all the feature change values ​​are smaller than the mutation threshold, determining that there is no counter violation.

[0010] Optionally, when any one of the feature change values ​​is greater than or equal to the mutation threshold, the Resnet network is used to analyze the first target key area and the first target key area, and before obtaining the analysis result, the method also includes: using a convolution kernel of a first set size and a maximum pooling layer of a second set size to construct the input part of the Resnet network; using multiple identical convolution kernels of a third set size or multiple different convolution kernels to construct the intermediate convolution part of the Resnet network; using global adaptive smooth pooling and fully connected layers in sequence to construct the output part of the Resnet network; forming the network structure of the Resnet network according to the input part, the intermediate convolution part and the output part; using training data to input into the untrained Resnet network for iterative training to obtain the Resnet network, the training data including sample key areas and sample analysis results corresponding to the sample key areas.

[0011] Optionally, when the analysis result indicates that there is a counter violation, after issuing an alarm message, the method further includes: when the second video frame is not the last video frame among all the video frames, repeating the acquisition step, the extraction step, the first analysis step, the second analysis step and the alarm step at least once in sequence until all the video frames have completed detection and analysis.

[0012] According to another aspect of the present application, a double-person counter detection device for a bank counter is provided, the device comprising: an acquisition unit, used to execute an acquisition step, to acquire a first video frame and a second video frame, the first video frame being the current video frame, and the second video frame being the next video frame of the first video frame; a first extraction unit, used to execute an extraction step, to extract a plurality of first key areas and a plurality of second key areas corresponding to the first video frame and the second video frame respectively according to a position feature set, the position feature set representing a set of features of the effective positions of the counter and the teller on the image, the effective position being a position that meets set requirements, the first key area being an area corresponding to when the effective positions of the counter and the teller in the first video frame meet the position feature set, the first key area being an area corresponding to when the effective positions of the counter and the teller in the second video frame meet the position feature set, and there is a one-to-one correspondence between the first key area and the second key area; a first analysis unit, used to execute a first analysis step, to perform feature change analysis on all the first key areas and the corresponding second key areas, to obtain to the corresponding feature change value, one of the first key areas and the corresponding one of the second key areas corresponds to one of the feature change values; a second analysis unit, used to perform the second analysis step, when any of the feature change values ​​is greater than or equal to the mutation threshold, using the Resnet network to analyze the first target key area and the second target key area to obtain an analysis result, the first target key area is the first key area in the first video frame where the feature change value is greater than or equal to the mutation threshold, the second target key area is the second key area in the second video frame where the feature change value is greater than or equal to the mutation threshold, the analysis result is one of the following: there is a counter violation or there is no counter violation, the counter violation is at least that there is only one teller or no teller on duty for a continuous time, the teller is in an invalid on-duty area, and the door is in a non-closed state when no one is on duty or during non-working hours; an alarm unit, used to perform the alarm step, when the analysis result is that there is a counter violation, an alarm message is issued, and the alarm message is used to prompt that there is a safety hazard in the current counter.

[0013] According to another aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute any one of the methods described.

[0014] According to another aspect of the present application, a computer program product is provided, comprising computer instructions, wherein when the computer instructions are executed by a processor, any one of the methods described above is implemented.

[0015] Applying the technical solution of the present application, in a method for detecting two people at a bank counter, first, an acquisition step is performed to acquire a first video frame and a second video frame, wherein the first video frame is a current video frame, and the second video frame is a video frame next to the first video frame; then, an extraction step is performed to extract a plurality of first key areas and a plurality of second key areas corresponding to the first video frame and the second video frame, respectively, according to a position feature set, wherein the position feature set represents a set of features of the effective positions of the counter and the teller on the image, wherein the effective positions are positions that meet set requirements, wherein the first key area is an area corresponding to when the effective positions of the counter and the teller in the first video frame meet the position feature set, wherein the first key area is an area corresponding to when the effective positions of the counter and the teller in the second video frame meet the position feature set, and there is a one-to-one correspondence between the first key area and the second key area; then, a first analysis step is performed to perform feature change analysis on all the first key areas and the corresponding second key areas to obtain corresponding feature change values. , one of the above-mentioned first key areas and the corresponding one of the above-mentioned second key areas correspond to one of the above-mentioned feature change values; then, in the second analysis step, when any of the above-mentioned feature change values ​​is greater than or equal to the mutation threshold, the Resnet network is used to analyze the first target key area and the second target key area to obtain an analysis result, the above-mentioned first target key area is the above-mentioned first key area in the above-mentioned first video frame where the above-mentioned feature change value is greater than or equal to the mutation threshold, the above-mentioned second target key area is the above-mentioned second key area in the above-mentioned second video frame where the above-mentioned feature change value is greater than or equal to the mutation threshold, and the above-mentioned analysis result is one of the following: there is a counter violation or there is no counter violation, and the above-mentioned counter violation is at least that there is only one of the above-mentioned teller or no teller is on duty for a continuous time, the above-mentioned teller is located in an invalid on-duty area, and the counter door is in a non-closed state when no one is on duty or during non-working hours; finally, in the alarm step, when the above-mentioned analysis result is that the above-mentioned counter violation exists, an alarm message is issued, and the above-mentioned alarm message is used to prompt that there is a safety hazard in the current counter. Before performing image analysis, this application first extracts the effective area of ​​the counter, uses the edge detection operator to extract the counter position features, takes the effective position area of ​​each counter and teller as the key area, splits the image into multiple key areas, and then uses the Resnet network to perform AI analysis on the key areas in each frame of the image. Rapidly compare the color and texture of the relative key areas of consecutive frames. If the feature change is less than the mutation threshold, it is considered that the change in the consecutive frames is not large, and no AI analysis is required, which improves the detection efficiency. It can also be combined with the number of people and the state of the cabinet door for analysis. If a violation of the two-person guarding occurs within the threshold time, an alarm will be issued.The present application solves the problem in the prior art that when analyzing video frames, the entire image is directly taken as input and each image of continuous video frames is directly analyzed, thereby increasing additional computing cost and time in non-critical areas and causing waste of resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A hardware structure block diagram of a mobile terminal for executing a two-person visit detection method at a bank counter provided in an embodiment of the present application is shown;

[0017] Figure 2 A schematic diagram of a process of a two-person counter detection method at a bank counter provided according to an embodiment of the present application is shown;

[0018] Figure 3 A schematic diagram of a flow chart of a method for detecting two people at a bank counter provided according to an embodiment of the present application is shown;

[0019] Figure 4 A schematic diagram of a residual learning structure of a Resnet network provided according to an embodiment of the present application is shown;

[0020] Figure 5 A schematic diagram of a residual learning structure of a specific Resnet network provided according to an embodiment of the present application is shown;

[0021] Figure 6 A structural block diagram of a double-person visit detection device for a bank counter provided according to an embodiment of the present application is shown.

[0022] The above drawings include the following reference numerals:

[0023] 102, processor; 104, memory; 106, transmission device; 108, input and output devices. DETAILED DESCRIPTION

[0024] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0025] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0027] As introduced in the background technology, in the prior art (1) when analyzing video frames, the entire image is directly taken as input, and additional computing cost and time are required for non-critical areas, which is not conducive to the recognition of video images in continuous time; (2) AI analysis is directly performed on each image of continuous video frames, which consumes a lot of computing resources for scenes with normal behavior most of the time, resulting in resource waste; (3) there is a lack of comprehensive consideration of the compliance status of tellers and counter doors, and the focus is mostly on human body recognition. In order to solve the problem in the prior art that when analyzing video frames, the entire image is directly taken as input and each image of continuous video frames is directly analyzed, thereby increasing additional computing cost and time in non-critical areas and causing resource waste, the embodiments of the present application provide a bank counter double-person counter detection method, a bank counter double-person counter detection device, a computer-readable storage medium and a computer program product.

[0028] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0029] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 1 is a hardware structure block diagram of a mobile terminal for a method for detecting two people at a bank counter according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is for illustration only and does not limit the structure of the mobile terminal. Figure 1More or fewer components as shown, or with Figure 1 Different configurations are shown.

[0030] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the double-person counter detection method of the bank counter in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof. The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0031] In this embodiment, a method for detecting two people visiting a bank counter is provided, which runs on a mobile terminal, a computer terminal or a similar computing device. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0032] Figure 2 FIG. 1 is a flow chart of a method for detecting two people at a bank counter according to an embodiment of the present application. Figure 2 As shown, the method comprises the following steps:

[0033] Step S201, an acquisition step, acquires a first video frame and a second video frame, wherein the first video frame is a current video frame, and the second video frame is a next video frame of the first video frame.

[0034] Specifically, the system acquires the video stream captured by the surveillance camera in real time and extracts the current video frame (first video frame) and the next video frame (second video frame) from it. It ensures that the first video frame and the second video frame are continuous in time, so that the changes between the two consecutive frames can be compared. By acquiring continuous video frames, the system can detect behavioral changes of the counter and the teller, providing a data basis for subsequent analysis.

[0035] Step S202, an extraction step, extracting a plurality of first key areas and a plurality of second key areas corresponding to the first video frame and the second video frame respectively according to a position feature set, wherein the position feature set represents a set of features of the effective positions of the counter and the teller on the image, wherein the effective positions are positions that meet the set requirements, and the first key area is an area corresponding to when the effective positions of the counter and the teller in the first video frame meet the position feature set, and the first key area is an area corresponding to when the effective positions of the counter and the teller in the second video frame meet the position feature set, and there is a one-to-one correspondence between the first key area and the second key area.

[0036] Specifically, key areas that match the location feature set are extracted from the first video frame and the second video frame, respectively, and these areas correspond to the effective positions of the counter and the teller. By extracting key areas, the system can focus on the most important parts of the video and improve the efficiency and accuracy of the analysis.

[0037] Step S203, a first analysis step, performs feature change analysis on all the first key areas and the corresponding second key areas to obtain corresponding feature change values. One first key area and one corresponding second key area correspond to one feature change value.

[0038] Specifically, Figure 3 As shown, firstly, the color and texture feature changes of the key areas in the continuous frames are compared, and the feature change analysis is performed on all the extracted first key areas and the corresponding second key areas, and the feature change value is calculated. Each first key area and the corresponding second key area has a feature change value, which indicates the degree of change between the first video frame and the second video frame. The feature change value can help the system identify significant changes between key areas and provide a basis for subsequent violation detection. Among them, the continuous frames represent two consecutive video frames, namely the first video frame and the second video frame mentioned above.

[0039] Step S204, the second analysis step, when any one of the above-mentioned feature change values ​​is greater than or equal to the mutation threshold, the Resnet network is used to analyze the first target key area and the second target key area to obtain an analysis result, the above-mentioned first target key area is the above-mentioned first key area in the above-mentioned first video frame where the above-mentioned feature change value is greater than or equal to the mutation threshold, the above-mentioned second target key area is the above-mentioned second key area in the above-mentioned second video frame where the above-mentioned feature change value is greater than or equal to the mutation threshold, and the above-mentioned analysis result is one of the following: there is a counter violation or there is no counter violation, and the above-mentioned counter violation is at least that there is only one of the above-mentioned teller or there is no teller on duty for a continuous time, the above-mentioned teller is located in an invalid on-duty area, and the counter door is in a non-closed state when there is no one on duty or during non-working hours.

[0040] Specifically, the color and texture feature changes of key areas in consecutive frames are compared. If the feature changes of a key area in a consecutive frame exceed the mutation threshold, it is necessary to further analyze the key area in depth and use the Resnet network for AI analysis to determine whether there is any violation and obtain the analysis results. The Resnet network can provide a deeper level of feature analysis and improve the accuracy of violation detection. By setting the mutation threshold, the system can filter out unimportant changes and reduce false positives.

[0041] Step S205, an alarm step, when the above analysis result shows that the above counter violation exists, an alarm message is issued, and the above alarm message is used to prompt that there is a safety hazard in the current counter.

[0042] Specifically, according to the analysis results of the Resnet network, it is determined whether there is any counter violation. If the analysis results show that there is a violation, the system will issue an alarm message. The alarm information can promptly notify relevant personnel to take corresponding security measures to improve the safety of the counter. Through real-time monitoring and alarms, violations can be effectively prevented and reduced, and the asset security of customers and banks can be protected.

[0043] In this embodiment, first, an acquisition step is performed to acquire a first video frame and a second video frame, wherein the first video frame is the current video frame, and the second video frame is the next video frame of the first video frame; then, an extraction step is performed to extract a plurality of first key areas and a plurality of second key areas corresponding to the first video frame and the second video frame respectively according to a position feature set, wherein the position feature set represents a set of features of the effective positions of the counter and the teller on the image, wherein the effective positions are positions that meet the set requirements, wherein the first key area is an area corresponding to when the effective positions of the counter and the teller in the first video frame meet the position feature set, and wherein the first key area is an area corresponding to when the effective positions of the counter and the teller in the second video frame meet the position feature set, and there is a one-to-one correspondence between the first key area and the second key area; then, a first analysis step is performed to perform feature change analysis on all the first key areas and the corresponding second key areas to obtain corresponding feature change values, wherein one of the first key areas And a corresponding second key area corresponds to one of the above-mentioned feature change values; then, a second analysis step, when any of the above-mentioned feature change values ​​is greater than or equal to the mutation threshold, the Resnet network is used to analyze the first target key area and the second target key area to obtain an analysis result, the above-mentioned first target key area is the above-mentioned first key area in the above-mentioned first video frame where the above-mentioned feature change value is greater than or equal to the mutation threshold, the above-mentioned second target key area is the above-mentioned second key area in the above-mentioned second video frame where the above-mentioned feature change value is greater than or equal to the mutation threshold, and the above-mentioned analysis result is one of the following: there is a counter violation or there is no counter violation, and the above-mentioned counter violation is at least that there is only one of the above-mentioned teller or no teller is on duty for a continuous time, the above-mentioned teller is located in an invalid on-duty area, and the counter door is in a non-closed state when no one is on duty or during non-working hours; finally, an alarm step, when the above-mentioned analysis result is that the above-mentioned counter violation exists, an alarm message is issued, and the above-mentioned alarm message is used to prompt that there is a safety hazard in the current counter. Before performing image analysis, this application first extracts the effective area of ​​the counter, uses the edge detection operator to extract the counter position features, takes the effective position area of ​​each counter and teller as the key area, splits the image into multiple key areas, and then uses the Resnet network to perform AI analysis on the key areas in each frame of the image. Rapidly compare the color and texture of the relative key areas of consecutive frames. If the feature change is less than the mutation threshold, it is considered that the change in the consecutive frames is not large, and no AI analysis is required, which improves the detection efficiency. It can also be combined with the number of people and the state of the cabinet door for analysis. If a violation of the two-person guarding occurs within the threshold time, an alarm will be issued.The present application solves the problem in the prior art that when analyzing video frames, the entire image is directly taken as input and each image of continuous video frames is directly analyzed, thereby increasing additional computing cost and time in non-critical areas and causing waste of resources.

[0044] In order to enable those skilled in the art to more clearly understand the technical solution of the present application, the implementation process of the double-person counter detection method at a bank counter of the present application will be described in detail below in conjunction with specific embodiments.

[0045] In order to determine the key areas and reduce the computational cost and time cost of identifying the non-key areas, in an optional implementation, after the above step S201, the method further includes:

[0046] Step S301, when the first video frame is the first video frame among all the video frames, an edge detection operator method is used to extract all valid position features, where the valid position features are features representing the valid positions of the counter and the teller on the image;

[0047] Step S302, dividing the area based on all the above-mentioned valid position features with the above-mentioned counters as units to obtain the above-mentioned position feature set, wherein a position feature in the above-mentioned position feature set corresponds to a key area in the above-mentioned video frame, and the number of the above-mentioned position features is the same as the number of the above-mentioned counters.

[0048] In the above embodiment, if Figure 3 As shown, it is determined whether the first video frame is the first frame. If so, the edge detection operator is used to extract the effective position features of the counter according to the first frame of the video frame. The edge detection operator is a discrete difference operator used for edge detection, which is used to calculate the grayscale approximation of the image brightness function. After calculating the image of the first frame, the effective position features are obtained, and the effective position features are divided into counters to obtain a set of position features with the same number as the counters. The set of position features is the key area and remains unchanged in subsequent video frame processing. Key area detection can narrow the image range and analyze the counter status more accurately and efficiently in consecutive frames.

[0049] In order to extract effective position features, reduce the amount of data for subsequent processing, and improve detection efficiency, in an optional implementation, the above step S301 includes:

[0050] Step S3011, performing a plane convolution operation on the first video frame with the horizontal matrix and the vertical matrix image, respectively, to obtain a first gradient approximation value and a second gradient approximation value, wherein the first gradient approximation value is a gradient approximation value of each pixel point in the first video frame in the horizontal direction, and the first gradient approximation value is a gradient approximation value of each pixel point in the first video frame in the vertical direction;

[0051] Step S3012, calculating the gradient size of each pixel point according to the first gradient approximation value and the second gradient approximation value of each pixel point according to the first formula, and the expression of the first formula is: G x represents the first gradient approximation mentioned above, G y represents the second gradient approximation, and G represents the gradient size;

[0052] Step S3013, calculating the gradient direction of each pixel point according to the second formula, the second formula is expressed as θ represents the above gradient direction;

[0053] Step S3014, identifying all edges in the first video frame according to the gradient magnitude and the gradient direction of each pixel point;

[0054] Step S3015: determining the position features formed by all the identified edges as the valid position features.

[0055] In the above embodiment, the principle of the operator is to use pixel points to calculate the corresponding gradient vector and the norm of the vector, and to realize the edge detection in the corresponding directions in the horizontal and vertical directions based on image convolution. The operator contains two sets of 3x3 matrices, one for horizontal and one for vertical. By performing a plane convolution with the image, the horizontal and vertical brightness difference approximations can be obtained respectively. It should be noted that the horizontal matrix and the vertical matrix: these two matrices are predefined and are used to perform convolution operations on the image to extract the gradient information in the horizontal and vertical directions. If A represents the original image, Gx and Gy represent the image grayscale values ​​after horizontal and vertical edge detection respectively (that is, the above-mentioned horizontal approximation and vertical gradient approximation), and the formula is as follows: and The approximate horizontal and vertical gradients of each pixel of the image can be combined using the following formula to calculate the magnitude of the gradient. Then calculate the gradient direction: If the angle θ is equal to zero, it means that the image has a vertical edge at that location, and the left side is darker than the right side. The first formula (usually the Sobel operator or other gradient operator) is used to combine the horizontal and vertical gradient approximations to calculate the gradient magnitude of each pixel. The gradient magnitude can quantify the strength of the edge, which is very useful for distinguishing strong edges from weak edges. Larger gradient values ​​usually correspond to real edges, while smaller values ​​may be noise, which helps with denoising. The second formula (such as the inverse tangent function) is used to calculate the gradient direction of each pixel. The gradient direction provides information about the edge direction, which is crucial for understanding the image structure and conducting further image analysis. Based on the gradient magnitude and direction, a threshold is set to determine which pixels belong to the edge. Connect adjacent edge pixels to form a continuous edge. By identifying pixels whose gradient magnitude exceeds the threshold, the edges in the image can be detected. Continuous edges can form the outline of objects in the image, and the position features formed by the identified edges are recorded. By identifying effective features, the amount of data for subsequent processing can be reduced, the efficiency of the algorithm can be improved, and useful structural information can be extracted from the video frame, laying the foundation for subsequent image analysis and processing.

[0056] In order to improve the recognition efficiency of two people visiting the counter, in an optional implementation, the above step S203 includes:

[0057] Step S2031, using Euclidean distance to calculate the similarity between the color features and texture features of all the first key areas and the corresponding second key areas, to obtain the corresponding feature change values.

[0058] In the above embodiment, in order to improve the recognition efficiency, the color and texture feature changes of the key areas of the video frames are compared. Euclidean distance is a similarity calculation method used to measure the similarity between two vectors. In the present invention, the Euclidean distance similarity is used to compare the pixel values ​​of two images to determine their similarity. The formula is as follows: k represents the total number of pixels in the key area, x i Represents the i-th pixel of the first video frame, y iRepresents the i-th pixel of the second video frame. First, we need to calculate the Euclidean distance of the color features and texture features between all the first key areas and the corresponding second key areas. Euclidean distance is a commonly used feature similarity measurement method, which can measure the distance between two feature vectors. The smaller the distance, the higher the similarity. For color features, color histograms or color moments can be used to represent the color distribution of each area. Then, the Euclidean distance of the color features between each first key area and the corresponding second key area is calculated. For texture features, we can use gray-level co-occurrence matrices or wavelet transforms to represent the texture features of each area. Similarly, the Euclidean distance of the texture features between each first key area and the corresponding second key area is calculated. The similarity of the color features and texture features between each first key area and the corresponding second key area can be obtained, and the average value of the similarity of the color features and texture features can be calculated to obtain the corresponding feature change value, or the feature with a larger similarity can be selected as the feature change value. By calculating the similarity and feature change values ​​between features, the relationship between the color and texture features of different areas in the image can be enhanced, and these feature change values ​​are used as reference indicators for counter detection. That is, if the changes in the key area of ​​consecutive frames exceed the mutation threshold, it is considered that the characteristics of the key area have changed significantly, and the current frame needs to be combined with deep learning and subsequent processing; if it does not exceed the mutation threshold, it is considered that there is no obvious change and no further analysis is required.

[0059] In order to improve the efficiency of handling illegal behaviors, in an optional implementation manner, after the above step S203, the method further includes:

[0060] Step S401: When all of the above characteristic change values ​​are less than the above mutation threshold, it is determined that the above counter violation does not exist.

[0061] In the above embodiment, if the threshold requirement is not met, it is considered that the key area has no obvious feature changes, indicating that there is no counter violation and no violation processing is required. The system can detect potential risks and abnormal behaviors in a timely manner, and take corresponding measures to control them.

[0062] In order to further improve recognition efficiency and accuracy, in an optional implementation, before the above step S204, the method further includes:

[0063] Step S501, constructing the input part of the above Resnet network using a convolution kernel of a first set size and a maximum pooling layer of a second set size;

[0064] Step S502, using a plurality of the same convolution kernels of the third set size or a plurality of different convolution kernels to construct the intermediate convolution part of the Resnet network;

[0065] Step S503, using global adaptive smooth pooling and fully connected layers in sequence to construct the output part of the above Resnet network;

[0066] Step S504, forming a network structure of the Resnet network according to the input part, the intermediate convolution part and the output part;

[0067] Step S505, using training data to input into the untrained Resnet network for iterative training to obtain the Resnet network, wherein the training data includes sample key areas and sample analysis results corresponding to the sample key areas.

[0068] In the above embodiment, Net neural network belongs to deep residual network, and it is proposed to use residual learning to solve the problem of degradation of deep learning network. For a structure composed of several layers, when the input is X, the learned feature is H(X). In order to learn the residual F(X)=H(X)-X, the original learning feature is F(X)=H(X)+X. The structure of residual learning is as follows Figure 4 As shown, the right side is a short-circuit connection, that is, X is also used as the input of the relu layer, and the left side is the learned residual F(X). The final output of the relu layer is F(X) = H(X) + X. The Net neural network structure is mainly divided into an input part, an intermediate convolution part, and an output part. After the data enters the input part, it passes through the convolution part, and finally passes through the average pooling and fully connected layer to get the output. The input part consists of a convolution kernel of size 7×7 and a stride of 2, and a maximum pooling of size 3×3 and a stride of 2, completing the first step of image feature extraction. The intermediate convolution part is a multi-layer stack of 3×3 convolutions to achieve information extraction, which also includes residual learning, such as Figure 5 As shown on the left side of , the input data is divided into two paths, one through two 3×3 convolutions, and the other through a direct short circuit. The two are added together and output through relu. Figure 5 As shown on the right side of , when the network is deeper, three layers of residual learning are used, and the convolution kernels are 1×1, 3×3, and 1×1 respectively. The 1×1 convolution kernel can realize the linear combination of multiple feature maps while maintaining the size of the original feature map, greatly reducing the computational complexity. In the network output part, global adaptive smoothing pooling is performed, and then the fully connected layer outputs. The number of output nodes is consistent with the number of predicted categories. For frames with obvious changes in key areas, the Resnet network is used to further analyze whether violations have occurred. The original image of the key area is used as the input of the Resnet network, and the analysis result is used as the output of the Resnet network. The test set and training set are constructed based on the teller and counter location area, and the network is built and trained to form a Resnet network recognition model. Then, the key areas that exceed the mutation threshold are judged to identify violations.

[0069] In order to ensure that no possible exception is missed, in an optional implementation, after the above step S205, the method further includes:

[0070] Step S601, when the second video frame is not the last one of all the video frames, repeat the acquisition step, the extraction step, the first analysis step, the second analysis step and the alarm step at least once in sequence until all the video frames have completed detection and analysis.

[0071] In the above embodiment, it is a cyclic process for continuously detecting and analyzing a series of video frames until all video frames are processed. Repeat the above acquisition step, the above extraction step, the above first analysis step, the above second analysis step and the above alarm step until all video frames are processed. This means that the system will continuously monitor the video stream and analyze each frame in real time to ensure that no possible anomalies are missed. It is possible to monitor the video stream in real time and promptly detect and respond to potential security threats or incidents. Reduce the need for manual monitoring and improve efficiency and accuracy through automated analysis. This cyclic process ensures the continuity and integrity of the video frames, allowing the system to conduct a comprehensive analysis of the entire video sequence, thereby improving the accuracy and efficiency of detection.

[0072] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0073] The embodiment of the present application also provides a two-person counter detection device for a bank counter. It should be noted that the two-person counter detection device for a bank counter in the embodiment of the present application can be used to execute the two-person counter detection method for a bank counter provided in the embodiment of the present application. The device is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.

[0074] The following is an introduction to the double-person visit detection device for a bank counter provided in an embodiment of the present application.

[0075] Figure 6 1 is a structural block diagram of a double-person counter detection device for a bank counter according to an embodiment of the present application. Figure 6 As shown, the device comprises:

[0076] The acquisition unit 10 is used to execute the acquisition step to acquire a first video frame and a second video frame, wherein the first video frame is a current video frame, and the second video frame is a next video frame of the first video frame.

[0077] Specifically, the system acquires the video stream captured by the surveillance camera in real time and extracts the current video frame (first video frame) and the next video frame (second video frame) from it. It ensures that the first video frame and the second video frame are continuous in time, so that the changes between the two consecutive frames can be compared. By acquiring continuous video frames, the system can detect behavioral changes of the counter and the teller, providing a data basis for subsequent analysis.

[0078] The first extraction unit 20 is used to perform the extraction step, and extract multiple first key areas and multiple second key areas corresponding to the first video frame and the second video frame respectively according to the position feature set, the position feature set represents a set of features of the effective positions of the counter and the teller on the image, the effective position is a position that meets the set requirements, the first key area is the area corresponding to the effective position of the counter and the teller in the first video frame when it meets the position feature set, the first key area is the area corresponding to the effective position of the counter and the teller in the second video frame when it meets the position feature set, and there is a one-to-one correspondence between the first key area and the second key area.

[0079] Specifically, key areas that match the location feature set are extracted from the first video frame and the second video frame, respectively, and these areas correspond to the effective positions of the counter and the teller. By extracting key areas, the system can focus on the most important parts of the video and improve the efficiency and accuracy of the analysis.

[0080] The first analysis unit 30 is used to execute the first analysis step, perform feature change analysis on all the above-mentioned first key areas and the corresponding above-mentioned second key areas, and obtain corresponding feature change values. One above-mentioned first key area and one corresponding above-mentioned second key area correspond to one above-mentioned feature change value.

[0081] Specifically, Figure 3 As shown, firstly, the color and texture feature changes of the key areas in the continuous frames are compared, and the feature change analysis is performed on all the extracted first key areas and the corresponding second key areas, and the feature change value is calculated. Each first key area and the corresponding second key area has a feature change value, which indicates the degree of change between the first video frame and the second video frame. The feature change value can help the system identify significant changes between key areas and provide a basis for subsequent violation detection. Among them, the continuous frames represent two consecutive video frames, namely the first video frame and the second video frame mentioned above.

[0082] The second analysis unit 40 is used to perform the second analysis step. When any one of the above-mentioned feature change values ​​is greater than or equal to the mutation threshold, the Resnet network is used to analyze the first target key area and the second target key area to obtain an analysis result. The above-mentioned first target key area is the above-mentioned first key area in the above-mentioned first video frame where the above-mentioned feature change value is greater than or equal to the mutation threshold, and the above-mentioned second target key area is the above-mentioned second key area in the above-mentioned second video frame where the above-mentioned feature change value is greater than or equal to the mutation threshold. The above-mentioned analysis result is one of the following: there is a counter violation or there is no counter violation. The above-mentioned counter violation is at least that there is only one of the above-mentioned teller or no teller is on duty for a continuous time, the above-mentioned teller is located in an invalid on-duty area, and the counter door is in a non-closed state when there is no one on duty or during non-working hours.

[0083] Specifically, the color and texture feature changes of key areas in consecutive frames are compared. If the feature changes of a key area in a consecutive frame exceed the mutation threshold, it is necessary to further analyze the key area in depth and use the Resnet network for AI analysis to determine whether there is any violation and obtain the analysis results. The Resnet network can provide a deeper level of feature analysis and improve the accuracy of violation detection. By setting the mutation threshold, the system can filter out unimportant changes and reduce false positives.

[0084] The alarm unit 50 is used to execute the alarm step, and when the above analysis result shows that the above counter violation exists, an alarm message is issued, and the above alarm message is used to prompt that there is a safety hazard in the current counter.

[0085] Specifically, according to the analysis results of the Resnet network, it is determined whether there is any violation at the counter. If the analysis results indicate that there is a violation, the system will issue an alarm message. The alarm message can promptly notify relevant personnel to take corresponding security measures to improve the safety of the counter. Through real-time monitoring and alarms, violations can be effectively prevented and reduced to protect the interests of customers and banks.

[0086] In this embodiment, an acquisition unit is used to execute an acquisition step to acquire a first video frame and a second video frame, wherein the first video frame is a current video frame, and the second video frame is a video frame next to the first video frame; a first extraction unit is used to execute an extraction step to extract a plurality of first key areas and a plurality of second key areas corresponding to the first video frame and the second video frame according to a position feature set, wherein the position feature set represents a set of features of the effective positions of the counter and the teller on the image, wherein the effective positions are positions that meet set requirements, wherein the first key area is an area corresponding to when the effective positions of the counter and the teller in the first video frame meet the position feature set, wherein the first key area is an area corresponding to when the effective positions of the counter and the teller in the second video frame meet the position feature set, and there is a one-to-one correspondence between the first key area and the second key area; a first analysis unit is used to execute a first analysis step to perform feature change analysis on all the first key areas and the corresponding second key areas to obtain corresponding feature change values, wherein one of the first key areas is a region corresponding to when the effective positions of the counter and the teller in the second video frame meet the position feature set. A key area and a corresponding second key area correspond to one of the above-mentioned feature change values; a second analysis unit is used to execute the second analysis step, and when any of the above-mentioned feature change values ​​is greater than or equal to the mutation threshold, the first target key area and the second target key area are analyzed by using the Resnet network to obtain an analysis result, the above-mentioned first target key area is the above-mentioned first key area in the above-mentioned first video frame where the above-mentioned feature change value is greater than or equal to the mutation threshold, and the above-mentioned second target key area is the above-mentioned second key area in the above-mentioned second video frame where the above-mentioned feature change value is greater than or equal to the mutation threshold, and the above-mentioned analysis result is one of the following: there is a counter violation or there is no counter violation, and the above-mentioned counter violation is at least that there is only one of the above-mentioned teller or no teller is on duty for a continuous time, the above-mentioned teller is located in an invalid on-duty area, and the counter door is in a non-closed state when no one is on duty or during non-working hours; an alarm unit is used to execute the alarm step, and when the above-mentioned analysis result is that the above-mentioned counter violation exists, an alarm message is issued, and the above-mentioned alarm message is used to prompt that there is a safety hazard in the current counter. Before performing image analysis, the application first extracts the effective area of ​​the counter, uses the edge detection operator to extract the counter position features, takes the common location area of ​​each counter and the teller as the key area, splits the image into multiple key areas, and then uses the Resnet network to perform AI analysis on the key areas in each frame. The color and texture of the relative key areas of consecutive frames are quickly compared. If the feature change is less than the mutation threshold, it is considered that the change in the consecutive frames is not large, and no AI analysis is required, which improves the detection efficiency. It can also be combined with the number of people and the state of the cabinet door for analysis. If a violation of the two-person guarding occurs within the threshold time, an alarm will be issued.The present application solves the problem in the prior art that when analyzing video frames, the entire image is directly taken as input and each image of continuous video frames is directly analyzed, thereby increasing additional computing cost and time in non-critical areas and causing waste of resources.

[0087] In order to determine the key area to reduce the computational cost and time cost of identifying the non-key area, in an optional implementation manner, the device further includes:

[0088] A second extraction unit is used for, after acquiring the first video frame and the second video frame, extracting all valid position features by using an edge detection operator method when the first video frame is the first video frame among all the video frames, wherein the valid position features are features representing the valid positions of the counter and the teller on the image;

[0089] The division unit is used to divide the area according to all the above-mentioned valid position features with the above-mentioned counters as units to obtain the above-mentioned position feature set, wherein a position feature in the above-mentioned position feature set corresponds to a key area in the above-mentioned video frame, and the number of the above-mentioned position features is the same as the number of the above-mentioned counters.

[0090] In the above embodiment, if Figure 3 As shown, it is determined whether the first video frame is the first frame. If so, the edge detection operator is used to extract the effective position features of the counter according to the first frame of the video frame. The edge detection operator is a discrete difference operator used for edge detection, which is used to calculate the grayscale approximation of the image brightness function. After calculating the image of the first frame, the effective position features are obtained, and the effective position features are divided into counters to obtain a set of position features with the same number as the counters. The set of position features is the key area and remains unchanged in subsequent video frame processing. Key area detection can narrow the image range and analyze the counter status more accurately and efficiently in consecutive frames.

[0091] In order to extract effective position features, reduce the amount of data for subsequent processing, and improve detection efficiency, in an optional implementation, the second extraction unit includes:

[0092] A convolution module, used for performing a plane convolution operation on the first video frame with a horizontal matrix image and a vertical matrix image, respectively, to obtain a first gradient approximation value and a second gradient approximation value, wherein the first gradient approximation value is a gradient approximation value of each pixel point in the first video frame in a horizontal direction, and the first gradient approximation value is a gradient approximation value of each pixel point in the first video frame in a vertical direction;

[0093] The first calculation module is used to calculate the gradient size of each pixel point according to the first gradient approximation value and the second gradient approximation value of each pixel point according to the first formula, and the expression of the first formula is: G x represents the first gradient approximation mentioned above, G y represents the second gradient approximation, and G represents the gradient size;

[0094] The second calculation module is used to calculate the gradient direction of each pixel point according to the second formula. The second formula is expressed as θ represents the above gradient direction;

[0095] An identification module, configured to identify all edges in the first video frame according to the gradient magnitude and the gradient direction of each pixel point;

[0096] The determination module is used to determine the position features formed by all the identified edges as the valid position features.

[0097] In the above embodiment, the principle of the operator is to use pixel points to calculate the corresponding gradient vector and the norm of the vector, and to realize the edge detection in the corresponding directions in the horizontal and vertical directions based on image convolution. The operator contains two sets of 3x3 matrices, one for horizontal and one for vertical. By performing a plane convolution with the image, the horizontal and vertical brightness difference approximations can be obtained respectively. It should be noted that the horizontal matrix and the vertical matrix: these two matrices are predefined and are used to perform convolution operations on the image to extract the gradient information in the horizontal and vertical directions. If A represents the original image, Gx and Gy represent the image grayscale values ​​after horizontal and vertical edge detection respectively (that is, the above-mentioned horizontal approximation and vertical gradient approximation), and the formula is as follows: and The approximate horizontal and vertical gradients of each pixel of the image can be combined using the following formula to calculate the magnitude of the gradient. Then calculate the gradient direction: If the angle θ is equal to zero, it means that the image has a vertical edge at that location, and the left side is darker than the right side. The first formula (usually the Sobel operator or other gradient operator) is used to combine the horizontal and vertical gradient approximations to calculate the gradient magnitude of each pixel. The gradient magnitude can quantify the strength of the edge, which is very useful for distinguishing strong edges from weak edges. Larger gradient values ​​usually correspond to real edges, while smaller values ​​may be noise, which helps with denoising. The second formula (such as the inverse tangent function) is used to calculate the gradient direction of each pixel. The gradient direction provides information about the edge direction, which is crucial for understanding the image structure and conducting further image analysis. Based on the gradient magnitude and direction, a threshold is set to determine which pixels belong to the edge. Connect adjacent edge pixels to form a continuous edge. By identifying pixels whose gradient magnitude exceeds the threshold, the edges in the image can be detected. Continuous edges can form the outline of objects in the image, and the position features formed by the identified edges are recorded. By identifying effective features, the amount of data for subsequent processing can be reduced, the efficiency of the algorithm can be improved, and useful structural information can be extracted from the video frame, laying the foundation for subsequent image analysis and processing.

[0098] In order to improve the recognition efficiency of two people visiting the counter, in an optional implementation manner, the first analysis unit includes:

[0099] The third calculation module is used to calculate the similarity between the color features and texture features of all the first key areas and the corresponding second key areas using Euclidean distance to obtain the corresponding feature change values.

[0100] In the above embodiment, in order to improve the recognition efficiency, the color and texture feature changes of the key areas of the video frames are compared. Euclidean distance is a similarity calculation method used to measure the similarity between two vectors. In the present invention, the Euclidean distance similarity is used to compare the pixel values ​​of two images to determine their similarity. The formula is as follows: k represents the total number of pixels in the key area, x i Represents the i-th pixel of the first video frame, y iRepresents the i-th pixel of the second video frame. First, we need to calculate the Euclidean distance of the color features and texture features between all the first key areas and the corresponding second key areas. Euclidean distance is a commonly used feature similarity measurement method, which can measure the distance between two feature vectors. The smaller the distance, the higher the similarity. For color features, color histograms or color moments can be used to represent the color distribution of each area. Then, the Euclidean distance of the color features between each first key area and the corresponding second key area is calculated. For texture features, we can use gray-level co-occurrence matrices or wavelet transforms to represent the texture features of each area. Similarly, the Euclidean distance of the texture features between each first key area and the corresponding second key area is calculated. The similarity of the color features and texture features between each first key area and the corresponding second key area can be obtained, and the average value of the similarity of the color features and texture features can be calculated to obtain the corresponding feature change value, or the feature with a larger similarity can be selected as the feature change value. By calculating the similarity and feature change values ​​between features, the relationship between the color and texture features of different areas in the image can be enhanced, and these feature change values ​​are used as reference indicators for counter detection. That is, if the changes in the key area of ​​consecutive frames exceed the mutation threshold, it is considered that the characteristics of the key area have changed significantly, and the current frame needs to be combined with deep learning and subsequent processing; if it does not exceed the mutation threshold, it is considered that there is no obvious change and no further analysis is required.

[0101] In order to improve the efficiency of handling illegal behaviors, in an optional implementation manner, the device further includes:

[0102] The determination unit is used to perform feature change analysis on all the above-mentioned first key areas and the corresponding above-mentioned second key areas to obtain corresponding feature change values, and then determine that the above-mentioned counter violation does not exist when all the above-mentioned feature change values ​​are less than the above-mentioned mutation threshold.

[0103] In the above embodiment, if the threshold requirement is not met, it is considered that the key area has no obvious feature changes, indicating that there is no counter violation and no violation processing is required. The system can detect potential risks and abnormal behaviors in a timely manner, and take corresponding measures to control them.

[0104] In order to further improve recognition efficiency and accuracy, in an optional implementation manner, the device further includes:

[0105] A first construction unit is used to use a Resnet network to analyze the first target key area and the second target key area when any of the above-mentioned feature change values ​​is greater than or equal to the mutation threshold, and before obtaining the analysis result, use a convolution kernel of a first set size and a maximum pooling layer of a second set size to construct an input part of the above-mentioned Resnet network;

[0106] A second construction unit is used to construct an intermediate convolution part of the Resnet network by using a plurality of the convolution kernels of the same third set size or a plurality of different convolution kernels;

[0107] The third construction unit is used to sequentially use global adaptive smooth pooling and a fully connected layer to construct the output part of the above Resnet network;

[0108] A fourth construction unit, used to form a network structure of the Resnet network according to the input part, the intermediate convolution part and the output part;

[0109] A training unit is used to use training data to input into the untrained Resnet network for iterative training to obtain the Resnet network, wherein the training data includes sample key areas and sample analysis results corresponding to the sample key areas.

[0110] In the above embodiment, Net neural network belongs to deep residual network, and it is proposed to use residual learning to solve the problem of degradation of deep learning network. For a structure composed of several layers, when the input is X, the learned feature is H(X). In order to learn the residual F(X)=H(X)-X, the original learning feature is F(X)=H(X)+X. The structure of residual learning is as follows Figure 4 As shown, the right side is a short-circuit connection, that is, X is also used as the input of the relu layer, and the left side is the learned residual F(X). The final output of the relu layer is F(X) = H(X) + X. The Net neural network structure is mainly divided into an input part, an intermediate convolution part, and an output part. After the data enters the input part, it passes through the convolution part, and finally passes through the average pooling and fully connected layer to get the output. The input part consists of a convolution kernel of size 7×7 and a stride of 2, and a maximum pooling of size 3×3 and a stride of 2, completing the first step of image feature extraction. The intermediate convolution part is a multi-layer stack of 3×3 convolutions to achieve information extraction, which also includes residual learning, such as Figure 5 As shown on the left side of , the input data is divided into two paths, one through two 3×3 convolutions, and the other through a direct short circuit. The two are added together and output through relu. Figure 5As shown on the right side of , when the network is deeper, three layers of residual learning are used, and the convolution kernels are 1×1, 3×3, and 1×1 respectively. The 1×1 convolution kernel can realize the linear combination of multiple feature maps while maintaining the size of the original feature map, greatly reducing the computational complexity. In the network output part, global adaptive smoothing pooling is performed, and then the fully connected layer outputs. The number of output nodes is consistent with the number of predicted categories. For frames with obvious changes in key areas, the Resnet network is used to further analyze whether violations have occurred. The original image of the key area is used as the input of the Resnet network, and the analysis result is used as the output of the Resnet network. The test set and training set are constructed based on the teller and counter location area, and the network is built and trained to form a Resnet network recognition model. Then, the key areas that exceed the mutation threshold are judged to identify violations.

[0111] In order to ensure that no possible anomalies are missed, in an optional embodiment, the device further includes:

[0112] A repeating unit is used for, after issuing an alarm message when the above-mentioned analysis result shows that the above-mentioned counter violation exists, and when the above-mentioned second video frame is not the last one of all the above-mentioned video frames, repeating the above-mentioned acquisition step, the above-mentioned extraction step, the above-mentioned first analysis step, the above-mentioned second analysis step and the above-mentioned alarm step at least once in sequence until all the above-mentioned video frames have completed detection and analysis.

[0113] In the above embodiment, it is a cyclic process for continuously detecting and analyzing a series of video frames until all video frames are processed. Repeat the above acquisition step, the above extraction step, the above first analysis step, the above second analysis step and the above alarm step until all video frames are processed. This means that the system will continuously monitor the video stream and analyze each frame in real time to ensure that no possible anomalies are missed. It is possible to monitor the video stream in real time and promptly detect and respond to potential security threats or incidents. Reduce the need for manual monitoring and improve efficiency and accuracy through automated analysis. This cyclic process ensures the continuity and integrity of the video frames, allowing the system to conduct a comprehensive analysis of the entire video sequence, thereby improving the accuracy and efficiency of detection.

[0114] The above-mentioned double-person counter detection device for bank counters includes a processor and a memory, and the above-mentioned acquisition unit, the first extraction unit and the first analysis unit are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize corresponding functions. The above-mentioned modules are all located in the same processor; or, the above-mentioned modules are located in different processors in the form of any combination.

[0115] The processor includes a kernel, which calls the corresponding program unit from the memory. One or more kernels can be set, and the kernel parameters are adjusted to solve the problem in the prior art that when analyzing video frames, the entire image is directly used as input and each image of continuous video frames is directly analyzed, thereby increasing the calculation cost and time in non-critical areas and causing resource waste.

[0116] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0117] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program. When the program is running, the device where the computer-readable storage medium is located is controlled to execute the two-person counter detection method at the bank counter.

[0118] Specifically, the double-person counter detection method at the bank counter includes:

[0119] Step S201, an acquisition step, acquiring a first video frame and a second video frame, wherein the first video frame is a current video frame, and the second video frame is a next video frame of the first video frame;

[0120] Step S202, an extraction step, extracting a plurality of first key areas and a plurality of second key areas corresponding to the first video frame and the second video frame respectively according to a position feature set, wherein the position feature set represents a set of features of the effective positions of the counter and the teller on the image, wherein the effective positions are positions that meet the set requirements, and the first key areas are areas corresponding to the effective positions of the counter and the teller in the first video frame when they meet the position feature set, and the first key areas are areas corresponding to the effective positions of the counter and the teller in the second video frame when they meet the position feature set, and there is a one-to-one correspondence between the first key areas and the second key areas;

[0121] Step S203, a first analysis step, performing feature change analysis on all the first key areas and the corresponding second key areas to obtain corresponding feature change values, wherein one first key area and one corresponding second key area correspond to one feature change value;

[0122] Step S204, the second analysis step, in the case where any one of the above-mentioned feature change values ​​is greater than or equal to the mutation threshold, the Resnet network is used to analyze the first target key area and the second target key area to obtain an analysis result, the above-mentioned first target key area is the above-mentioned first key area in the above-mentioned first video frame where the above-mentioned feature change value is greater than or equal to the mutation threshold, the above-mentioned second target key area is the above-mentioned second key area in the above-mentioned second video frame where the above-mentioned feature change value is greater than or equal to the mutation threshold, and the above-mentioned analysis result is one of the following: there is a counter violation or there is no counter violation, and the above-mentioned counter violation is at least that there is only one of the above-mentioned tellers or no tellers are on duty in a continuous period of time, the above-mentioned tellers are located in an invalid on-duty area, and the counter door is in a non-closed state when no one is on duty or during non-working hours;

[0123] Step S205, an alarm step, when the above analysis result shows that the above counter violation exists, an alarm message is issued, and the above alarm message is used to prompt that there is a safety hazard in the current counter.

[0124] An embodiment of the present invention provides a processor, which is used to run a program, wherein the program executes the two-person counter detection method at the bank counter when running.

[0125] An embodiment of the present invention provides a double-person counter detection system, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, at least the following steps are implemented:

[0126] Step S201, an acquisition step, acquiring a first video frame and a second video frame, wherein the first video frame is a current video frame, and the second video frame is a next video frame of the first video frame;

[0127] Step S202, an extraction step, extracting a plurality of first key areas and a plurality of second key areas corresponding to the first video frame and the second video frame respectively according to a position feature set, wherein the position feature set represents a set of features of the effective positions of the counter and the teller on the image, wherein the effective positions are positions that meet the set requirements, and the first key areas are areas corresponding to the effective positions of the counter and the teller in the first video frame when they meet the position feature set, and the first key areas are areas corresponding to the effective positions of the counter and the teller in the second video frame when they meet the position feature set, and there is a one-to-one correspondence between the first key areas and the second key areas;

[0128] Step S203, a first analysis step, performing feature change analysis on all the first key areas and the corresponding second key areas to obtain corresponding feature change values, wherein one first key area and one corresponding second key area correspond to one feature change value;

[0129] Step S204, the second analysis step, in the case where any one of the above-mentioned feature change values ​​is greater than or equal to the mutation threshold, the Resnet network is used to analyze the first target key area and the second target key area to obtain an analysis result, the above-mentioned first target key area is the above-mentioned first key area in the above-mentioned first video frame where the above-mentioned feature change value is greater than or equal to the mutation threshold, the above-mentioned second target key area is the above-mentioned second key area in the above-mentioned second video frame where the above-mentioned feature change value is greater than or equal to the mutation threshold, and the above-mentioned analysis result is one of the following: there is a counter violation or there is no counter violation, and the above-mentioned counter violation is at least that there is only one of the above-mentioned tellers or no tellers are on duty in a continuous period of time, the above-mentioned tellers are located in an invalid on-duty area, and the counter door is in a non-closed state when no one is on duty or during non-working hours;

[0130] Step S205, an alarm step, when the above analysis result shows that the above counter violation exists, an alarm message is issued, and the above alarm message is used to prompt that there is a safety hazard in the current counter.

[0131] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program for initializing at least the following method steps:

[0132] Step S201, an acquisition step, acquiring a first video frame and a second video frame, wherein the first video frame is a current video frame, and the second video frame is a next video frame of the first video frame;

[0133] Step S202, an extraction step, extracting a plurality of first key areas and a plurality of second key areas corresponding to the first video frame and the second video frame respectively according to a position feature set, wherein the position feature set represents a set of features of the effective positions of the counter and the teller on the image, wherein the effective positions are positions that meet the set requirements, and the first key areas are areas corresponding to the effective positions of the counter and the teller in the first video frame when they meet the position feature set, and the first key areas are areas corresponding to the effective positions of the counter and the teller in the second video frame when they meet the position feature set, and there is a one-to-one correspondence between the first key areas and the second key areas;

[0134] Step S203, a first analysis step, performing feature change analysis on all the first key areas and the corresponding second key areas to obtain corresponding feature change values, wherein one first key area and one corresponding second key area correspond to one feature change value;

[0135] Step S204, the second analysis step, in the case where any one of the above-mentioned feature change values ​​is greater than or equal to the mutation threshold, the Resnet network is used to analyze the first target key area and the second target key area to obtain an analysis result, the above-mentioned first target key area is the above-mentioned first key area in the above-mentioned first video frame where the above-mentioned feature change value is greater than or equal to the mutation threshold, the above-mentioned second target key area is the above-mentioned second key area in the above-mentioned second video frame where the above-mentioned feature change value is greater than or equal to the mutation threshold, and the above-mentioned analysis result is one of the following: there is a counter violation or there is no counter violation, and the above-mentioned counter violation is at least that there is only one of the above-mentioned tellers or no tellers are on duty in a continuous period of time, the above-mentioned tellers are located in an invalid on-duty area, and the counter door is in a non-closed state when no one is on duty or during non-working hours;

[0136] Step S205, an alarm step, when the above analysis result shows that the above counter violation exists, an alarm message is issued, and the above alarm message is used to prompt that there is a safety hazard in the current counter.

[0137] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0138] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0139] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0140] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0141] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0142] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0143] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0144] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0145] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0146] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:

[0147] 1) The double-person counter detection method for bank counters of the present application, before image analysis, first extracts the effective counter area, uses the edge detection operator to extract the counter position features, takes the effective position area of ​​each counter and teller as the key area, splits the image into multiple key areas, and then uses the Resnet network to perform AI analysis on the key areas in each frame of the image. Rapidly compare the color and texture of the relative key areas of consecutive frames. If the feature change is less than the mutation threshold, it is considered that the change in the consecutive frames is not large, and no AI analysis is required, which improves the detection efficiency. It can also be combined according to the number of people and the state of the counter door. If a violation of the double-person counter occurs within the threshold time, an alarm is issued. The present application solves the problem in the prior art that the entire image is directly used as input when analyzing video frames and each image of the continuous video frames is directly analyzed, thereby increasing the computing cost and time in non-critical areas and causing waste of resources.

[0148] 2) The double-person counter detection device for bank counters of the present application first extracts the effective counter area before image analysis, uses the edge detection operator to extract the counter position features, takes the common position area of ​​each counter and the teller as the key area, splits the image into multiple key areas, and then uses the Resnet network to perform AI analysis on the key areas in each frame of the image. Rapidly compare the color and texture of the relative key areas of consecutive frames. If the feature change is less than the mutation threshold, it is considered that the change in the consecutive frames is not large, and no AI analysis is required, which improves the detection efficiency. It can also be combined according to the number of people and the state of the counter door. If a violation of the double-person counter occurs within the threshold time, an alarm is issued. The present application solves the problem in the prior art that the entire image is directly used as input when analyzing video frames and each image of the continuous video frames is directly analyzed, thereby increasing the computing cost and time in non-critical areas and causing waste of resources.

[0149] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for detecting two people at a bank counter, characterized in that: include: An acquisition step of acquiring a first video frame and a second video frame, wherein the first video frame is a current video frame, and the second video frame is a video frame next to the first video frame; An extraction step, respectively extracting a plurality of first key areas and a plurality of second key areas corresponding to the first video frame and the second video frame according to a position feature set, wherein the position feature set represents a set of features of the effective positions of the counter and the teller on the image, wherein the effective position is a position that meets the set requirements, and the first key area is an area corresponding to when the effective positions of the counter and the teller in the first video frame meet the position feature set, and the first key area is an area corresponding to when the effective positions of the counter and the teller in the second video frame meet the position feature set, and there is a one-to-one correspondence between the first key area and the second key area; A first analysis step is to perform feature change analysis on all the first key areas and the corresponding second key areas to obtain corresponding feature change values, wherein one first key area and one corresponding second key area correspond to one feature change value; The second analysis step, when any one of the feature change values ​​is greater than or equal to the mutation threshold, uses the Resnet network to analyze the first target key area and the second target key area to obtain an analysis result, wherein the first target key area is the first key area in the first video frame where the feature change value is greater than or equal to the mutation threshold, and the second target key area is the second key area in the second video frame where the feature change value is greater than or equal to the mutation threshold, and the analysis result is one of the following: there is a counter violation or there is no counter violation, and the counter violation is at least that there is only one teller or no teller on duty for a continuous period of time, the teller is located in an invalid on-duty area, and the counter door is in a non-closed state when no one is on duty or during non-working hours; The alarm step is to issue an alarm message when the analysis result shows that there is a violation of the counter. The alarm message is used to prompt that there is a safety hazard in the current counter.

2. The method according to claim 1, characterized in that After acquiring the first video frame and the second video frame, the method further includes: In the case where the first video frame is the first video frame among all the video frames, an edge detection operator method is used to extract all effective position features, where the effective position features are features representing the effective positions of the counter and the teller on the image; The area is divided according to all the valid position features using the counter as a unit to obtain the position feature set, wherein a position feature in the position feature set corresponds to a key area in the video frame, and the number of the position features is the same as the number of the counters.

3. The method according to claim 2, characterized in that The edge detection operator method is used to extract all effective position features, including: Performing a planar convolution operation on the first video frame with the horizontal matrix image and the vertical matrix image, respectively, to obtain a first gradient approximation value and a second gradient approximation value, respectively, wherein the first gradient approximation value is a gradient approximation value of each pixel point in the first video frame in the horizontal direction, and the first gradient approximation value is a gradient approximation value of each pixel point in the first video frame in the vertical direction; The gradient magnitude of each pixel is calculated according to the first gradient approximation value and the second gradient approximation value of each pixel according to a first formula, and the expression of the first formula is: G x represents the first gradient approximation, G y represents the second gradient approximation, and G represents the gradient magnitude; The gradient direction of each pixel is calculated according to the second formula, where the second formula represents θ represents the gradient direction; Identify all edges in the first video frame according to the gradient magnitude and the gradient direction of each pixel point; The position features formed by all the identified edges are determined as valid position features.

4. The method according to claim 1, characterized in that: Performing feature change analysis on all the first key areas and the corresponding second key areas to obtain corresponding feature change values ​​includes: The similarities between the color features and the texture features of all the first key areas and the corresponding second key areas are calculated using the Euclidean distance to obtain the corresponding feature change values.

5. The method according to claim 1, characterized in that After performing feature change analysis on all the first key areas and the corresponding second key areas to obtain corresponding feature change values, the method further includes: When all of the feature change values ​​are smaller than the mutation threshold, it is determined that no counter violation occurs.

6. The method according to claim 1, characterized in that In the case where any one of the feature change values ​​is greater than or equal to the mutation threshold, the Resnet network is used to analyze the first target key area and the second target key area, and before obtaining the analysis result, the method further includes: Constructing an input portion of the Resnet network using a convolution kernel of a first set size and a maximum pooling layer of a second set size; Using a plurality of the convolution kernels of the same third set size or a plurality of different convolution kernels to construct an intermediate convolution part of the Resnet network; The output part of the Resnet network is constructed by using global adaptive smooth pooling and fully connected layers in sequence; Forming a network structure of the Resnet network according to the input part, the intermediate convolution part and the output part; The training data is input into the untrained Resnet network for iterative training to obtain the Resnet network, wherein the training data includes a sample key area and a sample analysis result corresponding to the sample key area.

7. The method according to claim 1, characterized in that When the analysis result shows that the counter violation occurs, after issuing an alarm message, the method further includes: In the case where the second video frame is not the last one of all the video frames, the acquisition step, the extraction step, the first analysis step, the second analysis step and the alarm step are repeated in sequence at least once until all the video frames have completed detection and analysis.

8. A double-person counter detection device for a bank counter, characterized in that: The device comprises: An acquisition unit, configured to execute an acquisition step to acquire a first video frame and a second video frame, wherein the first video frame is a current video frame, and the second video frame is a video frame next to the first video frame; a first extraction unit, configured to perform an extraction step, and extract a plurality of first key areas and a plurality of second key areas corresponding to the first video frame and the second video frame respectively according to a position feature set, wherein the position feature set represents a set of features of the effective positions of the counter and the teller on the image, wherein the effective position is a position that meets set requirements, and the first key area is an area corresponding to when the effective positions of the counter and the teller in the first video frame meet the position feature set, and the first key area is an area corresponding to when the effective positions of the counter and the teller in the second video frame meet the position feature set, and there is a one-to-one correspondence between the first key area and the second key area; A first analysis unit is used to execute a first analysis step, perform feature change analysis on all the first key areas and the corresponding second key areas, and obtain corresponding feature change values, where one first key area and one corresponding second key area correspond to one feature change value; A second analysis unit is used to perform a second analysis step, and in the case where any one of the feature change values ​​is greater than or equal to the mutation threshold, a Resnet network is used to analyze the first target key area and the second target key area to obtain an analysis result, wherein the first target key area is the first key area in the first video frame where the feature change value is greater than or equal to the mutation threshold, and the second target key area is the second key area in the second video frame where the feature change value is greater than or equal to the mutation threshold, and the analysis result is one of the following: there is a counter violation or there is no counter violation, and the counter violation is at least that there is only one teller or no teller on duty for a continuous period of time, the teller is located in an invalid on-duty area, and the counter door is in a non-closed state when no one is on duty or during non-working hours; The alarm unit is used to execute the alarm step, and when the analysis result shows that there is a violation of the counter, an alarm message is issued, wherein the alarm message is used to prompt that there is a safety hazard in the current counter.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.

10. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.