Method and device for improving gate passing people counting based on video sequence analysis

By deploying a video classification network model based on the spatio-temporal orthogonal timing attention mechanism in the self-service security check system, combined with the event-driven two-stage sampling strategy, the problems of poor real-time response capabilities and large computing overhead in the existing technology are solved, and high-precision and low-latency video number statistics are achieved.

CN120388327APending Publication Date: 2025-07-29RECONOVA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510365103.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing self-service security inspection system has poor real-time response capabilities, low analysis efficiency, large storage requirements, large calculation overhead, and poor video recognition accuracy, which cannot meet the real-time requirements.

Method used

The video classification network model based on the space-time orthogonal timing attention mechanism is adopted, combined with the event-driven two-stage sampling strategy, samples are performed from the real-time video stream of the gate image acquisition device, and the 2Dnet spatial feature extraction network and the OCA3Dnet timing feature extraction network are extracted, spatial and timing feature information are extracted, and the 2Dsnet semantic feature extraction network is spliced, and counting results are finally obtained through classification.

Benefits of technology

It improves the accuracy and speed of the video recognition algorithm, meets real-time requirements, reduces calculation delay, and realizes real-time monitoring and alarms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388327A_ABST
    Figure CN120388327A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for improving gate passing people counting based on video sequence analysis, and the method comprises the steps: obtaining image frames from a video stream through employing an event-driven two-stage sampling strategy, sampling an image once every K frames, inputting the image frames into a 2Dnet space extraction network, extracting the space features of the image frames, and generating feature representation; the feature representations are stacked to form a queue, and when the number of sampling frames accumulated in the queue reaches a threshold N, the number is input into an OCA3Dnet time sequence feature extraction network. Meanwhile, feature representations in the 2Dnet space extraction network can be spliced into new input, the new input is sent to the 2Dsnet semantic feature extraction network, time sequence feature information and semantic feature information are spliced, and finally a counting result is obtained through classification. The event-driven two-stage sampling strategy provided by the invention allows the model to reduce the calculation delay, and the algorithm output can be performed more quickly while the model precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of self-service security inspection systems, and particularly to a method and device for improving the statistics of the number of people passing through a turnstile based on video sequence analysis. Background Art

[0002] With the rapid development of artificial intelligence technology, the airport security inspection process has undergone a major transformation. The introduction of face recognition and object detection technologies has successfully realized the wide application of self-service security inspection turnstiles. This self-service security inspection system not only reduces labor costs but also improves the passing efficiency of passengers, so it has been widely used in the scenarios of passenger security inspection and boarding. In addition, the person-counting function equipped in the self-service security inspection system can ensure that the number of passing passengers matches the actual boarding passengers, thus enhancing the safety guarantee of passengers.

[0003] Currently, person counting mainly relies on video processing technology, that is, it is realized by combining image detection and tracking with a person-counting strategy. However, this method has many deficiencies: on the one hand, affected by the perspective deviation, in a specific perspective, it is difficult to accurately identify the target in a single-frame image, which is extremely likely to cause errors in tracking and counting, such as overcounting or undercounting; on the other hand, hats and some circular objects are easily misidentified as human heads, further increasing the counting error and thus triggering potential safety risks.

[0004] In order to effectively solve these problems, in recent years, relevant research has begun to analyze and predict videos by means of video recognition technology. Compared with single-frame images, video recognition technology can utilize richer temporal information, thus significantly improving the detection accuracy. However, most of the existing video processing methods need to process the entire video data after the video playback ends, which greatly limits the real-time response ability and analysis efficiency of the system, and at the same time increases the storage requirements of the system. Although two-stream methods and 3D convolutional techniques attempt to solve this problem by learning short video segments, for longer video streams, these techniques need to increase the length of video slices to improve the processing ability of temporal information, which in turn leads to a large computational overhead.

[0005] In view of this, the present invention deeply conceives and actively researches and develops improvements in view of the problems that when the existing self-service security inspection system processes videos, it needs to obtain complete video data, resulting in poor real-time response ability of the system, low analysis efficiency, large storage requirements, and large computational overhead due to the need to increase video slices for longer video streams, and thus poor video recognition accuracy, low speed, and inability to meet real-time requirements. Summary of the Invention

[0006] The object of the present invention is to overcome the deficiencies of the prior art and provide a method and device for improving the accuracy and speed of video recognition algorithms and meeting real-time requirements for enhancing the statistical count of the number of people passing through turnstiles based on video sequence analysis.

[0007] To achieve the above object, the solution of the present invention is as follows: A method for enhancing the statistical count of the number of people passing through turnstiles based on video sequence analysis, which includes the following steps: Step S1: Deploy a video classification network model based on a spatio-temporal orthogonal temporal attention mechanism in the turnstile system; Step S2: Adopt an event-driven two-stage sampling strategy to sample from the real-time video stream of the turnstile image acquisition device, including a coarse sampling stage and a fine sampling stage. The coarse sampling uses turnstile events as the sampling time period, configures the sampling interval K and the number of sampled frames N. After starting the video stream, monitor the turnstile events. The turnstile event is the sampling time period, the opening of the gate is taken as the starting point of a video segment, and the closing of the gate is taken as the ending point of a video segment; Step S3: When the gate is triggered to open, start fine sampling with a sliding window and send the sampling results into the video classification network model based on the spatio-temporal orthogonal temporal attention mechanism for inference. The inference process is as follows: Sample an image every K frames, input the image into the 2Dnet spatial extraction network. The 2Dnet spatial extraction network extracts the spatial features of the image, generates feature representations. These feature representations are stacked in the constructed time sequence to form a queue. When the number of sampled frames accumulated in the queue reaches the threshold N, the N feature maps are used as a feature group and input into the OCA3Dnet temporal feature extraction network to output temporal feature information; at the same time, the feature representations in the 2Dnet spatial extraction network are concatenated into a new input and input into the 2Dsnet semantic feature extraction network to output semantic feature information. Subsequently, the temporal feature information and the semantic feature information are concatenated, and finally the counting result is obtained through classification; Step S4: After each inference is completed, output the prediction result and store the prediction result in the prediction set ; Step S5: After the gate is triggered to close, end the sampling, and average the results of the prediction set as the final number counting result.

[0008] Further, in step S3, the specific steps of the sliding window sampling are as follows: Step A1: Initialize the image acquisition device of the turnstile. When the turnstile triggers the gate opening event, initialize the number of frames F to 0; Step A2: When the number of frames F satisfies F / K = 0, perform a frame sampling, and at the same time input the frame into the 2Dnet spatial extraction network for feature extraction, and store the spatially extracted feature representation obtained this time in the queue; Step A3: When the number of frames meets N, perform a queue sampling. The stacked spatial features in the queue are concatenated along the time dimension to form [c, N, h / 8, w / 8], which serves as the queue feature input to the OCA3Dnet temporal feature extraction network. Here, c is the number of feature channels, N is the number of sampled frames, h is the height of the feature map, and w is the width of the feature map. Step A4: Input the queue features into the OCA3Dnet temporal feature extraction network. The stacked feature representations in the queue construct a video feature map with a size of [N*C, h / 8, w / 8] and are input into the 2Dsnet for feature extraction operations to output feature vectors. Step A5: Concatenate the feature vectors extracted by the OCA3Dnet and the feature vectors extracted by the 2Dsnet, and finally output the classification result through softmax. Store the classification result in the prediction set. ; Step A6: Perform a loop in sequence until the gate triggers a door-closing event, ending the iteration. Take the average of the prediction results in the set. as the final result and report it.

[0009] Furthermore, in Step S3, the feature extraction steps of the 2Dnet spatial extraction network are as follows: Step B1: Input N video frames. The video frames are sampled from the video stream, and each frame is an input image with a size of {3, h, w}. Process each input image separately through Steps B2 - B4. Step B2: Perform downsampling through a 7x7 convolutional kernel. Step B3: Perform pooling operations through a 3x3 max-pooling layer. Step B4: Input into two residual modules. Each residual module contains two convolutional layers with 3x3 convolutional kernels, which are used to extract local features. The stride of the first 3x3 convolutional layer is set to 2 for downsampling to complete the downsampling. Step B5: Output N feature representations with a size of {c, h / 8, w / 8}.

[0010] Further, in step S3, the OCA3Dnet temporal feature extraction network uses a resnet network with 3D convolutional kernels. Besides the convolutional layers, this resnet network also defines an orthogonal attention mechanism module. This orthogonal attention mechanism module compresses and excites in three dimensions by constructing a set of orthogonal filters, that is, applying orthogonal filters respectively on t, h, and w. Here, t is the time dimension, h is the height, and w is the width. Then, an excitation mechanism is used to obtain the attention vector. The orthogonal attention mechanism module is integrated into the residual module of the temporal feature network and is called the orthogonal attention layer, so as to extract richer temporal information. After the OCA3Dnet temporal feature extraction network completes downsampling, finally, a feature vector is extracted through a pooling layer and a fully connected layer.

[0011] Further, the specific process of the orthogonal attention mechanism module is as follows: Step C1: Randomly initialize a set of filters according to the dimensions (c, t, h, w) of the input feature tensor. Step C2: If T×h×w < c, initialize t×h×w filters; otherwise, initialize c filters. Step C3: Use the Gram - Schmidt orthogonalization process to convert these filters into an orthogonal set. Step C4: Input the 3D feature tensor x = {B, c, t, h, w}, where B is the batch size, c is the number of channels, t is the time dimension, h is the height, and w is the width. Multiply the orthogonal filters with the D feature tensor x and sum them to obtain the compressed vector. Step C5: Flatten the compressed vector and pass it through two fully connected layers and the ReLU activation function, and then obtain the attention weights through the Sigmoid function. Step C6: Multiply the attention weights with the input feature tensor to obtain the weighted feature vector.

[0012] Further, in step S3, the specific process of the 2Dsnet semantic feature extraction network is as follows: Construct a video feature map with a size of {N×c, h / 8, w / 8}. The video feature map comes from the spatial features extracted by the 2Dnet spatial feature extraction network. N is the number of sampled frames, and c is the number of feature channels output in the 2Dnet spatial feature extraction network. Input the video feature map into the 2Dsnet semantic feature extraction network, and the 2Dsnet semantic feature extraction network uses two residual modules in resnet for feature extraction operations. Then, pass through a global average pooling layer, and finally output a 1024 - dimensional feature vector.

[0013] Further, the training of the video classification network model based on the spatio - temporal orthogonal temporal attention mechanism includes data and annotation and model training. Model training relies on labeled data, which is sourced from the video stream recorded by an image acquisition device. In the scenario of a turnstile passage, each video segment between one opening and the next closing of the turnstile is considered as a video. The annotator labels the number of people passing through the turnstile within the time period from the start to the end of one end of the video. After completion of the annotation, the video segment and the label file recording the number of people are used as the dataset for training. Using the already labeled dataset, it is divided into a training set and a test set. A classification loss function is defined to supervise the training of the model. During the training process, the model gradually updates the weight parameters of the already labeled tag information. After a certain number of iterations, the model is evaluated on the test set, and the model with the lowest classification loss value on the test set is saved. Finally, the model with the best performance is selected for deployment in the turnstile system.

[0014] A device for improving the statistics of the number of people passing through a turnstile based on video sequence analysis deploys a video classification network model based on a spatio-temporal orthogonal temporal attention mechanism in a turnstile system equipped with an image acquisition device. The video classification network model based on the spatio-temporal orthogonal temporal attention mechanism includes an image sampling module, a 2Dnet spatial feature extraction network, an OCA3Dnet temporal feature extraction network, a 2Dsnet semantic feature extraction network, and a splicing module. The image sampling module samples from the real-time video stream of the turnstile image acquisition device using an event-driven two-stage sampling strategy, including a coarse sampling stage and a fine sampling stage. Coarse sampling is based on turnstile events as the sampling time period. The time period of a turnstile event takes the opening of the turnstile as the starting point of a video segment and the closing of the turnstile as the ending point of the video segment. In the fine sampling stage, a K-frame sampling interval is set, and then an N-frame sampling number is set as a sliding window. In the first stage, within one round of the sliding window, an image is sampled every K frames, and the sampled images are input into the 2Dnet spatial extraction network. The 2Dnet spatial feature extraction network is used to extract the spatial features of the input image, thereby generating feature representations. These feature representations are sequentially constructed into a temporal stack in the program to form a queue. When the number of sampled frames accumulated in the queue reaches the threshold N, the N feature maps are used as a feature group and input into the OCA3Dnet temporal feature extraction network. The 2Dnet spatial extraction network is also used to splice the feature representations into a new input and send it into the 2Dsnet semantic feature extraction network. The OCA3Dnet temporal feature extraction network is used to extract the temporal features of the N input feature maps in the temporal feature extraction network and output a weighted feature vector with temporal feature information. The 2Dsnet semantic feature extraction network is used to extract the semantic information of the spatial feature sequence formed after splicing by the 2Dnet spatial feature extraction network and output a feature vector with semantic information. The splicing module is used to splice the feature vectors output by the 2Dsnet semantic feature extraction network and the weighted feature vectors output by the OCA3Dnet temporal feature extraction network, and finally output the classification result.

[0015] After adopting the above solution, the method and device for improving the statistics of the number of people passing through the turnstile based on video sequence analysis in the present invention take into account both spatial feature information and temporal feature information through the 2Dnet spatial feature extraction network and the OCA3Dnet temporal feature extraction network, improve the calculation efficiency, can obtain spatial features and temporal features with less computational cost, and improve the accuracy and speed of the video recognition algorithm.

[0016] The present invention proposes an event-driven two-stage sampling strategy, which samples and analyzes in parallel according to the real-time video stream, without the need to obtain all samples before recognition. This strategy clarifies the frame sequence required by the algorithm, thereby improving the algorithm accuracy, and can give an alarm under the condition of low latency, meeting the requirements of high real-time work.

[0017] Compared with the prior art, the present invention has the following advantages: The 2Dnet spatial feature extraction network of the present invention can extract spatial features and is insensitive to the time sequence of the video. Therefore, it can obtain long-term feature information through the sampling strategy for 3D fusion, and better capture the dynamic changes of the video.

[0018] The present invention proposes an event-driven two-stage sampling strategy, which allows the network to predict the real-time video stream, can meet the needs of real-time monitoring and alarm, without the need to obtain the entire video for analysis, and improves the real-time performance of the algorithm.

[0019] The orthogonal attention mechanism module of the OCA3Dnet temporal feature extraction network of the present invention can predict the number of people passing through the turnstile with a small amount of calculation; using the event-driven two-stage sampling strategy, it realizes the real-time inference of the model and can predict based on the online video. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is the overall flowchart of the present invention.

[0021] Figure 2 It is the architecture diagram of the 2D spatial feature extraction network of the present invention.

[0022] Figure 3 It is the flow schematic diagram of the residual module of the present invention.

[0023] Figure 4 It is the architecture diagram of the 3D temporal feature extraction network of the present invention.

[0024] Figure 5This is a schematic flowchart of the OCA layer module of the present invention.

[0025] Figure 6 This is a schematic flowchart of the OCA calculation process of the present invention.

[0026] Figure 7 This is a flowchart of the 2D semantic feature extraction network of the present invention.

[0027] Figure 8 This is a flowchart of the feature splicing and classification structure output of the present invention. Detailed implementation manners

[0028] In order to further explain the technical solution of the present invention, the present invention will be elaborated in detail below through specific embodiments.

[0029] As Figure 1 shown, the present invention proposes a method for improving the statistics of the number of people passing through a turnstile based on video sequence analysis, which includes the following steps: Step S1: Deploy a video classification network model based on a spatio-temporal orthogonal temporal attention mechanism in a turnstile system with an image acquisition device; Step S2: Adopt an event-driven two-stage sampling strategy to sample from the real-time video stream of the turnstile image acquisition device, including a coarse sampling stage and a fine sampling stage. The coarse sampling uses turnstile events as sampling time periods, configures a sampling interval K and a sampling number of frames N. After starting the video stream, turnstile events are monitored. The turnstile event is the sampling time period, the opening of the gate is used as the starting point of a video segment, and the closing of the gate is used as the ending point of a video segment; Step S3: When the gate is triggered to open, start the fine sampling of the sliding window, and send the sampling result into the video classification network model based on the spatio-temporal orthogonal temporal attention mechanism for inference. The inference process is as follows: Sample an image every K frames, input the image into the 2Dnet spatial extraction network. The 2Dnet spatial extraction network extracts the spatial features of the image and generates feature representations. These feature representations are stacked in the constructed time sequence to form a queue. When the accumulated number of sampled frames in the queue reaches the threshold N, the N feature maps are used as a feature group and input into the OCA3Dnet temporal feature extraction network to output temporal feature information; at the same time, the feature representations in the 2Dnet spatial extraction network are spliced into a new input and input into the 2Dsnet semantic feature extraction network to output semantic feature information. Subsequently, the temporal feature information and the semantic feature information are spliced, and finally the counting result is obtained through classification; Step S4: Output the prediction result after each inference is completed, and store the prediction result in the prediction set ; Step S5: After the gate is triggered to close, end the sampling, and the prediction set The results are averaged to obtain the final number count result.

[0030] The present invention also proposes a device for improving the statistics of the number of people passing through a turnstile based on video sequence analysis. In a turnstile system with an image acquisition device, a video classification network model based on a spatio-temporal orthogonal temporal attention mechanism is deployed. The video classification network model based on the spatio-temporal orthogonal temporal attention mechanism includes an image sampling module, a 2Dnet spatial feature extraction network, an OCA3Dnet temporal feature extraction network, a 2Dsnet semantic feature extraction network, and a splicing module.

[0031] The image sampling module uses an event-driven two-stage sampling strategy to sample from the real-time video stream of the turnstile camera, including a coarse sampling stage and a fine sampling stage. Coarse sampling is based on the turnstile event as the sampling time period. The time period of the turnstile event is that the opening of the gate is the starting point of a video segment, and the closing of the gate is the ending point of a video segment. In order to perform video analysis in real time, in the fine sampling stage, a K-frame sampling interval is set to obtain the changing part of the video. Then, an N-frame sampling number is set as a sliding window. In the first stage, within one round of the sliding window, an image is sampled every K frames, and the sampled image is input to the 2Dnet spatial extraction network.

[0032] The specific sampling steps are as follows: Step A1: Initialize the image acquisition device of the turnstile. When the turnstile triggers a gate opening event, initialize the frame number F to 0; Step A2: When the frame number F satisfies F / K = 0, perform a frame sampling, and at the same time input the frame into the 2Dnet spatial extraction network for feature extraction, and store the spatially extracted feature representation obtained each time into a queue; Step A3: When the frame number satisfies N, perform a queue sampling. The spatially extracted feature representations stacked in the queue are concatenated along the time dimension to form [c, N, h / 8, w / 8], which is used as the queue feature input to the OCA3Dnet temporal feature extraction network, where c is the number of feature channels, N is the number of sampled frames, h is the height of the feature map, and w is the width of the feature map; Step A4: Input the queue feature into the OCA3Dnet temporal feature extraction network. The stacked feature representations in the queue construct a video feature map with a size of [N*C, h / 8, w / 8], and input it into the 2Dsnet for feature extraction operations to output a feature vector; Step A5: Concatenate the feature vector extracted by the OCA3Dnet and the feature vector extracted by the 2Dsnet, and finally output the classification result through softmax, and store the classification result in the prediction set ; Step A6: Perform a loop in sequence until the gate triggers a closing event, end the iteration, and in the set The prediction results are averaged as the final result and reported.

[0033] The 2Dnet spatial feature extraction network is used to extract the spatial features of the image, thereby generating feature representations. These feature representations are sequentially constructed and stacked in time series in the program to form a queue. When the number of sampled frames accumulated in the queue reaches the threshold N, N feature maps are used as a feature group and input into the OCA3Dnet temporal feature extraction network; the 2Dnet spatial extraction network is also used to splice the feature representations into a new input and send it to the 2Dsnet semantic feature extraction network.

[0034] As Figure 2 shown, in this embodiment, the 2Dnet spatial extraction network uses the commonly used 2D vision feature extraction network resnet in the industry, and intercepts part of the network layers as the spatial feature information sampling scheme. Specifically, after the 2Dnet spatial extraction network receives an image input of size {3, h, w} (3 represents the number of spatial feature channels, h represents the image height, and w represents the image width), it performs downsampling through a 7x7 convolutional kernel, then performs pooling operations through a 3x3 max pooling layer, and then inputs it into two residual modules. As Figure 3 shown, each residual module contains two convolutional layers with 3x3 convolutional kernels. These convolutional layers are used to extract local features. The stride of the first 3x3 convolutional layer is set to 2 to achieve downsampling. After completing downsampling, the output size of the finally extracted features is {c, h / 8, w / 8}. Not all network layers are used in this stage to avoid information loss caused by too high abstraction dimensions. Since the 2Dnet spatial extraction network has nothing to do with temporal features, it is possible to sample image frames at any interval for spatial feature extraction. By adjusting the sampling interval, longer-term feature information can be obtained without affecting the model effect. The steps of feature extraction are as follows: Given N video frames, the video frames are sampled from the video stream. Each frame is an input image of size {3, h, w}; The video frames are input into the 2Dnet spatial extraction network and N feature maps are output. The size of the output feature maps should be {c, h / 8, w / 8}, where c is the number of spatial feature channels output by the 2Dnet spatial extraction network.

[0035] The OCA3Dnet temporal feature extraction network is used to extract the temporal features of the N input feature maps in the temporal feature extraction network and output feature maps with temporal feature information. The features extracted by the 2Dnet spatial feature extraction network are stacked in time order in the queue. When the number of accumulated sampled frames reaches the set N frames of the threshold, they are input into the OCA3Dnet temporal feature extraction network. As Figure 4As shown, the shape of the input image should be {c, N, h / 8, w / 8} at this time, where c is the number of spatial feature channels, w is the image width, and h is the image height. The OCA3Dnet temporal feature extraction network (Orthogonal Channel Attention 3D network) proposed in the present invention uses a resnet network with 3D convolutional kernels. In addition to the basic convolutional layers, an orthogonal attention mechanism module is defined. This orthogonal attention mechanism module compresses and excites in three dimensions by constructing a set of orthogonal filters. That is, orthogonal filters are applied respectively on T, h, and w, where T is the time dimension, h is the height, and w is the width, and then an excitation mechanism is used to obtain the attention vector. As Figure 5 shown, the orthogonal attention mechanism module is integrated into the residual module of the temporal feature network, called the orthogonal attention (OCA) layer, to extract richer temporal information. After the OCA3Dnet temporal feature extraction network completes downsampling, the feature vector is finally extracted through the pooling layer and the fully connected layer.

[0036] As Figure 6 shown, the specific process of the orthogonal attention mechanism module is as follows: Step C1: Randomly initialize a set of filters according to the dimensions (c, t, h, w) of the input feature tensor; Step C2: If T×h×w < c, then initialize t×h×w filters; otherwise, initialize c filters; Step C3: Use the Gram-Schmidt orthogonalization process to convert these filters into an orthogonal set; Step C4: Input the 3D feature tensor x = {B, c, t, h, w}, where B is the batch size, c is the number of channels, t is the time dimension, h is the height, and w is the width. Multiply the orthogonal filters with the D feature tensor x and sum them to obtain the compressed vector; Step C5: Flatten the compressed vector and pass it through two fully connected layers and the ReLU activation function, and then obtain the attention weights through the Sigmoid function; Step C6: Multiply the attention weights with the input feature tensor to obtain the weighted feature vector.

[0037] The 2Dsnet semantic feature extraction network is used to extract the semantic information of the spatial feature sequence formed after splicing the spatial features extracted by the 2Dnet spatial feature extraction network. The spatial features extracted from the 2Dnet spatial feature extraction network form a spatial feature sequence at the video level after splicing. This part of the spatial feature sequence usually contains semantic information at the video level. Therefore, the present invention proposes 2Dsnet as a parallel network to extract semantic feature information for enhancing video semantic understanding. As Figure 7As shown in the figure, first, a video feature map with the size of {N×c, h / 8, w / 8} needs to be constructed. The video feature map comes from the spatial features extracted by the 2Dnet spatial feature extraction network. N is the number of sampled frames, and c is the number of feature channels output in the 2Dnet spatial feature extraction network. The video feature map is input into the 2Dsnet semantic feature extraction network. The 2Dsnet semantic feature extraction network also uses two residual modules in the resnet for feature extraction operations. Then, through a global average pooling layer, a 1024-dimensional feature vector is finally output.

[0038] As Figure 8 shown in the figure, the splicing module is used to splice the feature vector output by the 2Dsnet semantic feature extraction network and the weighted feature vector output by the OCA3Dnet temporal feature extraction network, and finally, the classification result is output through the softmax layer.

[0039] The training of the video classification network model based on the spatio-temporal orthogonal temporal attention mechanism of the present invention includes dataset annotation and model training; Dataset annotation: The training of the model relies on the annotated data. The data comes from the video stream recorded by the image acquisition device. In the scenario of the turnstile passage, each video segment between each door opening and closing is used as a video. The annotator labels the number of people passing through the turnstile during the period from the start to the end of a video segment. After completion of the annotation, the video segment and the label file recording the number of people are used as the dataset for training.

[0040] The model is trained using the annotated dataset, which is divided into a training set and a test set. The classification loss function is defined to supervise the training of the model. During the training process of the model, the weight parameters of the annotated label information are gradually updated. After a certain number of iterations, the model is evaluated on the test set, and the model with the lowest classification loss value on the test set is saved. Finally, the model with the best effect is selected for deployment.

[0041] The general idea of the present invention is to adopt an event-driven two-stage sampling strategy to obtain image frames from a video stream. First, define K frames as the frame interval, sample an image every K frames, and input the image frames into the 2Dnet spatial extraction network. The 2Dnet spatial extraction network is responsible for extracting the spatial features of the images, thereby generating feature representations. These feature representations are sequentially constructed and stacked in time series in the program to form a queue. When the number of sampled frames accumulated in the queue reaches the threshold N (N is the number of sampled frames), the N feature maps are used as a feature group and input into the OCA3Dnet temporal feature extraction network to output temporal feature information. At the same time, the feature representations in the 2Dnet spatial extraction network are concatenated into a new input and fed into the 2Dsnet semantic feature extraction network to output semantic feature information. Subsequently, the temporal feature information and the semantic feature information are concatenated, and finally, the counting result is obtained through classification. The event-driven two-stage sampling strategy proposed by the present invention allows the model to reduce the computational delay, improve the model accuracy, and output the algorithm faster.

[0042] The above embodiments and diagrams do not limit the product form and style of the present invention. Any appropriate changes or modifications made by those of ordinary skill in the art shall be regarded as not departing from the patent scope of the present invention.

Claims

1. A method for improving the statistics of the number of people passing through a turnstile based on video sequence analysis, characterized in that, It includes the following steps: Step S1: Deploy a video classification network model based on a spatio-temporal orthogonal temporal attention mechanism in the turnstile system; Step S2: Adopt an event-driven two-stage sampling strategy to sample from the real-time video stream of the turnstile image acquisition device, including a coarse sampling stage and a fine sampling stage. The coarse sampling uses turnstile events as the sampling time period, configures the sampling interval K and the sampling number of frames N. After starting the video stream, monitor the turnstile events. The turnstile event is the sampling time period, the opening of the gate is used as the starting point of a video segment, and the closing of the gate is used as the ending point of a video segment; Step S3: When the gate is triggered to open, start the fine sampling of the sliding window and send the sampling result into the video classification network model based on the spatio-temporal orthogonal temporal attention mechanism for inference. The inference process is as follows: Sample an image every K frames, input the image into the 2Dnet spatial extraction network. The 2Dnet spatial extraction network extracts the spatial features of the image and generates feature representations. These feature representations are stacked according to the constructed time sequence to form a queue. When the accumulated number of sampled frames in the queue reaches the threshold N, the N feature maps are used as a feature group and input into the OCA3Dnet temporal feature extraction network to output temporal feature information; at the same time, the feature representations in the 2Dnet spatial extraction network are concatenated into a new input and input into the 2Dsnet semantic feature extraction network to output semantic feature information. Subsequently, the temporal feature information and the semantic feature information are concatenated, and finally the counting result is obtained through classification; Step S4: Output the prediction result after each inference is completed and store the prediction result in the prediction set ; Step S5, after triggering the gate to close, end the sampling, and average the results of the prediction set to obtain the final number count result.

2. The method for improving the statistical number of people passing through the turnstile based on video sequence analysis according to claim 1, wherein: In step S3, the specific steps of the sliding window sampling are as follows: Step A1: Initialize the image acquisition device of the turnstile. When the turnstile triggers the gate opening event, initialize the number of frames F to 0; Step A2: When the number of frames F satisfies F / K = 0, perform a frame sampling, and at the same time input the frame into the 2Dnet spatial extraction network for feature extraction, and store the spatially extracted feature representation obtained each time into the queue; Step A3: When the number of frames satisfies N, perform a queue sampling. The spatially extracted feature representations stacked in the queue are concatenated along the time dimension to form [c, N, h / 8, w / 8], which is used as the queue feature input of the OCA3Dnet temporal feature extraction network, where c is the number of feature channels, N is the number of sampled frames, h is the height of the feature map, and w is the width of the feature map; Step A4: Input the queue feature into the OCA3Dnet temporal feature extraction network. The stacked feature representations in the queue construct a video feature map with a size of [N*C, h / 8, w / 8] and input it into the 2Dsnet for feature extraction operations to output a feature vector; Step A5: Concatenate the feature vectors extracted by OCA3Dnet and the feature vectors extracted by 2Dsnet, and finally output the classification result through softmax, and store the classification result in the prediction set ; Step A6: Loop until the gate triggers the closing event, end the iteration, and The prediction results are averaged and reported as the final result.

3. The method for improving the statistics of the number of people passing through the turnstile based on video sequence analysis according to claim 1, wherein In step S3, the feature extraction steps of the 2Dnet spatial extraction network are as follows: Step B1: Input N video frames. The video frames are from the sampling of the video stream, and each frame is an input image with a size of {3, h, w}; each input image is processed separately in steps B2 - B4; Step B2: Perform downsampling through a 7x7 convolution kernel; Step B3: Perform a pooling operation through a 3x3 max pooling layer; Step B4: Input into two residual modules. Each residual module contains two convolutional layers with 3x3 convolutional kernels, which are used to extract local features. The stride of the first 3x3 convolutional layer is set to 2 to achieve downsampling and complete the downsampling. Step B5: Output N feature representations with the size of {c, h / 8, w / 8}.

4. The method for improving the statistical number of people passing through the turnstile based on video sequence analysis according to claim 1, characterized in that, In step S3, the OCA3Dnet temporal feature extraction network uses a resnet network with 3D convolutional kernels. In addition to convolutional layers, this resnet network also defines an orthogonal attention mechanism module. The orthogonal attention mechanism module compresses and excites in three dimensions by constructing a set of orthogonal filters, that is, applying orthogonal filters in t, h, and w respectively, where t is the time dimension, h is the height, and w is the width, and then uses an excitation mechanism to obtain the attention vector. The orthogonal attention mechanism module is integrated into the residual module of the temporal feature network and is called the orthogonal attention layer to extract richer temporal information. After the OCA3Dnet temporal feature extraction network completes downsampling, finally, a feature vector is extracted through a pooling layer and a fully connected layer.

5. The method for improving the statistical count of the number of people passing through the turnstile based on video sequence analysis according to claim 4, wherein The specific process of the orthogonal attention mechanism module is as follows: Step C1: Randomly initialize a set of filters according to the dimensions (c, t, h, w) of the input feature tensor. Step C2: If T×h×w < c, initialize t×h×w filters; otherwise, initialize c filters. Step C3: Use the Gram - Schmidt orthogonalization process to convert these filters into an orthogonal set. Step C4: Input the 3D feature tensor x = {B, c, t, h, w}, where B is the batch size, c is the number of channels, t is the time dimension, h is the height, and w is the width. Multiply the orthogonal filters with the D feature tensor x and sum to obtain the compressed vector. Step C5: Flatten the compressed vector and pass it through two fully connected layers and the ReLU activation function, and then obtain the attention weights through the Sigmoid function. Step C6: Multiply the attention weights with the input feature tensor to obtain the weighted feature vector.

6. The method for improving the statistical count of the number of people passing through a turnstile based on video sequence analysis according to claim 1, wherein In step S3, the specific process of the 2Dsnet semantic feature extraction network is as follows: Construct a video feature map with the size of {N×c, h / 8, w / 8}. The video feature map comes from the spatial features extracted by the 2Dnet spatial feature extraction network. N is the number of sampled frames, and c is the number of feature channels output in the 2Dnet spatial feature extraction network. Input the video feature map into the 2Dsnet semantic feature extraction network, and the 2Dsnet semantic feature extraction network uses two residual modules in resnet for feature extraction operations. Then, through a global average pooling layer, finally, a 1024 - dimensional feature vector is output.

7. The method for improving the statistical number of people passing through the turnstile based on video sequence analysis according to claim 1, characterized in that The training of the video classification network model based on the spatio - temporal orthogonal temporal attention mechanism includes data and annotation and model training. Model training relies on labeled data. The data comes from the video stream recorded by the image acquisition device. In the gate passage scenario, the video clip between each opening and closing of the gate is regarded as a video segment. The annotator uses the start and end time period of the video on one end to mark the number of people passing through the gate during this time period. After the labeling is completed, the video clip and the label file recording the number of people are used as the training data set. Use the labeled data set, divide it into training set and test set, define the classification loss function to supervise the model training, and gradually update the weight parameters of the labeled label information during the training process. After a certain number of iterations, the model is evaluated on the test set, and the model with the lowest classification loss value on the test set is saved. Finally, the model with the best effect is selected and deployed in the gate system.

8. An apparatus for improving the statistics of the number of people passing through a turnstile based on video sequence analysis, characterized in that, Deploy a video classification network model based on the spatiotemporal orthogonal temporal attention mechanism in a gate system with image acquisition equipment. The video classification network model based on the spatiotemporal orthogonal temporal attention mechanism includes an image sampling module, a 2Dnet spatial feature extraction network, an OCA3Dnet temporal feature extraction network, a 2Dsnet semantic feature extraction network, and a splicing module. The image sampling module adopts an event-driven two-stage sampling strategy to sample from the real-time video stream of the gate image acquisition device, including a coarse sampling stage and a fine sampling stage. The coarse sampling is based on the gate event as the sampling time period, and the gate event time period is the gate opening as the starting point of a video segment, and the gate closing as the end point of a video segment. In the fine sampling stage, a K-frame sampling interval is set, and then the N-frame sampling number is set as a sliding window. In the first stage, within a round of sliding window, an image is sampled every K frames, and the sampled image is input into the 2Dnet spatial extraction network. The 2Dnet spatial feature extraction network is used to extract the spatial features of the input image, thereby generating feature representations. These feature representations are sequentially constructed into a time series stack in the program to form a queue. When the number of sampled frames accumulated in the queue reaches a threshold N, the N feature maps are input as feature groups into the OCA3Dnet temporal feature extraction network. The 2Dnet spatial extraction network is also used to splice the feature representations into a new input and send it to the 2Dsnet semantic feature extraction network. The OCA3Dnet temporal feature extraction network is used to extract N input feature maps and output a weighted feature vector with temporal feature information; The 2Dsnet semantic feature extraction network is used to extract the semantic information of the spatial feature sequence formed after the 2Dnet spatial feature extraction network is spliced, and output a feature vector with semantic information; The splicing module is used to perform feature splicing on the feature vector output by the 2Dsnet semantic feature extraction network and the weighted feature vector output by the OCA3Dnet temporal feature extraction network, and finally output the classification result.