A method and system for detecting and classifying pedestrians on primary and secondary school campuses
By applying the pedestrian detection and classification method of multi-scale attention feature aggregation in primary and secondary school campuses, and using the improved ResNet-50 network and two-way feature pyramid network, the problem of insufficient applicability of pedestrian detection and fine-grained classification in campus scenarios in the prior art is solved, and high accuracy and high efficiency pedestrian detection and classification are achieved.
Patent Information
- Application Number
- CN202210044349.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-01-14
AI Technical Summary
The existing technology has insufficient applicability in pedestrian detection and fine-grained classification in primary and secondary school campuses. The traditional pedestrian detection algorithm has a high missed detection rate, and the fine-grained classification algorithm cannot be effectively applied to campus scenarios.
A pedestrian detection and classification method for primary and secondary school campuses based on multi-scale attention feature aggregation is adopted, and the improved ResNet-50 network and two-way feature pyramid network are used to extract the multi-layer attention characteristics and discriminant characteristics of pedestrians to achieve accurate detection and classification of students and off-campus personnel.
It improves the accuracy and efficiency of pedestrian testing, can effectively detect and distinguish students and off-campus personnel in primary and secondary school campuses in crowded environments, reduces the missed inspection rate, and enhances the practicality of campus safety management.
Smart Images

Figure CN114550206B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of campus security, and in particular relates to a method and system for detecting and classifying pedestrians in primary and secondary school campuses. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Pedestrian detection and fine-grained classification are research topics that have attracted much attention in computer vision in recent years. Pedestrian detection visualizes the physical characteristics and location information of target pedestrians by combining image processing and artificial intelligence technologies. Pedestrian detection can realize visual tasks such as pedestrian search, pedestrian re-identification, and pedestrian tracking, and plays a vital role in the fields of autonomous driving, intelligent video surveillance, and regional security. Fine-grained classification can correctly identify the type of target by distinguishing the details of the local area when the visual features of the target are roughly similar. Fine-grained classification is currently mainly used in real-life scenarios such as vehicle recognition and plant and animal classification, and is rarely used to distinguish between students and non-school personnel in primary and secondary school campuses.
[0004] Accurately detecting and classifying people on primary and secondary school campuses is an important method to maintain campus safety. Applying pedestrian detection and fine-grained classification algorithms to detect non-school personnel in campus areas can make use of the deep convolutional neural network's ability to learn high-level semantic information of images to detect pedestrians in campus areas and distinguish between students and non-school personnel. It effectively prevents non-school personnel from entering campuses and maintains campus safety, which has important application value in building a safe campus.
[0005] There are two main types of pedestrian detection frameworks: two-stage detectors and one-stage detectors. Compared with two-stage detectors, one-stage detectors have the advantages of simple detection process and fast detection speed. They can perform best when facing the changing environment and crowded crowds of campuses. The fine-grained classification method based on positioning-recognition first automatically locates the area with distinctive features without supervision, and then classifies it according to the distinctive features of the discrimination area. This fine-grained classification method shows strong applicability when applied in campus scenes. Therefore, this method uses a one-stage detector and a fine-grained classification method based on positioning-recognition when detecting and classifying pedestrians in primary and secondary school campuses.
[0006] The pedestrian detection and classification method in primary and secondary school campuses based on multi-scale attention feature aggregation performs well in crowded environments, which is conducive to accurately detecting and distinguishing students and non-school personnel in primary and secondary school campuses, effectively solving problems such as high missed detection rate of traditional pedestrian detection algorithms and inability to apply fine-grained classification algorithms to campus scenes, thereby improving detection and classification efficiency.
[0007] In the process of implementing the present invention, the inventors found that the current technology has the following problems:
[0008] 1. Traditional pedestrian detection algorithms are less applicable to pedestrian detection in primary and secondary school campuses.
[0009] 2. There are few reports on advanced methods of pedestrian detection in primary and secondary school campuses based on deep learning.
[0010] 3. Existing fine-grained classification methods are not used to distinguish between students in primary and secondary school campuses and people outside the school. Summary of the invention
[0011] In order to solve the technical problems existing in the above-mentioned background technology, the present invention provides a method and system for detecting and classifying pedestrians on primary and secondary school campuses. The pedestrian detection and classification results obtained by the method and system can detect students and non-school personnel in the campus area, which is helpful for applications in ensuring the personal safety of students while at school and building a safe campus.
[0012] In order to achieve the above object, the present invention adopts the following technical solution:
[0013] The first aspect of the present invention provides a method for detecting and classifying pedestrians in primary and secondary school campuses.
[0014] A method for detecting and classifying pedestrians in primary and secondary school campuses, comprising: a pedestrian feature map positioning part and a pedestrian recognition part;
[0015] The pedestrian location feature map positioning part includes: obtaining a campus surveillance image, using an improved ResNet-50 network to extract a feature map; based on the feature map, using a pedestrian attention network to obtain a multi-layer attention feature of the pedestrian location; using a bidirectional feature pyramid network to fuse the multi-layer attention features of the pedestrian location to obtain a pedestrian location feature; based on the pedestrian location feature, extracting a pedestrian location feature map;
[0016] The pedestrian recognition part includes: based on the feature map of the pedestrian, using an improved ResNet-50 network combined with a bidirectional feature pyramid network to extract several discriminative features of the pedestrian, and judging whether the pedestrian is an outsider based on the several discriminative features of the pedestrian.
[0017] A second aspect of the present invention provides a pedestrian detection and classification system for primary and secondary school campuses.
[0018] A pedestrian detection and classification system for primary and secondary school campuses, comprising: a pedestrian feature map positioning module and a pedestrian recognition module;
[0019] The pedestrian location feature map positioning module is configured as follows: obtaining a campus surveillance image, using an improved ResNet-50 network to extract a feature map; based on the feature map, using a pedestrian attention network to obtain a multi-layer attention feature of the pedestrian location; using a bidirectional feature pyramid network to fuse the multi-layer attention features of the pedestrian location to obtain a pedestrian location feature; based on the pedestrian location feature, extracting a pedestrian location feature map;
[0020] The pedestrian recognition module is configured as follows: based on the feature map of the pedestrian, an improved ResNet-50 network combined with a bidirectional feature pyramid network is used to extract several discriminative features of the pedestrian, and whether the pedestrian is an outsider is determined based on the several discriminative features of the pedestrian.
[0021] A third aspect of the present invention provides a computer-readable storage medium.
[0022] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for detecting and classifying pedestrians in primary and secondary school campuses as described in the first aspect above.
[0023] A fourth aspect of the present invention provides a computer device.
[0024] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the method for detecting and classifying pedestrians in primary and secondary school campuses as described in the first aspect above are implemented.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] 1. The depth expansion module can obtain deeper feature information while maintaining a relatively large feature map, increasing its perception field and effectively improving the problem of insufficient extraction of deep feature information in pedestrian detection.
[0027] 2. The pedestrian attention mechanism fully mines the relevant information of feature channels and spatial dimensions, selects the information that is more critical to the current task, extracts the correlation between channels and feature spaces, and improves the network detection accuracy.
[0028] 3. The feature aggregation module combines the rich high-level semantic features in the high-resolution feature map with the low-level precise positioning information to generate a richer feature map. The combination of low-level features and high-level features can enhance pedestrian feature extraction and further improve the accuracy of the algorithm.
[0029] 4. While considering real-time performance, nodes with relatively small contributions to feature network fusion are deleted to reduce computational costs. In terms of practicality, the present invention can better detect, locate and classify pedestrians. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0031] Figure 1 is a flow chart of a method for detecting and classifying pedestrians in primary and secondary school campuses shown in the present invention;
[0032] Figure 2 It is a network architecture diagram of a multi-scale attention feature aggregation pedestrian detection algorithm shown in the present invention;
[0033] Figure 3 This is a diagram of the network layer structure newly added in the sixth stage shown in the present invention;
[0034] Figure 4 is a structural diagram of the ABottleNeck submodule of the present invention;
[0035] FIG5( a ) is a structural diagram of a DBottleNeck submodule with a 1×1 convolution branch according to the present invention;
[0036] FIG5( b ) is a structural diagram of a DBottleNeck submodule without a 1×1 convolution branch shown in the present invention;
[0037] Figure 6 is a structural diagram of a channel attention module shown in the present invention;
[0038] Figure 7 It is a structural diagram of the dimensional attention module shown in the present invention;
[0039] Figure 8 It is a structural diagram of a feature aggregation module shown in the present invention;
[0040] Fig. 9 It is a structural diagram of a detection head shown in the present invention;
[0041] Fig.10 This is a diagram showing an application example of pedestrian detection and classification in primary and secondary school campuses according to the present invention. DETAILED DESCRIPTION
[0042] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0043] It should be noted that the following detailed descriptions are all illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0044] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0045] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the methods and systems according to various embodiments of the present disclosure. It should be noted that each box in the flowchart or block diagram can represent a module, a program segment, or a part of a code, and the module, program segment, or a part of a code may include one or more executable instructions for implementing the logical functions specified in each embodiment. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the flowchart and / or block diagram, and the combination of boxes in the flowchart and / or block diagram can be implemented using a dedicated hardware-based system that performs a specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.
[0046] Embodiment 1
[0047] The present embodiment provides a method for detecting and classifying pedestrians on primary and secondary school campuses. The present embodiment uses the method applied to a server as an example for illustration. It is understandable that the method can also be applied to a terminal, and can also be applied to a system including a terminal, a server, and a server, and is implemented through the interaction between the terminal and the server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network servers, cloud communications, middleware services, domain name services, security services CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in this application. In the present embodiment, the method includes:
[0048] The pedestrian feature map positioning part and the pedestrian recognition part;
[0049] The pedestrian location feature map positioning part includes: obtaining a campus surveillance image, using an improved ResNet-50 network to extract a feature map; based on the feature map, using a pedestrian attention network to obtain a multi-layer attention feature of the pedestrian location; using a bidirectional feature pyramid network to fuse the multi-layer attention features of the pedestrian location to obtain a pedestrian location feature; based on the pedestrian location feature, extracting a pedestrian location feature map;
[0050] The pedestrian recognition part includes: based on the feature map of the pedestrian, using an improved ResNet-50 network combined with a bidirectional feature pyramid network to extract several discriminative features of the pedestrian, and judging whether the pedestrian is an outsider based on the several discriminative features of the pedestrian.
[0051] The specific technical solution of this embodiment can be implemented by the following steps: Figure 1 As shown:
[0052] 1. The entire network is divided into five modules, namely, the depth expansion module, the pedestrian attention module, the feature aggregation module, the detection head, and the classification module. The school monitoring equipment is used to collect real-time monitoring images of the campus. When the image is input into the sampling network, the feature map is extracted. Figure 2 As shown in the figure, the sampling is divided into six stages, the network backbone is the improved ResNet-50, and then a new network layer is added after the 5th stage, that is, the 6th stage.
[0053] The backbone network of the multi-scale attention feature aggregation pedestrian detection algorithm adopts an improved ResNet-50 and adds a new module at the end, such as Figure 2 As shown in the figure, the ResNet-50 backbone network 1, the depth expansion module 2, the pedestrian attention module 3, the feature aggregation module 4, and the detection head 5 are connected in sequence.
[0054] like Figure 3 As shown in the figure. The downsampling factors of the first five stages are 2, 4, 8, 16, and 32 respectively. To ensure that the size of the feature map is the same as that of the fifth stage, the downsampling factor of the sixth stage is still 32. The sixth stage is called the deep expansion module, which consists of an ABottleNeck submodule and two DBottleNeck submodules. Among them, the first DBottleNeck submodule has a 1×1 convolution branch, and the second DBottleNeck submodule does not have a 1×1 convolution branch. The structure of the submodule is shown in the figure. Figure 4 , as shown in Figure 5.
[0055] Specifically, one ABottleNeck submodule and two DBottleNeck submodules are used in the new network layer, where the first DBottleNeck submodule has a 1×1 convolution branch and the second DBottleNeck submodule does not have a 1×1 convolution branch. The feature maps are connected at the end of each submodule through the Relu activation function, and finally all parameters are input into the attention module, such as Figure 3 shown.
[0056] The ABottleNeck submodule and the DBottleNeck submodule are obtained by making different adjustments to the bottleneck structure in ResNet-50, but this adjustment does not change the computational complexity. Based on the BottleNeck module, the ABottleNeck submodule adds a 1×1 average pooling layer with a stride of 2 before the right branch 1×1 convolution layer, and changes the stride of the original 1×1 convolution layer to 1, as shown in Figure 4 As shown in Figure 5(a), the DBottleNeck submodule is based on the BottleNeck module. The stride of the first 1×1 convolution layer on the left branch is changed to 1, and the stride of the 3×3 convolution layer is changed to 1. At this time, the expansion ratio is 2. In order to ensure the size of the feature map and increase the receptive field for pedestrian localization, the stride of the 1×1 convolution layer on the right branch is changed to 1, as shown in Figure 5(a); Figure 5(b) does not have a 1×1 convolution layer on the right branch.
[0057] 2. The pedestrian attention module is divided into a channel attention module and a dimension attention module, such as Figure 6 , Figure 7 As shown in the figure. It derives the attention weights along the channel and spatial dimensions in turn, multiplies the weights with the original feature map, and adaptively adjusts the features so that each branch learns the classification and positioning information on the channel axis and dimension axis respectively. The channel attention module and the dimension attention module use the maximum pooling method and the global average pooling method to utilize different information and effectively calculate the channel attention. The channel attention focuses on the meaningful features in the input image, and the dimension attention focuses on the location of the meaningful features in the input image.
[0058] When the characteristic F of H×W×C sx After being input into the channel attention module, global average pooling and global maximum pooling are performed, and the spatial dimension is compressed to obtain two 1×1×C vectors, which are input into a two-layer dense network. The two output features are added together and the weight coefficient F is obtained using the Sigmoid activation function. c Finally, F c With F sx Multiply to get the new output F p ,like Figure 6 As shown in the figure, 1 represents global average pooling and 2 represents global maximum pooling.
[0059] Output F of the channel attention module p Perform average pooling and maximum pooling to obtain two H×W×1 vectors. After the two vectors pass through the 3×3 convolution layer, they are activated by the Sigmoid function to obtain the weight coefficient F. d Finally, F d With F p Multiply to get the output attention feature F n ,like Figure 7 As shown in the figure, 1 represents average pooling, 2 represents maximum pooling, and 3 represents a 3×3 convolutional layer.
[0060] 3. The feature aggregation module uses a bidirectional feature pyramid network to fuse five layers of attention features, such as Figure 8 As shown. The feature aggregation module contains two paths, B and P. The B branch includes B2, B3, B4, B5, and B6, which represent the feature maps generated by the top-down path of the feature pyramid. In order to achieve more optimized cross-linking, B2 and B6, which contribute relatively little to the feature aggregation network, are deleted. The P branch includes P2, P3, P4, P5, and P6 from bottom to top to enhance the feature pyramid. Among them, P2 and B2 represent the same feature layer, and P3, P4, P5, and P6 are the results of the fusion of B3, B4, B5, and B6.
[0061] The feature aggregation module includes two paths, B and p, such as Figure 8 As shown. The B branch includes B2, B3, B4, B5, and B6, representing the feature graph generated by the top-down path of the feature pyramid. In order to achieve more optimized cross-linking, B2 and B6, which contribute relatively little to the feature aggregation network, are deleted. The P branch includes P2, P3, P4, P5, and P6, which are used to enhance the feature pyramid. Among them, P2 and B2 represent the same feature layer, and P3, P4, P5, and P6 are the results of the fusion of B3, B4, B5, and B6. In the figure, 1 represents the feature pyramid and 2 represents the feature aggregation module.
[0062] 4. After obtaining high-quality feature maps, convolution is used in the detection head to predict the center, size, and offset. First, a 3×3 convolution with a channel number of 256 is used to extract features to obtain feature maps. Three parallel convolution layers of 1×1, 1×1, and 2×2 are added to generate center point heat maps, size prediction maps, and offset prediction maps, respectively, as shown in Figure 2. Fig. 9 As shown in the figure, the center point heat map finds the center point of the pedestrian in the image, the size prediction map predicts the scale of the center point, and the offset prediction map reduces the error caused by inaccurate position information.
[0063] The detection head is used to predict the center, size, and offset. First, a 3×3 convolution layer with a channel number of 256 is used to extract features to obtain a feature map. Three parallel convolution layers of 1×1, 1×1, and 2×2 are added to generate the center point heat map, size prediction map, and offset prediction map, respectively, as shown in Fig. 9 As shown in the figure, 1 represents a 3×3 convolutional layer, 2 represents a 1×1 convolutional layer, and 3 represents a 2×2 convolutional layer.
[0064] Fig.10 The results of pedestrian detection and classification using this method on images collected from primary and secondary school campuses are shown on the left. The original photo is on the left, and the detection and classification results are on the right. Students in the campus area are marked with gray boxes, and people outside the school are marked with black boxes.
[0065] 5. The loss function includes center point position loss, dimension size loss and offset regression loss, respectively expressed as L point , L dimension and L offset It is represented as shown in formula (1).
[0066] L=λ p L point +λ d L dimension +λ o L offset (1)
[0067] Among them, λ p ,λ d ,λ o They are the weights of center point classification, dimension regression, and offset regression.
[0068] In the prediction process, a two-dimensional Gaussian mask is added to the positive sample to quickly reduce the error caused by the nearby negative samples and improve the network training advantage. The specific calculation formula is shown in the following formula (2).
[0069]
[0070] Among them, n is the number of targets in the image, x is the horizontal coordinate of the center point, y is the vertical coordinate of the center point, and the variance τ w and τ h Proportional to the width and height of the pedestrian.
[0071] The center point position loss function is shown in equation (3), equation (4) and equation (5).
[0072]
[0073]
[0074]
[0075] Among them, γ represents the focal hyperparameter, σ ij It is used to reduce the impact of negative sample points around positive samples on the overall loss function. ν represents the hyperparameter of the control parameter penalty, and h ij is the probability that the pixel is predicted as the center point, y ij is the ground truth coordinate, when y ij = 1, the sample is marked as a positive sample, and other cases are marked as negative samples.
[0076] The dimension size loss function is defined as shown in Equation (6) and Equation (7).
[0077]
[0078]
[0079] Among them, L teg represents the smooth L1 loss function, G n and R n Respectively represent the predicted value and true value of each positive sample point, and their role is to reduce the impact of the uncertainty of the center point.
[0080] The offset loss function uses a smooth L1 loss function.
[0081] 6. In order to verify the effectiveness of the pedestrian detection algorithm of this embodiment in a crowded environment, this embodiment conducted a large number of experiments under the Caltech dataset. The Caltech dataset includes 11 groups of images, the first six groups with a total of 42,782 images are used for training, and the last five groups with a total of 4,024 images are used for testing. The experimental results show that the pedestrian detection algorithm based on multi-scale attention feature aggregation proposed in this embodiment can achieve good results. This method can detect pedestrians in crowded and dense environments. This shows that the pedestrian detection algorithm proposed in this embodiment is effective, provides a better method for obtaining high-precision results, and has certain practical value.
[0082] 7. After the output pedestrian detection results are input into the classification network, several discriminative parts of the target pedestrian are first found through regional positioning. For example, the height, hairstyle, school uniform, school bag and the red scarf worn by primary school students. Then, the features of these parts are extracted through the same deep expansion module and feature aggregation module as the above pedestrian detection method, and then the extracted fine-grained features are learned and classified.
[0083] 8. To verify the effectiveness of the fine-grained classification algorithm of this embodiment, this embodiment first selected 100,000 photos of adults and 50,000 photos of minors for pre-training on the Seeprettyface dataset. The pre-trained model was then fine-tuned on a dataset made of 5,000 student photos taken in primary and secondary school campuses, and tested with 1,000 student photos. The experimental results show that the fine-grained algorithm proposed in this embodiment can achieve a good classification effect. The method can accurately distinguish between students and non-school personnel in primary and secondary school campuses, which shows that the fine-grained classification algorithm proposed in this embodiment is effective, provides a better method for obtaining high-accuracy results, and has certain practical value.
[0084] 9. The background sets adults appearing in the primary and secondary school campus area as suspicious targets. When a suspicious target appears in the monitoring area, the system will automatically detect the suspicious target and identify it through the school's existing face recognition system. If the face recognition result is an outsider, the system will record it in the background.
[0085] Embodiment 2
[0086] This embodiment provides a system for detecting and classifying pedestrians in primary and secondary school campuses.
[0087] A pedestrian detection and classification system for primary and secondary school campuses, comprising: a pedestrian feature map positioning module and a pedestrian recognition module;
[0088] The pedestrian location feature map positioning module is configured as follows: obtaining a campus surveillance image, using an improved ResNet-50 network to extract a feature map; based on the feature map, using a pedestrian attention network to obtain a multi-layer attention feature of the pedestrian location; using a bidirectional feature pyramid network to fuse the multi-layer attention features of the pedestrian location to obtain a pedestrian location feature; based on the pedestrian location feature, extracting a pedestrian location feature map;
[0089] The pedestrian recognition module is configured as follows: based on the feature map of the pedestrian, an improved ResNet-50 network combined with a bidirectional feature pyramid network is used to extract several discriminative features of the pedestrian, and whether the pedestrian is an outsider is determined based on the several discriminative features of the pedestrian.
[0090] It should be noted that the pedestrian feature map positioning module and the pedestrian recognition module are the same as the examples and application scenarios implemented by the steps in Embodiment 1, but are not limited to the contents disclosed in Embodiment 1. It should be noted that the above modules as part of the system can be executed in a computer system such as a set of computer executable instructions.
[0091] Embodiment 3
[0092] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of the method for detecting and classifying pedestrians in primary and secondary school campuses as described in the first embodiment above are implemented.
[0093] Embodiment 4
[0094] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the method for detecting and classifying pedestrians in primary and secondary school campuses as described in the first embodiment above are implemented.
[0095] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program codes.
[0096] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0097] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0098] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0099] A person skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0100] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for detecting and classifying pedestrians in primary and secondary school campuses, characterized in that: include: The pedestrian feature map positioning part and the pedestrian recognition part; The feature map positioning part of the pedestrian includes: obtaining a campus surveillance image, and using an improved ResNet-50 network to extract a feature map; the improved ResNet-50 network includes using an ABottleNeck submodule and two DBottleNeck submodules connected in sequence in the sixth stage; based on the feature map, a pedestrian attention network is used to obtain multi-layer attention features of the pedestrian position; the pedestrian attention network includes a channel attention module and a dimensional attention module, the channel attention module is used to concentrate on the features representing pedestrians in the feature map, and the dimensional attention module is used to concentrate on the features representing pedestrians. The method comprises the following steps: first, a multi-layer attention feature of the pedestrian position is used to fuse the multi-layer attention features of the pedestrian position, and second, a bidirectional feature pyramid network is used to fuse the multi-layer attention features of the pedestrian position to obtain the pedestrian position features; third, a feature map of the pedestrian is extracted based on the pedestrian position features, specifically including: using a convolutional layer to predict the center point heat map, size prediction map and offset prediction map of the pedestrian position features, and combining the loss function with the center point heat map, size prediction map and offset prediction map of the pedestrian position features to obtain the feature map of the pedestrian; wherein the center point heat map is used to find the center point of the pedestrian in the image, the size prediction map is used to predict the scale of the center point, and the offset prediction map is used to reduce the error caused by inaccurate position information; The pedestrian recognition part includes: based on the feature map of the pedestrian, using an improved ResNet-50 network combined with a bidirectional feature pyramid network to extract several discriminative features of the pedestrian, and judging whether the pedestrian is an outsider based on the several discriminative features of the pedestrian.
2. The method for detecting and classifying pedestrians in primary and secondary school campuses according to claim 1 is characterized in that: The downsampling factor of the sixth stage in the improved ResNet-50 network is the same as that of the fifth stage.
3. The method for detecting and classifying pedestrians in primary and secondary school campuses according to claim 1 is characterized in that: The first DBottleNeck submodule has a 1×1 convolution branch, and the second DBottleNeck submodule does not have a 1×1 convolution branch.
4. The method for detecting and classifying pedestrians in primary and secondary school campuses according to claim 1 is characterized in that: The loss functions include a center point position loss function, a dimension size loss function and an offset regression loss function.
5. A pedestrian detection and classification system for primary and secondary school campuses, characterized in that: include: Pedestrian feature map positioning module and pedestrian recognition module; A pedestrian feature map positioning module is configured to: obtain a campus surveillance image, and use an improved ResNet-50 network to extract a feature map; the improved ResNet-50 network includes an ABottleNeck submodule and two DBottleNeck submodules connected in sequence in the sixth stage; based on the feature map, a pedestrian attention network is used to obtain multi-layer attention features of the pedestrian position; the pedestrian attention network includes a channel attention module and a dimensional attention module, the channel attention module is used to concentrate the features representing the pedestrian in the feature map, and the dimensional attention module is used to concentrate the multi-layer attention features representing the pedestrian position; A bidirectional feature pyramid network is used to fuse the multi-layer attention features of the pedestrian position to obtain the pedestrian position feature; Based on the location features of pedestrians, extract the feature map of pedestrians, specifically including: using the convolution layer to predict the center point heat map, size prediction map and offset prediction map of the pedestrian location features, and combining the loss function with the center point heat map, size prediction map and offset prediction map of the pedestrian location features to obtain the feature map of pedestrians; wherein the center point heat map is used to find the center point of the pedestrian in the image, the size prediction map is used to predict the scale of the center point, and the offset prediction map is used to reduce the error caused by inaccurate location information; The pedestrian recognition module is configured as follows: based on the feature map of the pedestrian, an improved ResNet-50 network combined with a bidirectional feature pyramid network is used to extract several discriminative features of the pedestrian, and whether the pedestrian is an outsider is determined based on the several discriminative features of the pedestrian.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the method for detecting and classifying pedestrians in primary and secondary school campuses as described in any one of claims 1-4 are implemented.
7. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the method for detecting and classifying pedestrians in primary and secondary school campuses as described in any one of claims 1-4 are implemented.
Citation Information
Patent Citations
Multi-scale pedestrian re-identification method based on multi-granularity depth feature fusion
CN112818931A
Pedestrian attribute recognition and positioning method and convolutional neural network system
WO2019041360A1