Rail transit personnel detection method and device and storage medium

By using the image processing model of feature extraction network, neck network and dynamic detection head in rail transit scenarios, combined with coordinate attention and multiple perception mechanisms, the problems of low accuracy and poor real-time detection in rail transit scenarios are solved, and the personnel detection effect with high accuracy and high real-time is achieved.

CN119964090AInactive Publication Date: 2025-05-09INST OF COMPUTING TECH CHINA ACAD OF RAILWAY SCI +2
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510437529.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In rail transit scenarios, personnel detection has low accuracy and poor real-time performance for small targets, especially in low light and dark environments, which are difficult to detect intruders in a timely manner.

Method used

An image processing model including feature extraction network, neck network and dynamic detection head is adopted, and small target features are extracted through coordinate attention mechanisms, and classified positioning is performed in combination with multiple perception mechanisms to generate image information for personnel detection.

Benefits of technology

It improves the accuracy and real-time detection of small targets, and can accurately identify and locate intruders in complex environments to meet the automatic detection needs in rail transit scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964090A_ABST
    Figure CN119964090A_ABST
Patent Text Reader

Abstract

The invention provides a rail transit personnel detection method and device and a storage medium, and relates to the technical field of personnel detection, and the method comprises the steps: obtaining rail transit image information; inputting the rail transit image information into a preset image processing model for processing, and obtaining personnel detection image information; the image processing model comprises a feature extraction network, a neck network and a dynamic detection head. In a feature extraction network, small target features are effectively extracted by using coordinate attention, the degree of attention to a small target is improved, meanwhile, a dynamic detection head is fused by using multiple sensing mechanisms, the sensing ability to targets with different sizes is improved, the sensing recognition performance to the small target is further improved, people with the small target can be found in time, and the recognition efficiency is improved. In addition, the method is also beneficial for improving the success rate and the accuracy of small target recognition, and better meets the functional requirements of personnel detection in a rail transit scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of personnel detection, and in particular to a rail transit personnel detection method, device and storage medium. Background Art

[0002] With the rapid expansion and speed increase of rail transit, rail transit safety faces severe challenges. Although rail transit adopts fully enclosed management, there are still cases of human intrusion, which affects the order of rail transit transportation and poses a safety hazard. The traditional solution relies on manual inspections, which requires a lot of manpower and cannot guarantee safety around the clock. Therefore, monitoring is used instead. The monitoring method of monitoring by cameras in the monitoring room is popular, but it still relies on manual observation, which is susceptible to eye fatigue and poses a hidden danger. With the development of artificial intelligence technology, the monitoring-based method uses deep learning models to perform video image analysis and processing to achieve automatic detection of human intrusion in rail transit scenarios, which has gradually become the mainstream.

[0003] In rail transit scenarios, for security reasons, the real-time and accuracy of personnel detection are required to be high. However, due to the limitations of the layout of surveillance cameras, there are situations where intruders are far away from the surveillance cameras, resulting in smaller figures in the surveillance images. For small targets, especially in low-light and dark environments, the personnel detection method in the existing technology needs to wait for the intruder to get close to the surveillance camera before being detected. There is a problem of poor real-time performance due to the inability to detect in time, and the detection accuracy is also low.

[0004] Therefore, for personnel detection in rail transit scenarios, how to improve the real-time and accuracy of small target detection is an urgent problem that needs to be solved. Summary of the invention

[0005] The present invention provides a rail transit personnel detection method, device and storage medium, which are used to solve the defects of low small target detection accuracy and poor real-time performance in the prior art personnel detection in rail transit scenarios.

[0006] The present invention provides a rail transit personnel detection method, comprising: Obtain rail transit image information; Inputting the rail transit image information into a preset image processing model for processing to obtain personnel detection image information; Wherein, the image processing model includes a feature extraction network, a neck network and a dynamic detection head; The feature extraction network performs feature extraction processing based on the rail transit image information based on coordinate attention to obtain first feature information, wherein the coordinate attention is used to capture the features of the spatial position to improve the attention to small targets; The neck network performs feature fusion processing according to the first feature information to obtain second feature information; The dynamic detection head performs classification and positioning processing based on the fusion of multiple perception mechanisms and obtains the personnel detection image information according to the second feature information.

[0007] According to a rail transit personnel detection method provided by the present invention, the feature extraction network includes a feature extraction module group and a coordinate attention module, the feature extraction module group is used to perform feature extraction processing according to the rail transit image information to obtain feature extraction information, and the coordinate attention module is used to perform: According to the feature extraction information, encoding each channel along the horizontal direction and the vertical direction of the feature map respectively to obtain horizontal encoding feature information and vertical encoding feature information; Obtaining intermediate feature information based on the horizontal coding feature information and the vertical coding feature information, based on convolution transformation processing and coding processing; According to the intermediate feature information, height tensor information and width tensor information are obtained based on spatial dimension decomposition and convolution transformation processing; Acquire coordinate attention information according to the height tensor information, the width tensor information and the feature extraction information; Among them, the feature extraction information and the coordinate attention information are used to form the first feature information.

[0008] According to a rail transit personnel detection method provided by the present invention, the dynamic detection head includes at least one group of perception units, and the perception unit group includes a size perception attention unit, a space perception attention unit and a task perception attention unit; The size perception attention unit processes the input information based on size perception to obtain first perception attention information; The spatial perception attention unit processes the first perception attention information based on spatial perception to obtain second perception attention information; The task-aware attention unit processes the second perceived attention information based on task perception to obtain superimposed perceived attention information; Wherein, when the dynamic detection head includes at least two of the perception unit groups, the perception unit groups are nested, and the superimposed perception attention information is used to generate the personnel detection image information.

[0009] According to a rail transit personnel detection method provided by the present invention, the output expression of the coordinate attention module includes: ; ; ; ; ; ; in, Extract information for features, including the format The feature map of is the channel of the feature map, is the height of the feature map, is the width of the feature map; For Channel The level of encoding feature information; For Channel Vertically coded feature information; is the intermediate feature information; is the encoding function; is the convolution transformation function; is the height tensor information; is the width tensor information; for A tensor decomposed along the height dimension; for A tensor decomposed along the width dimension; is the encoding function; is the highly convolutional transformation function; is the width convolution transformation function; is the coordinate attention information.

[0010] According to a rail transit personnel detection method provided by the present invention, the output expression of the perception unit group includes: ; ; ; ; in, is the output function of the perception unit group; is the size-aware attention function; is the spatial perception attention function; For task-aware attention function; To input feature information, include A three-dimensional tensor of is the level of the feature map in the feature information, S is the product of the width and height of the feature map in the feature information, is the number of channels of the feature map in the feature information; is the activation function; is the linear function approximated by the convolutional layer; is the number of sparse sampling locations; is the sampling position at k; is the self-learning offset; For location The importance scalar of self-learning; is the maximum value function; is the characteristic slice of the Cth channel; and It is the self-learning parameter.

[0011] According to a rail transit personnel detection method provided by the present invention, the image processing model is determined by a model training process, and the model training process includes: Acquire a data set, the data set comprising basic image data and corresponding actual bounding box information; Inputting the basic image data into the initial image processing model to obtain predicted bounding box information; According to a preset loss function, based on the predicted bounding box information and the actual bounding box information, loss parameter information is calculated and obtained; According to the loss parameter information, adjusting the parameters of the image processing model to obtain new loss parameter information until a preset training completion condition is met; The loss function calculates the outlier degree based on the difference between the predicted bounding box information and the actual bounding box information to obtain the loss parameter information.

[0012] According to a rail transit personnel detection method provided by the present invention, the calculation expression of the loss function includes: ; ; ; ; in, is the bounding box loss; For distance attention; To adjust the bounding box loss; The horizontal coordinate of the center point of the predicted bounding box; The ordinate of the center point of the predicted bounding box; is the horizontal coordinate of the center point of the actual bounding box; is the ordinate of the center point of the actual bounding box; is the width of the minimum bounding rectangle of the predicted bounding box and the actual bounding box; is the height of the minimum bounding rectangle of the predicted bounding box and the actual bounding box; is the outlier degree; is the intersection-and-union loss value of the current predicted bounding box; is the average value of the intersection-over-union loss of the predicted bounding box; is the non-monotonic focusing coefficient; and is a hyperparameter.

[0013] According to a rail transit personnel detection method provided by the present invention, the initial image processing model is obtained in the following manner: Obtain a YOLO model, wherein the YOLO model includes a backbone network, a neck network, and an output layer; Adding a preset coordinate attention module to the backbone network to form the feature extraction network; Replacing the output layer with a preset dynamic detection head; The feature extraction network, the neck network and the dynamic detection head form the image processing model.

[0014] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, a rail transit personnel detection method as described in any one of the above is implemented.

[0015] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, a rail transit personnel detection method as described in any one of the above is implemented.

[0016] The present invention provides a rail transit personnel detection method, device and storage medium, which have at least beneficial effects: an image processing model processes rail transit image information to obtain personnel detection image information and realize the function of rail transit personnel intrusion detection. In the image processing model, feature extraction processing is performed on rail transit image information through a feature extraction network to obtain first feature information, and processing is performed based on coordinate attention during the feature extraction process. The coordinate attention is used to capture the characteristics of the spatial position, which can improve the attention and positioning ability of small targets, so that the first feature information more accurately reflects the characteristics of small targets; the neck network performs feature fusion processing based on the first feature information to fuse the features of different levels to obtain the second feature information, so that the second feature information can more accurately reflect the characteristics of targets of different sizes, which is conducive to accurately identifying targets in various complex environments; the dynamic detection head is based on the fusion of multiple perception mechanisms, such as size perception, space perception and task perception, and the second feature information is classified and positioned to identify the target object for classification and positioning, and generate personnel detection image information, so as to achieve the effect of detecting whether the rail transit image includes intruders and positioning the personnel. Therefore, in the feature extraction network, coordinate attention is used to effectively extract small target features, improve the degree of attention to small targets, and at the same time, the dynamic detection head uses a variety of perception mechanisms to enhance the perception of targets of different sizes, further improving the perception and recognition performance of small targets, which is conducive to timely detection of small target personnel, improving the real-time detection, and also conducive to improving the success rate and accuracy of small target recognition, so as to better meet the functional requirements of personnel detection in rail transit scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 It is a schematic diagram of the processing process of a rail transit personnel detection method provided by the present invention.

[0019] Figure 2 It is a structural diagram of the basic model when the basic model is the YOLO model.

[0020] Figure 3 It is a structural schematic diagram of an image processing model in one embodiment of a rail transit personnel detection method provided by the present invention.

[0021] Figure 4It is a schematic diagram of adding a coordinate attention module in one embodiment of a rail transit personnel detection method provided by the present invention.

[0022] Figure 5 It is a schematic diagram of the processing process of a coordinate attention module in one embodiment of a rail transit personnel detection method provided by the present invention.

[0023] Figure 6 It is a schematic diagram of the processing process of a sensing unit group in a dynamic detection head in one embodiment of a rail transit personnel detection method provided by the present invention.

[0024] Figure 7 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0026] The target detection algorithm based on deep learning can be divided into a two-stage detection algorithm and a single-stage detection algorithm. The two-stage detection algorithm divides the detection problem into two stages. First, candidate areas are generated, and then the candidate areas are classified and regressed. The two-stage detection algorithm has high recognition accuracy, but is slow. The single-stage detection algorithm does not need to generate candidate areas, and directly generates the category probability and position coordinate values ​​of the target. The final detection result can be directly obtained after a single detection, so it has a faster detection speed. Due to the high real-time requirements for rail transit personnel intrusion detection, single-stage detection algorithms are often used for detection to ensure detection speed.

[0027] Combine the following Figure 1-Figure 6 A rail transit personnel detection method of the present invention is described, comprising: Obtain rail transit image information; Inputting the rail transit image information into a preset image processing model for processing to obtain personnel detection image information; Wherein, the image processing model includes a feature extraction network, a neck network and a dynamic detection head; The feature extraction network performs feature extraction processing based on the rail transit image information based on coordinate attention to obtain first feature information, wherein the coordinate attention is used to capture the features of the spatial position to improve the attention to small targets; The neck network performs feature fusion processing according to the first feature information to obtain second feature information; The dynamic detection head performs classification and positioning processing based on the fusion of multiple perception mechanisms and obtains the personnel detection image information according to the second feature information.

[0028] The image processing model processes rail transit image information to obtain personnel detection image information and realize the function of rail transit personnel intrusion detection. In the image processing model, the rail transit image information is subjected to feature extraction processing through the feature extraction network to obtain the first feature information. In the feature extraction process, the coordinate attention is used to process the feature. The coordinate attention is used to capture the characteristics of the spatial position, which can improve the attention and positioning ability of small targets, so that the first feature information can more accurately reflect the characteristics of small targets; the neck network is used to perform feature fusion processing based on the first feature information to fuse the features of different levels to obtain the second feature information, so that the second feature information can more accurately reflect the characteristics of targets of different sizes, which is conducive to accurately identifying targets in various complex environments; the dynamic detection head is based on the fusion of multiple perception mechanisms, such as size perception, space perception, and task perception, and the second feature information is classified and positioned to identify the target object for classification and positioning, and generate personnel detection image information, so as to detect whether the rail transit image includes intruders and position the personnel.

[0029] Therefore, in the feature extraction network, coordinate attention is used to effectively extract small target features, improve the degree of attention to small targets, and at the same time, the dynamic detection head uses a variety of perception mechanisms to enhance the perception of targets of different sizes, further improving the perception and recognition performance of small targets, which is conducive to timely detection of small target personnel, improving the real-time detection, and also conducive to improving the success rate and accuracy of small target recognition, so as to better meet the functional requirements of personnel detection in rail transit scenarios.

[0030] The rail transit image information may be generated by surveillance cameras around the rails, and may include rail transit image data packets or rail transit image data streams, etc. The personnel detection image information may be a result of marking personnel identification on the rail transit image, and may be a bounding box for the personnel identification, etc. It is understandable that the bounding box for the personnel identification may not necessarily exist in the personnel detection image information, such as when there is no person in the rail transit image.

[0031] It should be noted that the first feature information may include features of multiple different levels. For example, the first feature information includes first-level features, second-level features and third-level features. Correspondingly, the second feature information includes features obtained by fusion of features of different levels and features fused again, such as the first fused feature obtained by fusion of the first-level features and the second-level features, the second fused feature obtained by fusion of the first-level features and the third-level features, the third fused feature obtained by fusion of the first fused features and the second fused features, etc.

[0032] In some embodiments of the present invention, an open source deep learning model can be selected as the basic model, and an attention mechanism and a dynamic detection head can be added to the basic model so that the basic model can further utilize the fusion of the attention mechanism and multiple perception mechanisms on the basis of the original performance to improve the detection performance of small targets, thereby forming the image processing model of the present invention, which is conducive to improving the acquisition efficiency of the image processing model.

[0033] refer to Figure 2 and Figure 3 In some embodiments of a rail transit personnel detection method of the present invention, the initial acquisition method of the image processing model includes: Obtain a YOLO model, wherein the YOLO model includes a backbone network, a neck network, and an output layer; Adding a preset coordinate attention module to the backbone network to form the feature extraction network; Replacing the output layer with a preset dynamic detection head; The feature extraction network, the neck network and the dynamic detection head form the image processing model.

[0034] The single-stage detection YOLO model is used as the basic model. The YOLO model has excellent performance in image processing and target recognition. The YOLO model includes a backbone network for feature extraction, a neck network for feature fusion, and an output layer for classification and positioning. On the basis of the YOLO model, a coordinate attention module is added to the backbone network so that during the feature extraction process, the coordinate attention module uses the coordinate attention mechanism to capture the characteristics of the spatial position and improve the attention and positioning ability of small targets; and the output layer is replaced with a dynamic detection head so that during the classification and positioning process, based on the fusion of multiple perception mechanisms, such as size perception, space perception, and task perception, the model's perception and recognition ability of targets of different sizes is improved, which is conducive to improving the real-time and accuracy of identifying small target personnel. In this way, a coordinate attention module is added to the backbone network to adapt to the feature extraction of small targets in rail transit scenarios. After the addition and modification, a feature extraction network is formed, and the output layer is replaced by a dynamic detection head to improve the classification and positioning performance of small targets in rail transit scenarios. The feature extraction network is combined with the original neck network and the dynamic detection head to form the required image processing model, which is conducive to targeted improvement of the performance of personnel detection in rail transit scenarios and meet application requirements.

[0035] It is understandable that after obtaining the initial image processing model, a model training process is required to allow the image processing model to perform deep learning to a level that can be applied in practice.

[0036] refer to Figure 2 , the figure is a schematic diagram of the structure of the YOLO model. After adding the coordinate attention module and replacing the output layer with the dynamic detection head, the structural diagram of the image processing model is as follows Figure 3 It should be noted that, in the image processing model of the present invention, multiple dynamic detection heads may be included to process feature information of different levels and different fusions, such as Figure 3 shown.

[0037] In some embodiments of the present invention, in addition to using the YOLO model as the basic model, other single-stage detection models may also be used as the basic model, such as the SSD model. The YOLO model has different versions and size complexities. For example, the performance YOLOv8n model may be selected as the basic model. For the convenience of description, the YOLO model is used as the basic model in the following description, but the basic model is not limited to the YOLO model.

[0038] refer to Figure 3 and Figure 4In some embodiments of a rail transit personnel detection method of the present invention, the feature extraction network includes a feature extraction module group and a coordinate attention module, the feature extraction module group is used to perform feature extraction processing according to the rail transit image information to obtain feature extraction information, and the coordinate attention module is used to perform: According to the feature extraction information, encoding each channel along the horizontal direction and the vertical direction of the feature map respectively to obtain horizontal encoding feature information and vertical encoding feature information; Obtaining intermediate feature information based on the horizontal coding feature information and the vertical coding feature information, based on convolution transformation processing and coding processing; According to the intermediate feature information, height tensor information and width tensor information are obtained based on spatial dimension decomposition and convolution transformation processing; Acquire coordinate attention information according to the height tensor information, the width tensor information and the feature extraction information; Among them, the feature extraction information and the coordinate attention information are used to form the first feature information.

[0039] The feature extraction module group performs feature extraction processing on the rail transit image information to obtain features of different sizes and levels in the rail transit image and form feature extraction information. Due to the personnel far away from the surveillance camera, the surveillance camera obtains small-sized targets, which occupy less information in the entire rail transit image and have a low resolution. In order to improve the attention to small targets and reduce the interference of irrelevant information, the coordinate attention module encodes each channel of the feature map in the horizontal and vertical directions according to the feature extraction information, and obtains horizontal encoding feature information and vertical encoding feature information. While reflecting the long dependency in one direction, it can retain the accurate position information in another direction, which is conducive to accurately locating important information. The attention module performs convolution transformation processing according to the horizontal encoding feature information and the vertical encoding feature information, and performs encoding processing in the horizontal and vertical directions to obtain the intermediate feature map as the intermediate feature information. The attention module performs spatial dimension decomposition along the height and width of the intermediate feature map according to the intermediate feature information, and performs convolution transformation processing to obtain height tensor information and width tensor information, which can reflect the attention weight of the coordinates, and then combines the feature extraction information to obtain the coordinate attention information. In this way, based on the coordinate attention mechanism, coordinate attention information is obtained, and feature extraction information and coordinate attention information form the first feature information, which can more accurately capture the spatial characteristics of rail transit images to improve the ability to reflect the characteristics of small targets, which is conducive to improving the attention to small targets, and thus improving the real-time and accuracy of small target personnel identification.

[0040] It can be understood that the feature extraction information may include feature maps of different levels.

[0041] refer to Figure 2 and Figure 4 The feature extraction module can be each module of the backbone network in the YOLO model, such as each Conv module, C2f module and SPPF module. The coordinate attention module is a new module added by the present invention for the personnel detection needs of the rail transit scene, so as to achieve the goal of adding the coordinate attention mechanism in the feature extraction process. The feature extraction module can output different levels of feature extraction information according to the rail transit image information, so as to facilitate the subsequent feature fusion processing, such as Figure 3 and Figure 4 shown.

[0042] In related technologies, the SE (Squeeze-and-Excitation) attention mechanism and the BAM (Bottleneck attention module) attention mechanism are usually used for personnel recognition using deep learning models. When obtaining channel attention, the maximum pooling or average pooling method is generally used, which leads to the loss of spatial details and poor recognition performance for small target personnel. In the present invention, the coordinate attention mechanism is used to more accurately capture the spatial features of the image by performing attention calculations on the height and width respectively, and the dependency relationship between features is more comprehensively obtained, which is conducive to accurately locating small target personnel.

[0043] refer to Figure 1 and Figure 6 , in some embodiments of a rail transit personnel detection method of the present invention, the dynamic detection head includes at least one set of perception unit groups, and the perception unit group includes a size perception attention unit, a space perception attention unit and a task perception attention unit; The size perception attention unit processes the input information based on size perception to obtain first perception attention information; The spatial perception attention unit processes the first perception attention information based on spatial perception to obtain second perception attention information; The task-aware attention unit processes the second perceived attention information based on task perception to obtain superimposed perceived attention information; Wherein, when the dynamic detection head includes at least two of the perception unit groups, the perception unit groups are nested, and the superimposed perception attention information is used to generate the personnel detection image information.

[0044] In the rail transit environment, personnel detection is performed. Since the distance between the personnel and the surveillance camera is variable, that is, the size span of the personnel in the rail transit image is large, in order to better adapt to the diversity of the characteristic sizes of the personnel targets, in the perception unit group of the dynamic detection head, the size perception attention unit, the space perception attention unit and the task perception attention unit are used to process the information in sequence, such as Figure 6 As shown in the figure, the fusion of multiple perception mechanisms of size perception, space perception and task perception is realized. The size perception attention unit dynamically processes features according to the semantic importance of different sizes. The space perception attention unit focuses on the area where there is consistency between space and features to process the features. The task perception attention unit realizes the task of personnel detection based on the perception of personnel and classifies and locates the features. The superimposed perception attention information obtained after being processed in sequence by the size perception attention unit, the space perception attention unit and the task perception attention unit superimposes the size attention, space attention and task attention together, so that personnel targets of different sizes and positions can be accurately identified. The superimposed perception attention information reflects the area of ​​attention, that is, the area where the identified personnel are located, which facilitates the generation of personnel detection image information.

[0045] In some embodiments of the present invention, based on the superimposed perceived attention information, a bounding box can be generated to reflect the area where attention is focused, that is, the area where the identified person is located. The original rail transit image information can be combined to generate personnel detection image information, and the rail transit intruders marked by the bounding box can be observed on the rail transit image.

[0046] It should be noted that the size-aware attention unit, the space-aware attention unit, and the task-aware attention unit process information in sequence to extract important information from the information in the gradually ascending order of local size, overall space, and target task, and logically integrate the three perception mechanisms to enhance the perception of targets of different sizes.

[0047] In some embodiments of the present invention, in order to further improve the detection performance, multiple groups of perception units may be nested in the dynamic detection head, that is, the multiple groups of perception units sequentially perform multiple perception mechanism fusion processing on the information to more accurately identify human targets.

[0048] The number of nested perception unit groups will affect the performance of the image processing model. A larger number of nested perception unit groups can improve the accuracy of personnel detection, but a larger number of perception unit groups will also make the image processing model more complex, thereby increasing the processing complexity and affecting the efficiency, resulting in poor real-time performance. Therefore, it is necessary to reasonably set the number of nested perception unit groups in the dynamic detection head, refer to the following table:

[0049] The table shows the test results of nesting 1-6 groups of perception unit groups. It can be seen that for each additional nested group of perception unit groups, the number of parameters of the image processing model will increase by about 0.5M, and the computing power will increase by about 0.5GFLOPs. At the same time, the size of the weight file will increase by about 1MB. In some embodiments of the present invention, after comprehensively considering the performance and complexity of the model, the number of nested perception unit groups is set to 2 to achieve a balance between detection accuracy and model complexity.

[0050] In the image processing model of the present invention, multiple dynamic detection heads can be included to process feature information of different levels and different fusions, such as Figure 3 shown.

[0051] refer to Figure 5 Schematic diagram of the processing process of the coordinate attention module. In some embodiments of the rail transit personnel detection method of the present invention, the output expression of the coordinate attention module includes: ; ; ; ; ; ; in, Extract information for features, including the format The feature map of is the channel of the feature map, is the height of the feature map, is the width of the feature map; For Channel The level of encoding feature information; For Channel Vertically coded feature information; is the intermediate feature information; is the encoding function; is the convolution transformation function; is the height tensor information; is the width tensor information; for A tensor decomposed along the height dimension; for A tensor decomposed along the width dimension; is the encoding function; is the highly convolutional transformation function; is the width convolution transformation function; is the coordinate attention information.

[0052] Height tensor information And the width tensor information , and feature extraction information It can be understood that the subscript C represents the C channel, and h in the formula corresponds to the height H of the feature map, and h represents a certain height value, such as represents the i-th feature information of the C channel and h height value in the feature extraction information; similarly, w corresponds to the width W of the feature map, and w represents a certain width value, such as Indicates the jth feature information of the C channel and w width in the feature extraction information. The superscripts and subscripts of other expressions can be understood according to the same rules.

[0053] In the feature extraction network, after performing feature extraction processing on the rail transit image information to obtain feature extraction information, the attention module processes the feature extraction information through the above expression to realize the coordinate attention mechanism, and then the feature extraction network achieves feature extraction processing based on the rail transit image information based on the coordinate attention, and obtains the feature extraction information and the coordinate attention information to form the first feature information.

[0054] refer to Figure 6 Schematic diagram of the processing process of the perception unit group. In some embodiments of the rail transit personnel detection method of the present invention, the output expression of the perception unit group includes: ; ; ; ; in, is the output function of the perception unit group; is the size-aware attention function; is the spatial perception attention function; For task-aware attention function; To input feature information, include A three-dimensional tensor of is the level of the feature map in the feature information, S is the product of the width and height of the feature map in the feature information, is the number of channels of the feature map in the feature information; is the activation function; is the linear function approximated by the convolutional layer; is the number of sparse sampling locations; is the sampling position at k; is the self-learning offset; For location The importance scalar of self-learning; is the maximum value function; is the characteristic slice of the Cth channel; and It is the self-learning parameter.

[0055] In the dynamic detection head, size perception, space perception and task perception are combined, and the attention mechanism is deployed in each perception function respectively, and finally nested into an attention function , to achieve the purpose of integrating multiple perception mechanisms.

[0056] Enter feature information It can be the second feature information of the neck network. When nesting multiple groups of perception units, input feature information It can also be the feature information output by the previous perception unit group, that is, the superimposed perception attention information.

[0057] In some embodiments of a rail transit personnel detection method of the present invention, the image processing model is determined by a model training process, and the model training process includes: Acquire a data set, the data set comprising basic image data and corresponding actual bounding box information; Inputting the basic image data into the initial image processing model to obtain predicted bounding box information; According to a preset loss function, based on the predicted bounding box information and the actual bounding box information, loss parameter information is calculated and obtained; According to the loss parameter information, adjusting the parameters of the image processing model to obtain new loss parameter information until a preset training completion condition is met; The loss function calculates the outlier degree based on the difference between the predicted bounding box information and the actual bounding box information to obtain the loss parameter information.

[0058] According to the basic model, such as the YOLO model, after adding the coordinate attention module and replacing the dynamic detection head, the initial structure of the image processing model is formed. The image processing model needs to be trained to perform a deep learning process, so that the image processing model can learn to adjust the model parameters to achieve the personnel detection function. The data set used for model training includes basic image data and corresponding actual bounding box information. The basic image data is input into the image processing model to obtain the predicted bounding box information output by the image processing model, and the actual bounding box information is compared. The loss parameter information is calculated based on the set loss function, and then the image processing model is adjusted to make the predicted bounding box information of the image processing model more accurate. Repeat the process of inputting basic image data, calculating and obtaining loss parameter information, and adjusting the parameters of the image processing model until the training completion conditions are met, such as the output of the image processing model converges or the set number of training iterations is reached. The model training is considered to be completed, and the image processing model can be put into use to obtain rail transit image information for processing and obtain personnel detection image information, so as to achieve the effect of automatic personnel detection in rail transit scenarios.

[0059] During the model training process, the loss function uses the outlier degree to characterize the quality of the predicted bounding box, reflecting the difference between the predicted bounding box and the actual bounding box, and then obtains the loss parameter information to optimize the parameters of the image processing model. The loss function is based on the outlier degree. For low-quality bounding boxes, its impact is reduced, thereby reducing interference with model training, which is conducive to improving the model's adaptability to complex situations in rail transit scenarios, realizing dynamic focus learning, and improving the generalization performance of the image processing model.

[0060] In some embodiments of the present invention, the loss function is a Wise-IoU function.

[0061] In some embodiments of a rail transit personnel detection method of the present invention, the calculation expression of the loss function includes: ; ; ; ; in, is the bounding box loss; For distance attention; To adjust the bounding box loss; The horizontal coordinate of the center point of the predicted bounding box; The ordinate of the center point of the predicted bounding box; is the horizontal coordinate of the center point of the actual bounding box; is the ordinate of the center point of the actual bounding box; is the width of the minimum bounding rectangle of the predicted bounding box and the actual bounding box; is the height of the minimum bounding rectangle of the predicted bounding box and the actual bounding box; is the outlier degree; is the intersection-and-union loss value of the current predicted bounding box; is the average value of the intersection-over-union loss of the predicted bounding box; is the non-monotonic focusing coefficient; and is a hyperparameter.

[0062] Hyperparameters and , which remains unchanged during training and can be set by the user according to actual conditions. The loss parameter information can include adjusting the bounding box loss , distance attention , bounding box loss , outlier , nonmonotonic focusing coefficient wait.

[0063] There is a direct relationship between the loss function and the detection performance of the image processing model. If the basic model adopts the YOLOv8n model, the YOLOv8n model adopts the CIoU loss function. CIoU adds the aspect ratio of the bounding box to the quantization standard of the loss function, but uses too many complex functions, consumes a lot of computing power during the calculation process, and has limitations for the regression of low-quality bounding boxes.

[0064] In this regard, the present invention uses the outlier degree to characterize the quality of the predicted bounding box through the expression of the above-mentioned loss function, reflects the difference between the anchor box and the real target, and then calculates the loss parameter information to optimize the parameters of the image processing model. The calculation complexity of the loss function is low, which is conducive to saving computing power, and can reduce the impact of low-quality bounding boxes, reduce interference with model training, improve the model's adaptability to rail transit scenarios, realize dynamic focus learning, and improve the generalization performance of the image processing model.

[0065] In the present invention, based on the basic model, by adding a coordinate attention module, replacing the dynamic detection head, and the loss function based on the outlier degree to obtain the loss parameter information, the image processing model is obtained through three major improvements. The basic model + coordinate attention module, the basic model + coordinate attention module + dynamic detection head, the basic model + coordinate attention module + dynamic detection head + loss function improvement are used as three test schemes, and ablation tests are performed to reflect the impact of the three improvements on the model performance. The basic model uses YOLOv8n, and the test results are as follows:

[0066] As can be seen from the table, the three improvements all improved the detection effect of the model, with the final accuracy increased from 89.4% to 91.6%, the recall rate increased from 81.4% to 83.7%, the mAP50 increased from 87.3% to 89.9%, and the mAP50-95 increased from 58.0% to 60.9%. Compared with the FPS index of the model, the processing efficiency of the improved model did not decrease significantly, and a good processing speed was maintained. The results show that the three improvements can significantly improve performance without sacrificing processing efficiency, that is, the image processing model can be better applied to rail transit personnel detection scenarios.

[0067] In order to test the performance of the image processing model provided by the present invention, different models are tested using the same test data. The results are shown in the following table:

[0068] The image processing model of the present invention is compared with target detection models such as SSD, Faster R-CNN, YOLOv5, YOLOv6, and YOLOv7. It can be seen from the table that the average accuracy of the image processing model of the present invention is higher than that of other target detection models, which further verifies the effect of the improved model.

[0069] Figure 7 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 7 As shown, the electronic device may include: a processor 810, a communications interface 820, a memory 830 and a communication bus 840, wherein the processor 810, the communications interface 820 and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the above-mentioned rail transit personnel detection method.

[0070] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0071] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute a rail transit personnel detection method provided by the above-mentioned methods.

[0072] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0073] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0074] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0075] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.

[0076] All actions to obtain signals, information or data in this application are carried out in compliance with the relevant data protection laws and policies of the location and with the authorization given by the owner of the corresponding device.

[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A rail transit personnel detection method, characterized in that: include: Obtain rail transit image information; Inputting the rail transit image information into a preset image processing model for processing to obtain personnel detection image information; Wherein, the image processing model includes a feature extraction network, a neck network and a dynamic detection head; The feature extraction network performs feature extraction processing based on the rail transit image information based on coordinate attention to obtain first feature information, wherein the coordinate attention is used to capture the features of the spatial position to improve the attention to small targets; The neck network performs feature fusion processing according to the first feature information to obtain second feature information; The dynamic detection head performs classification and positioning processing based on the fusion of multiple perception mechanisms and obtains the personnel detection image information according to the second feature information.

2. A rail transit personnel detection method according to claim 1, characterized in that: The feature extraction network includes a feature extraction module group and a coordinate attention module. The feature extraction module group is used to perform feature extraction processing according to the rail transit image information to obtain feature extraction information. The coordinate attention module is used to perform: According to the feature extraction information, encoding each channel along the horizontal direction and the vertical direction of the feature map respectively to obtain horizontal encoding feature information and vertical encoding feature information; Obtaining intermediate feature information based on the horizontal coding feature information and the vertical coding feature information, based on convolution transformation processing and coding processing; According to the intermediate feature information, height tensor information and width tensor information are obtained based on spatial dimension decomposition and convolution transformation processing; Acquire coordinate attention information according to the height tensor information, the width tensor information and the feature extraction information; Among them, the feature extraction information and the coordinate attention information are used to form the first feature information.

3. A rail transit personnel detection method according to claim 1, characterized in that: The dynamic detection head includes at least one set of perception unit groups, and the perception unit group includes a size perception attention unit, a space perception attention unit and a task perception attention unit; The size perception attention unit processes the input information based on size perception to obtain first perception attention information; The spatial perception attention unit processes the first perception attention information based on spatial perception to obtain second perception attention information; The task-aware attention unit processes the second perceived attention information based on task perception to obtain superimposed perceived attention information; Wherein, when the dynamic detection head includes at least two of the perception unit groups, the perception unit groups are nested, and the superimposed perception attention information is used to generate the personnel detection image information.

4. A rail transit personnel detection method according to claim 2, characterized in that: The output expression of the coordinate attention module includes: ; ; ; ; ; ; in, Extract information for features, including the format The feature map of is the channel of the feature map, is the height of the feature map, is the width of the feature map; For Channel The level of encoding feature information; For Channel Vertically coded feature information; is the intermediate feature information; is the encoding function; is the convolution transformation function; is the height tensor information; is the width tensor information; for A tensor decomposed along the height dimension; for A tensor decomposed along the width dimension; is the encoding function; is the highly convolutional transformation function; is the width convolution transformation function; is the coordinate attention information.

5. A rail transit personnel detection method according to claim 3, characterized in that: The output expression of the perception unit group includes: ; ; ; ; in, is the output function of the perception unit group; is the size-aware attention function; is the spatial perception attention function; For task-aware attention function; To input feature information, include A three-dimensional tensor of is the level of the feature map in the feature information, S is the product of the width and height of the feature map in the feature information, is the number of channels of the feature map in the feature information; is the activation function; is the linear function approximated by the convolutional layer; is the number of sparse sampling locations; is the sampling position at k; is the self-learning offset; For location The importance scalar of self-learning; is the maximum value function; is the characteristic slice of the Cth channel; and It is the self-learning parameter.

6. A rail transit personnel detection method according to claim 1, characterized in that: The image processing model is determined by a model training process, and the model training process includes: Acquire a data set, the data set comprising basic image data and corresponding actual bounding box information; Inputting the basic image data into the initial image processing model to obtain predicted bounding box information; According to a preset loss function, based on the predicted bounding box information and the actual bounding box information, loss parameter information is calculated and obtained; According to the loss parameter information, adjusting the parameters of the image processing model to obtain new loss parameter information until a preset training completion condition is met; The loss function calculates the outlier degree based on the difference between the predicted bounding box information and the actual bounding box information to obtain the loss parameter information.

7. A rail transit personnel detection method according to claim 6, characterized in that: The calculation expression of the loss function includes: ; ; ; ; in, is the bounding box loss; For distance attention; To adjust the bounding box loss; The horizontal coordinate of the center point of the predicted bounding box; The ordinate of the center point of the predicted bounding box; is the horizontal coordinate of the center point of the actual bounding box; is the ordinate of the center point of the actual bounding box; is the width of the minimum bounding rectangle of the predicted bounding box and the actual bounding box; is the height of the minimum bounding rectangle of the predicted bounding box and the actual bounding box; is the outlier degree; is the intersection-and-union loss value of the current predicted bounding box; is the average value of the intersection-over-union loss of the predicted bounding box; is the non-monotonic focusing coefficient; and is a hyperparameter.

8. A rail transit personnel detection method according to claim 1, characterized in that: The initial acquisition method of the image processing model includes: Obtain a YOLO model, wherein the YOLO model includes a backbone network, a neck network, and an output layer; Adding a preset coordinate attention module to the backbone network to form the feature extraction network; Replacing the output layer with a preset dynamic detection head; The feature extraction network, the neck network and the dynamic detection head form the image processing model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, a rail transit personnel detection method as described in any one of claims 1 to 8 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, a rail transit personnel detection method as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Multi-scale attention-fused traffic helmet small target detection system and method

    CN116665156A

  • Personnel target detection method and system

    CN117037211A

  • Submarine cable target detection method based on improved deep network

    CN118447363A

  • Pedestrian target detection model training method, device and equipment

    CN118865433A

  • Unmanned aerial vehicle aerial image target detection method based on improved YOLOv8

    CN119295969A