Human target detection method in infrared environment based on improved yolov3-tinier network
By improving the yolov3-tinier network and optimizing the human target detection model in infrared environments, the problems of low detection accuracy and slow speed in infrared environments are solved, efficient detection on the ZYNQ side is achieved, and the real-time and portability in infrared environments are improved.
Patent Information
- Application Number
- CN202510103470.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-01-22
AI Technical Summary
Existing human target detection methods have limited performance in infrared environments, with low detection accuracy and slow speed. In addition, the network size is large and cannot be effectively deployed on hardware, which limits the real-time and portability in infrared environments.
An improved YOLOv3-Tinier network is adopted, including the backbone module, feature extraction module and detection module. By optimizing the network structure and loss function, an infrared personnel target detection model suitable for the ZYNQ terminal is trained and deployed using the infrared imaging dataset.
While maintaining detection accuracy, the speed and real-time performance of human target detection in infrared environments are improved, making it easier to perform detection on smaller computing carriers and improving recognition efficiency.
Smart Images

Figure CN119810872B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning and target detection, and more specifically to a method for detecting human targets in an infrared environment based on an improved YOLOv3-Tinier network. Background Art
[0002] In infrared environments, due to the loss of effective target area, the task of detecting human targets on visible light images with limited quality is very challenging. This makes human target detection in infrared environments more difficult, resulting in false detections, missed detections, and even misjudgment of human targets to be rescued. Therefore, there is an urgent need for a system that can quickly detect human targets in infrared environments.
[0003] Human target detection methods are widely used under visible light conditions, but the performance of existing human target detection methods in infrared environments is very limited. In addition, the existing human target detection network is very large, with too many calculation parameters and the required calculation carrier is also very large, which cannot be transplanted to hardware for processing. As a result, the accuracy of human target detection is reduced and the detection speed is too slow in infrared environments, which limits the real-time and portability of human target detection in infrared environments. Summary of the Invention
[0004] In order to address the shortcomings of the above-mentioned prior art, the present invention proposes a method for detecting human targets in infrared environments based on an improved YOLOv3-Tinier network, in order to quickly and effectively detect human targets in infrared environments by deploying an optimized model with faster recognition speed and better recognition performance on the ZYNQ side, thereby improving the speed of human target detection in infrared environments without losing accuracy.
[0005] In order to achieve the above-mentioned object, the present invention adopts the following technical solutions:
[0006] The present invention provides a method for detecting human targets in an infrared environment based on an improved Yolov3-Tinier network, which comprises the following steps:
[0007] Step 1: Get the infrared environment personnel detection dataset Image, and let any infrared personnel image in Image be recorded as ,Will The positions and categories of the K infrared personnel in the image are marked to obtain the position labels and category labels. The location tag of the kth infrared person in is recorded as ,and ,in, Indicates the Infrared personnel in The coordinates of the center point of the rectangular box at the position in , and =( ), represents the horizontal coordinate of the center point of the rectangular frame, represents the vertical coordinate of the center point of the rectangular frame, Indicates the Infrared personnel in The size of the rectangular box at the position in , and = , Indicates the Infrared personnel in The width of the rectangular frame at the position in Indicates the Personnel in The height of the rectangular box at the position in The category label of the kth infrared person in is recorded as ;
[0008] Step 2: Establish an improved Yolov3-Tinier network, which includes: backbone module, feature extraction module, and detection module;
[0009] Step 2.1, the backbone modules are alternately indivual Unit and The maximum pooling layer is composed of indivual Unit consists of convolutional layers and The activation function layers are alternately composed and Processing to obtain infrared activation features ;
[0010] Step 2.2, the feature extraction module is composed of a first infrared feature extraction layer and a second infrared feature extraction layer in parallel, and Processing is performed to obtain the final infrared activation features of the first layer and the final infrared activated features of the second layer ;
[0011] Step 2.3, the detection module consists of two CR units, two decoding layers and a non-maximum suppression layer, and and Process and obtain Infrared suppression characteristics ;
[0012] Step 3: Based on , construct the loss function ;
[0013] Step 4: Based on Image, use the gradient descent method to train the improved yolov3-tinier network and calculate the loss function To adjust the network parameters, when the number of training iterations reaches the set number or the loss function When convergence occurs, training stops, and the optimal human target detection model in infrared environment is obtained. It is deployed on the ZYNQ development board for detecting human targets in infrared environment.
[0014] The method for detecting human targets in an infrared environment based on an improved Yolov3-Tinier network according to the present invention is also characterized in that step 2.1 includes the following steps:
[0015] When m=1,c=1, Input into the backbone module and pass through the mth Unit No. The mth convolution layer is processed to obtain The cth infrared convolution feature output by the unit ;Will Then enter the mth Unit No. activation function layer, and use formula (1) to get the mth The cth infrared activation feature output by the unit ;
[0016] (1)
[0017] When m=1,c=2,3,…,C, the c-1th infrared activation feature After the mth Unit No. convolutional layers and the The activation function layer is processed, so that the mth Unit No. The activation function layer outputs infrared activation features ;
[0018] When m=1, c=C, Input the mth maximum pooling layer for processing to obtain the mth maximum pooling infrared feature ;
[0019] When m=2,3,…,M, the m-1th maximum pooled infrared feature Enter the mth Processing is performed in the unit to obtain the Mth The unit outputs the final infrared activation signature .
[0020] Furthermore, step 2.2 includes the following steps:
[0021] Step 2.2.1, the first infrared feature extraction layer consists of a maximum pooling layer, a feature fusion layer and F CR units, and Processing is performed to obtain the final infrared activation features of the first layer ;
[0022] Step 2.2.1.1, the maximum pooling layer in the first infrared feature extraction layer Processing is performed to obtain the maximum pooled infrared eigenvalue ;
[0023] Step 2.2.1.2, After inputting into the first CR unit for convolution and activation operations, the first infrared activation feature of the first layer is obtained. ;Will After inputting into the second CR unit for convolution and activation operations, the second infrared activation feature of the first layer is obtained ;
[0024] Step 2.2.1.3, the feature fusion layer in the first infrared feature extraction layer will and After splicing in the channel number dimension, the first infrared fusion feature is obtained ;
[0025] Step 2.2.1.4, The convolution and activation operations are performed in the remaining CR units in turn to obtain the final infrared activation features of the first layer. ;
[0026] Step 2.2.2, the second infrared feature extraction layer consists of a feature fusion layer and D CR units, and Processing is performed to obtain the final infrared activation features of the second layer ;
[0027] Step 2.2.2.1: The first CR unit pair in the second infrared characteristic extraction layer After convolution and activation operations, the first infrared activation feature of the second layer is obtained ;
[0028] Step 2.2.2.2, After inputting the convolution and activation operations into the second CR unit in the second infrared feature extraction layer, the second infrared activation feature of the second layer is obtained. ;
[0029] Step 2.2.2.3, the feature fusion layer in the second infrared feature extraction layer will and Splicing is performed in the dimension of channel number to obtain the second infrared fusion feature ;
[0030] Step 2.2.2.4, After the convolution and activation operations of the remaining CR units in the second infrared feature extraction layer, the final infrared activation features of the second layer are obtained. .
[0031] Furthermore, step 2.3 includes the following steps:
[0032] Step 2.3.1, and Input into two CR units for processing respectively, and the first infrared activation feature is obtained accordingly and a second infrared activated feature ;
[0033] Step 2.3.2, and Input into two decoding layers for processing respectively, and the first infrared decoding feature is obtained accordingly and the second infrared decoded feature ;
[0034] Step 2.3.3, and Input into the non-maximum suppression layer for processing, and get Infrared suppression characteristics ,in, represent The predicted rectangular frame of the kth infrared person's location in ; The horizontal coordinate of the center point of the predicted rectangular box representing the location of the kth infrared person in I, The vertical coordinate of the center point of the predicted rectangular box representing the location of the kth infrared person in I, Represents the width of the predicted rectangular box of the kth infrared person's location in I, Represents the height of the predicted rectangular box of the kth infrared person's location in I; represent The prediction confidence of the kth infrared personnel in ; represent The predicted category of the kth infrared person in .
[0035] Furthermore, in step 3, the loss function is constructed using formula (2) :
[0036] (2)
[0037] In formula (2), for The center point loss of the rectangular box where the kth infrared person is located, for The width and height loss of the rectangular box where the kth infrared person is located, for The confidence loss of the kth infrared person in , for The category loss of the k-th infrared person in .
[0038] The electronic device of the present invention includes a memory and a processor, and is characterized in that the memory is used to store a program that supports the processor to execute the person target detection method, and the processor is configured to execute the program stored in the memory.
[0039] The present invention provides a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program executes the steps of the person target detection method when the computer program is executed by a processor.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] 1. The present invention uses human target detection images in infrared environments as the network data set. Infrared imaging highlights human targets through temperature field imaging on the surface of objects. It is not restricted by lighting conditions and is less affected by the environment. It has a strong ability to penetrate fog, rain, snow and other weather conditions. It can overcome the phenomenon that it is difficult to distinguish objects under low light conditions in visible light images, resulting in missed detection of human targets, and plays a role in supplementing information.
[0042] 2. The improved YOLOv3-Tinier network of the present invention includes a backbone module, a feature extraction module, and a detection module. The three modules include convolution, pooling, etc., which overcome the defects of the original YOLOv3 network such as large size, large number of parameters, slow calculation speed, and large calculation carrier. While maintaining detection accuracy, the number of parameters can be reduced, the calculation speed can be accelerated, and the calculation process can be performed on a smaller carrier ZYNQ end, thereby improving the real-time and portability of human target detection in infrared environments. Human target detection can be performed in a wider range of infrared environments, thereby improving recognition efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 This is a flow chart of the improved human target detection method in the yolov3-tinier network infrared environment in the present invention;
[0044] Figure 2This is the improved yolov3-tinier network structure diagram in the present invention.
[0045] Figure 3 This is the architecture diagram of the improved yolov3-tinier network infrared environment personnel target detection system in the present invention. DETAILED DESCRIPTION
[0046] In this embodiment, a method for detecting human targets in an infrared environment based on an improved yolov3-tinier network is described. Figure 1 , including the following steps:
[0047] Step 1: Get the infrared environment personnel detection dataset Image, and let any infrared personnel image in Image be recorded as ,Will The positions and categories of the K infrared personnel in the image are marked to obtain the position labels and category labels. The location tag of the kth infrared person in is recorded as ,and ,in, Indicates the Infrared personnel in The coordinates of the center point of the rectangular box at the position in , and =( ), Indicates the horizontal coordinate of the center point of the rectangular box, Indicates the vertical coordinate of the center point of the rectangular box, Indicates the Infrared personnel in The size of the rectangular box at the position in , and = , Indicates the Infrared personnel in The width of the rectangular frame at the position in Indicates the Personnel in The height of the rectangular box at the position in The category label of the kth infrared person in is recorded as In this example, the LLVIP low-light vision infrared dataset is used for processing. The model's input images are uniformly resized using bilinear interpolation, with the final resized value set to 416×416.
[0048] Step 2: Establish an improved yolov3-tinier network, such as Figure 2 As shown, it includes: backbone module, feature extraction module, and detection module in sequence;
[0049] Step 2.1, the backbone modules are alternately indivual Unit and The maximum pooling layer is composed of indivual Unit consists of convolutional layers and activation function layers are alternately composed; in this embodiment, Take 5, Taking 1, the backbone module can ensure a certain feature acquisition capability while making the inference speed faster, the model size smaller, the memory usage smaller, and the risk of overfitting lower.
[0050] When m=1,c=1, Input into the backbone module and pass through the mth Unit No. The mth convolution layer is processed to obtain The cth infrared convolution feature of the unit ; Then enter the mth Unit No. activation function layer, and use formula (1) to get the mth The cth infrared activation feature of the unit ;
[0051]
[0052] When m=1,c=2,3,…,C, the c-1th infrared activation feature After the mth Unit No. convolutional layers and the The activation function layer is processed, so that the mth Unit No. The activation function layer outputs infrared activation features In this embodiment, the first indivual The convolutional layer of the unit is composed of filters, each filter size is 3×3, and each filter has a total of 2× convolution kernels, and the stride and padding value of each convolution kernel is 1.
[0053] When m=1, c=C, Input the mth maximum pooling layer for processing to obtain the mth maximum pooling infrared feature In this embodiment, the size of any pooling layer in the backbone module is 2×2, and the stride is 2.
[0054] When m=2,3,…,M, the m-1th maximum pooled infrared feature Enter the mth The Mth The unit outputs the final infrared activation signature .
[0055] Step 2.2, the feature extraction module is composed of a first infrared feature extraction layer and a second infrared feature extraction layer in parallel;
[0056] Step 2.2.1, the first infrared feature extraction layer consists of a maximum pooling layer, a feature fusion layer and F CR units, and Processing is performed to obtain the final infrared activation features of the first layer ; In this embodiment, the value of F is 3;
[0057] Step 2.2.1.1, the maximum pooling layer in the first infrared feature extraction layer Processing is performed to obtain the maximum pooled infrared eigenvalue In this embodiment, infrared features are extracted through the maximum pooling operation, thereby enhancing the generalization ability of the model and retaining the most significant infrared maximum pooling features.
[0058] Step 2.2.1.2, After inputting into the first CR unit for convolution and activation operations, the first infrared activation feature of the first layer is obtained. ;Will After inputting into the second CR unit for convolution and activation operations, the second infrared activation feature of the first layer is obtained In this embodiment, while performing infrared feature extraction, dimensionality reduction and information concentration can be performed to construct a hierarchical infrared feature representation;
[0059] Step 2.2.1.3, the feature fusion layer in the first infrared feature extraction layer will and After splicing in the channel number dimension, the first infrared fusion feature is obtained ; In this embodiment, feature fusion aims to capture richer contextual infrared information and multi-scale infrared details.
[0060] Step 2.2.1.4, The convolution and activation operations are performed in the remaining CR units in turn to obtain the final infrared activation features of the first layer. In this embodiment, the convolution and activation operations are performed after the feature fusion operation, which can further Extract key infrared features while preserving important structural information.
[0061] Step 2.2.2, the second infrared feature extraction layer consists of a feature fusion layer and D CR units, and Processing is performed to obtain the final infrared activation features of the second layer ; In this embodiment, the value of D is 3;
[0062] Step 2.2.2.1: The first CR unit pair in the second infrared characteristic extraction layer After convolution and activation operations, the first infrared activation feature of the second layer is obtained ;
[0063] Step 2.2.2.2, After inputting the convolution and activation operations into the second CR unit in the second infrared feature extraction layer, the second infrared activation feature of the second layer is obtained. .
[0064] Step 2.2.2.3, the feature fusion layer in the second infrared feature extraction layer will and Splicing is performed in the dimension of channel number to obtain the second infrared fusion feature ;
[0065] Step 2.2.2.4, After the convolution and activation operations of the remaining CR units in the second infrared feature extraction layer, the final infrared activation features of the second layer are obtained. .
[0066] Step 2.3: The detection module consists of two CR units, two decoding layers, and a non-maximum suppression layer. In this embodiment, each filter of the CR unit of the detection module contains 255×256 convolution kernels, each of which has a size of 1×1, a stride value of 1, and a padding value of 0.
[0067] Step 2.3.1, and Input into two CR units for processing respectively, and the first infrared activation feature is obtained accordingly and a second infrared activated feature ;
[0068] Step 2.3.2, and Input into two decoding layers for processing respectively, and the first infrared decoding feature is obtained accordingly and the second infrared decoded feature .
[0069] Step 2.3.3, and Input into the non-maximum suppression layer for processing, and get Infrared suppression characteristics ,in, represent The predicted rectangular frame of the kth infrared person's location in ; The horizontal coordinate of the center point of the predicted rectangular box representing the location of the kth infrared person in I, The vertical coordinate of the center point of the predicted rectangular box representing the location of the kth infrared person in I, Represents the width of the predicted rectangular box of the kth infrared person's location in I, Represents the height of the predicted rectangular box of the kth infrared person's location in I; represent The prediction confidence of the kth infrared personnel in ; represent The predicted category of the kth infrared person in .
[0070] Step 3: Use formula (2) to construct the loss function :
[0071]
[0072] In formula (2) for The center point loss of the rectangular box where the kth infrared person is located, for The width and height loss of the rectangular box where the kth infrared person is located, for The confidence loss of the kth infrared person in , for The category loss of the k-th infrared person in .
[0073] In this embodiment, formula (3), formula (4), formula (5) and formula (6) are used to construct 、 、 and :
[0074] (3)
[0075] (4)
[0076] (5)
[0077] (6)
[0078] In formula (3), It is a coordination function used to coordinate the inconsistent contributions of rectangular boxes of different sizes to the error function; Indicates whether the predicted rectangle of the k-th infrared personnel position in I successfully predicts the k-th infrared personnel target. If the k-th infrared personnel target is successfully predicted, let =1, otherwise, let =0;
[0079] In formula (5), Indicates that the predicted rectangular box at the location of the kth infrared person in I contains the true value of the infrared person; is a weight value, which indicates the weight of the confidence error in the loss function when the predicted rectangle does not predict the infrared person target; If the predicted rectangle of the kth infrared person location in I fails to successfully predict the kth infrared person target, then let =1, otherwise, let =0;
[0080] Step 4: Based on Image, use the gradient descent method to train the improved yolov3-tinier network and calculate the loss function To adjust the network parameters, when the number of training iterations reaches the set number or the loss function When convergence occurs, training stops, thereby obtaining the optimal human target detection model in an infrared environment, which is deployed on the ZYNQ development board for detecting human targets in an infrared environment. In this embodiment, the number of training rounds is set to 1000, and the trained model is obtained after 300 epochs of training iterations.
[0081] In this embodiment, an electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0082] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are executed.
[0083] In this embodiment, a method for detecting human targets in an infrared environment based on an improved yolov3-tinier network is based on a hardware system. Figure 3The system includes an infrared camera, a ZYNQ development board, an HDMI data stream transmission cable, and an HDMI display. The infrared camera, mounted on the ZYNQ development board, captures images of people in an infrared environment and stores them in the DDR on the PS for detection. The trained improved YOLOv3-Tinier network is loaded onto the ZYNQ development board to detect the infrared image data in the DDR. The HDMI display, connected to the ZYNQ development board via the HDMI data stream transmission cable, displays the results of person detection.
Claims
1. A method for detecting human targets in infrared environment based on improved yolov3-tinier network, characterized in that: The steps include: Step 1: Get the infrared environment personnel detection dataset Image, and let any infrared personnel image in Image be recorded as I, mark the positions and categories of K infrared personnel in I, obtain the position label and category label, and record the position label of the kth infrared personnel in I as ,and ,in, represents the center coordinate of the rectangular box where the kth infrared person is located in I, and , represents the horizontal coordinate of the center point of the rectangular frame, represents the vertical coordinate of the center point of the rectangular frame, represents the size of the rectangular box where the kth infrared person is located in I, and , The width of the rectangular box representing the position of the kth infrared person in I, The height of the rectangular box representing the position of the kth person in I; the category label of the kth infrared person in I is recorded as ; Step 2: Establish an improved Yolov3-Tinier network, which includes: backbone module, feature extraction module, and detection module; Step 2.1, the backbone module is composed of alternating M CR units and M-1 maximum pooling layers, wherein any m-th CR unit is composed of C convolutional layers and C activation function layers alternately, and I is processed to obtain the infrared activation feature ; Step 2.2, the feature extraction module is composed of a first infrared feature extraction layer and a second infrared feature extraction layer in parallel, and Processing is performed to obtain the final infrared activation features of the first layer and the final infrared activation features of the second layer ; Step 2.3, the detection module consists of two CR units, two decoding layers and a non-maximum suppression layer, and and Processing is performed to obtain the infrared suppression characteristics of I ; Step 3: Based on , construct the loss function Loss; Step 4: Based on Image, the improved YOLOv3-Tinier network is trained using the gradient descent method, and the loss function Loss is calculated to adjust the network parameters. When the number of training iterations reaches the set number or the loss function Loss converges, the training stops, thereby obtaining the optimal human target detection model in the infrared environment, and deploying it on the ZYNQ development board for detecting human targets in the infrared environment.
2. A method for detecting human targets in an infrared environment based on an improved yolov3-tinier network according to claim 1, characterized in that: Step 2.1 includes the following steps: When m=1, c=1, I is input into the backbone module and processed by the cth convolution layer of the mth CR unit to obtain the cth infrared convolution feature output by the mth CR unit. ;Will Then input it into the cth activation function layer of the mth CR unit, and use formula (1) to get the cth infrared activation feature output by the mth CR unit ; (1) When m=1, c=2,3,…,C, the c-1th infrared activation feature After being processed by the cth convolution layer and the cth activation function layer of the mth CR unit in sequence, the infrared activation feature is output by the Cth activation function layer of the mth CR unit. ; When m=1, c=C, Input the mth maximum pooling layer for processing to obtain the mth maximum pooling infrared feature ; When m=2,3,…,M, the m-1th maximum pooled infrared feature Input into the mth CR unit for processing, and the Mth CR unit outputs the final infrared activation feature .
3. A method for detecting human targets in an infrared environment based on an improved yolov3-tinier network according to claim 2, characterized in that: Step 2.2 includes the following steps: Step 2.2.1, the first infrared feature extraction layer consists of a maximum pooling layer, a feature fusion layer and F CR units, and Processing is performed to obtain the final infrared activation features of the first layer ; Step 2.2.1.1, the maximum pooling layer in the first infrared feature extraction layer Processing is performed to obtain the maximum pooled infrared eigenvalue ; Step 2.2.1.2, After inputting into the first CR unit for convolution and activation operations, the first infrared activation feature of the first layer is obtained. ;Will After inputting into the second CR unit for convolution and activation operations, the second infrared activation feature of the first layer is obtained ; Step 2.2.1.3, the feature fusion layer in the first infrared feature extraction layer will and After splicing in the channel number dimension, the first infrared fusion feature is obtained ; Step 2.2.1.4, The convolution and activation operations are performed in the remaining CR units in turn to obtain the final infrared activation features of the first layer. ; Step 2.2.2, the second infrared feature extraction layer consists of a feature fusion layer and D CR units, and Processing is performed to obtain the final infrared activation features of the second layer ; Step 2.2.2.1: The first CR unit pair in the second infrared characteristic extraction layer After convolution and activation operations, the first infrared activation feature of the second layer is obtained ; Step 2.2.2.2, After inputting the convolution and activation operations into the second CR unit in the second infrared feature extraction layer, the second infrared activation feature of the second layer is obtained. ; Step 2.2.2.3, the feature fusion layer in the second infrared feature extraction layer will and Splicing is performed in the dimension of channel number to obtain the second infrared fusion feature ; Step 2.2.2.4, After the convolution and activation operations of the remaining CR units in the second infrared feature extraction layer, the final infrared activation features of the second layer are obtained. .
4. A method for detecting human targets in an infrared environment based on an improved yolov3-tinier network according to claim 3, characterized in that: Step 2.3 includes the following steps: Step 2.3.1, and Input into two CR units for processing respectively, and the first infrared activation feature is obtained accordingly and a second infrared activated feature ; Step 2.3.2, and Input into two decoding layers for processing respectively, and the first infrared decoding feature is obtained accordingly and the second infrared decoded feature ; Step 2.3.3, and Input into the non-maximum suppression layer for processing to obtain the infrared suppression feature of I ,in, Represents the predicted rectangular box of the kth infrared person's location in I; The horizontal coordinate of the center point of the predicted rectangular box representing the location of the kth infrared person in I, The vertical coordinate of the center point of the predicted rectangular box representing the location of the kth infrared person in I, Represents the width of the predicted rectangular box of the kth infrared person's location in I, Represents the height of the predicted rectangular box of the kth infrared person's location in I; Represents the prediction confidence of the kth infrared person in I; where, Represents the predicted category of the kth infrared person in I.
5. A method for detecting human targets in an infrared environment based on an improved yolov3-tinier network according to claim 4, characterized in that: In step 3, the loss function Loss is constructed using formula (2): (2) In formula (2), is the center point loss of the rectangular box where the kth infrared person is located in I, is the width and height loss of the rectangular box where the kth infrared person is located in I, is the confidence loss of the kth infrared person in I, is the category loss of the kth infrared person in I.
6. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the person target detection method according to any one of claims 1 to 5, and the processor is configured to execute the program stored in the memory.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for detecting a person target according to any one of claims 1 to 5 are executed.
Citation Information
Patent Citations
Infrared image weak and small target detection method based on improved YOLO v3
CN112101434A
Infrared target detection method based on improved YOLOv3
CN112949633A