Target detection method and system applied to rescue
Target detection is carried out through the YOLOv5 model, which solves the problem of many steps and long time for personnel going out when the rescue vehicle or equipment, and realizes the identification and connection of targets in the rescue vehicle, improving the rescue efficiency.
Patent Information
- Application Number
- CN202410467141.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-18
- Publication Date
- 2025-07-25
AI Technical Summary
When rescuing large vehicles or equipment, the prior art requires personnel to operate outside the vehicle, with many steps and long time, resulting in inefficiency.
The YOLOv5 model is used for object detection, and four feature extractions are performed through the backbone network. The neck network enhances high-level semantics and low-level detailed features, the head network performs regression prediction, and images are obtained using depth cameras and target recognition and connection are completed in the rescue vehicle.
The target detection is completed in the rescue vehicle, reducing the operation steps of personnel going out, and improving the rescue efficiency.
Smart Images

Figure CN120374930A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of rescue technologies, and in particular to a target detection method and system applied to a rescue device. Background Art
[0002] Currently, when rescuing large vehicles or equipment, most rescue vehicles are started. After arriving at the rescue location, rescue personnel need to get out of the vehicle to operate. First, the towing connecting rod and connecting piece are removed from the rescue vehicle, and then the connecting rod and connecting piece are installed in the correct order of operations. Then, the rescue vehicle towes the vehicle or equipment away from the scene.
[0003] This operation method requires personnel to operate outside the vehicle, and there are problems such as many operation steps and long time consumption.
[0004] Therefore, there is an urgent need for a target detection method applied to rescue to solve the problem that personnel need to operate outside the vehicle, which not only has many steps but also takes a long time. Summary of the Invention
[0005] The present invention provides a target detection method and system applied to a rescue device to solve the problem that personnel need to operate outside the vehicle, which not only has many steps but also takes a long time.
[0006] The present invention provides a target detection method applied to rescue in a first aspect. The method includes:
[0007] Obtain an initial input image, process the initial input image to obtain a target input image;
[0008] Utilize the backbone network in the YOLOv5 model to perform four times of feature extraction on the target input image in sequence, and obtain a first feature map, a second feature map, a third feature map, and a fourth feature map in sequence;
[0009] Utilize the neck network in the YOLOv5 model to enhance the high-level semantic feature information from top to bottom and enhance the low-level detail features from bottom to top for the first feature map, the second feature map, the third feature map, and the fourth feature map respectively, and obtain a first prediction layer, a second prediction layer, a third prediction layer, and a fourth prediction layer in sequence;
[0010] Utilize the head network in the YOLOv5 model to perform regression prediction on the first prediction layer, the second prediction layer, the third prediction layer, and the fourth prediction layer, and obtain a first target layer, a second target layer, a third target layer, and a fourth target layer in sequence;
[0011] Utilize the first target layer, the second target layer, the third target layer, and the fourth target layer to predict the target of the rescue to obtain a prediction result.
[0012] In some implementable ways, the steps of obtaining an initial input image and processing the initial input image to obtain a target input image include:
[0013] Performing depth processing on the initial input image to obtain an input grayscale image;
[0014] Performing size scaling processing on the input grayscale image to obtain the target input image.
[0015] In some implementable ways, the steps of using the backbone network in the YOLOv5 model to sequentially perform four feature extractions on the target input image to obtain a first feature map, a second feature map, a third feature map, and a fourth feature map include:
[0016] Using the backbone network in the YOLOv5 model to perform convolution processing on the target input image and performing feature extraction on the target input image after convolution processing to obtain the first feature map;
[0017] Using the backbone network in the YOLOv5 model to perform convolution processing on the first feature map and performing feature extraction on the first feature map after convolution processing to obtain the second feature map;
[0018] Using the backbone network in the YOLOv5 model to perform convolution processing on the second feature map and performing feature extraction on the second feature map after convolution processing to obtain the third feature map;
[0019] Using the backbone network in the YOLOv5 model to perform convolution processing on the third feature map and performing feature extraction on the third feature map after convolution processing to obtain the fourth feature map.
[0020] In some implementable ways, the steps of using the neck network in the YOLOv5 model to respectively perform top-down enhancement of high-level semantic feature information and bottom-up enhancement of low-level detail features on the first feature map, the second feature map, the third feature map, and the fourth feature map to sequentially obtain a first prediction layer, a second prediction layer, a third prediction layer, and a fourth prediction layer include:
[0021] Using the neck network in the YOLOv5 model to sequentially perform convolution and upsampling on the fourth feature map, and then performing splicing processing with the third feature map to obtain a first enhanced high-level semantic feature information map;
[0022] Using the neck network in the YOLOv5 model to sequentially perform convolution and upsampling on the first enhanced high-level semantic feature information map, and then splicing it with the second feature map to obtain second enhanced high-level semantic feature information;
[0023] Using the neck network in the YOLOv5 model, after performing convolution and upsampling on the second enhanced high-level semantic feature information map in sequence, and then splicing it with the first feature map, a third enhanced high-level semantic feature information map is obtained. Among them, the first enhanced high-level semantic feature information map, the second enhanced high-level semantic feature information map, and the third enhanced high-level semantic feature information map form enhanced high-level semantic features, and the semantic information contained in the first enhanced high-level semantic feature information map, the second enhanced high-level semantic feature information map, and the third enhanced high-level semantic feature information map is enhanced in sequence.
[0024] In some implementable ways, the steps of using the neck network in the YOLOv5 model to respectively perform top-down enhanced high-level semantic features and bottom-up enhanced low-level detail features on the first feature map, the second feature map, the third feature map, and the fourth feature map, and sequentially obtaining the first prediction layer, the second prediction layer, the third prediction layer, and the fourth prediction layer further include:
[0025] Using the backbone network in the YOLOv5 model, after convolving the third enhanced high-level semantic feature information map, and then splicing it with the convolved second enhanced high-level semantic feature information map, the second prediction layer is obtained, where the third enhanced high-level semantic feature information represents the first prediction layer;
[0026] Using the backbone network in the YOLOv5 model, after convolving the second prediction layer, and then splicing it with the convolved first enhanced high-level semantic feature information map, the third prediction layer is obtained;
[0027] Using the backbone network in the YOLOv5 model, after the third prediction layer, and then splicing it with the convolved fourth feature map, the fourth prediction layer is obtained.
[0028] In some implementable ways, the steps of using the head network in the YOLOv5 model to perform regression prediction on the first prediction layer, the second prediction layer, the third prediction layer, and the fourth prediction layer, and sequentially obtaining the first target layer, the second target layer, the third target layer, and the fourth target layer include:
[0029] Using the head network in the YOLOv5 model, and using the linear regression algorithm to perform regression prediction on the first prediction layer to obtain the first target layer;
[0030] Using the head network in the YOLOv5 model, and using the linear regression algorithm to perform regression prediction on the second prediction layer to obtain the second target layer;
[0031] Using the head network in the YOLOv5 model, perform regression prediction on the third prediction layer using the linear regression algorithm to obtain the third target layer;
[0032] Using the head network in the YOLOv5 model, perform regression prediction on the fourth prediction layer using the linear regression algorithm to obtain the fourth target layer.
[0033] In some implementable ways, the step of using the first target layer, the second target layer, the third target layer, and the fourth target layer to predict the target of the rescue and obtain a prediction result includes:
[0034] Using the first target layer, the second target layer, the third target layer, and the fourth target layer to predict the target of the rescue to obtain the confidence level and coordinate information of the detected target;
[0035] According to the confidence level and coordinate information of the detected target, obtain the prediction result.
[0036] The second aspect of the present invention provides a target detection system for rescue, which is applied to the aforementioned target detection method for rescue. The system includes:
[0037] A depth camera, configured to acquire an initial input image, process the initial input image to obtain a target input image;
[0038] A processing unit, configured to use the backbone network in the YOLOv5 model to perform four times of feature extraction on the target input image in sequence to obtain a first feature map, a second feature map, a third feature map, and a fourth feature map in sequence;
[0039] The processing unit is further configured to use the neck network in the YOLOv5 model to enhance the high-level semantic feature information from top to bottom and enhance the low-level detail features from bottom to top for the first feature map, the second feature map, the third feature map, and the fourth feature map respectively, and obtain a first prediction layer, a second prediction layer, a third prediction layer, and a fourth prediction layer in sequence;
[0040] The processing unit is further configured to use the head network in the YOLOv5 model to perform regression prediction on the first prediction layer, the second prediction layer, the third prediction layer, and the fourth prediction layer to obtain a first target layer, a second target layer, a third target layer, and a fourth target layer in sequence;
[0041] An output unit, configured to use the first target layer, the second target layer, the third target layer, and the fourth target layer to predict the target of the rescue to obtain a prediction result.
[0042] The third aspect of the present invention provides a connection structure of an equipment rescue device, which is applied to the aforementioned target detection system for rescue, and includes a connection structure main body. A depth camera, a processing unit, and an output unit are arranged on the connection structure main body, and the processing unit is respectively connected to the camera and the output unit.
[0043] The fourth aspect of the present invention provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the aforementioned target detection method for rescue.
[0044] Advantages of the present invention:
[0045] The present invention provides a target detection method for rescue. First, an initial input image is obtained and processed to obtain a target input image; then, the backbone network in the YOLOv5 model is used to perform four times of feature extraction on the target input image in sequence to obtain a first feature map, a second feature map, a third feature map, and a fourth feature map respectively; the neck network in the YOLOv5 model is used to enhance the high-level semantic feature information from top to bottom and the low-level detail features from bottom to top for the first feature map, the second feature map, the third feature map, and the fourth feature map respectively to obtain a first prediction layer, a second prediction layer, a third prediction layer, and a fourth prediction layer respectively; the head network in the YOLOv5 model is used to perform regression prediction on the first prediction layer, the second prediction layer, the third prediction layer, and the fourth prediction layer to obtain a first target layer, a second target layer, a third target layer, and a fourth target layer respectively; finally, the first target layer, the second target layer, the third target layer, and the fourth target layer are used to predict the rescue target to obtain a prediction result. Through the above method, the detection of the rescue target is realized, so that the personnel in the rescue vehicle can complete the rescue of large vehicles or equipment by using the corresponding connection structure of the equipment rescue device through the identified rescue target information. Description of the drawings
[0046] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 It is a flowchart of a target detection method for rescue according to the present invention;
[0048] Figure 2 For a target detection method applied to rescue in the present invention;
[0049] Figure 3 It is a schematic structural diagram of a connection structure of a rescue device equipped with the present invention.
[0050] Explanation of reference numerals:
[0051] 1. Upper oil cylinder seat; 2. Connection seat; 3. Main oil cylinder; 4. Side oil cylinder; 5. Upper oil cylinder; 6. Binocular camera; 7. Bracket; 8. Shackle; 9. Ball eye joint; 10. Multi-way solenoid valve; 11. Edge computing box. Specific implementation manners
[0052] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0053] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0054] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined. In addition, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0055] Next, some terms involved in the present application will be explained to facilitate the understanding of the present invention:
[0056] YOLOv5 is a real-time object detection algorithm characterized by high accuracy and fast inference speed.
[0057] As Figure 1 and Figure 2 shown, the first aspect of this application provides an object detection method for rescue, and the method includes:
[0058] S100, obtain an initial input image, process the initial input image to obtain a target input image.
[0059] Specifically, first, perform depth processing on the initial input image to obtain an input grayscale image; then, perform size scaling processing on the input grayscale image to obtain the target input image.
[0060] Performing grayscale processing on the input image, that is, converting a color image into a grayscale image, can effectively reduce the amount of calculation and improve the operation speed. In addition, for the input grayscale image, perform size adjustment. Exemplarily, the pixels of the initial input image are 1920x1080. For the convenience of operation and unified annotation, the pixels of the input image are adjusted to 640x640.
[0061] S200, use the backbone network in the YOLOv5 model to perform four times of feature extraction on the target input image in sequence, and obtain a first feature map, a second feature map, a third feature map, and a fourth feature map in sequence.
[0062] Among them, use the backbone network in the YOLOv5 model to perform convolution processing on the target input image, and perform feature extraction on the target input image after convolution processing to obtain the first feature map.
[0063] The convolution processing method mentioned in this application is a conventional convolution processing method. This application does not limit how to implement convolution processing. In addition, for feature extraction, this application also adopts a conventional image feature extraction method. This application does not limit how to implement image feature extraction.
[0064] In addition, the convolution processing of the target input image can be once or twice. Exemplarily, in the convolution processing of the target input image using the backbone network in the YOLOv5 model, two convolution operations are adopted to improve the operation speed in subsequent steps.
[0065] Use the backbone network in the YOLOv5 model to perform convolution processing on the first feature map, and perform feature extraction on the first feature map after convolution processing to obtain the second feature map.
[0066] Use the backbone network in the YOLOv5 model to perform convolution processing on the second feature map, and extract features from the second feature map after convolution processing to obtain the third feature map.
[0067] Use the backbone network in the YOLOv5 model to perform convolution processing on the third feature map, and extract features from the third feature map after convolution processing to obtain the fourth feature map.
[0068] It should be noted that by continuously processing the input image, the recognition effect of image features is enhanced so that more image features can be recognized.
[0069] S300, use the neck network in the YOLOv5 model to enhance the high-level semantic feature information from top to bottom and the low-level detail features from bottom to top for the first feature map, the second feature map, the third feature map, and the fourth feature map respectively, and sequentially obtain the first prediction layer, the second prediction layer, the third prediction layer, and the fourth prediction layer.
[0070] Among them, after sequentially performing convolution and upsampling on the fourth feature map using the neck network in the YOLOv5 model, it is concatenated with the third feature map to obtain the first enhanced high-level semantic feature information map.
[0071] After sequentially performing convolution and upsampling on the first enhanced high-level semantic feature information map using the neck network in the YOLOv5 model, it is then concatenated with the second feature map to obtain the second enhanced high-level semantic feature information.
[0072] After sequentially performing convolution and upsampling on the second enhanced high-level semantic feature information map using the neck network in the YOLOv5 model, it is then concatenated with the first feature map to obtain the third enhanced high-level semantic feature information map.
[0073] Among them, the first enhanced high-level semantic feature information map, the second enhanced high-level semantic feature information map, and the third enhanced high-level semantic feature information map constitute the enhanced high-level semantic feature information, and the semantic information contained in the first enhanced high-level semantic feature information map, the second enhanced high-level semantic feature information map, and the third enhanced high-level semantic feature information map is enhanced sequentially.
[0074] Specifically, after obtaining the first feature map, the second feature map, the third feature map, and the fourth feature map in the aforementioned steps, in a top-down manner, the fourth feature map, the third feature map, the second feature map, and the first feature map are sequentially enhanced for high-level semantic feature information recognition, gradually increasing the high-level semantic information, so as to help improve the accuracy and robustness of object detection through the high-level semantic feature information. The high-level semantic information realizes the deep feature representation of the feature map, and these features contain a more abstract and high-level understanding of the feature map, which helps the model better understand the semantic information in the scene. By utilizing the high-level semantic feature information, YOLOv5 can better understand the rescue targets in the image and can perform object detection and localization more accurately. These high-level semantic feature information can help the YOLOv5 model identify information such as the shapes and textures of complex rescue targets, thereby improving the performance of rescue object detection.
[0075] It should be noted that for the first enhanced high-level semantic feature information map, the second enhanced high-level semantic feature information map, and the third enhanced high-level semantic feature information map, after convolving the third enhanced high-level semantic feature information map using the backbone network in the YOLOv5 model, it is then concatenated with the convolved second enhanced high-level semantic feature information map to obtain the second prediction layer.
[0076] Among them, the third enhanced high-level semantic feature information represents the first prediction layer and can be directly sent to the head network.
[0077] Using the backbone network in the YOLOv5 model, after convolving the second prediction layer, it is then concatenated with the convolved first enhanced high-level semantic feature information map to obtain the third prediction layer.
[0078] Using the backbone network in the YOLOv5 model, after the third prediction layer, it is then concatenated with the convolved fourth feature map to obtain the fourth prediction layer.
[0079] Exemplarily, 8-fold, 16-fold, and 32-fold downsampling are respectively performed in the Neck network. The feature map sizes of the corresponding detection layers (conventional units existing in the YOLOv5 model) are 80×80, 40×40, and 20×20, which are respectively used to detect small targets, medium targets, and large targets. On this basis, to improve the recognition accuracy of small targets at long distances, a detection layer is added to the YOLOv5 model network. The detection layer adds a first feature map. That is to say, one upsampling is added in the Neck network. After the third upsampling (upsampling the first feature map), it is fused with the second layer of the backbone network to obtain a newly added prediction layer of 160×160 for detecting small targets. A detection layer with 4 prediction scales (corresponding to four feature maps respectively) is adopted, making full use of the high-resolution of the underlying features and the high semantic information of the deep features, and at the same time not significantly increasing the complexity of the YOLOv5 model network.
[0080] S400, using the head network in the YOLOv5 model, perform regression prediction on the first prediction layer, the second prediction layer, the third prediction layer, and the fourth prediction layer, and sequentially obtain the first target layer, the second target layer, the third target layer, and the fourth target layer.
[0081] Among them, using the head network in the YOLOv5 model, use the linear regression algorithm to perform regression prediction on the first prediction layer to obtain the first target layer.
[0082] Using the head network in the YOLOv5 model, use the linear regression algorithm to perform regression prediction on the second prediction layer to obtain the second target layer.
[0083] Using the head network in the YOLOv5 model, use the linear regression algorithm to perform regression prediction on the third prediction layer to obtain the third target layer.
[0084] Using the head network in the YOLOv5 model, use the linear regression algorithm to perform regression prediction on the fourth prediction layer to obtain the fourth target layer.
[0085] Specifically, the linear regression algorithm used by the head network in the YOLOv5 model is a conventional algorithm, and this application does not improve it. Use the linear regression algorithm to perform regression prediction on the four detection layers in the Neck network respectively, so as to obtain the first target layer, the second target layer, the third target layer, and the fourth target layer.
[0086] S500, use the first target layer, the second target layer, the third target layer, and the fourth target layer to predict the rescue target and obtain the prediction result.
[0087] Among them, first, the first target layer, the second target layer, the third target layer, and the fourth target layer are used to predict the target of rescue, and the confidence and coordinate information of the detected target are obtained. Then, according to the confidence and coordinate information of the detected target, the prediction result is obtained. That is to say, in the YOLOv5 model, the input is an image (usually a three-dimensional array, specifically a tensor with a shape of (height, width, channels), where channels are usually 3 (corresponding to the RGB color channels)), and the output is a tensor containing the detection results. This tensor usually contains information about each detected object, including: Bounding box coordinates: including the pixel coordinates of the upper left corner, lower right corner, and center point of the bounding box; Confidence: the confidence of the model that the detected object is indeed a certain category; Class name.
[0088] The second aspect of the present invention provides a target detection system for rescue, which is applied to the aforementioned target detection method for rescue. The system includes:
[0089] A depth camera 6, configured to obtain an initial input image, process the initial input image, and obtain a target input image;
[0090] A processing unit, configured to use the backbone network in the YOLOv5 model to perform four times of feature extraction on the target input image in sequence, and obtain a first feature map, a second feature map, a third feature map, and a fourth feature map in sequence;
[0091] The processing unit is further configured to use the neck network in the YOLOv5 model to enhance the high-level semantic feature information from top to bottom and the low-level detailed features from bottom to top for the first feature map, the second feature map, the third feature map, and the fourth feature map respectively, and obtain a first prediction layer, a second prediction layer, a third prediction layer, and a fourth prediction layer in sequence;
[0092] The processing unit is further configured to use the head network in the YOLOv5 model to perform regression prediction on the first prediction layer, the second prediction layer, the third prediction layer, and the fourth prediction layer, and obtain a first target layer, a second target layer, a third target layer, and a fourth target layer in sequence;
[0093] An output unit, configured to use the first target layer, the second target layer, the third target layer, and the fourth target layer to predict the target of rescue and obtain a prediction result.
[0094] As Figure 3As shown in the figure, the third aspect of the present invention provides a connection structure for an equipment rescue device, which is applied to the aforementioned target detection system for rescue. It includes a connection structure main body, on which a depth camera, a processing unit, and an output unit are provided. The processing unit is respectively connected to the camera and the output unit.
[0095] Specifically, the processing unit and the output unit are loaded into the edge computing box 11. After obtaining the prediction result, the edge computing box 11 is used for output. By performing inverse operations, the target position that the magnetostrictive displacement sensor needs to reach is obtained. The edge computing box 11 issues an action instruction, and the multi-way solenoid valve 10 controls the on-off of the oil pipes in the main cylinder 3, the upper cylinder 5, and the side cylinder 4, so that the main cylinder 3, the upper cylinder 5, and the side cylinder 4 make corresponding stretching or retracting movements. After a series of telescopic movements, finally, the shackle 8 moves to the vicinity of the target, and the shackle 8 is used to connect with the rescue target.
[0096] It should be noted that one end of the fish-eye joint 9 is connected to the telescopic end of the main cylinder 3, and the other end of the fish-eye joint 9 is rotatably connected to the shackle 8. The upper cylinder seat 1 is arranged at the fixed end of the upper cylinder 5, and the upper cylinder seat 1 is rotatably connected to the rescue vehicle. The through hole on the connecting seat 2 is connected to the tail hook of the rescue vehicle. When the hydraulic pump is started, the fixed end of the main cylinder 3 is subjected to the acting force generated by the telescopic end of the upper cylinder 5 and moves horizontally on the connecting seat 2; the fixed end of the main cylinder 3 is subjected to the acting force generated by the telescopic end of the side cylinder 4, so as to move vertically on the connecting seat 2. Synchronously, the telescopic end of the main cylinder 3 also expands and contracts to connect with the equipment to be rescued. The depth camera 6 is arranged on the bracket, and the position of the bracket can be placed according to needs.
[0097] The fourth aspect of the present invention provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the aforementioned target detection method for rescue.
[0098] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A target detection method applied to rescue, characterized in that The method includes: Obtain an initial input image, process the initial input image to obtain a target input image; Utilize the backbone network in the YOLOv5 model to perform four times of feature extraction on the target input image successively, and obtain a first feature map, a second feature map, a third feature map and a fourth feature map in sequence; Utilize the neck network in the YOLOv5 model to enhance the high-level semantic feature information from top to bottom and enhance the low-level detail features from bottom to top for the first feature map, the second feature map, the third feature map and the fourth feature map respectively, and obtain a first prediction layer, a second prediction layer, a third prediction layer and a fourth prediction layer in sequence; Utilize the head network in the YOLOv5 model to perform regression prediction on the first prediction layer, the second prediction layer, the third prediction layer and the fourth prediction layer, and obtain a first target layer, a second target layer, a third target layer and a fourth target layer in sequence; Utilize the first target layer, the second target layer, the third target layer and the fourth target layer to predict the target of the rescue and obtain a prediction result.
2. The object detection method for rescue according to claim 1, wherein The step of obtaining an initial input image, processing the initial input image to obtain a target input image includes: Perform depth processing on the initial input image to obtain an input grayscale image; Perform size scaling processing on the input grayscale image to obtain the target input image.
3. The object detection method for rescue according to claim 1, characterized in that, The step of utilizing the backbone network in the YOLOv5 model to perform four times of feature extraction on the target input image successively, and obtain a first feature map, a second feature map, a third feature map and a fourth feature map in sequence includes: Utilize the backbone network in the YOLOv5 model to perform convolution processing on the target input image, and perform feature extraction on the target input image after convolution processing to obtain the first feature map; Utilize the backbone network in the YOLOv5 model to perform convolution processing on the first feature map, and perform feature extraction on the first feature map after convolution processing to obtain the second feature map; Utilize the backbone network in the YOLOv5 model to perform convolution processing on the second feature map, and perform feature extraction on the second feature map after convolution processing to obtain the third feature map; Utilize the backbone network in the YOLOv5 model to perform convolution processing on the third feature map, and perform feature extraction on the third feature map after convolution processing to obtain the fourth feature map.
4. The object detection method for rescue according to claim 1, wherein The step of utilizing the neck network in the YOLOv5 model to enhance the high-level semantic feature information from top to bottom and enhance the low-level detail features from bottom to top for the first feature map, the second feature map, the third feature map and the fourth feature map respectively, and obtain a first prediction layer, a second prediction layer, a third prediction layer and a fourth prediction layer in sequence includes: Utilize the neck network in the YOLOv5 model to perform convolution and upsampling on the fourth feature map successively, and then perform splicing processing with the third feature map to obtain a first enhanced high-level semantic feature information map; Using the neck network in the YOLOv5 model, after performing convolution and upsampling on the first enhanced high-level semantic feature information map in sequence, and then splicing it with the second feature map, the second enhanced high-level semantic feature information is obtained; Using the neck network in the YOLOv5 model, after performing convolution and upsampling on the second enhanced high-level semantic feature information map in sequence, and then splicing it with the first feature map, the third enhanced high-level semantic feature information map is obtained, where the first enhanced high-level semantic feature information map, the second enhanced high-level semantic feature information map, and the third enhanced high-level semantic feature information map form the enhanced high-level semantic feature information, and the semantic information contained in the first enhanced high-level semantic feature information map, the second enhanced high-level semantic feature information map, and the third enhanced high-level semantic feature information map is enhanced in sequence.
5. The object detection method for rescue according to claim 4, wherein, The step of using the neck network in the YOLOv5 model to perform enhanced high-level semantic features from top to bottom and enhanced low-level detail features from bottom to top on the first feature map, the second feature map, the third feature map, and the fourth feature map respectively, and obtaining the first prediction layer, the second prediction layer, the third prediction layer, and the fourth prediction layer in sequence, further includes: Using the backbone network in the YOLOv5 model, after convolving the third enhanced high-level semantic feature information map, and then splicing it with the convolved second enhanced high-level semantic feature information map, the second prediction layer is obtained, where the third enhanced high-level semantic feature information represents the first prediction layer; Using the backbone network in the YOLOv5 model, after convolving the second prediction layer, and then splicing it with the convolved first enhanced high-level semantic feature information map, the third prediction layer is obtained; Using the backbone network in the YOLOv5 model, after the third prediction layer, and then splicing it with the convolved fourth feature map, the fourth prediction layer is obtained.
6. The object detection method for rescue according to claim 1, characterized in that, The step of using the head network in the YOLOv5 model to perform regression prediction on the first prediction layer, the second prediction layer, the third prediction layer, and the fourth prediction layer, and obtaining the first target layer, the second target layer, the third target layer, and the fourth target layer in sequence, includes: Using the head network in the YOLOv5 model, using the linear regression algorithm to perform regression prediction on the first prediction layer, and obtaining the first target layer; Using the head network in the YOLOv5 model, using the linear regression algorithm to perform regression prediction on the second prediction layer, and obtaining the second target layer; Using the head network in the YOLOv5 model, using the linear regression algorithm to perform regression prediction on the third prediction layer, and obtaining the third target layer; Using the head network in the YOLOv5 model, using the linear regression algorithm to perform regression prediction on the fourth prediction layer, and obtaining the fourth target layer.
7. The object detection method for rescue according to claim 1, characterized in that, The step of using the first target layer, the second target layer, the third target layer, and the fourth target layer to predict the target of the rescue and obtain the prediction result, includes: Predict the target of rescue using the first target layer, the second target layer, the third target layer, and the fourth target layer to obtain the confidence level and coordinate information of the detected target; Obtain the prediction result according to the confidence level and coordinate information of the detected target.
8. A target detection system applied to rescue, characterized in that, Applied to the target detection method for rescue described in any one of claims 1-7, the system includes: A depth camera for obtaining an initial input image, processing the initial input image to obtain a target input image; A processing unit for sequentially performing four feature extractions on the target input image using the backbone network in the YOLOv5 model to obtain a first feature map, a second feature map, a third feature map, and a fourth feature map in sequence; The processing unit is further configured to use the neck network in the YOLOv5 model to enhance the high-level semantic feature information from top to bottom and the low-level detail feature from bottom to top for the first feature map, the second feature map, the third feature map, and the fourth feature map respectively, and obtain a first prediction layer, a second prediction layer, a third prediction layer, and a fourth prediction layer in sequence; The processing unit is further configured to use the head network in the YOLOv5 model to perform regression prediction on the first prediction layer, the second prediction layer, the third prediction layer, and the fourth prediction layer to obtain a first target layer, a second target layer, a third target layer, and a fourth target layer in sequence; An output unit for predicting the target of rescue using the first target layer, the second target layer, the third target layer, and the fourth target layer to obtain a prediction result.
9. A connection structure of an equipment rescue device, characterized in that, Applied to the target detection system for rescue described in claim 8, including a connection structure body, on which a depth camera, a processing unit, and an output unit are provided, and the processing unit is respectively connected to the camera and the output unit.
10. An electronic device, characterized in that, Including: At least one processor; And a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the target detection method for rescue described in any one of claims 1-7.