Construction site multi-target intelligent identification and detection method and device based on yo9
By applying the multi-objective intelligent identification and detection method based on yolo9 in construction site safety supervision, the problems of insufficient small-objective detection performance and insufficient robustness under complex lighting conditions are solved, and efficient and accurate target detection and flexible target category expansion are achieved to adapt to the diversity and dynamics of construction site safety needs.
Patent Information
- Application Number
- CN202411961626.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-30
AI Technical Summary
In the safety supervision of construction sites, the existing technology has problems such as insufficient small target detection performance, insufficient robustness under complex lighting conditions, low real-time processing efficiency of targets, and limited expansion capabilities of target categories.
A multi-object intelligent identification detection method for construction sites based on yolo9 is proposed, and a scheduling algorithm is used to optimize process processing, and an object detection algorithm combined with multi-scale feature fusion and attention mechanism is used to improve the detection ability of small targets and complex environments and achieve good scalability.
Improve the efficiency and accuracy of site target detection, enhance the robustness under complex lighting conditions, and support flexible target category expansion to adapt to the diversity and dynamics of site safety needs.
Smart Images

Figure CN120071237A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target recognition, involving various deep learning technologies and computational resource optimization methods, and particularly relates to a method and device for intelligent multi-target recognition and detection on construction sites based on YOLO9. Background Art
[0002] Construction site safety supervision is crucial for the safety of construction sites. However, the traditional manual inspection method has many deficiencies, such as low inspection efficiency, easy omission of potential risks, and inability to achieve real-time supervision. For this reason, computer vision technology has been widely applied to the field of construction site safety supervision, and automatic identification of safety hazards through target detection technology has become an important means to solve these problems. Currently, target detection technology is mainly implemented through deep learning models. Single-stage target detection algorithms represented by the YOLO series (You Only Look Once) have received extensive attention due to their fast speed and high accuracy. The single-stage architecture of YOLO can simultaneously predict the positions and categories of multiple targets through a single forward propagation, thus significantly improving the detection efficiency. However, the existing YOLO models still have some limitations in the construction site safety supervision scenario:
[0003] 1. Insufficient performance in detecting small targets: Small targets on construction sites (such as safety helmet buckles, safety signs) are easily overlooked in complex environments, and the existing technology has a low detection accuracy for these small targets.
[0004] 2. Insufficient robustness under complex lighting conditions: The construction site environment usually has the problem of drastic lighting changes (such as strong light, backlight, etc.), which poses higher requirements for the robustness of the target detection model.
[0005] 3. Efficiency problem in real-time processing of targets: Low-configured computing devices are difficult to quickly complete target detection tasks when facing real-time data from a large number of cameras, affecting practical applications.
[0006] 4. Limited ability to expand target categories: Traditional methods are difficult to flexibly add new detection targets and cannot meet the diversity and dynamics of construction site safety requirements.
[0007] In response to the above problems, the existing technology has not provided a comprehensive and effective solution. Especially in terms of real-time and efficient processing on low-configured devices, multi-target detection, and complex environment adaptation ability, there is still a large room for improvement. Summary of the Invention
[0008] The purpose of the present invention is to propose a method and device for intelligent multi-target recognition and detection on construction sites based on YOLO9 in view of the deficiencies of the existing technology, which realizes the detection of common targets on construction sites such as safety helmets and safety helmet buckles, and is deployed efficiently, ultimately achieving high target detection efficiency.
[0009] The object of the present invention is achieved by the following technical solutions: In the first aspect, the present invention provides a method for intelligent multi-target recognition and detection on a construction site based on YOLO9, and this method includes two parts: a scheduling stage and a target detection stage;
[0010] Scheduling stage: Receive camera information based on a scheduling algorithm, allocate an independent process for the detection task of each camera, and set the priority queue of the process; after the process is processed, perform target detection in a step-by-step matching form based on the target detection algorithm;
[0011] Target detection stage: Obtain the dataset for target detection, input it into the target detection algorithm built by the YOLO9 framework to obtain the weight file of the target detection algorithm; the target detection algorithm includes a backbone network and an auxiliary branch. The backbone network obtains three original features through multiple aggregation networks, and then through multiple sampling splicing and aggregation networks, three detection feature maps are obtained respectively; the auxiliary branch takes the three original features as inputs respectively, performs batch linear processing respectively and then performs feature fusion, and then passes through an aggregation network and a dual attention residual module to obtain three detection feature maps; the six detection feature maps are converged to perform final target detection, and the bounding box, category and confidence are output.
[0012] Further, the camera information includes the camera device number, the target type to be detected, and the detection video image frame.
[0013] Further, the scheduling algorithm sets the priority queue of the process based on the judgment process algorithm; the judgment process algorithm first establishes a process queue, arranges all the processes to be processed in the order of priority or arrival, and then sets a parameter N, which represents the maximum number of processes that can be processed simultaneously. The algorithm will take the first N processes from the queue for processing first, and then process another batch of processes in the queue after a set time.
[0014] Further, the set time is a fixed time or dynamically adjusted according to the system load.
[0015] Further, the step-by-step matching form is specifically as follows: Define in advance the target types to be judged. The target detection algorithm detects the input video image frame, analyzes the image, matches and filters the targets that meet the characteristics of the predefined types multiple times and performs annotation; the annotated image is then transmitted back to the front end through the URL interface for visual display.
[0016] Furthermore, the dual attention residual module is specifically as follows: The original input passes through the first average pooling and the first max pooling respectively, and then they are concatenated and put into the first fully connected layer. Together with the input, it passes through the first activation function to generate Output 1. Then the original input is put into the fifth convolution, passes through the second activation function, and then performs pixel multiplication with the original input. The generated result is pixel-added to Output 1 to generate Output 2.
[0017] Furthermore, the object detection algorithm model is specifically as follows:
[0018] The input image of the backbone network first passes through the initial silent layer, and then is processed by the first convolution and the second convolution. The features are processed by the repeated asymmetric convolution spatial and channel attention network blocks of the first aggregation network and downsampled using the first advanced downsampling; the features are processed again by the second aggregation network, the third aggregation network's repeated asymmetric convolution spatial and channel attention network blocks and the second downsampling, the third downsampling; then, the features enter the first spatial pyramid pooling after being processed by the fourth aggregation network, and then pass through the first upsampling and are concatenated with the previous feature map, and are processed again by the fifth aggregation network. This upsampling and processing process is repeated again to form the fourth detection feature map; Next, use the fourth advanced downsampling and connect it with the previous features, and form the fifth detection feature map again through the seventh aggregation network, and then process it again through the fifth advanced downsampling, connection and the eighth aggregation network to form the sixth detection feature map;
[0019] The auxiliary branch first performs three batch linear processes on the early features, and then repeats the early steps of the backbone network, including the third and fourth convolution downsamplings and the ninth aggregation network. Then, use the sixth advanced downsampling and fuse it with the three features processed by the convolutional batch linear process. Then, it is processed again by the tenth aggregation network, the eleventh aggregation network and the twelfth aggregation network, and three attention residual modules are applied; This process is repeated three times, and each time it is fused with the convolutional batch linear features at different levels to obtain the first detection feature map, the second detection feature map and the third detection feature map respectively; Finally, the six feature maps from the backbone network and the auxiliary branch converge, and final object detection is performed here, and the output includes the bounding box, category and confidence.
[0020] In a second aspect, the present invention provides a multi-object intelligent recognition and detection device for construction sites based on YOLO9, including a memory and one or more processors. An executable code is stored in the memory. When the processor executes the executable code, it implements the multi-object intelligent recognition and detection method for construction sites based on YOLO9 described above.
[0021] In a third aspect, the present invention provides a computer-readable storage medium with a program stored thereon. When the program is executed by a processor, the described multi-object intelligent recognition and detection method for construction sites based on YOLO9 is implemented.
[0022] In a fourth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the described multi-object intelligent recognition and detection method for construction sites based on YOLO9 is implemented.
[0023] Compared with the prior art, the advantages of the present invention are as follows:
[0024] 1. A new scheduling algorithm is proposed, enabling low-configuration computing devices to process a large amount of image information. This algorithm can transmit the camera footage to the server for processing and upload the processing results to the front-end for visualization, avoiding the waste of resources caused by each camera requiring a separate chip for image processing.
[0025] 2. Based on the YOLO9 framework, an algorithm model specifically designed for construction site target detection is developed. This model can efficiently identify various safety-related targets on construction sites, such as whether a safety helmet is worn, whether the safety helmet buckle is used, and whether the workers' clothing is standardized.
[0026] 3. The detection algorithm improves the detection ability of small targets and the robustness in different environments through multi-scale feature fusion and attention mechanisms.
[0027] 4. It has good scalability and can easily add new detection target categories to adapt to the changing safety requirements of construction sites. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0029] Figure 1 For the overall scheduling of the multi-object intelligent recognition and detection algorithm for construction sites.
[0030] Figure 2 For the dual attention residual module.
[0031] Figure 3 For the target detection algorithm model diagram.
[0032] Figure 4 For the safety helmet detection result, where hat indicates that the safety helmet is worn and person indicates that the safety helmet is not worn.
[0033] Figure 5 For the detection result of safety helmet buckle, wu_buckle means no safety helmet button, and you_buckle means there is a safety helmet button.
[0034] Figure 6 For the detection result of clothing, vest means short sleeves, and h_heel represents high heels, etc.
[0035] Figure 7 For the detection result of personnel behavior, fall means the person falls to the ground, and cigarette means the person smokes.
[0036] Figure 8 For the detection result of the environment, fire means flame, and smoke means smoke.
[0037] Figure 9 It is a structural diagram of a multi-target intelligent recognition and detection device for construction sites based on YOLO9 provided by the present invention. Detailed implementation manners
[0038] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described below with reference to the accompanying drawings and implementation cases. It should be understood that the specific implementation cases described herein are only used to explain the present invention and are not used to limit the present invention.
[0039] A multi-object intelligent recognition and detection method for construction sites based on YOLO9 proposed by the present invention. The detection method of the present invention is developed based on the YOLO9 (You Only Look Once version 9) architecture and is specifically used to detect key safety targets on construction sites, such as whether workers wear safety helmets, whether the safety helmets are properly fastened, and whether the workers' clothing complies with the regulations (for example, detecting non-compliant clothing such as wearing high heels or short sleeves). The present invention adopts a multi-scale feature fusion technology, enabling the algorithm to process targets of different sizes simultaneously, thereby improving the detection accuracy of various sized objects on construction sites. In addition, an attention mechanism is introduced to help the network better focus on key areas, such as the head and feet, further enhancing the detection accuracy of clothing such as safety helmets and shoes. Through the training of a large amount of construction site scene data, including images of various weather conditions, lighting environments, and construction stages, the algorithm can accurately identify various safety hazards, including not wearing a safety helmet, the safety helmet not being properly worn, wearing non-compliant shoes or tops, etc. At the same time, the present invention applies advanced data augmentation techniques, such as mixed sampling and adaptive data augmentation, to further improve the generalization ability of the model. This automated detection method not only improves the efficiency of safety supervision but also can provide real-time warnings to help managers discover and solve safety problems in a timely manner. The present invention has developed a complete construction site safety management system, integrating the detection results with the construction site management system, realizing the full-process automation from detection to warning, from record to statistical analysis. The detection method of the present invention includes a scheduling stage and an object detection stage; specifically as follows:
[0040] Scheduling stage:
[0041] Step 1: Preprocessing stage of the scheduling algorithm
[0042] Such as Figure 1As shown in the figure, the front end passes information such as the targets to be detected by the camera and the device number (including the detected video image frames) to the scheduling algorithm through the URL interface. The scheduling algorithm allocates different processes by parsing the URL information. The information transmitted includes the camera device number, the type of target to be detected, the detected video image frames, and other relevant parameters. After receiving this information, the scheduling algorithm will allocate an independent process for the detection task of each camera. To prevent a large number of camera devices to be detected from causing system jams due to the simultaneous processing of a large amount of process information, the scheduling algorithm designs a special process judgment algorithm to solve such problems. The process judgment algorithm first establishes a process queue and arranges all the processes to be processed according to the priority or arrival order. Then a parameter N is set, which represents the maximum number of processes that the system can handle simultaneously. The algorithm will take the first N processes from the queue for processing first, and then process another batch of processes in the queue after a certain period of time (which can be a fixed time or adjusted dynamically according to the system load). This method can well solve the problem that low-configuration computing devices can also process a large amount of image information. By controlling the number of processes processed simultaneously, it is possible to avoid excessive occupation of system resources, thus ensuring the stability and response speed of the system, so that it can run smoothly even on computing devices with lower configurations.
[0043] Step 2: The scheduling algorithm calls the target detection algorithm processing stage
[0044] After the scheduling algorithm finishes the process processing, it will call the target detection algorithm to perform target detection again. The target detection judges the target type through a step-by-step matching form. First, the system will pre-define the target types to be judged, such as safety helmets, high heels, smoking, etc. Then the target detection algorithm detects the input video image frames, analyzes the images, and searches for targets that meet the characteristics of the pre-defined types. This process may involve multiple matches and screenings to improve the detection accuracy. After detecting the corresponding target, the algorithm will mark the target. Finally, the marked image is passed back to the front end through the URL interface for visual display. In this way, users can view the processed and marked images in real time on the front-end interface and intuitively understand the detection results. This two-step process design allows the system to efficiently process a large amount of image information from multiple cameras, and at the same time, through the intelligent management of the scheduling algorithm, it avoids excessive occupation of system resources. Finally, the goal of being able to process a large amount of image information even on low-configuration devices is achieved.
[0045] Target detection stage:
[0046] Step 1: Target detection preprocessing stage
[0047] The present invention has established datasets for safety helmets, safety helmet buckles, clothing wear, personnel behavior, and fireworks. Select Q TraThe original color image RGB Tra and the final result image GT of saliency detection Tra , and they form the training set Tra. Select Q in the same way val (Q Tra :Q val = 7:3) color images R different from the training set Tra Tra and the final result image GT of saliency detection corresponding to each original color image val and they form the validation set Val. The color images mainly consist of RGB images and multispectral images taken in different scenes. The RGB images record the spectral information of the red, green, and blue bands, and the multispectral images record the spectral information of other different bands. Each band information is equivalent to a channel component, that is, each original scene image with salient objects contains the R channel component, G channel component, and B channel component of the RGB image. Input the training set and the validation set into the object detection algorithm built by the yolo9 framework to obtain the weight file of the object detection algorithm.
[0048] Step 2: Detailed process of the object detection algorithm
[0049] As Figure 2 shown in the dual attention residual module. At the beginning of the input, it goes through the first average pooling and the first max pooling respectively, and then they are concatenated and put into the first fully connected layer to generate output one together with the input through the first activation function. Then the original input is put into the fifth convolution, goes through the second activation function, and then is pixel-multiplied with the original input. The generated result is pixel-added to output one to generate output two.
[0050] As Figure 3The diagram of the target detection algorithm model shown. The input image first passes through the initial silent layer, and then is processed through the first convolution and the second convolution. Next, the features are processed through the repeated asymmetric convolution spatial and channel attention network blocks of the first aggregation network, and downsampled using the first advanced downsampling. This process is repeated, and the features are again processed through the repeated asymmetric convolution spatial and channel attention network blocks of the second aggregation network and the third aggregation network, and the second downsampling and the third downsampling. After that, the features are processed through the fourth aggregation network and then enter the first spatial pyramid pooling, and then pass through the first upsampling and are concatenated with the previous feature map, and are again processed through the fifth aggregation network. This upsampling and concatenation process enters the sixth aggregation network to form the P4 feature map. Next, the fourth advanced downsampling is used and connected to the output of the fifth aggregation network, and is again processed through the seventh aggregation network to form the P5 feature map, and then is again processed through the fifth advanced downsampling, connection, and the eighth aggregation network to form the P6 feature map. At the same time, the auxiliary branch starts to work. First, three batch linear processes are performed on the outputs of the second, third, and fourth aggregation networks, and then the early steps of the backbone network are repeated, including the third and fourth convolutions and the ninth aggregation network. Subsequently, the sixth advanced downsampling is used and fused with the three feature maps of the convolution batch linear process, and is again processed through the tenth aggregation network, the eleventh aggregation network, and the twelfth aggregation network, and three attention residual modules are applied. This process is repeated three times, each time fused with the convolution batch linear features at different levels. Finally, the six feature maps from the backbone network and the auxiliary branch converge, and the final target detection is performed here, and the output includes the bounding box, category, and confidence level.
[0051] The algorithm of the present invention inherits the single-stage detection framework of YOLO, and can simultaneously predict the positions and categories of multiple targets through one forward propagation, greatly improving the detection speed. In order to adapt to the special environment of the construction site, the network structure is optimized, the detection ability for small targets (such as safety helmet buckles) is enhanced, and the robustness under different lighting conditions is improved.
[0052] As Figures 4 - 8 shown, they are respectively the detection results of the detection method of the present invention on the datasets of safety helmets (whether worn), safety helmet buckles (whether there are), clothing (short sleeves or high heels), human behaviors (falling to the ground or smoking), and the environment (flames or smoke); it shows that the present invention can achieve good detection on these datasets.
[0053] The present invention fine-tunes the YOLO9 algorithm to generalize it to the collected construction site detection data set, and uses the process rotation execution method to solve the problem of high hardware resource occupation when YOLO9 is actually deployed. And the detection algorithm is deployed on a low-profile server to achieve more flexible real-time monitoring. By optimizing the model structure and quantization technology, the model size and computational complexity are significantly reduced, so that it can run smoothly on a low-profile server. Corresponding to the aforementioned embodiment of a construction site multi-target intelligent recognition detection method based on yolo9, the present invention also provides an embodiment of a construction site multi-target intelligent recognition detection device based on yolo9.
[0054] See also Figure 9 An embodiment of the present invention provides a construction site multi-target intelligent recognition and detection device based on yolo9, including a memory and one or more processors, wherein the memory stores executable code, and when the processor executes the executable code, it is used to implement a construction site multi-target intelligent recognition and detection method based on yolo9 in the above embodiment.
[0055] The embodiment of the construction site multi-target intelligent identification and detection device based on yolo9 provided by the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities in which it is located reading the corresponding computer program instructions in the non-volatile memory into the internal memory for execution. From the hardware level, such as Figure 9 As shown, it is a hardware structure diagram of a construction site multi-target intelligent identification and detection device based on yolo9 provided by the present invention, in which any device with data processing capability is located, except Figure 9 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities in which the apparatus in the embodiments is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.
[0056] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0057] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions in the method embodiments. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. A person of ordinary skill in the art can understand and implement it without creative work.
[0058] An embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a method for intelligent multi-target recognition and detection on a construction site based on YOLO9 in the above embodiment.
[0059] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store the data that has been output or will be output.
[0060] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the method for intelligent multi-target recognition and detection on a construction site based on YOLO9.
[0061] The above embodiments are used to explain the present invention, rather than limit the present invention. Any modifications and changes made to the present invention within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.
Claims
1. A construction site multi-target intelligent recognition and detection method based on yolo9, characterized in that: The method includes two parts: scheduling phase and target detection phase; Scheduling stage: Receive camera information based on the scheduling algorithm, assign an independent process to each camera's detection task, and set the priority queue of the process; After the process is processed, the target is detected by step-by-step matching based on the target detection algorithm; Target detection stage: obtain the target detection data set, input it into the target detection algorithm built by the yolo9 framework, and obtain the weight file of the target detection algorithm; The target detection algorithm includes a backbone network and auxiliary branches. The backbone network obtains three original features after multiple aggregation networks, and then obtains three detection feature maps after multiple sampling splicing and aggregation networks respectively; the auxiliary branches take the three original features as input, perform feature fusion after batch linear processing, and then obtain three detection feature maps after aggregation networks and dual attention residual modules; the six detection feature maps are converged for final target detection, and the bounding box, category and confidence are output.
2. A construction site multi-target intelligent recognition and detection method based on yolo9 according to claim 1, characterized in that: The camera information includes the camera device number, the target type to be detected, and the detection video image frame.
3. A construction site multi-target intelligent recognition and detection method based on yolo9 according to claim 1, characterized in that: The scheduling algorithm sets the priority queue of the process based on the judgment process algorithm; the judgment process algorithm first establishes a process queue, arranges all the processes to be processed according to priority or arrival order, and then sets a parameter N, which represents the maximum number of processes that can be processed at the same time. The algorithm will take out the first N processes from the queue for processing first, and then process another batch of processes in the queue after a set time.
4. A construction site multi-target intelligent recognition and detection method based on yolo9 according to claim 3, characterized in that: The set time is a fixed time or is dynamically adjusted according to the system load.
5. A construction site multi-target intelligent recognition and detection method based on yolo9 according to claim 1, characterized in that: The specific form of step-by-step matching is as follows: the target type to be judged is defined in advance, the target detection algorithm detects the input video image frame, analyzes the image, matches and filters the targets that meet the predefined type characteristics multiple times and marks them; the marked image is then passed back to the front end through the URL interface for visual display.
6. A construction site multi-target intelligent recognition and detection method based on yolo9 according to claim 1, characterized in that: The dual attention residual module is specifically as follows: the original input undergoes the first average pooling and the first maximum pooling respectively, and then they are concatenated and put into the first fully connected layer together with the input to generate output one through the first activation function, and then the original input is put into the fifth convolution, passes through the second activation function, and then multiplied pixel by pixel with the original input, and the generated result is added to the output one pixel to generate output two.
7. A construction site multi-target intelligent recognition and detection method based on yolo9 according to claim 1, characterized in that: The target detection algorithm model is as follows: The input image of the backbone network first passes through the initial silent layer, and then is processed by the first convolution and the second convolution. The features are processed by the repeated asymmetric convolution space and channel attention network blocks of the first aggregation network, and downsampled using the first advanced downsampling; the features are again processed by the repeated asymmetric convolution space and channel attention network blocks of the second and third aggregation networks, and the second downsampling and the third downsampling; after that, the features are processed by the fourth aggregation network and enter the first spatial pyramid pooling, and then go through the first upsampling and splicing with the previous feature map, and then processed by the fifth aggregation network again. This upsampling and processing process is repeated again to form the fourth detection feature map; next, the fourth advanced downsampling is used and connected with the previous features, and the fifth detection feature map is formed again through the seventh aggregation network, and then the fifth advanced downsampling, connection and eighth aggregation network are processed again to form the sixth detection feature map; The auxiliary branch first performs three batch linear processing on the early features, and then repeats the early steps of the backbone network, including the third and fourth convolutional downsampling and the ninth aggregation network. Subsequently, the sixth advanced downsampling is used and fused with the three features processed by the convolutional batch linear processing, and then processed again by the tenth aggregation network, the eleventh aggregation network and the twelfth aggregation network, and the attention residual module is applied three times; this process is repeated three times, each time fused with the convolutional batch linear features of different levels, and the first detection feature map, the second detection feature map and the third detection feature map are obtained respectively; finally, the six feature maps from the backbone network and the auxiliary branch are converged, and the final target detection is performed here, and the output includes the bounding box, category and confidence.
8. A construction site multi-target intelligent recognition and detection device based on yolo9, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the processor executes the executable code, a construction site multi-target intelligent recognition and detection method based on YOLO9 as described in any one of claims 1-7 is implemented.
9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by the processor, a construction site multi-target intelligent recognition and detection method based on YOLO9 as described in any one of claims 1-7 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by the processor, a construction site multi-target intelligent recognition and detection method based on YOLO9 is implemented as described in any one of claims 1-7.