Bimodal construction scene target detection and tracking method
By using a dual-modal video acquisition and fuzzy control system, combined with visible light and infrared images, a target detection and tracking model is constructed. This solves the problem of accuracy in target detection and tracking under complex construction site conditions, enables precise identification of mechanical collision risks, and ensures construction safety.
Patent Information
- Application Number
- CN202511078289.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-11-11
AI Technical Summary
Existing target detection and tracking technologies at construction sites have low accuracy in complex environments, especially at night, in bad weather, or in low light conditions. Single-modal video data cannot provide clear and accurate target information, affecting the assessment of mechanical collision risks.
By employing dual-modal video acquisition technology and combining visible light and infrared images, a target detection model and a tracking model are constructed. A fuzzy control system is used to comprehensively consider multiple factors to determine the risk of mechanical collision, including indicators such as target proximity, approach speed, and spatial congestion.
It improves the accuracy of target detection and tracking, can work stably in complex environments, provides more reliable mechanical collision risk level identification, and ensures the safety of construction sites.
Smart Images

Figure CN120931907A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of construction scene detection technology, and in particular to a dual-modal construction scene target detection and tracking method. Background Technology
[0002] In the field of modern construction, ensuring construction safety and efficient management are crucial tasks. With the continuous expansion of construction scale and increasing complexity, numerous potential safety risks exist on construction sites. Among these, mechanical collision accidents are relatively common and highly dangerous. For example, in large construction sites, various construction machines such as cranes, excavators, and concrete pump trucks operate simultaneously, their spatial ranges intersecting. Once a collision occurs, it can not only damage the machinery and equipment, affecting construction progress, but also pose a serious threat to the personal safety of on-site construction workers.
[0003] Some existing studies have attempted to apply target detection and tracking technologies to construction sites, but most of them are based on single-modal video data, such as using only visible light video. However, the construction site environment is complex and variable. At night, in severe weather (such as fog, rain, snow, etc.), or in insufficient light, single-modal video data often cannot provide clear and accurate target information, resulting in a significant decrease in the accuracy of target detection and tracking, which in turn affects the judgment of mechanical collision risks. Summary of the Invention
[0004] In view of this, the present invention proposes a dual-modal construction scene target detection and tracking method, which can effectively solve the defects of low accuracy in target detection and tracking in the existing technology.
[0005] The technical solution of this invention is implemented as follows:
[0006] A dual-modal target detection and tracking method for construction scenarios, specifically including:
[0007] Collect dual-modal video of the construction site target during on-site operations;
[0008] A target detection model and a target tracking model are constructed based on the dual-modal video of the target at the construction site;
[0009] Real-time acquisition of video stream data from construction site operations;
[0010] Real-time video stream data is input into the target detection model, and the target detection model outputs the target detection result;
[0011] Real-time video stream data and target detection results are input into the target tracking model to obtain the target trajectory result;
[0012] The target detection results and target trajectory results are input into the fuzzy control system, which then determines the mechanical collision risk level at the construction site.
[0013] As a further optional solution to the aforementioned dual-modal construction scene target detection and tracking method, the step of constructing a target detection model based on the dual-modal video of the target at the construction site specifically includes:
[0014] The dual-modal video of the target at the construction site is processed to obtain the target detection dataset;
[0015] Based on the object detection dataset, a Dual-YOLO object detection model was built and trained using GPU in the PyTorch framework environment, thus constructing the object detection model.
[0016] As a further optional solution to the aforementioned dual-modal construction scene target detection and tracking method, the processing of the dual-modal video of the construction site target to obtain the target detection dataset specifically includes:
[0017] Images are extracted from bimodal video using video frame extraction to obtain visible light and infrared images;
[0018] The visible light image and the infrared image are converted into intermediate mode images respectively using a mean filter, resulting in visible light intermediate mode images and infrared intermediate mode images;
[0019] Feature point detection and directional gradient histogram description of feature information are performed on visible light intermediate mode images and infrared intermediate mode images to obtain visible light target feature maps and infrared target feature maps;
[0020] The visible light target feature map and the infrared target feature map are labeled to obtain the target detection dataset.
[0021] As a further optional solution to the dual-modal construction scene target detection and tracking method, the construction of the target tracking model specifically includes:
[0022] The dual-modal video of the target at the construction site is input into the target detection model to obtain the target detection results of the dual-modal video.
[0023] The target detection results of the bimodal video are input into the bimodal target tracking model based on the ByteTrack algorithm for training, and the target tracking model is constructed.
[0024] As a further optional solution to the aforementioned dual-modal construction scene target detection and tracking method, the construction of the fuzzy control system specifically includes:
[0025] Based on the target detection results output by the target detection model and the target trajectory results output by the target tracking model, the target proximity index, target approach speed index, and spatial congestion index are selected as input variables of the fuzzy control system, and the mechanical collision risk level at the construction site is used as the output variable.
[0026] For each input and output variable, construct a fuzzy set and set the membership function of the fuzzy set;
[0027] Build a fuzzy rule base;
[0028] Construct the Mamdani inference mechanism and the centroid method for fuzzy resolution.
[0029] As a further optional solution to the aforementioned dual-modal construction scene target detection and tracking method, the fuzzy control system determines the mechanical collision risk level at the construction site, specifically including:
[0030] The Mamdani inference mechanism is used to perform fuzzy inference on the input variables to obtain the degree of membership of the output variables in each fuzzy set.
[0031] The centroid method is used to defuzzify the membership degree of the output variable in each fuzzy set, and the fuzzy set of the output variable is converted into a quantitative risk value, which corresponds to the mechanical collision risk level at the construction site.
[0032] As a further optional solution to the aforementioned dual-modal construction scene target detection and tracking method, the use of the Mamdani inference mechanism to perform fuzzy inference on the input variables specifically includes:
[0033] The degree to which an input variable belongs to each fuzzy set is determined based on its actual value and membership function.
[0034] Reasoning is performed based on the rules in the fuzzy rule base to obtain the degree of membership of the output variable in each fuzzy set.
[0035] A dual-modal construction scene target detection and tracking system includes:
[0036] The dual-modal video acquisition module is used to acquire dual-modal video of targets at the construction site during on-site operations.
[0037] The model building module is used to build target detection and target tracking models based on the dual-modal video of the target at the construction site.
[0038] The real-time data acquisition module is used to acquire video stream data of construction site operations in real time.
[0039] The target detection module is used to input real-time video stream data into the target detection model, and the target detection model outputs the target detection results;
[0040] The target tracking module is used to input real-time video stream data and target detection results into the target tracking model to obtain the target trajectory result;
[0041] The fuzzy control module is used to input the target detection results and target trajectory results into the fuzzy control system, which then determines the mechanical collision risk level at the construction site.
[0042] A computing device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the dual-modal construction scene target detection and tracking method described above.
[0043] A computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the dual-modal construction scene target detection and tracking method described above.
[0044] The beneficial effects of this invention are as follows: By acquiring dual-modal video of targets at the construction site, the two modal data complement each other, greatly enriching the information features of the targets and providing a more comprehensive and accurate data foundation for target detection and tracking. This effectively avoids target omissions or false detections caused by the limitations of single-modal data, significantly improving the accuracy of target detection. Based on the dual-modal video, target detection and target tracking models are constructed separately. This independent construction method allows each model to focus on its own task, avoiding the mixing of model functions, making the detection and tracking process more professional and efficient, and improving the stability and accuracy of target tracking. Real-time acquisition of video stream data from construction site operations and timely input of it into the target detection and target tracking models... This real-time capability ensures that the system can acquire the latest dynamic information from the construction site immediately, enabling timely detection and tracking of targets. By inputting the target detection results and target trajectory results into the fuzzy control system, and selecting indicators closely related to mechanical collision risk, such as target proximity, target approach speed, and spatial congestion, as input variables, and constructing reasonable fuzzy sets, membership functions, and fuzzy rule bases, as well as employing appropriate inference and defuzzification methods, the system can comprehensively consider the impact of multiple factors on mechanical collision risk. This allows for a more accurate determination of the mechanical collision risk level at the construction site, avoiding the one-sidedness of relying solely on a single factor. This provides a more reliable and accurate decision-making basis for construction safety management, further ensuring the safety of the construction site. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 This is a flowchart of a dual-modal construction scene target detection and tracking method according to the present invention;
[0047] Figure 2 This is a schematic diagram of the composition of a dual-modal construction scene target detection and tracking system according to the present invention;
[0048] Figure 3 This is a schematic diagram of the composition of a computing device according to the present invention. Detailed Implementation
[0049] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0050] refer to Figures 1 to 3 A dual-modal target detection and tracking method for construction scenarios, specifically including:
[0051] The video data collected by drones included dual-modal videos of on-site operations under different shooting conditions. The dual-modal videos included visible light videos and infrared videos. The video data included nine types of construction site targets: workers, cars, concrete pump trucks, trucks, concrete mixer trucks, excavators, rotary drilling rigs, cranes, and trenchers. The shooting conditions included daytime conditions with sufficient natural ambient light, nighttime conditions with insufficient natural ambient light but sufficient artificial light sources, and nighttime conditions with insufficient light.
[0052] Specifically, drones are used to collect dual-modal videos of on-site operations under different shooting conditions, covering visible light and infrared video. Visible light video can clearly present the visual characteristics of the target, such as appearance and color, while infrared video can capture the target's thermal radiation information. The combination of the two can obtain more comprehensive and richer information on construction site targets, including nine types of construction targets such as workers, cars, and concrete pump trucks, providing a sufficient data foundation for subsequent detection and tracking.
[0053] Video acquisition under different shooting conditions enables this method to adapt to diverse environmental conditions at construction sites. For example, in different scenarios such as normal daytime lighting, nighttime with artificial light sources but no natural light, and nighttime with no light, dual-modal video can effectively acquire target information, overcoming the limitations of single shooting conditions or single-modal video in complex environments, and ensuring the stability and reliability of data acquisition.
[0054] Based on bimodal video of the construction site target, a target detection model and a target tracking model are constructed. In some embodiments, the construction of the target detection model based on the bimodal video of the construction site target specifically includes:
[0055] The dual-modal video of the target at the construction site is processed to obtain the target detection dataset;
[0056] Based on an object detection dataset, a Dual-YOLO object detection model was built and trained using GPUs within the PyTorch framework. Specifically, the Dual-YOLO object detection model employs a dual-branch backbone network. The two branches extract visible light and infrared modal features through convolution, respectively. The dual-modal features at different scale levels are processed and fused by the Bimodal Feature Fusion (BFFM) module and then input into the C2f module at the corresponding scale level of the neck network for extraction, enhancement, and fusion. Finally, the output feature tensor is fed to the Dyhead detection head of the head network to generate prediction results. During training, the Wise-IoU loss function is used to continuously update the network parameters and backpropagate the parameters. After multiple rounds of training, the accuracy, recall, and comprehensive metrics AP and mAP are evaluated. The optimal model weight file of the final trained parameters is saved, thus constructing the object detection model.
[0057] Specifically, the Dual-YOLO object detection model built and trained using GPU based on the PyTorch framework adopts a dual-branch backbone network. This design can effectively extract features of visible light and infrared modes respectively. Through convolution operations, the model can deeply explore the target features under different modes, providing a good foundation for subsequent feature fusion and detection.
[0058] The dual-modal features at different scales are processed and fused by the bimodal feature fusion module BFFM, which can fully combine the feature information of the two modalities. This fusion method takes into account the feature representation at different scales, enabling the model to understand the characteristics of the target more comprehensively and improve the target recognition ability.
[0059] The fused features are input into the C2f module at the corresponding scale level of the neck network for extraction, enhancement, and fusion. Finally, the output feature tensor is fed into the Dyhead detection head of the head network to generate prediction results. Further processing by the C2f module can improve the expressive power of the features, while the Dyhead detection head can more accurately generate the location and category prediction of the target, thus improving the accuracy of target detection.
[0060] During training, the Wise-IoU loss function is used to continuously update network parameters and backpropagate the parameters. The Wise-IoU loss function can more accurately measure the difference between the predicted box and the ground truth box, guiding the model to learn in a more accurate direction and helping to improve the model's prediction accuracy of the target bounding box. After multiple rounds of training, precision, recall, and comprehensive metrics AP and mAP are evaluated. These evaluation metrics can comprehensively evaluate the model's performance from different perspectives, ensuring the model's accuracy and reliability in the target detection task. By saving the optimal model weight file of the final training parameters, a high-performance and stable target detection model can be built, providing accurate target detection results for subsequent target tracking and construction safety management.
[0061] In some embodiments, the processing of the bimodal video of the construction site target to obtain the target detection dataset specifically includes:
[0062] Images are extracted from bimodal video using video frame extraction, and the extracted images are cropped to obtain unregistered visible light and infrared images.
[0063] The visible light image and the infrared image are smoothed by applying a mean filter. Then, the smoothed image pixel value is subtracted from the original image pixel value to obtain the intermediate mode image of the high frequency feature component, thus obtaining the visible light intermediate mode image and the infrared intermediate mode image.
[0064] For visible light intermediate mode images and infrared intermediate mode images, the ORB algorithm is used to calculate the gray-level centroid coordinates and directions of candidate feature points by calculating the gray-level moments of the candidate feature points' neighborhoods to confirm the feature points. The suppression radius is determined based on the local density of each candidate feature point. The feature point with the highest responsivity is selected to suppress the expression of other feature points within the corresponding suppression radius. At the same time, the gradient field of the image pixels is calculated using the Sobel matrix operator to obtain the gradient direction and amplitude features of each pixel. The image is divided into several small units, and the gradient direction distribution is statistically analyzed to construct the unit HOG. After normalizing the contrast of the units, multiple unit HOGs are combined to form a block feature vector. Finally, all block feature vectors are concatenated in spatial order to output the HOG features of the image, thus obtaining the visible light target feature map and the infrared target feature map.
[0065] Nine types of targets, including workers, cars, concrete pump trucks, trucks, concrete mixers, excavators, rotary drilling rigs, cranes, and trenchers, were labeled in the visible light target feature map and infrared target feature map, and a target detection dataset in TXT format was created.
[0066] Specifically, images are extracted from bimodal video using video frame extraction and then cropped. This process removes redundant non-target areas from the video, allowing subsequent processing to focus on key content, reducing computational load, and increasing the proportion of the target in the image, thus laying the foundation for accurate target detection.
[0067] The visible light and infrared images are smoothed by using a mean filter, which effectively reduces noise interference in the images. In the construction site environment, video acquisition may be affected by various factors and generate noise. Smoothing makes the image clearer, highlights the features of the target, and helps to extract target information more accurately in the future.
[0068] By subtracting the pixel values of the smoothed image from the pixel values of the original image, an intermediate modal image with high-frequency feature components is obtained. This processing method enhances the high-frequency features in the image, such as the edge and contour information of the target, making the distinction between the target and the background more obvious and further improving the accuracy of target detection.
[0069] The ORB algorithm is used to detect feature points in visible light intermediate mode images and infrared intermediate mode images. The gray centroid coordinates and directions of the feature points are calculated by the gray moment of the candidate feature points to confirm the feature points. The suppression radius is determined based on the local density to select the feature point with the highest responsivity and suppress the expression of other feature points. This method can accurately locate the key feature points of the target, avoid feature point redundancy and false detection, and improve the quality and reliability of feature points.
[0070] By using the Sobel matrix operator to calculate the gradient field of image pixels at the edges, the gradient direction and magnitude features of each pixel are obtained. The image is divided into several small units and the gradient direction distribution is statistically analyzed to construct the unit HOG. After normalization, multiple unit HOGs that constitute a spatial block are combined to form a block feature vector. This multi-level feature description method, from pixel gradient to unit and block feature vectors, comprehensively describes the appearance features of the target, can better distinguish different types of targets, and provides rich feature information for target detection.
[0071] Nine types of targets, including workers, cars, and concrete pump trucks, were labeled in the visible light target feature map and infrared target feature map to create a target detection dataset in TXT format. The clear labeling makes the target information in the dataset clear and accurate, providing reliable labeled data for model training and helping the model learn the feature patterns of different targets.
[0072] In some embodiments, the construction of the target tracking model specifically includes:
[0073] The dual-modal video of the target at the construction site is input into the target detection model to obtain the target detection results of the dual-modal video.
[0074] The target detection results from the bimodal video are input into a bimodal target tracking model based on the ByteTrack algorithm for training. In this model, the target bounding box confidence scores are processed: targets with a confidence score less than 0.1 are discarded; predicted boxes with a confidence score greater than 0.1 but less than 0.6 are classified as low-confidence targets; and predicted boxes with a confidence score greater than 0.6 are classified as high-confidence targets. Kalman filtering is used to predict the motion of existing trajectories in the current frame, resulting in trajectory prediction boxes. Three matching operations are performed: the first matching matches high-confidence target boxes with trajectory prediction boxes; if a match is successful, the target trajectory is updated to a tracking and activated state; if a match fails, the unmatched target boxes and prediction boxes are additionally cached; the second matching matches the remaining targets with low-confidence target boxes. The tracking state prediction boxes in the remaining trajectory prediction boxes are matched. If the match is successful, the target trajectory is updated to the tracking and activated state. If the match fails, the low-confidence target boxes are discarded, the prediction boxes are updated to the lost state and cached. When the lost state of the prediction box exceeds 30 frames, the target trajectory is removed. The third match is to match the remaining high-confidence target boxes with the inactive state trajectory boxes in the remaining trajectory prediction boxes. If the match is successful, the target trajectory is updated to the tracking and activated state. If the match fails, the prediction box and its trajectory are removed, and a new target trajectory is created based on the target box. The inactive state prediction boxes only come from the newly created trajectory in the previous frame, and the trajectory prediction box at this time is the original target detection box. After multiple rounds of training, the optimal model weight file of the final trained parameters is saved to build the target tracking model.
[0075] Specifically, the bimodal video of the construction site target is input into the constructed target detection model to obtain the target detection result of the bimodal video, and then used as the training input of the ByteTrack algorithm bimodal target tracking model. This step ensures that the tracking model is trained based on accurate target detection information, laying a solid data foundation for subsequent accurate target tracking, because accurate target detection results contain key information such as the target's location and category, enabling the tracking model to better learn the target's features and motion patterns.
[0076] In the ByteTrack algorithm model, targets with a confidence level less than 0.1 are discarded, predicted boxes with a confidence level greater than 0.1 but less than 0.6 are classified as low-confidence targets, and predicted boxes with a confidence level greater than 0.6 are classified as high-confidence targets. This classification method can effectively filter out reliable targets and reduce the impact of noise and false detections on the tracking model. High-confidence targets usually correspond to more accurate target detection results, while low-confidence targets may have some uncertainty. Through classification, subsequent matching operations can be performed more effectively.
[0077] Kalman filtering is used to predict the motion of the existing trajectory in the current frame and obtain the trajectory prediction box. Kalman filtering can predict the position of the target in the current frame based on the target's historical motion information, which improves the accuracy of trajectory prediction. Even when the target is briefly occluded or its motion state changes, Kalman filtering can still provide relatively reliable prediction results, which helps to maintain continuous tracking of the target.
[0078] Three matching operations are performed. The first matching matches high-confidence target boxes with trajectory prediction boxes. If a match is successful, the target trajectory is updated to a tracking and activated state. If a match fails, the unmatched target boxes and prediction boxes are cached separately. This strategy prioritizes high-confidence target boxes, ensuring the tracking accuracy of high-reliability targets. The second matching matches low-confidence target boxes with tracking state prediction boxes in the remaining trajectory prediction boxes, further exploring possible target matches and improving the tracking capability for low-confidence but real targets. The third matching matches the remaining high-confidence target boxes with inactive state trajectory boxes in the remaining trajectory prediction boxes, making full use of all possible target box information, maximizing target tracking and reducing target loss.
[0079] During the matching process, the state of the target trajectory is updated reasonably according to the matching results, such as the tracking and activated state, the lost state, etc., and the target trajectory is removed when the predicted box is lost for more than 30 frames. This state management and trajectory removal mechanism can clean up invalid trajectory information in a timely manner, avoid interference with subsequent tracking, and ensure the efficiency and accuracy of the tracking model.
[0080] Real-time acquisition of video stream data from construction site operations.
[0081] Real-time video stream data is input into the target detection model, and the target detection model outputs the target detection results.
[0082] Real-time video stream data and target detection results are input into the target tracking model to obtain the target trajectory result.
[0083] The target detection results and target trajectory results are input into the fuzzy control system, which determines the mechanical collision risk level at the construction site. In some embodiments, the construction of the fuzzy control system specifically includes:
[0084] Based on the target detection results output by the target detection model and the target trajectory results output by the target tracking model, the target proximity index, target approach speed index, and spatial congestion index are selected as input variables of the fuzzy control system, and the mechanical collision risk level at the construction site is used as the output variable.
[0085] For each input and output variable, construct a fuzzy set and set the membership function of the fuzzy set;
[0086] Build a fuzzy rule base;
[0087] Construct the Mamdani inference mechanism and the centroid method for fuzzy resolution.
[0088] Specifically, based on the results output by the target detection model and the target trajectory results output by the target tracking model, target proximity index, target approach speed index, and spatial congestion index are selected as input variables. Target proximity can intuitively reflect the spatial distance relationship between machines, approach speed can reflect the speed at which machines approach each other, and spatial congestion can describe the density of machine distribution on the construction site as a whole. These three indicators comprehensively and accurately cover the key factors affecting the risk of machine collisions, providing a reliable data foundation for accurately judging the risk level. The construction site environment is complex and changeable, and the motion state and spatial position of the machines are constantly changing. These input variables can dynamically reflect the relative relationship between machines and the real-time situation of the construction site. Whether the machines are stationary, moving slowly, or moving quickly, they can be accurately described through these indicators, enabling the fuzzy control system to adapt to various complex construction scenarios.
[0089] For each input and output variable, a fuzzy set is constructed, and a membership function is set. By reasonably defining the range of the fuzzy set and the shape of the membership function, the degree of membership of each variable under different states can be accurately described. For example, for the target proximity index, fuzzy sets such as "far," "medium," and "near" can be defined, and the degree to which a specific distance value belongs to each set can be determined by the membership function, thereby more finely characterizing the variable features and providing accurate input information for fuzzy inference. The construction of fuzzy sets and membership functions enables the system to handle uncertain and fuzzy information. In the construction site, many situations are difficult to describe with precise numerical values, but fuzzy sets and membership functions can quantify this fuzzy information, enhancing the system's adaptability and flexibility to complex situations.
[0090] The constructed fuzzy rule base, based on expert experience and actual construction site conditions, comprehensively considers various combinations of input variables in the form of "IF-THEN". For example, when the target proximity is "near", the target approach speed is "fast", and the space congestion is "high", the mechanical collision risk level is determined to be "high risk". These rules cover all possible input scenarios, ensuring that the system can provide reasonable risk assessment results under various circumstances. The fuzzy rule base systematically organizes and expresses expert experience, enabling this valuable knowledge to be effectively utilized in the system. At the same time, it also facilitates subsequent modification and improvement of the rules based on actual conditions, continuously improving the system's performance and accuracy.
[0091] The Mamdani inference method is used for fuzzy inference. It determines the degree to which an input variable belongs to each fuzzy set based on its actual value and membership function. Then, it infers the degree of membership of the output variable in each fuzzy set according to rules in the fuzzy rule base. This inference method fully considers the interaction between various factors and accurately infers the fuzzy result of the mechanical collision risk level. The centroid method is used to defuzzify the fuzzy inference result, converting the fuzzy set of the output variable into a quantitative risk value. This quantitative risk value corresponds to the mechanical collision risk level at the construction site, providing a clear and explicit basis for construction safety early warning and decision-making. This facilitates timely implementation of appropriate measures by construction and management personnel to ensure the safety of the construction site.
[0092] In some embodiments, the fuzzy control system determines the level of mechanical collision risk at the construction site, specifically including:
[0093] The Mamdani inference mechanism is used to perform fuzzy inference on the input variables to obtain the degree of membership of the output variables in each fuzzy set.
[0094] The centroid method is used to defuzzify the membership degree of the output variable in each fuzzy set, and the fuzzy set of the output variable is converted into a quantitative risk value, which corresponds to the mechanical collision risk level at the construction site.
[0095] Specifically, the Mamdani inference mechanism is used to perform fuzzy inference on the input variables. Since the risk of mechanical collisions at construction sites is influenced by a combination of factors, such as the distance between machines, relative speed, and direction of movement, these factors are often fuzzy and uncertain. The Mamdani inference mechanism can effectively handle this fuzzy information. Through pre-defined fuzzy rules, it comprehensively considers the interactions between various input variables, fully assesses the probability of mechanical collisions, and thus obtains the membership degree of the output variable in each fuzzy set, providing a basis for accurately determining the risk level.
[0096] The centroid method is used to defuzzify the membership degree of the output variable in each fuzzy set, and the fuzzy set is converted into a quantitative risk value. The quantitative risk value has a clear numerical meaning and can more intuitively reflect the level of mechanical collision risk. For example, the risk value can be divided into different intervals, corresponding to different risk levels, such as low risk, medium risk, and high risk, which makes it easier for construction personnel and managers to understand and judge quickly.
[0097] By organically combining fuzzy reasoning and defuzzification, this technical solution can more accurately determine the level of mechanical collision risk at construction sites. Compared with traditional risk assessment methods based on precise mathematical models, it can better handle fuzzy and uncertain information at construction sites, reduce misjudgments and omissions, and improve the reliability of risk identification.
[0098] In some embodiments, the use of the Mamdani inference mechanism to perform fuzzy inference on the input variables specifically includes:
[0099] The degree to which an input variable belongs to each fuzzy set is determined based on its actual value and membership function.
[0100] Reasoning is performed based on the rules in the fuzzy rule base to obtain the degree of membership of the output variable in each fuzzy set.
[0101] Specifically, the construction site environment is complex, and the input variables related to the risk of mechanical collisions (such as the distance between machines, relative speed, etc.) are often fuzzy and uncertain, making it difficult to describe them accurately with precise values. The Mamdani inference mechanism can handle this fuzzy information well. It determines the degree to which the input variable belongs to each fuzzy set based on the actual value of the input variable and the membership function, quantifies the uncertain information, and provides an accurate basis for subsequent inference.
[0102] Reasoning is performed based on the rules in the fuzzy rule base, which is established based on expert experience and actual construction site conditions. It contains the logical relationship between various combinations of input variables and output results. The Mamdani reasoning mechanism can reasonably infer the degree of membership of input variables in various fuzzy sets according to these rules, and obtain the degree of membership of output variables in various fuzzy sets. This rule-based reasoning method can make full use of prior knowledge and improve the efficiency and accuracy of reasoning.
[0103] The mechanical movement and spatial position at the construction site are constantly changing, and the input variables also change in real time. The Mamdani inference mechanism can quickly infer based on the new input variable values and rule base, update the output results in a timely manner, realize the real-time assessment of mechanical collision risks, and provide timely decision-making basis for construction safety management.
[0104] By accurately processing fuzzy information and efficiently performing rule-based reasoning, the Mamdani inference mechanism helps to more accurately determine the risk level of mechanical collisions at construction sites. It can comprehensively consider the influence of multiple factors, avoid misjudgments caused by the limitations of a single factor or precise numerical model, and improve the reliability of risk assessment.
[0105] A dual-modal construction scene target detection and tracking system includes:
[0106] The dual-modal video acquisition module is used to acquire dual-modal video of targets at the construction site during on-site operations.
[0107] The model building module is used to build target detection and target tracking models based on the dual-modal video of the target at the construction site.
[0108] The real-time data acquisition module is used to acquire video stream data of construction site operations in real time.
[0109] The target detection module is used to input real-time video stream data into the target detection model, and the target detection model outputs the target detection results;
[0110] The target tracking module is used to input real-time video stream data and target detection results into the target tracking model to obtain the target trajectory result;
[0111] The fuzzy control module is used to input the target detection results and target trajectory results into the fuzzy control system, which then determines the mechanical collision risk level at the construction site.
[0112] A computing device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the dual-modal construction scene target detection and tracking method described above.
[0113] A computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the dual-modal construction scene target detection and tracking method described above.
[0114] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A dual-modal target detection and tracking method for construction scenarios, characterized in that, Specifically, it includes: Collect dual-modal video of the construction site target during on-site operations; A target detection model and a target tracking model are constructed based on the dual-modal video of the target at the construction site; Real-time acquisition of video stream data from construction site operations; Real-time video stream data is input into the target detection model, and the target detection model outputs the target detection result; Real-time video stream data and target detection results are input into the target tracking model to obtain the target trajectory result; The target detection results and target trajectory results are input into the fuzzy control system, which then determines the mechanical collision risk level at the construction site.
2. The dual-modal construction scene target detection and tracking method according to claim 1, characterized in that, The construction of the target detection model based on the dual-modal video of the construction site target specifically includes: The dual-modal video of the target at the construction site is processed to obtain the target detection dataset; Based on the object detection dataset, a Dual-YOLO object detection model was built and trained using GPU in the PyTorch framework environment, thus constructing the object detection model.
3. The dual-modal construction scene target detection and tracking method according to claim 2, characterized in that, The process of processing the bimodal video of the construction site target to obtain the target detection dataset specifically includes: Images are extracted from bimodal video using video frame extraction to obtain visible light and infrared images; The visible light image and the infrared image are converted into intermediate mode images respectively using a mean filter, resulting in visible light intermediate mode images and infrared intermediate mode images; Feature point detection and directional gradient histogram description of feature information are performed on visible light intermediate mode images and infrared intermediate mode images to obtain visible light target feature maps and infrared target feature maps; The visible light target feature map and the infrared target feature map are labeled to obtain the target detection dataset.
4. The dual-modal construction scene target detection and tracking method according to claim 3, characterized in that, The construction of the target tracking model specifically includes: The dual-modal video of the target at the construction site is input into the target detection model to obtain the target detection results of the dual-modal video. The target detection results of the bimodal video are input into the bimodal target tracking model based on the ByteTrack algorithm for training, and the target tracking model is constructed.
5. The dual-modal construction scene target detection and tracking method according to claim 4, characterized in that, The construction of the fuzzy control system specifically includes: Based on the target detection results output by the target detection model and the target trajectory results output by the target tracking model, the target proximity index, target approach speed index, and spatial congestion index are selected as input variables of the fuzzy control system, and the mechanical collision risk level at the construction site is used as the output variable. For each input and output variable, construct a fuzzy set and set the membership function of the fuzzy set; Build a fuzzy rule base; Construct the Mamdani inference mechanism and the centroid method for fuzzy resolution.
6. The dual-modal construction scene target detection and tracking method according to claim 5, characterized in that, The fuzzy control system determines the level of mechanical collision risk at the construction site, specifically including: The Mamdani inference mechanism is used to perform fuzzy inference on the input variables to obtain the degree of membership of the output variables in each fuzzy set. The centroid method is used to defuzzify the membership degree of the output variable in each fuzzy set, and the fuzzy set of the output variable is converted into a quantitative risk value, which corresponds to the mechanical collision risk level at the construction site.
7. The dual-modal construction scene target detection and tracking method according to claim 6, characterized in that, The use of the Mamdani inference mechanism to perform fuzzy inference on the input variables specifically includes: The degree to which an input variable belongs to each fuzzy set is determined based on its actual value and membership function. Reasoning is performed based on the rules in the fuzzy rule base to obtain the degree of membership of the output variable in each fuzzy set.
8. A dual-modal construction scene target detection and tracking system, characterized in that, include: The dual-modal video acquisition module is used to acquire dual-modal video of the construction site target during on-site operations; The model building module is used to build target detection and target tracking models based on the dual-modal video of the target at the construction site. The real-time data acquisition module is used to acquire video stream data of construction site operations in real time. The target detection module is used to input real-time video stream data into the target detection model, and the target detection model outputs the target detection results; The target tracking module is used to input real-time video stream data and target detection results into the target tracking model to obtain the target trajectory result; The fuzzy control module is used to input the target detection results and target trajectory results into the fuzzy control system, which then determines the mechanical collision risk level at the construction site.
9. A computing device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the dual-modal construction scene target detection and tracking method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the steps of the dual-modal construction scene target detection and tracking method according to any one of claims 1-7.
Citation Information
Patent Citations
A Method and System for Worker and Machinery Collision Avoidance Prediction Based on Computer Vision
CN114937240A
Power grid safety early warning method and device, computer equipment and storage medium
CN115620208A