Behavior detection method of engineering project, electronic equipment and program product
By uploading the image to be reviewed to the large model for review when the confidence level of the small model is uncertain, and adjusting and training parameters based on the case library, the problems of false detection and missed detection of the small model in the construction site environment are solved, and efficient construction safety monitoring is achieved.
Patent Information
- Application Number
- CN202511178988.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-09-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing engineering site safety monitoring, behavior detection based on small models of edge devices is susceptible to environmental interference, resulting in a high misjudgment rate, while full upload of large models to the cloud for review is limited by resources and real-time performance. Existing technologies have failed to effectively solve the problems of false detection and missed detection of small models in different environments.
When the confidence level of the small model continuously detects the same target category within the target threshold range, the image to be reviewed is uploaded to the large model for review. Based on the comparison between the detection results of the large model and the small model, the parameters of the small model are adjusted. Incremental training and fine-tuning are performed in combination with the dispute case library and the false detection case library to optimize the model performance.
It improves the detection accuracy and robustness of small models in different environments, reduces false detections and missed detections, realizes self-optimization and continuous improvement of the detection system, and improves the real-time and effectiveness of construction safety monitoring.
Smart Images

Figure CN120673482A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a behavior detection method, electronic equipment, and program product for an engineering project. Background Art
[0002] In construction site safety monitoring, lightweight models (small models) based on edge devices can detect targets in real time, but their output confidence is easily affected by environmental interference (such as occlusion and lighting), leading to a dilemma in handling low-confidence targets. Blindly trusting the small model will result in a high rate of misjudgment, triggering invalid alerts. Uploading the entire data to the cloud for review is limited by resources and real-time performance. Summary of the Invention
[0003] The present disclosure provides a behavior detection method for an engineering project, electronic equipment, and a program product.
[0004] According to one aspect of the present disclosure, a behavior detection method for an engineering project is provided, including: acquiring an image to be reviewed, the image to be reviewed being determined by a small model when the confidence level of continuously detecting entity targets of the same target category is within a target threshold range; identifying the image to be reviewed based on the large model to determine an image detection result; generating an alarm message when the image detection result of the large model is that an entity target of the target category exists in the image to be reviewed; and adjusting the parameters of the small model based on a comparison result between the image detection result of the large model and the image detection result of the small model.
[0005] According to one aspect of the technical solution, when the small model continuously detects entity targets of the same target category and the confidence level is within the target threshold range, the image to be reviewed is determined and uploaded, and the large model reviews the detection results of the small model based on the image to be reviewed. This can reduce the ineffective occupation of resources while ensuring the accuracy of the recognition results.
[0006] Moreover, by adjusting the parameters of the small model based on the comparison results between the image detection results of the large model and the image detection results of the small model, the small model can be continuously adapted to different construction scenarios and environmental changes, thereby improving the accuracy and robustness of detection of various risk behaviors and reducing false detection and missed detection problems caused by fixed model parameters.
[0007] According to the behavior detection method of the engineering project of at least one embodiment of the present disclosure, the parameters of the small model are adjusted according to the comparison result between the image detection result of the large model and the image detection result of the small model, including: when the image detection result of the large model is that the entity target of the target category exists in the image to be reviewed, and the image detection result of the small model is that the entity target of the target category does not exist in the image to be reviewed, lowering the detection threshold of the small model for the target category, and the detection threshold is the lower limit of confidence for judging whether the entity target of the corresponding target category exists in the image; when the image detection result of the large model is that the entity target of the target category does not exist in the image to be reviewed, and the image detection result of the small model is that the entity target of the target category exists in the image to be reviewed, increasing the detection threshold of the small model for the target category.
[0008] According to the technical solution of this embodiment, the parameters of the small model are automatically adjusted according to the comparison results, so that the small model can continuously adapt to different construction scenarios and environmental changes, improve the detection accuracy and robustness of various risk behaviors, reduce false detection and missed detection problems caused by fixed model parameters, and realize self-optimization and continuous improvement of the detection system.
[0009] According to the behavior detection method of the engineering project of at least one embodiment of the present disclosure, the method also includes: acquiring a site image collected at the construction site; identifying the site image based on the large model to determine the scene category of the construction site, wherein the scene category is used to characterize the lighting conditions of the construction site; and adjusting the parameters of the small model according to the scene category.
[0010] According to the technical solution of this embodiment, by dynamically adjusting the small model parameters according to the lighting conditions of the construction site, the small model can better adapt to different environmental changes, improve its detection accuracy and robustness of target categories in various lighting scenarios, and reduce false detection and missed detection problems caused by lighting factors.
[0011] According to the behavior detection method for engineering projects of at least one embodiment of the present disclosure, the method further includes: adding case data in which there is a judgment conflict between the large model and the small model to a dispute case library; and performing incremental training on the small model based on the case data in the dispute case library.
[0012] According to the technical solution of this embodiment, by using the data in the dispute case library to perform incremental training on the small model, it is possible to specifically solve the problem of inconsistency between the small model and the large model in the actual detection process, and continuously improve the detection accuracy and adaptability of the small model to various targets, making it more stable and reliable in complex and changeable construction site environments.
[0013] According to the behavior detection method for engineering projects of at least one embodiment of the present disclosure, the method also includes: adding case data in which the large model makes incorrect judgments to a false detection case library; and fine-tuning parameters of the large model based on the case data in the false detection case library.
[0014] According to the technical solution of this embodiment, by using the data in the false detection case library to perform targeted parameter fine-tuning on the large model, the false detection problem of the large model in specific target categories or scenarios can be effectively corrected, and its detection accuracy and robustness for complex situations on the construction site can be continuously improved, thereby reducing the occurrence of false detections and improving the quality of construction safety monitoring.
[0015] According to the behavior detection method for engineering projects of at least one embodiment of the present disclosure, after fine-tuning the large model according to the case data in the false detection case library, the method further includes: retraining the small model according to the fine-tuned large model, wherein, when retraining the small model, logical loss and distillation loss are added to the loss function corresponding to the target category related to the target risk behavior, the logical loss is the cross entropy between the predicted probability output by the small model and the judgment result of the large model, and the distillation loss is the distance between the probability distribution of the small model and the large model.
[0016] According to the technical solution of this embodiment, by introducing logical loss and distillation loss, the small model can learn the judgment results and probability distribution characteristics of the large model, thereby more accurately identifying target risk behaviors, reducing the occurrence of false detection and missed detection, and improving the detection accuracy and reliability of the small model.
[0017] According to at least one embodiment of the present disclosure, a behavior detection method for an engineering project is provided, in which the image to be reviewed is identified based on the large model to determine an image detection result, including: guiding the large model to make a directional judgment on the image to be reviewed based on a structured prompt word template, wherein the structured prompt word template includes one or more target category inspection items; and determining the image detection result based on the output result of the large model for the target category inspection item.
[0018] The technical solution of this embodiment provides clear guidance and standards for the large model's judgment through a structured prompt word template, enabling it to more accurately identify target categories and risky behaviors in images. Furthermore, the unified prompt word template ensures consistent judgment across different images and scenarios, reducing biases caused by unclear or inconsistent prompts.
[0019] According to at least one embodiment of the present disclosure, the behavior detection method for engineering projects further includes: generating multiple prompt word templates with semantic differences for the same target category; deploying the multiple prompt word templates in parallel into the large model; and selecting one of the multiple prompt word templates as the target structured prompt word template based on the inference results of the large model under different prompt word templates.
[0020] According to the technical solution of this embodiment, by generating multiple prompt word templates with semantic differences and performing comparative evaluation, the prompt word template with the best performance in terms of accuracy and stability can be screened out, thereby improving the judgment quality of the large model in image detection tasks and enabling it to more accurately identify target risk behaviors.
[0021] According to at least one embodiment of the present disclosure, a behavior detection method for an engineering project selects one of multiple prompt word templates as a target structured prompt word template based on the inference results of the large model under different prompt word templates, including: determining the accuracy and stability index of the inference results corresponding to each prompt word template based on the inference results of the large model under different prompt word templates; selecting a predetermined number of prompt word templates with the best performance from the multiple prompt word templates as candidate prompt word templates based on the accuracy and stability index; and conducting an A / B test based on the candidate prompt word templates to determine the target structured prompt word template.
[0022] According to the technical solution of this embodiment, through the comprehensive evaluation of accuracy and stability indicators, the performance of different prompt word templates under various conditions can be comprehensively and objectively measured, ensuring that the screened prompt word templates have high accuracy and reliability, and providing high-quality candidates for subsequent A / B testing.
[0023] According to another aspect of the present disclosure, an electronic device is provided, comprising: a memory storing execution instructions; and a processor executing the execution instructions stored in the memory, so that the processor executes the method of any embodiment of the present disclosure.
[0024] According to another aspect of the present disclosure, a readable storage medium is provided, wherein the readable storage medium stores execution instructions, and when the execution instructions are executed by a processor, the execution instructions are used to implement the method of any embodiment of the present disclosure.
[0025] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the method of any embodiment of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings illustrate exemplary embodiments of the present disclosure and together with the description serve to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification.
[0027] Figure 1 A schematic diagram of an application scenario that can be applied to the present disclosure is shown.
[0028] Figure 2 A flow chart of a behavior detection method for an engineering project according to an embodiment of the present disclosure is shown.
[0029] Figure 3 A flow chart of step S240 in a behavior detection method for an engineering project according to an embodiment of the present disclosure is shown.
[0030] Figure 4 A schematic diagram of a process for adjusting small model parameters based on illumination, which is also included in a behavior detection method for an engineering project according to an embodiment of the present disclosure, is shown.
[0031] Figure 5 A schematic diagram of a process for incrementally training a small model, which is also included in a behavior detection method for an engineering project according to an embodiment of the present disclosure, is shown.
[0032] Figure 6 A schematic diagram of a process for fine-tuning a large model, which is also included in a behavior detection method for an engineering project according to an embodiment of the present disclosure, is shown.
[0033] Figure 7 A flow chart of step S220 in a behavior detection method for an engineering project according to an embodiment of the present disclosure is shown.
[0034] Figure 8 A schematic diagram of a process for determining a structural prompt word template is shown in a behavior detection method for an engineering project according to an embodiment of the present disclosure.
[0035] Figure 9 A flow chart of step S830 in a method for detecting a behavior of an engineering project according to an embodiment of the present disclosure is shown.
[0036] Figure 10 A schematic diagram of a framework of a behavior detection system for an engineering project according to an embodiment of the present disclosure is shown.
[0037] Figure 11 A schematic structural block diagram of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0038] The present disclosure is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific examples described herein are intended only to illustrate the relevant content and are not intended to limit the present disclosure. It should also be noted that, for ease of description, only the portions relevant to the present disclosure are shown in the accompanying drawings.
[0039] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in the present disclosure can be combined with each other. The technical solution of the present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0040] With the rapid development of high-risk industries such as industrial construction, transportation, and the electric power and petrochemical industries, on-site operating environments are complex and fraught with numerous risk factors. Safety supervision has long relied on manual inspections and passive management, resulting in inadequate oversight and delayed risk response. To improve the real-time and automated nature of safety management, artificial intelligence (AI) technology, particularly vision-based object detection, is being gradually introduced into the field of safety monitoring.
[0041] In recent years, lightweight object detection algorithms such as the YOLO (You Only Look Once) series have been widely used in edge devices, enabling real-time detection on low-computing devices. At the same time, large language models (such as GPT-4 and Qwen) possess powerful semantic understanding and multimodal perception capabilities, demonstrating significant advantages in handling complex and abstract scene judgment tasks. Combining these two types of models to achieve collaborative judgment through a "small edge model + large backend model" approach has become a new trend in intelligent security monitoring.
[0042] Currently, a mature solution for edge devices like smart helmets is to use detection models such as YOLOv5 and YOLOv8 to detect key safety equipment or behavioral targets (such as not wearing a helmet, fire extinguishers, ladders, etc.), deployed on embedded terminals (such as smart helmets and surveillance cameras). Large model technologies (such as CLIP, BLIP-2, and Qwen-VL) are primarily deployed in the cloud to perform more complex task judgments, such as scene attribute analysis (whether it is a high-altitude working environment), behavior recognition (whether a fire is being started), and text and image understanding.
[0043] However, in the above scheme, the small model is prone to misjudging low-confidence targets, and because its threshold is fixed and lacks dynamic adjustment capabilities, it is easy to cause over-detection or missed detection.
[0044] To this end, the present disclosure proposes the following technical solution, in which when the small model continuously detects entity targets of the same target category and the confidence level is within the target threshold range, the image to be reviewed is determined and uploaded, and the large model reviews the detection results of the small model based on the image to be reviewed. This can reduce the ineffective occupation of resources while ensuring the accuracy of the recognition results.
[0045] Furthermore, by comparing the image detection results of the large model with those of the small model, generating an alert or adjusting the parameters of the small model can ensure the effectiveness of the alert generation. Furthermore, if there is a conflict between the two detection results, adjusting the parameters of the small model can improve the accuracy of the detection results output by the small model.
[0046] Figure 1 A schematic diagram of an application scenario that can be applied to the present disclosure is shown.
[0047] like Figure 1 As shown, the application scenario may include a terminal device 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the terminal device 101 and the server 103, which may include various connection types, such as a wired communication link, a wireless communication link, and the like.
[0048] It should be understood that Figure 1 The number of terminal devices 101, networks 102, and servers 103 in the embodiment is merely illustrative. Depending on the implementation requirements, there may be any number of terminal devices 101, networks 102, and servers 103. For example, the server 103 may be a server cluster consisting of multiple servers.
[0049] In actual use, construction workers can carry terminal devices 101 (such as smart helmets) while working at the construction site. Terminal devices 101 can capture images of the construction site or work process. A small model configured on terminal device 101 can perform real-time detection of the captured images to determine whether risky behavior exists. When the small model continuously detects physical objects of the same target category with a confidence level within a target threshold, it identifies an image for review and sends it via network 102 to server 103, which is equipped with a large model.
[0050] The server can identify the image to be reviewed based on the large model and determine an image detection result. The image detection result indicates whether an object of the target category is present in the image to be reviewed. The image detection result of the large model is compared with the image detection result of the small model. Based on the comparison result, an alert is generated or the parameters of the small model are adjusted. The alert is used to describe the risk behavior in the image to be reviewed.
[0051] Figure 2 FIG. 1 shows a flow chart of a behavior detection method for an engineering project according to an embodiment of the present disclosure. Figure 2 As shown, the behavior detection method of the engineering project includes at least steps S210 to S240, which are described in detail as follows.
[0052] In step S210, an image to be reviewed is obtained, where the image to be reviewed is determined by the small model when the confidence level of entity targets of the same target category is continuously detected to be within a target threshold range.
[0053] Entity targets refer to specific objects or behaviors that can be directly observed in construction scene images or videos, with clear shapes, outlines, and positions. These targets typically possess clear visual features, such as shape, color, and texture, enabling computer vision models to distinguish them from the background through image processing and analysis techniques.
[0054] Target categories can be abstract concept sets that classify various physical targets or abstract concepts based on their characteristics, uses, or the types of safety hazards they represent, in accordance with the needs of construction safety supervision. These target categories serve as the basis for classifying targets and are used to organize and manage target detection tasks, enabling the model to adopt appropriate detection strategies and algorithms for different target categories. In behavioral detection of engineering projects, targets can be divided into categories such as "not wearing a hard hat," "wooden ladders," "fire extinguishers," "edges and holes," "height work," and "hot work." For example, the target category "not wearing a hard hat" encompasses all physical targets such as personnel who fail to wear a hard hat as required. "Height work" is an abstract category that may encompass various construction operations performed at height and related physical targets, such as personnel working at height and scaffolding used for height work.
[0055] The small model can be a lightweight target detection algorithm model deployed in edge devices (such as smart helmets, surveillance cameras, etc.), such as YOLOv5 and YOLOv8 in the YOLO series.
[0056] Confidence can be the degree to which the small model believes that the detected entity target belongs to a specific category (such as "person", "safety helmet", etc.) in target detection. It is usually expressed as a value between 0 and 1. The closer it is to 1, the more confident the model is in judging the target category.
[0057] The target threshold interval can be a confidence range interval pre-set by technical personnel in this field based on prior experience. According to the target threshold interval, detection targets that are less certain about the small model and may have the risk of false detection or missed detection can be screened out, and then the detection results that need further review can be determined.
[0058] In this embodiment, small models in edge devices such as front-end smart helmets (such as detection models based on the YOLO series) can perform target detection on real-time collected images or video streams, identify various types of entity targets (such as "not wearing a helmet", "ladder", "fire extinguisher", etc.), and assign a confidence value to each detected entity target.
[0059] Continuously monitor the detection results and their corresponding confidence levels output by the small model. When the confidence level of a detected entity target of the same target category (e.g., "person") continuously falls within a pre-set target threshold range (e.g., [0.4, 0.6]), the system determines that the situation requires further review. It should be noted that "continuous" can be a reasonable number or time interval set based on the actual application scenario, such as three consecutive detection results meeting the confidence level within the range, or multiple such occurrences within a minute.
[0060] When the above-mentioned confidence trigger conditions are met, the image associated with the detection result (i.e., the frame or multiple frames of images containing the physical target whose confidence is within the target threshold range) can be determined as the image to be reviewed and transmitted from the edge device to the back-end server or cloud for further in-depth analysis and judgment by the large model.
[0061] In this way, by setting the target threshold range, the parts with higher uncertainty in the small model detection results can be effectively screened out. These parts are often key points that are prone to false detection or missed detection, thereby avoiding sending all detection results directly to the large model for review, reducing unnecessary waste of computing resources, and improving the operating efficiency of the entire system.
[0062] In step S220, the image to be reviewed is identified based on the large model to determine an image detection result.
[0063] Large models, such as the Qwen2.5-VL-7B, can be equipped with powerful semantic understanding and multimodal perception capabilities. They demonstrate significant advantages in handling complex and abstract scene judgment tasks and can be deployed on backend servers or in the cloud to review small model detection results and handle more complex task judgments.
[0064] The image detection result of the large model can be the judgment information output by the large model after recognizing the input image to be reviewed, indicating whether an object of the target category is present in the image. It is usually presented as a binary result (such as presence / absence, 0 / 1, etc.) and can also include a brief explanation or reasoning for the judgment result.
[0065] In this embodiment, a large model on a backend server or cloud receives images for review transmitted from a frontend edge device (such as a smart helmet). The large model leverages its multimodal perception capabilities to extract features from the input images for review, analyzing the various visual elements, scene information, and possible semantic associations within the images. These include the presence of specific physical objects (such as people, helmets, ladders, etc.), the relative positions of these objects, and the environmental context (e.g., whether the work environment is elevated, or whether there are any edges or holes).
[0066] Based on the above analysis and reasoning process, the large model obtains the image detection results, clearly points out whether there is an entity target corresponding to the target category in the image to be reviewed, and outputs the results in the set format.
[0067] Please continue to refer to Figure 2 In step S230, when the image detection result of the large model is that an entity target of the target category exists in the image to be reviewed, an alarm information is generated, and the alarm information is used to describe the risk behavior existing in the image to be reviewed.
[0068] When a system determines that a risky behavior (i.e., a physical object of the target category) exists in an image to be reviewed, an alert message is generated describing the risky behavior. This alert message may include a detailed description of the risky behavior type, the location at which the image was captured (typically represented by the edge device's positioning information), and the severity of the risk (e.g., high, medium, or low risk). This allows personnel to quickly identify potential safety hazards on-site and take appropriate countermeasures.
[0069] In this embodiment, after the large model determines that there is an entity target of the target category in the image to be reviewed, corresponding warning information can be generated according to the category and severity of the corresponding risk behavior.
[0070] In one example, a corresponding alarm template can be retrieved from a pre-set alarm template library. The alarm template is then populated and customized, combining the image detection results of the large model with relevant information, such as the location at the time of image acquisition and a description of the risk scenario, to generate a specific alarm message. For example, the alarm message could read: "In surveillance images at [specific time], a person was detected working at height in [specific area] without a helmet. This poses a high-risk safety hazard. Please immediately visit the site for verification and resolution."
[0071] After the corresponding alarm information is generated, it can be promptly sent to designated security supervisors or relevant responsible persons through pre-set channels (such as SMS, email, system push, etc.). At the same time, the alarm information can also be displayed in a prominent manner on the front-end interface of the monitoring system (such as the display screen of the security monitoring center, the mobile device app of the manager, etc.), making it convenient for relevant personnel to view and handle it in a timely manner.
[0072] In this way, by generating detailed alarm information and sending it to relevant personnel in a timely manner, it can ensure that construction safety supervisors are aware of the risk behaviors on site at the first time, take timely measures to intervene and deal with them, effectively avoid the occurrence of safety accidents, and improve the timeliness and effectiveness of construction safety management.
[0073] Please continue to refer to Figure 2 In step S240, the parameters of the small model are adjusted according to the comparison result between the image detection result of the large model and the image detection result of the small model.
[0074] In this embodiment, after the large model outputs the corresponding image detection result for the image to be reviewed, the image detection result of the large model can be compared with the image detection result of the corresponding small model (for the image to be reviewed) to determine whether the image detection results of the two are consistent.
[0075] If the comparison results indicate that the small model's image detection results contain false positives or missed detections, the small model's parameter adjustment mechanism can be triggered according to pre-set parameter adjustment rules. For example, if the large model determines that a certain type of object (such as "edges and holes") exists but the small model fails to detect it, the small model's detection threshold for that type of object can be lowered to improve recall. Conversely, if the small model frequently falsely detects a certain type of object, the corresponding detection threshold can be increased to reduce false positives.
[0076] Therefore, the parameters of the small model are automatically adjusted according to the comparison results, so that the small model can continuously adapt to different construction scenarios and environmental changes, improve the accuracy and robustness of detection of various risk behaviors, reduce false detection and missed detection problems caused by fixed model parameters, and realize self-optimization and continuous improvement of the detection system.
[0077] Regarding step S240, in some embodiments of the present disclosure, it may include the following: Figure 3 Steps S241 to S242 are shown.
[0078] In step S241, when the image detection result of the large model is that an entity target of the target category exists in the image to be reviewed, and the image detection result of the small model is that an entity target of the target category does not exist in the image to be reviewed, the detection threshold of the small model for the target category is lowered. The detection threshold is the lower limit of confidence for judging whether an entity target of the corresponding target category exists in the image.
[0079] The detection threshold can be the lower confidence limit for determining whether a physical object of the corresponding target category exists in an image during object detection. When the small model's detection confidence for a particular target category exceeds the threshold, the image is deemed to contain a physical object of that category; otherwise, it is deemed to be absent. It should be noted that the detection thresholds for different target categories can be the same or different, and this is not specifically limited.
[0080] In this embodiment, based on the comparison results between the image detection results of the two models, if the large model determines that there is an entity target of a certain target category in the image to be reviewed, but the small model determines that it does not exist, it can be determined that the detection threshold of the small model for the target category needs to be adjusted at this time, so as to optimize the detection performance of the small model.
[0081] In one example, the magnitude of the detection threshold reduction for a particular target class can be determined based on a pre-defined threshold adjustment strategy. For example, the threshold can be lowered by a fixed step size (e.g., 0.05 or 0.1) each time it is adjusted, or the adjustment step can be dynamically determined based on the confidence level of the large model or historical adjustment data. The detection threshold for the target class in the small model is then lowered by the determined adjustment step. The updated detection threshold is then applied to the small model, replacing the original threshold setting, completing the local update of the model parameters.
[0082] In this way, by lowering the detection threshold, the small model can more easily determine the presence of physical targets of this target category in the image in subsequent detections, thereby identifying more potential risk behaviors that may have been missed, improving the detection recall rate of construction violations, and reducing the omission of safety hazards.
[0083] In step S242, when the image detection result of the large model is that there is no entity target of the target category in the image to be reviewed, and the image detection result of the small model is that there is an entity target of the target category in the image to be reviewed, the detection threshold of the small model for the target category is increased.
[0084] In this embodiment, based on the comparison results, if the large model determines that a physical object of a certain target category is not present in the image to be reviewed, while the small model determines that it is present, the small model's detection threshold for that target category can be adjusted. Specifically, this can increase the small model's detection threshold for that target category, making it less likely for the small model to determine that a physical object of that target category is present in the image during subsequent detections. This reduces false alarms caused by misdetection by the small model and improves the accuracy and reliability of the detection results.
[0085] Figure 4 A flow chart of adjusting small model parameters based on illumination, which is also included in the behavior detection method for an engineering project according to an embodiment of the present disclosure, is shown. The flow chart includes at least steps S410 to S430, which are described in detail as follows.
[0086] In step S410, a site image collected at the construction site is acquired.
[0087] In this embodiment, image acquisition equipment (such as cameras, smart helmets, etc.) installed at the construction site is used to capture images of the construction site according to set time intervals or trigger conditions (such as detection of human activity, equipment operation, etc.), and the captured image data is transmitted to the back-end system.
[0088] In step S420, the scene image is recognized based on the large model to determine the scene category of the construction site, where the scene category is used to characterize the lighting conditions of the construction site.
[0089] The scene category can be classified based on the lighting conditions at the construction site, for example, "insufficient lighting," "strong direct light," "weak light and shadows," etc., or simply divided into "indoor scene" or "outdoor scene," etc. The lighting environment characteristics of the construction site can be characterized based on the scene category.
[0090] In this embodiment, after receiving the on-site image, the large model can use its powerful multimodal perception and semantic understanding capabilities to analyze the on-site image and identify the scene category of the construction site, that is, to determine which of the above-mentioned predefined categories the lighting conditions at the construction site belong to.
[0091] In step S430, the parameters of the small model are adjusted according to the scene category.
[0092] In this embodiment, the corresponding adjustment strategy can be determined based on pre-defined small model parameter adjustment rules for different scene categories. For example, for "low-light" scenes, the small model's detection thresholds for certain target categories (such as "hardhats" and "people") can be pre-set to be lowered to improve detection of these targets in low-light environments. For "strong direct sunlight" scenes, the small model's image preprocessing parameters can be adjusted, such as increasing contrast and adjusting brightness, to improve image quality and, therefore, detection effectiveness.
[0093] Based on the determined adjustment strategy, specific adjustments are made to the relevant parameters of the small model. Parameter adjustments can include modifying detection thresholds, optimizing image preprocessing parameters, and adjusting certain hyperparameters in the object detection algorithm. Once the adjustments are complete, the updated small model parameters are applied to the small model on the edge device, replacing the original parameter settings. This allows the small model to better adapt to the lighting conditions of the construction site based on the current scene type in subsequent detections, thereby improving detection performance.
[0094] In this way, by dynamically adjusting the small model parameters according to the lighting conditions at the construction site, the small model can better adapt to different environmental changes, improve its detection accuracy and robustness of target categories in various lighting scenarios, and reduce false detection and missed detection problems caused by lighting factors.
[0095] Figure 5 FIG. 1 shows a flow chart of incremental training of a small model in a behavior detection method for an engineering project according to an embodiment of the present disclosure. Figure 5 As shown, incremental training of the small model includes at least steps S510 to S520, which are described in detail below.
[0096] In step S510, case data in which there is a judgment conflict between the large model and the small model is added to the dispute case library.
[0097] The dispute case database may be a database for storing case data of judgment conflicts between the large model and the small model. The case data may include disputed images, detection results of the large model, detection results of the small model, and related confidence information.
[0098] In this embodiment, each time the large and small model's detection results for the same image are compared, if they differ in their judgments about whether a certain object category exists in the image, a conflict exists. This case can be identified as a dispute. Detailed information about the disputed case is then extracted and added to the disputed case database for storage.
[0099] In one example, in a dispute case library, dispute cases can be classified and labeled, for example, they can be labeled according to dimensions such as target category, scenario characteristics, and conflict type, so that these data can be used more specifically in subsequent incremental training.
[0100] In step S520, incremental training is performed on the small model based on the case data in the dispute case library.
[0101] In this embodiment, the need to initiate incremental training of the small model is determined based on pre-set incremental training trigger conditions (e.g., the amount of data in the dispute case database reaches a certain threshold, or the number of dispute cases accumulated within a certain period of time reaches a certain threshold). For example, incremental training can be triggered every time 1,000 pieces of data in the dispute case database are added.
[0102] During incremental training, a data subset that meets the current incremental training requirements can be selected from the dispute case library. Based on the incremental training objectives (such as optimization for specific target categories, improved adaptability to specific scenarios, etc.), relevant dispute cases are selected from the library and their data is preprocessed, including image enhancement and data standardization, to meet the data requirements for small model training.
[0103] During the training process, the parameters of the small model can be tuned by adjusting training strategies such as learning rate and optimizer parameters while maintaining the original parameters of the small model, thereby improving the accuracy of the detection results of the small model.
[0104] In this way, by using the data in the dispute case library to incrementally train the small model, it is possible to specifically solve the problem of inconsistency between the small model and the large model during the actual detection process, and continuously improve the small model's detection accuracy and adaptability for various targets, making it more stable and reliable in complex and changeable construction site environments.
[0105] Moreover, the incremental training method does not require training a small model from scratch. Instead, it only needs to optimize and adjust the original model, which greatly saves training time and computing resource costs, improves the efficiency of model updates, and enables the improved model to be applied to actual inspections more quickly, responding to new situations and new problems at the construction site in a timely manner.
[0106] Figure 6 A schematic diagram of a process for fine-tuning a large model, which is also included in a method for detecting behavior of an engineering project according to an embodiment of the present disclosure, is shown. The process includes at least steps S610 to S620, which are described in detail below.
[0107] In step S610, the case data in which the large model makes an erroneous judgment is added to the false detection case library.
[0108] Among them, the false detection case library can be a database used to store case data of incorrect judgments made by the large model during the image detection process. The case data may include incorrectly judged images, incorrect detection results given by the large model (such as incorrect target category judgment, incorrect scene description, etc.) and correct annotation information (for subsequent model correction and optimization).
[0109] In this embodiment, after the large model detects and outputs results from a live image or an image to be reviewed, it compares these results with manually annotated accurate results or verified correct results to identify errors in the large model's judgment. If the large model's detection results are inconsistent with the correct results, the case is identified as a false positive, and its detailed data (including the image, the false positive detection result, the correct annotated result, etc.) is extracted and added to the false positive case library for storage.
[0110] In one embodiment, false detection cases and Xining classifications and detailed annotations can be made according to error types (e.g., falsely judging that a target category exists when it actually does not exist, falsely judging that the target category is wrong, etc.), scene characteristics, target category and other dimensions, so that these data can be used more specifically in the subsequent parameter fine-tuning process.
[0111] In step S620, the parameters of the large model are fine-tuned according to the case data in the false detection case library.
[0112] In this embodiment, the need to initiate parameter fine-tuning for the large model can be determined based on pre-set parameter fine-tuning trigger conditions (e.g., the amount of data in the false positive case database reaches a certain threshold, the cumulative number of false positive cases within a certain period of time reaches a certain threshold, or the concentration of false positive cases in a specific target category or scenario reaches a set ratio). For example, parameter fine-tuning can be triggered every time 500 pieces of data in the false positive case database are added.
[0113] When fine-tuning parameters, you can filter a data subset from the false positive case library to meet the current fine-tuning requirements. Based on the fine-tuning goal (such as addressing false positives for specific target categories or improving adaptability to specific scenarios), select appropriate false positive cases from the library and perform data preprocessing, including image enhancement and data normalization, to meet the data requirements for fine-tuning large models. At the same time, prepare corresponding prompt word templates (such as optimized prompt words for different target categories) and fine-tuning parameter configurations (such as learning rate and batch size).
[0114] During fine-tuning, most parameters of the large model can be kept fixed, and only the parameters of specific layers or modules related to false detection issues can be adjusted and optimized. A cue-word-based learning approach can be employed to optimize cue-word templates to guide the model in more accurately understanding the object categories and scene information in the image, thereby reducing false detections. Furthermore, by combining the correctly labeled information in the fine-tuning data and using an appropriate loss function (such as a classification loss function) to guide the update of model parameters, the model's detection accuracy can be gradually improved.
[0115] In this way, by using the data in the false detection case library to fine-tune the parameters of the large model in a targeted manner, it is possible to effectively correct the false detection problems of the large model in specific target categories or scenarios, continuously improve its detection accuracy and robustness for complex situations on the construction site, reduce the occurrence of false detections, and improve the quality of construction safety monitoring.
[0116] Moreover, the parameter fine-tuning method does not require large-scale retraining of the large model. It only requires local optimization and adjustment based on the original model to address the false detection problem. This greatly saves training time and computing resource costs, improves the efficiency of model updates, and enables the improved model to be applied to actual inspections in a timely manner, quickly responding to changes and needs at the construction site.
[0117] In some embodiments of the present disclosure, after fine-tuning the large model based on case data in the false positive case library, the method further includes: The small model is retrained based on the fine-tuned large model, wherein, when retraining the small model, logical loss and distillation loss are added to the loss function corresponding to the target category related to the target risk behavior, the logical loss is the cross entropy between the predicted probability output by the small model and the judgment result of the large model, and the distillation loss is the distance between the probability distribution of the small model and the large model.
[0118] Target risk behaviors can be pre-specified abstract behaviors, such as working near edges or holes, working at heights, and working with hot objects. Compared to targets with clear and easily identifiable visual features like "not wearing a hard hat" or "fire extinguisher," targets corresponding to abstract behaviors have fuzzy outlines and a variety of forms. Therefore, relying solely on small models to identify these behaviors will not yield effective results.
[0119] In this embodiment, after fine-tuning the parameters of the large model using the false positive case library, the latest version of the fine-tuned large model is used as the basis for retraining the small model. When retraining the small model, in addition to the original loss functions (such as category loss and bounding box loss), the logistic loss and distillation loss are added to the loss function corresponding to the target category associated with the target risk behavior. Specifically, the logistic loss is calculated by calculating the cross-entropy between the predicted probability output by the small model and the judgment result (0 / 1 binarization) of the large model for that category of target. The distillation loss is calculated by calculating the KL divergence between the probability distributions predicted by the small model and the large model for the target category. For risk behaviors other than the target risk behavior, training continues using the original loss function.
[0120] In one example, the loss function for the target risk behavior can be as follows: in, is an indicator function (if sample i belongs to the category set C (i.e., target risk behavior) that requires customized loss, the value is 1, otherwise it is 0. For example, the modified loss is used for abstract targets such as edges and holes, while the original loss is used for targets that are easy to judge such as helmets). are the logistic loss and distillation loss corresponding to the i-th sample respectively.
[0121] , is the predicted probability of the small model for the i-th type of target, is the judgment result of the large model on this type of target (0 / 1 binarization). , KL divergence is used to measure the distance between the probability distributions of two models.
[0122] is the loss function of the original small model, including category loss and border loss. 、 、 They are adjustable hyperparameters. In addition, the soft labels output by the large model are used as guidance to provide more supervisory signals.
[0123] Using the prepared dataset and constructed loss function, the small model is fully retrained. During training, the small model not only learns the characteristics and annotation information of the data itself, but also learns the judgment results and probability distribution information of the large model through logistic loss and distillation loss terms. This integrates the knowledge and experience of the large model into the small model, improving its detection performance.
[0124] In this way, by introducing logical loss and distillation loss, the small model can learn the judgment results and probability distribution characteristics of the large model, thereby more accurately identifying target risk behaviors, reducing the occurrence of false detections and missed detections, and improving the detection accuracy and reliability of the small model.
[0125] Regarding step S220, in some embodiments of the present disclosure, as Figure 7 As shown, step S220 at least includes steps S221 to S222, which are described in detail as follows.
[0126] In step S221 , the large model is guided to perform orientation judgment on the image to be reviewed according to a structured prompt word template, wherein the structured prompt word template includes one or more target category inspection items.
[0127] The structured prompt template can be a list of guiding questions designed according to a specific format and logic, used to guide the large model in examining and judging the target categories in the image. Each inspection item clearly identifies the target category or risk behavior characteristic to be judged, usually in the form of "whether...", and requires the model to provide a concise "yes" or "no" answer in sequence, along with a brief justification.
[0128] In one example, the structured prompt word template may be as follows: Please carefully examine each area of the photo, scanning the details from left to right and from top to bottom, and determine whether the following nine types of targets are clearly present in the image. Ignore any targets that are incomplete due to truncation at the edge of the image: 1. There are people who are not wearing helmets; 2. People who wear hard hats but do not wear hard hat lanyards; 3. Wooden stepladder; 4. Metal ladders, such as aluminum alloy; 5. Fire extinguisher; 6. Personnel smoking; 7. Personnel working with hot objects; 8. Personnel are working outside on high floors; 9. There are places without guardrails near edges or holes, such as the edges of stairs, elevator shafts, windows, balconies, deep pits, etc.
[0129] Please answer "yes" or "no" for each element in order. For example, if none of the elements exist, please answer "1. No\n2. No\n3. No\n4. No\n5. No\n6. No\n7. No\n8. No\n9. No", for a total of 9 items, and explain your reasons. In this embodiment, a structured prompt word template suitable for the current construction scenario and inspection requirements is retrieved from the system's preset prompt word template library. This template covers a variety of target category inspection items related to construction safety, such as not wearing a hard hat and not wearing a safety belt when working at height.
[0130] The image to be reviewed and the structured prompt word template are fed into the large model. Leveraging its powerful multimodal perception capabilities and the guiding questions in the prompt word template, the large model conducts a detailed analysis of the image. Following the order of the prompt word template, the large model analyzes and determines whether the image contains the features or behaviors described by each target category check item. For example, for the check item "Is there a person without a helmet?", the large model carefully identifies the features of the person's head in the image to determine whether a helmet is present.
[0131] Please continue to refer to Figure 7 In step S222, the image detection result is determined based on the output result of the large model for the target category inspection item.
[0132] In this example, based on the results of each item's analysis, the large model outputs a judgment result in the format required by the prompt word template. For each check item, a clear answer of "yes" or "no" is output, along with a brief justification, such as "The person's head is obscured in the image, making it impossible to determine whether they are not wearing a helmet." This provides a basis for subsequent review and processing.
[0133] In this way, the structured prompt word template provides clear guidance and standards for the large model's judgment, enabling it to more accurately identify target categories and risky behaviors in images. Furthermore, the unified prompt word template ensures consistent judgment across different images and scenarios, reducing biases caused by unclear or inconsistent prompts.
[0134] Furthermore, by requiring large models to output justification for their decisions, each judgment result is supported by evidence, enhancing the transparency and explainability of the judgment process. This helps human reviewers better understand the model's judgment logic, facilitates review and analysis of the results, and improves the credibility of the entire system.
[0135] Figure 8 FIG. 1 shows a flow chart of determining a structure prompt word template in a behavior detection method for an engineering project according to an embodiment of the present disclosure. Figure 8 As shown, determining the structured prompt word template at least includes steps S810 to S830, which are described in detail below.
[0136] In step S810 , a plurality of prompt word templates with semantic differences are generated for the same target category.
[0137] In this embodiment, for the same target category, multiple prompt word templates with different semantics are generated based on manual experience and historical judgment results. For example, for the target category "height work", there may be the following three prompt word templates: Template Version A: "Does this image pertain to working at height?" Template Version B: "Is there a risk of working at height in the image?" Template version C: "Please determine whether the person in the picture is working in an area more than 2 meters high?" In step S820, multiple prompt word templates are deployed in parallel into the large model.
[0138] In this embodiment, the generated multiple prompt word templates are deployed in parallel to the large model. When performing image detection, the large model simultaneously receives the same image to be reviewed and analyzes and judges it according to different prompt word templates, generating multiple sets of inference results.
[0139] In step S830 , based on the inference results of the large model under different prompt word templates, one of the multiple prompt word templates is selected as a target structured prompt word template.
[0140] In this embodiment, based on the inference results of the large model under different prompt word templates, the performance of each template can be compared in terms of judgment accuracy, consistency, and degree of agreement with manual annotation results. The prompt word template with the best performance is selected as the target structured prompt word template.
[0141] In this way, by generating multiple prompt word templates with semantic differences and conducting comparative evaluations, we can screen out the prompt word template that performs best in terms of accuracy and stability, thereby improving the judgment quality of the large model in image detection tasks and enabling it to more accurately identify target risk behaviors.
[0142] Regarding step S830, in some embodiments of the present disclosure, as Figure 9 As shown, it at least includes steps S831 to S833, which are described in detail as follows.
[0143] In step S831, based on the inference results of the large model under different prompt word templates, the accuracy and stability index of the inference results corresponding to each prompt word template are determined.
[0144] The accuracy rate can be the ratio of the number of samples correctly predicted by the large model under the guidance of the prompt word template to the total number of samples, which is used to measure the accuracy of the model judgment under the prompt word template.
[0145] The stability index can measure the judgment consistency of the prompt word template under different scenes, lighting, image quality and other conditions, and is used to evaluate the robustness and reliability of the prompt word template.
[0146] In this embodiment, a large model can collect inference results from a large number of sample images using different cue word templates, along with the corresponding accurate annotation information. For each cue word template, the corresponding inference results are compared with the accurate results from manual annotations. The number of correctly predicted samples is counted, and the corresponding accuracy rate is calculated. In one example, the accuracy rate can be calculated as follows: Accuracy = (number of correctly predicted samples / total number of samples) × 100%.
[0147] Next, the consistency of the inference results for each cue word template under different scenario conditions is analyzed. For example, stability can be assessed by calculating the fluctuations in the model's accuracy under different lighting intensities, angles, and background complexities. In one example, statistical methods can be used, such as calculating the standard deviation of the accuracy. The smaller the standard deviation, the more stable the model's performance across various scenarios, and the better the stability indicator.
[0148] In step S832, based on the accuracy rate and the stability index, a predetermined number of prompt word templates with the best performance are selected from the plurality of prompt word templates as candidate prompt word templates.
[0149] In this embodiment, based on preset accuracy and stability thresholds, a predetermined number (e.g., the top three) of the best-performing prompt word templates are screened from multiple prompt word templates to be selected as candidate prompt word templates. In one example, a weighted sum calculation can be performed based on the accuracy and stability index corresponding to each prompt word template to obtain a corresponding evaluation score, thereby determining the top predetermined number of best-performing prompt word templates as candidate prompt word templates.
[0150] In step S833, an A / B test is performed based on the candidate prompt word templates to determine a target structured prompt word template.
[0151] A / B testing can be a process in which different prompt word templates are deployed on different business branches or equipment batches, and the performance data of each group is collected and analyzed to determine which prompt word template has a better effect.
[0152] In this embodiment, the candidate prompt word templates are deployed on different business branches or equipment batches for A / B testing. The sample size of each test group is ensured to be sufficiently large and representative, and the operating environment, data distribution, and other conditions are kept as consistent as possible between the groups to accurately compare the effectiveness of different prompt word templates.
[0153] During the A / B test, we continuously collected performance data for each group, including accuracy, false positive rate, user feedback, and other indicators. We conducted statistical analysis on the collected data and compared the performance of different prompt word templates in actual applications.
[0154] Based on the A / B test results, the best-performing prompt word template is selected as the target prompt word template, taking into account factors such as accuracy, false positive rate, and user feedback. For example, if a prompt word template performs best in both accuracy and user satisfaction, it will be determined as the target structured prompt word template and applied to the actual image detection task.
[0155] In this way, through the comprehensive evaluation of accuracy and stability indicators, we can comprehensively and objectively measure the performance of different prompt word templates under various conditions, ensure that the selected prompt word templates have high accuracy and reliability, and provide high-quality candidates for subsequent A / B testing.
[0156] In addition, A / B testing can compare and verify multiple candidate prompt word templates in the actual operating environment, fully considering various factors in actual business scenarios (such as user feedback, the impact of misjudgment, etc.), so as to more accurately determine the optimal prompt word template and improve the performance of the entire detection system and user experience.
[0157] Based on the technical solutions of the above embodiments, a specific application scenario of the embodiments of the present application is introduced below: Figure 10 FIG. 1 shows a schematic diagram of a behavior detection system for an engineering project according to an embodiment of the present disclosure. Figure 10 As shown, the behavior detection system may include a small model 1010 , a large model 1020 , and a database 1030 .
[0158] Specifically, small model 1010 can perform real-time target detection based on image data collected at the construction site, identifying specific physical objects such as unworn hardhats, ladders, and fire extinguishers, and assigning a confidence score to each detected object. When the confidence score for a specific target category is continuously detected within a target threshold, a corresponding warning image (i.e., an image to be reviewed) is determined and sent to large model 1020 for further identification and judgment.
[0159] In one embodiment, the small model 1010 can be trained based on Yolov8. Specifically, the end-side small model can be trained based on existing collected data, where the detection categories include {not wearing a hardhat, wooden ladder, fire extinguisher, etc.} and {near edge or hole, working at height, hot work, etc.}. The former category is a specific physical target and can be trained with the small model to achieve relatively good results. The second category is a more abstract target and cannot be effectively trained with the small model alone.
[0160] To this end, we can abstract related entities from "edge and hole, high-altitude work, and hot work." For example, high-altitude work and hot work must be performed by a person, so the small model's goal is to detect the "person," and then let the large model understand the scene. Edge and hole scenes primarily include "windows, balconies, elevator shafts, stairs, and deep pits." Therefore, we can first convert these abstract objects to be identified into the basic objects they contain, and then use the large model's multimodal understanding capabilities to make judgments.
[0161] According to the above training objectives, data collection and data labeling are carried out, Yolov8 is used for model training, and the trained model is converted into ncnn format and implanted into the front-end edge device (such as smart helmet).
[0162] It should be noted that during the algorithm reasoning process, the threshold settings for different entity targets can be different. For targets that are easier to detect, since the algorithm itself can achieve better results, a relatively high threshold can be set to reduce false detections of the algorithm, thereby reducing the transmission of warning data.
[0163] For more abstract targets or targets that are difficult to identify, a lower threshold can be set to improve the algorithm's recall rate.
[0164] Large model 1020 receives the warning image sent by small model 1010, identifies it, and determines an image detection result. This image detection result indicates whether a physical object of the target category exists in the image to be reviewed. Simultaneously, it combines this with a prompt word template to perform a targeted judgment, outputting a detailed judgment result and reasoning. Based on the image detection results of large model 1020, it can determine whether small model 1010 has made false detections or missed detections (i.e., a secondary warning determination), allowing for dynamic adjustment of the detection threshold of small model 1010 for the target category (i.e., dynamic threshold adjustment).
[0165] In one example, the large model 1020 can be fine-tuned based on Qwen2.5-VL-7B to improve the accuracy of hidden danger classification, where the fine-tuning adopts the lora method. Specifically, different types of typical hidden danger pictures can be collected for small and scattered projects. The image data can be taken on-site from multiple small construction sites, covering indoor and outdoor, multi-angle, occlusion, night and other complex conditions, and the aforementioned structured prompt word template is used to annotate the collected pictures. The annotation process is reviewed by two people and reviewed by one person to ensure the consistency of the annotation. The large model is fine-tuned based on Qwen2.5-VL-7B.
[0166] In addition, the system can collect case data where there are judgment conflicts between the small model 1010 and the large model 1020 and add them to the dispute case library A in the database, and collect case data where the large model 1020 makes incorrect judgments and add them to the large model misdetection library B in the database.
[0167] The system can fine-tune the parameters of the large model 1020 based on the case data in the large model false detection library B. The small model can then be retrained based on the fine-tuned large model 1020, and the small model 1010 can be incrementally trained based on the case data in the dispute case library A, thereby continuously optimizing the system's detection performance.
[0168] The present disclosure also provides an electronic device. Figure 11 A schematic diagram showing a hardware implementation using a processing system is shown.
[0169] like Figure 11 As shown, the hardware structure of electronic device 1000 can be implemented using a bus architecture. The bus architecture can include any number of interconnecting buses and bridges, depending on the specific application and overall design constraints of the hardware. Bus 1100 connects various circuits including one or more processors 1200, memory 1300, and / or hardware modules. Bus 1100 can also connect various other circuits 1400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc. Bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Component Architecture (EISA) bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of illustration, the figure only uses a single connecting line, but this does not mean that there is only one bus or only one type of bus.
[0170] The present disclosure also provides a readable storage medium, in which a computer program is stored, and the computer program is used to implement the above-mentioned method when executed by a processor. "Readable storage medium" can be any device that can contain storage, communication, dissemination or transmission programs for use in an instruction execution system, device or equipment or in combination with these instruction execution systems, devices or equipment. More specific examples of readable storage media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and editable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM), etc.
[0171] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented in whole or in part using software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, the process or function of the present disclosure is performed in whole or in part.
[0172] A computer program or instruction can be stored in a readable storage medium or transferred from one readable storage medium to another. For example, the computer program or instruction can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The readable storage medium can be any accessible medium or a data storage device such as a server or data center that integrates one or more accessible media. The accessible medium can be a magnetic medium such as a floppy disk, hard disk, or magnetic tape; an optical medium such as a digital video disk; or a semiconductor medium such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.
[0173] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0174] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, apparatuses, electronic devices, and computer program products according to the present disclosure. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0175] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0176] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0177] In the description of this specification, the description with reference to the terms "one embodiment / method", "some embodiments / methods", "example", "specific example", or "some examples" means that the specific features, structures, or characteristics described in conjunction with the embodiment / method or example are included in at least one embodiment / method or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment / method or example. Moreover, the specific features, structures, or characteristics described may be combined in a suitable manner in any one or more embodiments / methods or examples. In addition, those skilled in the art may combine and combine different embodiments / methods or examples described in this specification and the features of different embodiments / methods or examples, unless they are contradictory.
[0178] Those skilled in the art will appreciate that the above embodiments are merely intended to clearly illustrate the present disclosure and are not intended to limit the scope of the present disclosure. Other changes or modifications may be made based on the above disclosure, and such changes or modifications are still within the scope of the present disclosure.
Claims
1. A behavior detection method for an engineering project, characterized in that: include: Acquire an image to be reviewed, wherein the image to be reviewed is determined by a small model when the confidence level of entity targets of the same target category is continuously detected within a target threshold range, and the small model is a lightweight target detection algorithm model deployed in the edge device; Identify the image to be reviewed based on a large model to determine an image detection result, wherein the large model is a model with multimodal perception capabilities deployed on a backend server or cloud; When the image detection result of the large model indicates that an entity target of the target category exists in the image to be reviewed, generating warning information, wherein the warning information is used to describe the risk behavior existing in the image to be reviewed; According to the comparison result between the image detection result of the large model and the image detection result of the small model, the parameters of the small model are adjusted.
2. The method according to claim 1, wherein Adjusting parameters of the small model according to a comparison result between the image detection result of the large model and the image detection result of the small model includes: When the image detection result of the large model indicates that an entity target of the target category exists in the image to be reviewed, and the image detection result of the small model indicates that an entity target of the target category does not exist in the image to be reviewed, lowering the detection threshold of the small model for the target category, where the detection threshold is a lower confidence limit value for determining whether an entity target of the corresponding target category exists in the image; When the image detection result of the large model is that there is no entity target of the target category in the image to be reviewed, and the image detection result of the small model is that there is an entity target of the target category in the image to be reviewed, the detection threshold of the small model for the target category is increased.
3. The method according to claim 1, wherein The method further comprises: Acquire on-site images collected at the construction site; Recognizing the on-site image based on the large model to determine a scene category of the construction site, where the scene category is used to characterize lighting conditions at the construction site; According to the scene category, the parameters of the small model are adjusted.
4. The method according to claim 1, wherein The method further comprises: Adding case data in which there is a judgment conflict between the large model and the small model to a dispute case library; The small model is incrementally trained based on the case data in the dispute case library.
5. The method according to claim 1, wherein The method further comprises: Adding case data judged incorrectly by the large model to a false positive case database; Parameters of the large model are fine-tuned according to case data in the false detection case library.
6. The method according to claim 5, wherein After fine-tuning the large model based on the case data in the false positive case library, the method further includes: The small model is retrained based on the fine-tuned large model, wherein, when retraining the small model, logical loss and distillation loss are added to the loss function corresponding to the target category related to the target risk behavior, the logical loss is the cross entropy between the predicted probability output by the small model and the judgment result of the large model, and the distillation loss is the distance between the probability distribution of the small model and the large model.
7. The method according to claim 1, wherein Identifying the image to be reviewed based on the large model and determining an image detection result includes: Guiding the large model to perform directional judgment on the image to be reviewed according to a structured prompt word template, wherein the structured prompt word template includes one or more target category inspection items; An image detection result is determined based on the output result of the large model for the target category inspection item.
8. The method according to claim 7, wherein The method further comprises: For the same target category, multiple prompt word templates with different semantics are generated; Deploying multiple prompt word templates in parallel into the large model; According to the inference results of the large model under different prompt word templates, one of the multiple prompt word templates is selected as the target structured prompt word template.
9. An electronic device, characterized in that: include: a memory storing execution instructions; A processor, wherein the processor executes the execution instructions stored in the memory, so that the processor executes the method according to any one of claims 1 to 8.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Model deployment method, end side equipment and storage medium
CN118135377A
Target detection method and device, equipment, storage medium and product
CN118968038A
Safety monitoring method based on mixing of large model and neural network algorithm
CN119380166A
Object category recognition model training method and apparatus, and object category recognition method and apparatus
WO2025167876A1