Security detection method and device, electronic equipment and storage medium
Through the multimodal large model combined with the object detection algorithm, it analyzes whether workers wear safety protective equipment in compliance, solves the problems of insufficient data dependence and rough processing logic in the existing technology, and achieves more efficient and accurate security detection.
Patent Information
- Application Number
- CN202510487610.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, the safety detection method based on the relative position of workers and safety protection tools has insufficient data dependence and rough data processing logic, resulting in insufficient accuracy of the safety detection results.
The multimodal large model is used to combine the object detection algorithm to obtain the target image frame, detect the target object and obtain its characteristic description information, and use the multimodal large language model to analyze whether the organism is compliant to wear safety protection tools, improving the accuracy of data dependence and data processing.
It improves the accuracy of security detection results, reduces the consumption of computing resources of electronic devices, enhances detection efficiency, and supports diversified detection needs.
Smart Images

Figure CN120279490A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular to application fields such as artificial intelligence, large models, security detection, video processing, etc. Specifically, the present disclosure relates to a security detection method, apparatus, electronic device, and storage medium. Background Art
[0002] In an industrial production environment and / or an engineering construction environment, accidents often occur due to workers' neglect of safety measures, and wearing safety protection equipment is one of the most basic and important protection measures. Safety protection equipment can effectively protect workers from reducing the risk of injury when encountering unexpected situations such as falling objects and impacts. Summary of the Invention
[0003] The present disclosure provides a security detection method, apparatus, electronic device, and storage medium.
[0004] According to a first aspect of the present disclosure, there is provided a security detection method, including:
[0005] Obtaining a target image frame;
[0006] Detecting the target image frame to determine at least one target object and feature description information of at least one target object;
[0007] When it is determined that there is a living body among at least one target object, using a multi-modal large model, based on the target image frame and the feature description information of at least one target object, obtaining a security detection result for the living body; wherein the security detection result is used to represent whether the living body wears safety protection equipment in compliance.
[0008] According to a second aspect of the present disclosure, there is provided a security detection apparatus, including:
[0009] An image acquisition unit for obtaining a target image frame;
[0010] An image detection unit for detecting the target image frame to determine at least one target object and feature description information of at least one target object;
[0011] A result acquisition unit for, when it is determined that there is a living body among at least one target object, using a multi-modal large model, based on the target image frame and the feature description information of at least one target object, obtaining a security detection result for the living body; wherein the security detection result is used to represent whether the living body wears safety protection equipment in compliance.
[0012] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0013] At least one processor;
[0014] A memory communicatively connected to the at least one processor;
[0015] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided in the first aspect of the present disclosure.
[0016] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method provided in the first aspect of the present disclosure.
[0017] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the method provided in the first aspect of the present disclosure.
[0018] Adopting the present disclosure can improve the accuracy of security detection results.
[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0021] Figure 1 is a schematic flowchart of a security detection method provided by an embodiment of the present disclosure;
[0022] Figure 2 is a schematic diagram of a method for obtaining multiple target detection results provided by an embodiment of the present disclosure;
[0023] Figure 3 is a complete flowchart illustration of a security detection method provided by an embodiment of the present disclosure;
[0024] Figure 4 is a schematic diagram of an application scenario of a security detection method provided by an embodiment of the present disclosure;
[0025] Figure 5 is a schematic structural block diagram of a security detection device provided by an embodiment of the present disclosure;
[0026] Figure 6 is a schematic structural block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The following describes exemplary embodiments of the present disclosure in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.
[0028] As previously mentioned, in an industrial production environment and / or an engineering construction environment, accidents often result from workers' neglect of safety measures, and wearing safety protection equipment is one of the most basic and important protection measures. Safety protection equipment can effectively protect workers from reducing the risk of injury when encountering unexpected situations such as falling objects and impacts. Currently, the safety detection results for workers are usually obtained in the following ways:
[0029] Obtain a target environment image;
[0030] Determine the relative positions of the worker and the safety protection equipment in the target environment image;
[0031] Based on the relative positions of the worker and the safety protection equipment, obtain the safety detection results for the worker to be used to characterize whether the worker wears the safety protection equipment compliantly.
[0032] However, through research by the inventor, it is found that in the above method, the safety detection results for the worker are obtained only based on the relative positions of the worker and the safety protection equipment. This process has insufficient data dependence and relatively rough data processing logic. Therefore, the accuracy of the safety detection results cannot be ensured.
[0033] To address the above problems, the embodiments of the present disclosure provide a safety detection method that can be applied to an electronic device. Among them, the electronic device can be a server, a workbench, a mainframe computer, or other similar computing devices. The following will be combined with Figure 1 the following flow schematic diagram to describe a safety detection method provided by the embodiments of the present disclosure. It should be noted that although the logical order is shown in the flow schematic diagram, in some cases, the steps shown or described in the flowchart can also be executed in other orders.
[0034] Step S101, obtain a target image frame.
[0035] Among them, the target image frame can be a video image frame intercepted from the video data stream after photographing the detection environment to obtain the video data stream. Here, the detection environment can be an industrial production environment or an engineering construction environment, and the embodiments of the present disclosure do not limit this.
[0036] Step S102: Detect the target image frame to determine at least one target object and the feature description information of at least one target object.
[0037] Among them, at least one target object may include organisms, may also include safety protection equipment, or may include both organisms and safety protection equipment at the same time. Here, the organism may be a person or other living beings; the safety protection equipment may be a safety helmet, protective clothing, safety shoes, face protection mask, goggles, etc.
[0038] In addition, it should be noted that in the embodiments of the present disclosure, for each target object among at least one target object, its feature description information may include the target category, target location, and fusion confidence level related to the target object. Among them, the target category is used to characterize whether the target object is an organism or safety protection equipment (specifically, it can be used to characterize which category of safety protection equipment); the target location is used to characterize the coordinate position of the target object in the target image frame; the fusion confidence level is used to characterize the credibility of the target category and target location of the target object.
[0039] Step S103: In the case where it is determined that there is an organism among at least one target object, use the multimodal large model to obtain the safety detection result for the organism based on the target image frame and the feature description information of at least one target object.
[0040] Among them, the multimodal large model is the Multimodal Large Language Models (MLLM), which is a neural network model that combines the natural language processing capabilities of the large language model (LLM) and the understanding and generation capabilities of data in other modalities (such as the visual modality). Here, the LLM can be a pre-trained neural network model (such as an autoregressive generation model with a Transformer architecture), which has general language knowledge, world knowledge, and professional knowledge in various fields (such as professional knowledge in the field of computer technology).
[0041] In one example, when it is determined that there is an organism in at least one target object, a multimodal large model can be used to obtain a safety detection result for the organism based on the target image frame and the feature description information of each target object in the at least one target object, so as to characterize whether the organism wears safety protection equipment compliantly. Among them, the organism wearing safety protection equipment compliantly can be that the organism wears the safety protection equipment correctly and the actual color of the safety protection equipment worn by the organism is the same as the target color, and the target color is a color matched with the organism; correspondingly, the organism not wearing safety protection equipment compliantly can be that the organism does not wear safety protection equipment, or that the organism does not wear the safety protection equipment correctly, or that the organism wears the safety protection equipment correctly, but the actual color of the safety protection equipment worn by the organism is different from the target color. Here, the organism wearing the safety protection equipment correctly can be that the organism wears the safety protection equipment at the correct position and at the correct angle; correspondingly, the organism not wearing the safety protection equipment correctly can be that the organism does not wear the safety protection equipment at the correct position and / or does not wear the safety protection equipment at the correct angle.
[0042] By using the safety detection method provided in the embodiments of the present disclosure, first, a target image frame will be obtained, and the target image frame will be detected to determine at least one target object and the feature description information of the at least one target object. After that, the target image frame and the feature description information of the at least one target object will be jointly used as rich and reliable data. When it is determined that there is an organism in the at least one target object, a multimodal large model is used to obtain a safety detection result for the organism based on the rich and reliable data, so as to characterize whether the organism wears safety protection equipment compliantly. Compared with the prior art, on the one hand, the problem of insufficient data dependence is solved; on the other hand, a multimodal large model (that is, a neural network model that combines the natural language processing ability of the LLM and the understanding and generation ability of data in other modalities (such as the visual modality)) is used to obtain a safety detection result for the organism based on the rich and reliable data, which also solves the problem of relatively rough data processing logic. Therefore, the accuracy of the safety detection result can be improved.
[0043] In some alternative embodiments, the target image frame can be obtained in the following ways:
[0044] In a detection environment, the video data stream captured by a monitoring camera is processed to obtain target image frames. Specifically, the video data stream captured by the monitoring camera can be segmented to obtain multiple candidate image frames, and the multiple candidate image frames are de-duplicated to obtain multiple target image frames, and then the image data is fed back. For example, according to the arrangement order of the multiple target image frames in the video data stream, each target image frame in the multiple target image frames is sent to an electronic device, that is, sent to a server, a workbench, a mainframe computer or other similar computing devices as an intelligent cloud. Among them, when de-duplicating the multiple candidate image frames, an image similarity comparison algorithm can be used to analyze and compare each candidate image frame in the multiple candidate image frames, determine multiple duplicate image frames with relatively high similarity and multiple non-duplicate image frames other than the multiple duplicate image frames from the multiple candidate image frames, select one duplicate image frame from the multiple duplicate image frames as a target image frame, and use the multiple non-duplicate image frames as target image frames to achieve de-duplication of the multiple candidate image frames.
[0045] Through the above method, in the embodiments of the present disclosure, de-duplication of multiple candidate image frames can be achieved to obtain multiple target image frames. In this way, the data processing volume of the electronic device can be reduced, thereby improving the execution efficiency of the security detection method. At the same time, the consumption of computing resources of the electronic device can be reduced.
[0046] In some optional embodiments, step S102, that is, "detect the target image frame to determine at least one target object and the feature description information of at least one target object" may include:
[0047] Step S102-1, determine multiple target detection algorithms.
[0048] Among them, the multiple target detection algorithms may include a first target detection algorithm that is good at detecting large targets and a second target detection algorithm that is good at detecting small targets. Here, the first target detection algorithm may be an improved You Only Look Once (YOLO) target detection algorithm, for example, the PP-YOLOE Plus target detection algorithm; the second target detection algorithm may be the first Real-Time Detection Transformer (RT-DETR) target detection algorithm.
[0049] Step S102-2, respectively use the multiple target detection algorithms to detect the target image frame to obtain multiple target detection results with relevance.
[0050] Among them, the multiple target detection results correspond to the multiple target detection algorithms one by one.
[0051] In addition, it should be noted that in the embodiments of the present disclosure, the target detection result may include a preliminary detection object. At the same time, it may also include the preliminary detection category, preliminary detection position of the preliminary detection object, and the algorithm confidence related to the preliminary detection object. Among them, the preliminary detection category is used to characterize whether the preliminary detection object is a living organism or a safety protection device (specifically, which category of safety protection device it is); the preliminary detection position is used to characterize the coordinate position of the preliminary detection object in the target image frame; the algorithm confidence is used to characterize the credibility of the preliminary detection category and preliminary detection position of the preliminary detection object.
[0052] It should also be noted that in the embodiments of the present disclosure, multiple target detection results with correlation may be target detection results with substantially the same preliminary detection position.
[0053] Please combine Figure 2 , for example, using the first target detection algorithm to detect the target image frame, the obtained first target detection result includes:
[0054] The first preliminary detection object;
[0055] The preliminary detection category of the first preliminary detection object: white safety helmet;
[0056] The preliminary detection position of the first preliminary detection object: (X11, Y11)&(X12, Y12);
[0057] The algorithm confidence related to the first preliminary detection object: 60%;
[0058] Among them, the preliminary detection position of the first preliminary detection object can be represented by a detection frame. Specifically, (X11, Y11) can be used to represent the coordinate position of the upper left corner point of the detection frame of the first preliminary detection object; (X12, Y12) can be used to represent the coordinate position of the lower right corner point of the detection frame of the first preliminary detection object.
[0059] Using the first target detection algorithm to detect the target image frame, the obtained second target detection result includes:
[0060] The second preliminary detection object;
[0061] The preliminary detection category of the second preliminary detection object: living organism;
[0062] The preliminary detection position of the second preliminary detection object: (X21, Y21)&(X22, Y22);
[0063] The algorithm confidence related to the second preliminary detection object: 60%;
[0064] Among them, the initial inspection position of the second initial inspection object can be represented by a detection frame. Specifically, (X21, Y21) can be used to represent the coordinate position of the upper left corner point of the detection frame of the second initial inspection object; (X22, Y22) can be used to represent the coordinate position of the lower right corner point of the detection frame of the second initial inspection object.
[0065] Using the second object detection algorithm to detect the target image frame, the obtained third object detection results include:
[0066] The third initial inspection object;
[0067] The initial inspection category of the third initial inspection object: black safety helmet;
[0068] The initial inspection position of the third initial inspection object: (X31, Y31)&(X32, Y32);
[0069] The algorithm confidence related to the third initial inspection object: 90%.
[0070] Among them, the initial inspection position of the third initial inspection object can be represented by a detection frame. Specifically, (X31, Y31) can be used to represent the coordinate position of the upper left corner point of the detection frame of the third initial inspection object; (X32, Y32) can be used to represent the coordinate position of the lower right corner point of the detection frame of the third initial inspection object.
[0071] Using the second object detection algorithm to detect the target image frame, the obtained fourth object detection results include:
[0072] The fourth initial inspection object;
[0073] The initial inspection category of the fourth initial inspection object: organism;
[0074] The initial inspection position of the fourth initial inspection object: (X41, Y41)&(X42, Y42);
[0075] The algorithm confidence related to the fourth initial inspection object: 90%.
[0076] Among them, the initial inspection position of the fourth initial inspection object can be represented by a detection frame. Specifically, (X41, Y41) can be used to represent the coordinate position of the upper left corner point of the detection frame of the fourth initial inspection object; (X42, Y42) can be used to represent the coordinate position of the lower right corner point of the detection frame of the fourth initial inspection object.
[0077] As described above, in the embodiments of the present disclosure, multiple object detection results with relevance may be object detection results with basically the same initial detection positions. Therefore, the first object detection result and the third object detection result belong to multiple object detection results with relevance, that is, the first object detection result and the third object detection result belong to a group of object detection results; the second object detection result and the fourth object detection result belong to multiple object detection results with relevance, that is, the second object detection result and the fourth object detection result belong to a group of object detection results.
[0078] Step S102-3: Fuse multiple object detection results to determine the target object and the feature description information of the target object.
[0079] In one example, multiple object detection results may be fused in the following manner to determine the target object and the feature description information of the target object:
[0080] Determine the initially detected object included in the multiple object detection results as the target object;
[0081] Based on the initially detected category, initially detected position, and algorithm confidence included in the multiple object detection results, determine the target category, target position of the target object, and the fusion confidence related to the target object;
[0082] Determine the target category, target position of the target object, and the fusion confidence related to the target object as the feature description information of the target object.
[0083] That is to say, in the embodiments of the present disclosure, the initially detected objects included in the multiple object detection results may be determined as the same target object. For example, the first initially detected object included in the first object detection result and the third initially detected object included in the third object detection result are determined as the same target object; the second initially detected object included in the second object detection result and the fourth initially detected object included in the fourth object detection result are determined as the same target object.
[0084] After determining the initially detected objects included in the multiple object detection results as the target object, based on the initially detected category, initially detected position, and algorithm confidence included in the multiple object detection results, the target category, target position of the target object, and the fusion confidence related to the target object may be determined. This process may include at least one of the following:
[0085] (1) Based on the algorithm confidence included in the multiple object detection results, determine the target category of the target object from the initially detected categories included in the multiple object detection results.
[0086] In one example, for the nth object detection result among multiple object detection results, the following data processing logic can be used to obtain the confidence parameter value of the nth object detection result among multiple object detection results:
[0087] Mn = Pn * CLn
[0088] Among them, M is used to represent the confidence parameter value of the nth object detection result among multiple object detection results; P is used to represent the weight value related to the nth object detection result among multiple object detection results; CL is used to represent the algorithm confidence related to the initial inspection object in the nth object detection result among multiple object detection results.
[0089] In one specific example, the following method can be used to determine the weight value related to the nth object detection result among multiple object detection results:
[0090] Determine the area of the detection frame represented by the initial inspection position included in the nth object detection result among multiple object detection results as the reference area;
[0091] When the object detection algorithm corresponding to the nth object detection result among multiple object detection results is the first object detection algorithm that is good at detecting large objects, obtain a weight value that is positively correlated with the reference area as the weight value related to the nth object detection result among multiple object detection results;
[0092] Or, when the object detection algorithm corresponding to the nth object detection result among multiple object detection results is the second object detection algorithm that is good at detecting small objects, obtain a weight value that is negatively correlated with the reference area as the weight value related to the nth object detection result among multiple object detection results.
[0093] Among them, obtaining a weight value that is positively correlated with the reference area can be: the larger the reference area, the larger the obtained weight value; the smaller the reference area, the smaller the obtained weight value; obtaining a weight value that is negatively correlated with the reference area can be: the larger the reference area, the smaller the obtained weight value; the smaller the reference area, the larger the obtained weight value.
[0094] After obtaining the confidence parameter value of each object detection result among multiple object detection results, the object detection result with a larger confidence parameter value can be determined from multiple object detection results as the reference detection result, and the initial inspection category included in the reference detection result can be used as the target category of the target object. Specifically, when the confidence parameter value of the reference detection result is greater than the preset parameter threshold, the initial inspection category included in the reference detection result can be used as the target category of the target object. Among them, the preset parameter threshold can be set according to actual application requirements, and the embodiments of the present disclosure do not limit this.
[0095] (2) Determine the target position of the target object based on the preliminary detection positions included in multiple target detection results.
[0096] As described above, in the embodiments of the present disclosure, for each target detection result among multiple target detection results, the preliminary detection position included therein can be represented by a detection box. Based on this, in the embodiments of the present disclosure, the detection boxes represented by the preliminary detection positions included in each target detection result among multiple target detection results can be merged to obtain a relatively larger detection box for representing the target position of the target object.
[0097] (3) Determine the fusion confidence related to the target object based on the algorithm confidence levels included in multiple target detection results.
[0098] In one example, the following data processing logic can be used to determine the fusion confidence related to the target object based on the algorithm confidence levels included in multiple target detection results:
[0099]
[0100] Wherein, CL' is used to represent the fusion confidence related to the target object; Pn is used to represent the weight value related to the nth target detection result among multiple target detection results; CLn is used to represent the algorithm confidence level related to the preliminary detection object in the nth target detection result among multiple target detection results; k is used to represent the total number of multiple target detection results. Here, multiple target detection results belong to the same set of target detection results.
[0101] Please combine Figure 2 , continuing with the foregoing example, for multiple target detection results (i.e., the first target detection result and the third target detection result) included in the first set of target detection results, the first preliminary detection object included in the first target detection result and the third preliminary detection object included in the third target detection result can be determined as the first target object. After determining the first preliminary detection object included in the first target detection result and the third preliminary detection object included in the third target detection result as the first target object, the first target category, the first target position, and the first fusion confidence related to the first target object can be determined based on the first preliminary detection category, the first preliminary detection position, and the first algorithm confidence level included in the first target detection result, and the third preliminary detection category, the third preliminary detection position, and the third algorithm confidence level included in the third target detection result.
[0102] Specifically, for the first target detection result among the multiple target detection results included in the first set of target detection results, that is, the first target detection result, the following data processing logic can be used to obtain the credible parameter values of the first target detection result:
[0103] M1 = P1 * CL1 = 0.3 * 60% = 0.18
[0104] Among them, M1 is used to represent the credibility parameter value of the first target detection result, specifically 0.18; P1 is used to represent the weight value related to the first target detection result, specifically 0.3; CL1 is used to represent the first algorithm confidence related to the first initial inspection object in the first target detection result, specifically 60%.
[0105] For the second target detection result included in the first group of target detection results, that is, the third target detection result, the credibility parameter value of the third target detection result can be obtained through the following data processing logic:
[0106] M2 = P2 * CL2 = 0.7 * 90% = 0.63
[0107] Among them, M2 is used to represent the credibility parameter value of the third target detection result, specifically 0.63; P2 is used to represent the weight value related to the third target detection result, specifically 0.7; CL2 is used to represent the third algorithm confidence related to the third initial inspection object in the third target detection result, specifically 90%.
[0108] After obtaining the credibility parameter values of the first target detection result and the third target detection result, the target detection result with the larger credibility parameter value can be determined from the first target detection result and the third target detection result, that is, the third target detection result, as the reference detection result, and when the credibility parameter value of the reference detection result (that is, 0.63) is greater than the preset parameter threshold (for example, 0.60), the initial inspection category included in the reference detection result (that is, the black safety helmet) is used as the first target category of the first target object.
[0109] After that, the detection frame represented by the first initial inspection position included in the first target detection result can be merged with the detection frame represented by the third initial inspection position included in the third target detection result to obtain a relatively larger detection frame for representing the first target position of the first target object, which can be specifically represented as (X11', Y11') & (X12', Y12').
[0110] Then, through the following data processing logic, the first fusion confidence related to the first target object can be obtained by determining based on the first algorithm confidence included in the first target detection result and the third algorithm confidence included in the third target detection result:
[0111]
[0112] Among them, CL1' is used to represent the first fusion confidence related to the first target object, specifically 81%; P1 is used to represent the weight value related to the first target detection result, specifically 0.3; CL1 is used to represent the first algorithm confidence in the first target detection result related to the first preliminary detection object, specifically 60%; P2 is used to represent the weight value related to the third target detection result, specifically 0.7; CL2 is used to represent the third algorithm confidence in the third target detection result related to the third preliminary detection object, specifically 90%.
[0113] In summary, after determining the first target category, the first target position of the first target object, and the first fusion confidence related to the first target object as the feature description information of the first target object, the feature description information of the first target object may include:
[0114] The first target category of the first target object: black safety helmet;
[0115] The first target position of the first target object: (X11', Y11') & (X12', Y12');
[0116] The first fusion confidence related to the first target object: 81%.
[0117] Similarly, for the multiple target detection results included in the second group of target detection results (that is, the second target detection result and the fourth target detection result), the second preliminary detection object included in the second target detection result and the fourth preliminary detection object included in the fourth target detection result can be determined as the second target object. After determining the second preliminary detection object included in the second target detection result and the fourth preliminary detection object included in the fourth target detection result as the second target object, the second target category, the second target position of the second target object, and the second fusion confidence related to the second target object can be determined based on the second preliminary detection category, the second preliminary detection position, and the second algorithm confidence included in the second target detection result, as well as the fourth preliminary detection category, the fourth preliminary detection position, and the fourth algorithm confidence included in the fourth target detection result.
[0118] Assume that in the foregoing example, after determining the second target category, the second target position of the second target object, and the second fusion confidence related to the second target object as the feature description information of the second target object, the feature description information of the second target object may include:
[0119] The second target category of the second target object: organism;
[0120] The second target position of the second target object: (X21', Y21') & (X22', Y22');
[0121] Second fusion confidence related to the second target object: 81%.
[0122] In the above manner, in the embodiments of the present disclosure, multiple target detection algorithms can be determined, and the multiple target detection algorithms are respectively used to detect the target image frame to obtain multiple relevant target detection results, and then the multiple target detection results are fused to determine the target object and the feature description information of the target object. That is to say, in the embodiments of the present disclosure, the target object and the feature description information of the target object are not determined only by using a single target detection algorithm, but the target object and the feature description information of the target object are determined by combining the target detection results of multiple target detection algorithms, so as to ensure the reliability of the determined target object and the feature description information of the target object.
[0123] Moreover, when fusing multiple target detection results to determine the target object and the feature description information of the target object, the initially detected objects included in the multiple target detection results are determined as the target object, and based on the initially detected category, initially detected position, and algorithm confidence included in the multiple target detection results, the target category, target position of the target object, and the fusion confidence related to the target object are determined, and then the target category, target position of the target object, and the fusion confidence related to the target object are determined as the feature description information of the target object. In this way, the richness of the feature description information can be further ensured. Furthermore, when determining the target category, target position of the target object, and the fusion confidence related to the target object based on the initially detected category, initially detected position, and algorithm confidence included in the multiple target detection results, the target category of the target object can be determined from the initially detected categories included in the multiple target detection results based on the algorithm confidence included in the multiple target detection results, the target position of the target object can also be determined based on the initially detected position included in the multiple target detection results, and the fusion confidence related to the target object can also be determined based on the algorithm confidence included in the multiple target detection results. In this process, no complex data processing logic is involved, so the execution efficiency of the security detection method can be further improved.
[0124] In some alternative embodiments, "using the multi-modal large model to obtain the security detection result for the organism based on the target image frame and the feature description information of at least one target object" in step S103 may include:
[0125] Step S103-1, determining multiple detection dimensions.
[0126] Among them, the multiple detection dimensions may include whether the safety protection equipment is worn at the correct position, or may also include whether the safety protection equipment is worn at the correct angle.
[0127] Step S103-2: Generate detection prompt information based on multiple detection dimensions.
[0128] Exemplarily, the detection prompt information may include: You are a security detection expert. Starting from the multiple detection dimensions provided for you, based on the target image frame and the feature description information of at least one target object, obtain a preliminary detection result for the organism.
[0129] Step S103-3: Input the detection prompt information into the multimodal large model so that the multimodal large model, according to the detection prompt information, based on the target image frame and the feature description information of at least one target object, obtains a preliminary detection result for the organism.
[0130] Among them, the preliminary detection result is used to characterize whether the organism correctly wears safety protection equipment. Here, the organism correctly wearing safety protection equipment can be that the organism wears the safety protection equipment at the correct position and at the correct angle; correspondingly, the organism not correctly wearing safety protection equipment can be that the organism does not wear the safety protection equipment at the correct position and / or does not wear it at the correct angle.
[0131] Step S103-4: Obtain a safety detection result based on the preliminary detection result.
[0132] Among them, the safety detection result may include at least partial identity information of the organism, specific violation behaviors, and violation evidence. Here, the violation evidence can be the target image frame.
[0133] In one example, the safety detection result can be obtained based on the preliminary detection result in the following way:
[0134] Determine the identity information of the organism;
[0135] Based on the preliminary detection result and the identity information of the organism, obtain a safety detection result.
[0136] Among them, the identity information of the organism may include name, gender, age, affiliated department, position, rank, number, etc.
[0137] In a specific example, face recognition technology can be used to identify the organism, extract the biological characteristics of the organism, and compare them with a pre-constructed identity database to determine the identity information of the organism. Among them, the biological characteristics can be facial contour, proportion of facial features, etc.
[0138] After determining the identity information of the organism, it is possible to determine whether it is necessary to supplement the preliminary test results based on the preliminary test results and the identity information of the organism. In the case where it is determined that the preliminary test results need to be supplemented, the preliminary test results are supplemented to obtain the safety test results; or, in the case where it is determined that the preliminary test results do not need to be supplemented, the preliminary test results are used as the safety test results. Based on this, in a specific example, the safety test results can be obtained based on the preliminary test results and the identity information of the organism in the following manner:
[0139] When the preliminary test results indicate that the organism is wearing safety protection equipment, determine the actual color of the safety protection equipment worn by the organism;
[0140] Based on the identity information of the organism, determine the target color that matches the organism;
[0141] When the actual color and the target color are different, supplement the preliminary test results to obtain the safety test results.
[0142] It can be understood that in the embodiments of the present disclosure, when the preliminary test results indicate that the organism is wearing safety protection equipment, this safety protection equipment actually also belongs to a target object. Therefore, the multimodal large model can determine the actual color of this safety protection equipment based on its understanding ability of data in other modalities (for example, the visual modality) and in combination with the feature description information of this safety protection equipment. Specifically, the multimodal large model can determine the actual color of this safety protection equipment based on its understanding ability of data in other modalities (for example, the visual modality) and in combination with the target category included in the feature description information of this safety protection equipment and the fusion confidence level related to this safety protection equipment.
[0143] In addition, taking the safety protection equipment as a safety helmet as an example, its target category can be:
[0144] White safety helmet, matching management personnel, supervision personnel, senior technical personnel, etc.;
[0145] Red safety helmet, matching safety supervisors, technical personnel, management personnel, etc.;
[0146] Blue safety helmet, matching technical personnel, patrol personnel, etc.;
[0147] Yellow safety helmet, matching ordinary workers, construction personnel, etc.;
[0148] Black safety helmet, matching personnel with special protection needs.
[0149] Based on this, in the embodiments of the present disclosure, after determining the actual color of the safety protection equipment worn by the organism and determining the target color matching the organism based on the identity information of the organism, when the actual color is different from the target color, it can be determined that the preliminary detection result needs to be supplemented, and the first supplementary information is used to supplement the preliminary detection result to obtain the safety detection result; or, when the actual color is the same as the target color, it can be determined that the preliminary detection result does not need to be supplemented, and the preliminary detection result is used as the safety detection result.
[0150] Exemplarily, when the actual color of the safety protection equipment worn by the organism is black and the target color matching the organism is blue, the first supplementary information can be: wearing the safety helmet incorrectly, it should be equipped with a blue safety helmet but is wrongly worn as a black safety helmet. Among them, the first supplementary information will be used as a specific violation behavior and included in the safety detection result.
[0151] Through the above method, in the embodiments of the present disclosure, multiple detection dimensions can be determined, and based on the multiple detection dimensions, detection prompt information is generated, and then the detection prompt information is input into the multimodal large model, so that the multimodal large model obtains a preliminary detection result for the organism according to the detection prompt information, based on the target image frame and the feature description information of at least one target object, so as to obtain a safety detection result based on the preliminary detection result. That is to say, in the embodiments of the present disclosure, the multimodal large model can execute the data processing logic according to the detection prompt information to clarify the data processing logic, thereby improving the accuracy of the preliminary detection result for the organism and further improving the accuracy of the safety detection result.
[0152] Moreover, when obtaining the safety detection result based on the preliminary detection result, the identity information of the organism can be determined, and the safety detection result can be obtained based on the preliminary detection result and the identity information of the organism. Specifically, when the preliminary detection result indicates that the organism wears safety protection equipment, the actual color of the safety protection equipment worn by the organism can be determined, and based on the identity information of the organism, the target color matching the organism can be determined, and then when the actual color is different from the target color, the preliminary detection result is supplemented to obtain the safety detection result. In this way, the detection surface of the safety detection method can be improved, so as to meet more diversified detection requirements.
[0153] In addition, in another specific example, the safety detection result can also be obtained based on the preliminary detection result and the identity information of the organism in the following way:
[0154] Determine the actual area where the organism is located;
[0155] Based on the identity information of the organism, determine the accessible area matching the organism;
[0156] When the actual location area of the organism is not within the allowed appearance area, supplement the preliminary detection result to obtain the safety detection result.
[0157] Among them, the actual location area of the organism can be determined based on the area label carried in the target image frame, and the embodiments of the present disclosure do not limit this.
[0158] In the embodiments of the present disclosure, the actual location area of the organism can be determined, and based on the identity information of the organism, the allowed appearance area matching the organism can be determined. Then, when the actual location area is not within the allowed appearance area, it can be determined that the preliminary detection result needs to be supplemented, and the second supplementary information is used to supplement the preliminary detection result to obtain the safety detection result; or, when the actual location area is within the allowed appearance area, the preliminary detection result is used as the safety detection result.
[0159] Exemplarily, when the actual location area of the organism is Area A, the allowed appearance area matching the organism is Area B, and Area A is not within Area B, the second supplementary information can be: Strayed into the non-allowed appearance area, please exit Area A. Among them, the second supplementary information will be used as a specific violation behavior and included in the safety detection result.
[0160] In an optional implementation manner, the safety detection method may further include:
[0161] When the safety detection result indicates that the organism does not wear the safety protection device in compliance, generate an alarm message based on the safety detection result;
[0162] Determine the target terminal related to the identity information of the organism;
[0163] Send the alarm message to the target terminal.
[0164] In an example, an alarm message including the safety detection result can be generated, or an alarm message including the identity information of the organism and the safety detection result can be generated.
[0165] In a specific example, the alarm message includes the safety detection result, and the target terminal related to the identity information of the organism can be the terminal device of the organism itself. In a specific example, the alarm message includes the identity information of the organism and the safety detection result, and the target terminal related to the identity information of the organism can be the terminal device of the superior person in charge of the organism.
[0166] In addition, in the embodiments of the present disclosure, the alarm message can be sent to the target terminal through various methods such as in-station pop-up windows, SMS push, emails, etc., and the embodiments of the present disclosure will not elaborate on this.
[0167] In the above manner, in the embodiments of the present disclosure, when the safety detection result indicates that the organism does not wear the safety protection device in compliance, an alarm message can be generated based on the safety detection result, and a target terminal related to the identity information of the organism can be determined, and then the alarm message can be sent to the target terminal. In this way, when the organism does not wear the safety protection device in compliance, the organism can be corrected in time to reduce the probability of the organism being injured.
[0168] In an alternative embodiment, the safety detection method may further include:
[0169] Determine a regulation knowledge base;
[0170] Determine the target regulation knowledge related to the safety detection result from the regulation knowledge base;
[0171] Using a large model, based on the safety detection result and the target regulation knowledge, obtain the reward and punishment information for the organism.
[0172] Wherein, the large model is the LLM, that is, it can be a pre-trained neural network model (for example, an autoregressive generation model with a Transformer architecture), which has general language knowledge, world knowledge, and professional knowledge in various fields (for example, professional knowledge in the field of computer technology), etc.
[0173] In addition, in the embodiments of the present disclosure, the reward and punishment information may include at least part of the identity information of the organism, specific violation behaviors, violation evidence, specific violated regulations, and reward and punishment measures. Among them, the violation evidence may be the target image frame.
[0174] In the above manner, in the embodiments of the present disclosure, a regulation knowledge base can be determined, the target regulation knowledge related to the safety detection result can be determined from the regulation knowledge base, and then using the large model, based on the safety detection result and the target regulation knowledge, the reward and punishment information for the organism can be obtained. That is to say, in the embodiments of the present disclosure, the Retrieval-augmented Generation (RAG) technology can be used to efficiently retrieve and intelligently analyze a large amount of and complex regulation knowledge articles to quickly locate the target regulation knowledge related to the safety detection result, and then use the large model to efficiently obtain the reward and punishment information for the organism based on the safety detection result and the target regulation knowledge.
[0175] Next, in combination with Figure 3 , a complete process of a safety detection method provided by the embodiments of the present disclosure will be described.
[0176] First, in the detection environment, the video data stream captured by the monitoring camera is processed to obtain target image frames. Specifically, the video data stream captured by the monitoring camera can be segmented to obtain multiple candidate image frames, and the multiple candidate image frames are de-duplicated to obtain multiple target image frames, and then the image data is backflow processed. For example, according to the arrangement order of the multiple target image frames in the video data stream, each target image frame in the multiple target image frames is sent to an electronic device, that is, sent to a server, a workbench, a mainframe computer or other similar computing devices as an intelligent cloud. Among them, when de-duplicating the multiple candidate image frames, an image similarity comparison algorithm can be used to analyze and compare each candidate image frame in the multiple candidate image frames, determine multiple duplicate image frames with relatively high similarity and multiple non-duplicate image frames other than the multiple duplicate image frames from the multiple candidate image frames, select one duplicate image frame from the multiple duplicate image frames as the target image frame, and use the multiple non-duplicate image frames as the target image frames to achieve de-duplication of the multiple candidate image frames.
[0177] Each time an electronic device receives a target image frame, it stores the target image frame and simultaneously enters the following process:
[0178] (1) Use the security detection module to determine multiple target detection algorithms, and respectively use the multiple target detection algorithms to detect the target image frame to obtain multiple relevant target detection results, and then fuse the multiple target detection results to determine the target object and the feature description information of the target object.
[0179] Among them, the multiple target detection results correspond one by one to the multiple target detection algorithms.
[0180] In addition, it should be noted that in the embodiments of the present disclosure, for the specific functions and examples of "determining multiple target detection algorithms", reference can be made to the relevant descriptions of the corresponding steps (that is, step S102-1) in the embodiments of the security detection method, which will not be elaborated here; for the specific functions and examples of "respectively using multiple target detection algorithms to detect the target image frame to obtain multiple relevant target detection results", reference can be made to the relevant descriptions of the corresponding steps (that is, step S102-2) in the embodiments of the security detection method, which will not be elaborated here; for the specific functions and examples of "fusing the multiple target detection results to determine the target object and the feature description information of the target object", reference can be made to the relevant descriptions of the corresponding steps (that is, step S102-3) in the embodiments of the security detection method, which will not be elaborated here.
[0181] (2) Use the event judgment module to determine multiple detection dimensions, generate detection prompt information based on the multiple detection dimensions, and then input the detection prompt information into the multimodal large model, so that the multimodal large model obtains a preliminary detection result for the organism according to the detection prompt information, based on the target image frame and the feature description information of at least one target object, and obtain a safety detection result based on the preliminary detection result.
[0182] Among them, the preliminary detection result is used to characterize whether the organism wears safety protection equipment correctly.
[0183] In addition, it should be noted that in the embodiments of the present disclosure, for the specific functions and examples of "determining multiple detection dimensions", reference can be made to the relevant descriptions of the corresponding steps (that is, step S103-1) in the embodiments of the safety detection method, which will not be elaborated here; for the specific functions and examples of "generating detection prompt information based on multiple detection dimensions", reference can be made to the relevant descriptions of the corresponding steps (that is, step S103-2) in the embodiments of the safety detection method, which will not be elaborated here; for the specific functions and examples of "inputting the detection prompt information into the multimodal large model, so that the multimodal large model obtains a preliminary detection result for the organism according to the detection prompt information, based on the target image frame and the feature description information of at least one target object", reference can be made to the relevant descriptions of the corresponding steps (that is, step S103-3) in the embodiments of the safety detection method, which will not be elaborated here; for the specific functions and examples of "obtaining a safety detection result based on the preliminary detection result", reference can be made to the relevant descriptions of the corresponding steps (that is, step S104-3) in the embodiments of the safety detection method, which will not be elaborated here.
[0184] After obtaining the safety detection result, the process can enter (3) and (4).
[0185] (3) Use the violation warning module to generate a warning message based on the safety detection result when the safety detection result indicates that the organism does not wear safety protection equipment in compliance, determine the target terminal related to the identity information of the organism, and then send the warning message to the target terminal.
[0186] It should be noted that in the embodiments of the present disclosure, for the specific functions and examples of process (3), reference can be made to the relevant descriptions of the corresponding steps in the embodiments of the safety detection method, which will not be elaborated here.
[0187] (4) Use the reward and punishment information acquisition module to determine the regulation knowledge base, determine the target regulation knowledge related to the safety detection result from the regulation knowledge base, and then use the large model to obtain the reward and punishment information for the organism based on the safety detection result and the target regulation knowledge.
[0188] Exemplarily, the reward and punishment information may include at least partial identity information of the organism, specific violations, evidence of violations, specific regulations violated, and reward and punishment measures.
[0189] Among them, at least partial identity information may include:
[0190] Name: Zhang San;
[0191] Number: A0001;
[0192] Rank: Head of AA Department.
[0193] Specific violations may include: wearing the wrong safety helmet. Instead of wearing a blue safety helmet as required, a black safety helmet was wrongly worn.
[0194] Specific regulations violated may include: violating the XXXX regulations.
[0195] Reward and punishment measures may include: imposing a fine of YY yuan.
[0196] In addition, it should be noted that in the embodiments of the present disclosure, for the specific functions and examples of process (4), reference may be made to the relevant descriptions of the corresponding steps in the embodiments of the security detection method, which will not be elaborated here.
[0197] The security detection method provided by the embodiments of the present disclosure at least has the following effects:
[0198] (1) Through the cooperation of a large-parameter model (e.g., a multimodal large model and a large model) and a small-parameter model (e.g., an image similarity comparison algorithm and multiple object detection algorithms), computing resources are saved. Moreover, by using the image similarity comparison algorithm, duplicate processing of multiple candidate image frames is achieved, and multiple target image frames are obtained. In this way, the number of calls to the multimodal large model and the large model can be reduced, that is, the data processing volume of the electronic device is reduced, thereby improving the execution efficiency of the security detection method. At the same time, the consumption of computing resources of the electronic device can be reduced.
[0199] (2) Through the interaction between the user and the large-parameter model, the purpose of intelligent operation is achieved.
[0200] Please refer to Figure 4 , which is a schematic diagram of an application scenario of a security detection method provided by the embodiments of the present disclosure.
[0201] The security detection method provided by the embodiments of the present disclosure is applied to an electronic device. Among them, the electronic device may be a server, a workbench, a mainframe computer, or other similar computing devices.
[0202] Here, the electronic device is used for:
[0203] Obtain target image frames;
[0204] Detect the target image frame to determine at least one target object and the feature description information of at least one target object.
[0205] When it is determined that there is a living organism among at least one target object, use a multimodal large model to obtain a safety detection result for the living organism based on the target image frame and the feature description information of at least one target object; wherein, the safety detection result is used to characterize whether the living organism wears safety protection equipment compliantly.
[0206] It should be noted that in the embodiments of the present disclosure, Figure 4 The schematic diagram of the application scenario shown is only illustrative and not restrictive. Those skilled in the art can make various obvious changes and / or substitutions based on Figure 4 the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.
[0207] To better implement the foregoing method, the embodiments of the present disclosure also provide a safety detection device, which can be integrated into an electronic device. Among them, the electronic device can be a server, a workbench, a mainframe computer or other similar computing devices. Hereinafter, a safety detection device 500 provided by the embodiments of the disclosure will be described with reference to Figure 5 the schematic structural block diagram shown.
[0208] The safety detection device 500 includes:
[0209] An image acquisition unit 501, configured to acquire a target image frame;
[0210] An image detection unit 502, configured to detect the target image frame to determine at least one target object and the feature description information of at least one target object;
[0211] A result acquisition unit 503, configured to, when it is determined that there is a living organism among at least one target object, use a multimodal large model to obtain a safety detection result for the living organism based on the target image frame and the feature description information of at least one target object; wherein, the safety detection result is used to characterize whether the living organism wears safety protection equipment compliantly.
[0212] In some alternative embodiments, the result acquisition unit 503 is configured to:
[0213] Determine multiple detection dimensions;
[0214] Generate detection prompt information based on the multiple detection dimensions;
[0215] Input the detection prompt information into the multimodal large model, so that the multimodal large model, according to the detection prompt information, based on the target image frame and the feature description information of at least one target object, obtains a preliminary detection result for the organism; wherein, the preliminary detection result is used to characterize whether the organism is wearing safety protection equipment correctly;
[0216] Based on the preliminary detection result, obtain the safety detection result.
[0217] In some optional embodiments, the result acquisition unit 503 is configured to:
[0218] Determine the identity information of the organism;
[0219] Based on the preliminary detection result and the identity information of the organism, obtain the safety detection result.
[0220] In some optional embodiments, the result acquisition unit 503 is configured to:
[0221] When the preliminary detection result indicates that the organism is wearing safety protection equipment, determine the actual color of the safety protection equipment worn by the organism;
[0222] Based on the identity information of the organism, determine the target color that matches the organism;
[0223] When the actual color is different from the target color, supplement the preliminary detection result to obtain the safety detection result.
[0224] In some optional embodiments, the safety detection device 500 further includes a reward and punishment information acquisition unit, configured to:
[0225] When the safety detection result indicates that the organism is not wearing the safety protection equipment in compliance, generate an alarm information based on the safety detection result;
[0226] Determine the target terminal related to the identity information of the organism;
[0227] Send the alarm information to the target terminal.
[0228] In some optional embodiments, the safety detection device 500 further includes an alarm information sending unit, configured to:
[0229] Determine the regulation knowledge base;
[0230] Determine the target regulation knowledge related to the safety detection result from the regulation knowledge base;
[0231] Using the large model, based on the safety detection result and the target regulation knowledge, obtain the reward and punishment information for the organism.
[0232] In some alternative embodiments, the image detection unit 502 is configured to:
[0233] Determine a plurality of target detection algorithms;
[0234] Respectively use the plurality of target detection algorithms to detect the target image frame, and obtain a plurality of target detection results with relevance; wherein, the plurality of target detection results correspond one-to-one to the plurality of target detection algorithms;
[0235] Fuse the plurality of target detection results to determine the target object and the feature description information of the target object.
[0236] In some alternative embodiments, the image detection unit 502 is configured to:
[0237] Determine the preliminary detection object included in the plurality of target detection results as the target object;
[0238] Based on the preliminary detection category, preliminary detection position, and algorithm confidence included in the plurality of target detection results, determine the target category, target position of the target object, and the fusion confidence related to the target object;
[0239] Determine the target category, target position of the target object, and the fusion confidence related to the target object as the feature description information of the target object.
[0240] In some alternative embodiments, the image detection unit 502 is used for at least one of the following:
[0241] Based on the algorithm confidence included in the plurality of target detection results, determine the target category of the target object from the preliminary detection categories included in the plurality of target detection results;
[0242] Based on the preliminary detection position included in the plurality of target detection results, determine the target position of the target object;
[0243] Based on the algorithm confidence included in the plurality of target detection results, determine the fusion confidence related to the target object.
[0244] In the embodiments of the present disclosure, for the specific functions and examples of each unit in the security detection device 500, reference may be made to the relevant descriptions of the corresponding steps in the embodiments of the security detection method, which will not be elaborated here.
[0245] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0246] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0247] Figure 6FIG. 0 shows a schematic structural block diagram of an exemplary electronic device 600 that can be used to implement embodiments of the present disclosure. The electronic device 600 is intended to represent various forms of digital computers, such as in-vehicle computing devices, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 600 can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0248] As Figure 6 shown, the electronic device 600 includes a computing unit 601 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0249] A plurality of components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of renderers, speakers, etc.; a storage unit 608, such as a disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0250] The computing unit 601 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the security detection method. For example, in some embodiments, the security detection method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the security detection method described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured as the security detection method by any other suitable means (e.g., by means of firmware).
[0251] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0252] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data optimization devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0253] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0254] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a rendering device for rendering information to the user (e.g., a cathode ray tube (CRT) renderer or a liquid crystal display (LCD) renderer); and a keyboard and a pointing device (e.g., a mouse or a trackball), through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including, acoustic input, voice input, or tactile input).
[0255] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: a Local Area Network (LAN), a Wide Area Network (WAN), and the Internet.
[0256] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs that run on the respective computers and have a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server that incorporates blockchain.
[0257] Embodiments of the present disclosure also provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute a security detection method.
[0258] Embodiments of the present disclosure also provide a computer program product, including a computer program that implements a security detection method when executed by a processor.
[0259] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and this is not limited herein. In addition, in the present disclosure, relational terms such as "first", "second", "third", etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. In addition, in the present disclosure, "a plurality of" can be understood to mean at least two.
[0260] The foregoing specific embodiments do not limit the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A security detection method, comprising: Obtaining a target image frame; Detecting the target image frame to determine at least one target object and feature description information of the at least one target object; When it is determined that there is a living body among the at least one target object, using a multimodal large model, based on the target image frame and the feature description information of the at least one target object, obtaining a security detection result for the living body; wherein, the security detection result is used to characterize whether the living body wears safety protection equipment in compliance; 2. The method according to claim 1, wherein, The step of obtaining a security detection result for the living body by using the multimodal large model based on the target image frame and the feature description information of the at least one target object includes: Determining a plurality of detection dimensions; Generating detection prompt information based on the plurality of detection dimensions; Inputting the detection prompt information into the multimodal large model so that the multimodal large model obtains a preliminary detection result for the living body according to the detection prompt information, based on the target image frame and the feature description information of the at least one target object; wherein, the preliminary detection result is used to characterize whether the living body wears the safety protection equipment correctly; Obtaining the security detection result based on the preliminary detection result.
3. The method according to claim 2, wherein The step of obtaining the security detection result based on the preliminary detection result includes: Determining the identity information of the living body; Obtaining the security detection result based on the preliminary detection result and the identity information of the living body.
4. The method according to claim 3, wherein, The step of obtaining the security detection result based on the preliminary detection result and the identity information of the living body includes: When the preliminary detection result indicates that the living body wears the safety protection equipment, determining the actual color of the safety protection equipment worn by the living body; Determining a target color matching the living body based on the identity information of the living body; When the actual color is different from the target color, supplementing the preliminary detection result to obtain the security detection result.
5. The method according to claim 3, further comprising: When the security detection result indicates that the living body does not wear the safety protection equipment in compliance, generating an alarm information based on the security detection result; Determining a target terminal related to the identity information of the living body; Sending the alarm information to the target terminal.
6. The method according to any one of claims 1 to 5, further comprising: Determining a regulation knowledge base; Determining target regulation knowledge related to the security detection result from the regulation knowledge base; Using a large model to obtain reward and punishment information for the living body based on the security detection result and the target regulation knowledge.
7. According to the method of any one of claims 1 to 5, wherein The step of detecting the target image frame to determine at least one target object and feature description information of the at least one target object includes: Determining a plurality of target detection algorithms; Use the multiple target detection algorithms respectively to detect the target image frame, and obtain multiple target detection results with relevance; wherein, the multiple target detection results correspond one-to-one to the multiple target detection algorithms; Fuse the multiple target detection results to determine the target object and the feature description information of the target object.
8. The method according to claim 7, wherein The fusing the multiple target detection results to determine the target object and the feature description information of the target object includes: Determine the initial detection object included in the multiple target detection results as the target object; Based on the initial detection category, initial detection position and algorithm confidence included in the multiple target detection results, determine the target category, target position of the target object, and the fusion confidence related to the target object; Determine the target category, target position of the target object, and the fusion confidence related to the target object as the feature description information of the target object.
9. The method according to claim 8, wherein The determining the target category, target position of the target object, and the fusion confidence related to the target object based on the initial detection category, initial detection position and algorithm confidence included in the multiple target detection results includes at least one of the following: Based on the algorithm confidence included in the multiple target detection results, determine the target category of the target object from the initial detection categories included in the multiple target detection results; Based on the initial detection position included in the multiple target detection results, determine the target position of the target object; Based on the algorithm confidence included in the multiple target detection results, determine the fusion confidence related to the target object.
10. A safety detection device, comprising: An image acquisition unit for acquiring a target image frame; An image detection unit for detecting the target image frame to determine at least one target object and the feature description information of the at least one target object; A result acquisition unit for, when it is determined that there is a living body among the at least one target object, using a multimodal large model, based on the target image frame and the feature description information of the at least one target object, obtaining a safety detection result for the living body; wherein, the safety detection result is used to characterize whether the living body wears safety protection equipment compliantly.
11. An electronic device, comprising: At least one processor; A memory communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 9.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.
13. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 9.