Rule violation identification method and device, electronic equipment, storage medium and program product
By using a pre-trained violation identification model and a multimodal large model to identify and re-judge violations in monitoring video images at power operation sites, the problems of missed and false detections in violation identification at power operation sites have been solved, and the accuracy and efficiency of identification have been improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies for identifying violations at power operation sites suffer from high rates of missed and false detections, as well as low efficiency in manual verification.
A pre-trained violation identification model is used to perform preliminary identification of surveillance video images, and then a target detection model and a multimodal large model are used for further judgment. The system is also reviewed using preset safety regulations to reduce false alarms and missed alarms.
It improved the accuracy of violation identification and the efficiency of review, reduced the false alarm rate, reduced the workload of manual review, and provided rectification suggestions.
Smart Images

Figure CN121767916A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of power operation safety management technology, specifically to a method, device, electronic equipment, storage medium, and program product for identifying violations. Background Technology
[0002] With increasingly stringent safety requirements in the power industry, the safety management of high-risk operations in power plants, such as working at heights, hot work, and hoisting operations, faces severe challenges. Currently, many systems employ small-scale model algorithms for real-time safety monitoring and violation identification of on-site personnel, environment, and equipment, sending alarms to the system for manual review. This approach not only suffers from high rates of missed and false alarms but also suffers from low efficiency in manual review. Summary of the Invention
[0003] In order to overcome the problems existing in the related technologies, this disclosure provides a method, device, electronic device, storage medium and program product for identifying violations.
[0004] In a first aspect, this disclosure provides a method for identifying violations, the method comprising: acquiring a monitoring video image of a power operation site; using a pre-trained violation identification model to identify violations in the monitoring video image to obtain a first violation identification result, wherein the violation identification model is used to identify violations of a preset violation type; using a pre-trained target detection model to detect safety targets in the monitoring video image to obtain a target detection result; acquiring a preset safety regulation corresponding to the preset violation type; and, based on the preset safety regulation and the target detection result, using a pre-trained multimodal large model to re-evaluate the first violation identification result to obtain a second violation identification result.
[0005] Secondly, this disclosure provides a violation identification device, comprising: an image acquisition module, a violation identification module, a target detection module, a regulation acquisition module, and a violation review module. The image acquisition module is used to acquire monitoring video images of a power operation site; the violation identification module is used to identify violations in the monitoring video images using a pre-trained violation identification model to obtain a first violation identification result, wherein the violation identification model is used to identify violations of a preset violation type; the target detection module is used to perform safety target detection in the monitoring video images using a pre-trained target detection model to obtain a target detection result; the regulation acquisition module is used to acquire preset safety regulations corresponding to the preset violation type; and the violation review module is used to review the first violation identification result based on the preset safety regulations and the target detection result, and using a pre-trained multimodal large model to obtain a second violation identification result.
[0006] Thirdly, this disclosure provides an electronic device comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to implement the steps of the first aspect.
[0007] Fourthly, this disclosure provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method described in the first aspect.
[0008] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.
[0009] This disclosure acquires monitoring video images of power operation sites; uses a pre-trained violation identification model to identify violations in the monitoring video images, obtaining a first violation identification result, wherein the violation identification model is used to identify violations of preset violation types; uses a pre-trained target detection model to detect safety targets in the monitoring video images, obtaining target detection results; acquires preset safety regulations corresponding to the preset violation types; and, based on the preset safety regulations and the target detection results, uses a pre-trained multimodal large model to re-judge the first violation identification result, obtaining a second violation identification result. In other words, for monitoring video images of power operation sites, a preliminary violation identification is performed using a violation identification model, and then, combined with the target detection results and preset safety regulations, a multimodal large model is used for re-judgment, reducing false alarms and missed alarms. Furthermore, the re-judgment of violation identification by the multimodal large model significantly improves the efficiency and accuracy of re-judgment compared to manual re-judgment.
[0010] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0011] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the following detailed description to explain the present disclosure, but do not constitute a limitation thereof. In the drawings: Figure 1 A flowchart illustrating a traffic violation identification method provided in an embodiment of this application is shown.
[0012] Figure 2 It shows Figure 1 A flowchart illustrating a sub-step of step S150 in one embodiment.
[0013] Figure 3A flowchart illustrating a violation identification method provided in another embodiment of this application is shown.
[0014] Figure 4 A schematic diagram of the architecture of a traffic violation identification system provided in an embodiment of this application is shown.
[0015] Figure 5 This is a block diagram of a traffic violation identification device according to an embodiment of this application.
[0016] Figure 6 This is a block diagram of an electronic device used to perform a violation identification method according to an embodiment of this application.
[0017] Figure 7 This is a storage unit in this application embodiment for storing or carrying program code that implements the violation identification method according to this application embodiment. Detailed Implementation
[0018] The specific embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit this disclosure.
[0019] It should be noted that all actions involving the acquisition of signals, information, or data in this disclosure are carried out in compliance with the relevant data protection laws and policies of the country where the location is situated, and with authorization from the owner of the relevant device.
[0020] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0021] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0022] The inventors have proposed a method, device, electronic device, storage medium, and program product for identifying traffic violations. The method for identifying traffic violations provided in the embodiments of this application will be described in detail below.
[0023] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a traffic violation identification method according to an embodiment of this application. The following will be combined with... Figure 1The method for identifying traffic violations provided in this application is described in detail. This method may include the following steps: Step S110: Acquire monitoring video images of the power operation site.
[0024] In this embodiment, to ensure the safety of workers at power work sites, cameras are typically installed at the sites to monitor their work behavior in real time. Based on this, an image frame can be extracted from the real-time monitoring video stream at preset intervals as the monitoring video image for subsequent violation identification. In other words, this application can perform safety monitoring and violation identification of workers at power work sites at preset intervals. The preset interval can be a pre-set value, such as 5 minutes, and can be adjusted according to different actual needs.
[0025] Step S120: Using a pre-trained violation recognition model, identify violations in the surveillance video image to obtain a first violation recognition result. The violation recognition model is used to identify violations of a preset violation type.
[0026] Optionally, the preset violation type may include at least one of the following: not wearing a seatbelt, climbing over a safety fence, not wearing a safety helmet, not wearing work clothes, sleeping on duty, making a phone call, and smoking. The violation recognition model may be obtained through supervised training of a YOLOv8s base model. The specific training process of the violation recognition model may include: acquiring a first sample image set, wherein each first sample image in the first sample image set carries a corresponding preset violation type label, and the first sample image set includes at least 5000 first sample images; inputting the first sample images into an initial recognition model built based on the YOLOv8s base model to obtain predicted violation type labels; calculating a prediction loss value based on the preset violation type label and the predicted violation type label of each first sample image; and finally, iteratively training the initial recognition model based on the prediction loss value until a first preset condition is met, obtaining the trained initial recognition model as the aforementioned violation recognition model.
[0027] The first preset condition can be: the predicted loss value is less than a preset value, the predicted loss value no longer changes, or the number of training iterations reaches a preset number, etc. It is understood that after iteratively training the initial recognition model on the first sample image set for multiple training cycles—each training cycle including multiple iterations—the parameters in the initial recognition model are continuously optimized, causing the predicted loss value to decrease until it reaches a fixed value or is less than the preset value. At this point, it indicates that the initial recognition model has converged. Alternatively, it can be determined that the initial recognition model has converged after the number of training iterations reaches a preset number. In this case, the trained initial recognition model can be used as the aforementioned violation recognition model. The preset value and the preset number of iterations are pre-set and can be adjusted according to different application scenarios; this embodiment does not impose any restrictions on this.
[0028] In other words, after acquiring the surveillance video image, it is first input into a pre-trained violation recognition model for preliminary identification of the violation, resulting in a first violation recognition result. Taking the preset violation type of not wearing a seatbelt as an example, the first violation recognition result would be either "wearing a seatbelt" or "not wearing a seatbelt." The first violation recognition result can be output as a first violation recognition result image, which includes a violation type label for the identified first violation and the coordinate information of the first violation in the surveillance video image. This coordinate information can be represented by a bounding box in the surveillance video image. The first violation is matched with violations of the preset violation types that the violation recognition model can identify.
[0029] For example, taking the preset violation type as whether or not a seatbelt is worn, the violation type label for this preset violation type can include a violation label for "not wearing a seatbelt" and a non-violation label for "wearing a seatbelt". If the first violation identification result is that there is a situation of not wearing a seatbelt in the surveillance video image, the image will be marked with a "not wearing a seatbelt" violation at the location where a seatbelt should be worn; if the first violation identification result is that there is no situation of not wearing a seatbelt in the surveillance video image, the image will be marked with a "wearing a seatbelt" non-violation label at the location where a seatbelt should be worn.
[0030] Step S130: Using a pre-trained target detection model, perform security target detection on the surveillance video image to obtain the target detection result.
[0031] Optionally, safety targets can be understood as common safety management-related targets at power operation sites, and may include at least workers, safety fences, scaffolding, reflective clothing, safety warning signs, hole covers, and crane booms. Similarly, the target detection model can be obtained through supervised training of an initial detection model built on the YOLOv8s basic model. The specific training process of the target detection model may include: acquiring a second sample image set, which includes multiple second sample images, and at least 5000 second sample images, each labeled with a preset bounding box and a preset target name for the safety target; inputting the second sample images into the initial detection model to obtain the target result image output by the initial detection model, which contains the predicted bounding box and predicted target name of the detected safety target; determining the target detection loss value based on the degree of difference between the predicted bounding box and the preset bounding box, and the degree of difference between the preset target name and the predicted target name; finally, iteratively training the initial detection model based on the target detection loss value until the second preset condition is met, obtaining the trained initial detection model as the aforementioned target detection model.
[0032] The second preset condition can be: the first preset condition can be: the target detection loss value is less than a preset value, the target detection loss value no longer changes, or the number of training iterations reaches a preset number, etc. It is understood that after iteratively training the initial detection model on the second sample image set for multiple training cycles, where each training cycle includes multiple iterations, the parameters in the initial detection model are continuously optimized, making the target detection loss value smaller and smaller until it becomes a fixed value or less than the preset value. At this point, it indicates that the initial detection model has converged. Alternatively, it can be determined that the initial detection model has converged after the number of training iterations reaches a preset number. In this case, the trained initial detection model can be used as the target detection model. The preset value and the preset number of iterations are pre-set and can be adjusted according to different application scenarios; this embodiment does not impose any restrictions on this.
[0033] To improve the accuracy of violation identification, a target detection model can be used to supplement the identification of security targets in surveillance video images, resulting in corresponding target detection results. These results can include a target detection image, which is labeled with the name of the identified security target and its coordinates within the surveillance video image.
[0034] In some implementations, the violation recognition image output by the violation recognition model and the detection result image output by the target detection model can be merged, that is, the information of the violation recognition result and the target detection result can be integrated into a single image to facilitate subsequent multimodal large model to re-judge violations.
[0035] Optionally, based on the coordinate information of the first violation in the first violation recognition result image, the violation type label of the first violation can be annotated to the corresponding position in the target result detection image to obtain a first target image. It should be noted that the coordinate information of the first violation will also be annotated in the first target image.
[0036] Optionally, the target name of the safety target can be annotated in the first violation recognition result image based on the coordinate information of the safety target in the target detection result image, thus obtaining the first target image. It should be noted that the coordinate information of the safety target will also be annotated in the first target image accordingly.
[0037] Optionally, considering that the target detection results often contain a large number of security targets, that is, the target detection result image contains a large amount of data, in order to improve the efficiency of image merging, the violation type label of the first violation behavior can be marked to the corresponding position in the target result detection image based on the coordinate information of the first violation behavior in the first violation recognition result image, so as to obtain the first target image.
[0038] Optionally, it can also be determined whether the number of safe targets in the target detection result image is greater than the number of first violations in the first violation recognition result image. If it is greater, it indicates that there is more data in the target detection result image. To improve image merging efficiency, the violation type label of the first violation can be marked at the corresponding position in the target detection result image based on the coordinate information of the first violation in the first violation recognition result image, thus obtaining the first target image. If it is less than or equal to, it indicates that there is more data in the first violation recognition result image. To improve image merging efficiency, the target name of the safe target can be marked in the first violation recognition result image based on the coordinate information of the safe target in the target detection result image, thus obtaining the first target image.
[0039] Step S140: Obtain the preset safety regulations corresponding to the preset violation type.
[0040] In this embodiment, a safety regulation knowledge base can be pre-set, containing various preset safety regulations, each of which corresponds to at least one preset violation type. Based on this, the preset safety regulations corresponding to the preset violation types can be obtained.
[0041] Step S150: Based on the preset safety regulations and the target detection results, and using a pre-trained multimodal large model to re-judge the first violation identification result, a second violation identification result is obtained.
[0042] In some implementations, please refer to Figure 2 Step S150 may include the contents of steps S151 to S153: Step S151: Based on the preset security regulations and the target detection results, generate target question text, which is used to ask whether the security target complies with the preset security regulations.
[0043] Specifically, based on preset safety regulations and target detection results, target problem text is generated according to a preset text template. For example, the preset text template could be: "Is the violation label marked in the image [Preset Violation Type] accurate? Does the [Safety Target] comply with the provisions of the [Preset Safety Regulations]?"; the preset safety regulations could be: "The safety helmets, safety belts, safety ropes, climbing self-locking devices, fall arrestors, etc., equipped by personnel working at heights should be inspected and qualified and meet the requirements. During operations, safety belts and safety ropes must be secured to a sturdy object to prevent them from falling off. Safety belts should be used in a high-hanging, low-use manner, and double-hook safety belts and fall arrestors should be used correctly"; the safety target could be "workers and safety belts"; and the "Preset Violation Type" could be "whether a safety helmet is worn." At this point, the generated target question text could be: "Is the violation sign in the image indicating whether a safety helmet is being worn accurate? Do the workers and safety belts comply with the regulations regarding the safety helmets, safety belts, safety ropes, climbing self-locking devices, and fall arresters used by workers at height? Safety belts and safety ropes must be secured to a sturdy object during operations to prevent them from falling off. Safety belts should be used in a high-hanging, low-use manner, and double-hook safety belts and fall arresters should be used correctly."
[0044] Step S152: Input the target question text and the first target image into the multimodal large model to obtain the output target answer text.
[0045] Step S153: Based on the target answer text, the first violation identification result is re-evaluated to obtain the second violation identification result.
[0046] Furthermore, the generated target question text and the first target image are input into the multimodal large model for review of preset violation types and supplementary identification of other violation types, resulting in the target answer text output by the multimodal large model. Optionally, the target answer text may include review identification text and rectification suggestion text, that is, the multimodal large model can output the corresponding rectification suggestion text according to preset safety regulations.
[0047] For example, the target answer text could be: "The person in the picture is wearing a safety helmet while working at height, so the first violation identification result is inaccurate. However, the person is not wearing a safety belt as required by regulations, which does not comply with the preset safety regulations. Safety education and training must be strengthened immediately, the requirement to wear safety belts must be strictly enforced, and qualified safety equipment must be provided and on-site supervision must be strengthened to ensure the safety of personnel working at height."
[0048] In some implementations, the second violation identification result includes a second violation identification result image, in which a violation type label of the identified second violation is marked, and the coordinate information of the second violation in the surveillance video image.
[0049] Based on this, firstly, the violation identifier in the violation type label of the first violation can be labeled to the corresponding position in the target result detection image according to the first preset labeling method, thus obtaining the aforementioned first target image. If the second violation does not include the first violation, then the violation identifier in the first target image is replaced with the non-violation identifier in the violation type label of the first violation according to the second preset labeling method, thus obtaining the second target image. The second preset labeling method is different from the first preset labeling method. The first preset labeling method can be understood as the labeling method for violations, while the second preset labeling method can be understood as the labeling method for violations. The purpose of the two different labeling methods is to enable safety supervisors to intuitively and quickly identify which violations exist in the image simply by viewing it.
[0050] For example, the first preset annotation method can be to annotate with red font and red annotation box, while the second preset annotation method can be to annotate with green font and green annotation box.
[0051] Furthermore, based on the coordinate information of the second violation in the second violation recognition result image, and according to the first preset annotation method, the violation mark in the violation type label of the second violation is annotated to the corresponding position in the second target image to obtain the third target image. That is, for violation types that the violation recognition model misses, supplementary markings will be made accordingly. For example, if the second violation is not wearing a helmet, then the violation mark "not wearing a helmet" can be annotated in red in the second target image to obtain the third target image.
[0052] Finally, the third target image and security alert information can be output, the security alert information being used to indicate that the security target has committed the second violation.
[0053] Optionally, the third target image and safety warning information can be directly output to the electronic device used by the safety supervisor. This allows the safety supervisor to intuitively and quickly identify any violations at the work site based on the output third target image and safety warning information, and promptly notify the corresponding workers to rectify the issues, thereby ensuring the safety of workers during electrical work.
[0054] Optionally, loudspeakers can be installed at the power work site. In addition to outputting the third-party target image and safety reminder information to the electronic equipment used by safety supervisors, the safety reminder information can also be output in the form of voice broadcasts through loudspeakers to promptly remind workers at the power work site of any violations and to rectify them as soon as possible, thereby ensuring the personal safety of the workers.
[0055] Optionally, each worker can carry a corresponding electronic device while performing electrical work. Based on this, safety alerts can be sent to the corresponding worker's electronic device, allowing for timely and targeted prompting of corrective action for workers exhibiting violations, thus ensuring worker safety. Of course, safety alerts and third-party target images can also be sent to the electronic device used by safety supervisors for storage.
[0056] In this embodiment, a collaborative approach using large and small models is employed to identify violations. First, a small model quickly identifies violations at the power work site. Then, a large multimodal model, based on the identification results, queries relevant safety regulations to make a professional and accurate re-evaluation. Simultaneously, the results of target detection are used for supplementary identification. This violation re-evaluation method reduces the false identification rate of violations and significantly improves re-evaluation efficiency compared to manual re-evaluation.
[0057] Please refer to Figure 3 , Figure 3 This is a flowchart illustrating a traffic violation identification method according to another embodiment of this application. The following will be combined with... Figure 3 The method for identifying traffic violations provided in this application is described in detail. This method may include the following steps: Step S210: Acquire monitoring video images of the power operation site.
[0058] In this embodiment, the specific implementation of step S210 can be found in the content of the foregoing embodiments, and will not be repeated here.
[0059] Step S220: Using multiple pre-trained violation recognition models, the monitoring video images are used to identify violations, resulting in multiple first violation recognition results. Different violation recognition models are used to identify different preset violation types.
[0060] In this embodiment, for different preset violation types, corresponding violation recognition models can be trained specifically, thereby enabling more targeted and accurate identification of violations in surveillance video images. For example, if the preset violation types include seven categories such as not wearing a seatbelt, climbing over a safety fence, not wearing a safety helmet, not wearing work clothes, employees sleeping on duty, employees making phone calls, and smoking, then seven violation recognition models need to be pre-trained, with each of the seven violation recognition models corresponding one-to-one with these seven preset violation types.
[0061] Based on this, such as Figure 4 As shown, the surveillance video images are input into multiple pre-trained violation recognition models, resulting in a first violation recognition result output by each model. In other words, multiple first violation recognition result images are obtained, each labeled with the identified first violation and its coordinates within the surveillance video image. It should be noted that the first violation in the first violation recognition results output by different violation recognition models is actually different.
[0062] The training method of each violation recognition model and the specific implementation method of each violation recognition model to identify the corresponding violation behavior in the monitoring and recognition images can be found in the above embodiments, and will not be repeated here.
[0063] Optionally, such as Figure 4 Multiple violation recognition result images output by multiple violation recognition models can be stored in a database.
[0064] Step S230: Using a pre-trained target detection model, perform security target detection on the surveillance video image to obtain the target detection result.
[0065] In this embodiment, the specific implementation of step S230 can be found in the content of the foregoing embodiments, and will not be repeated here.
[0066] In some implementations, considering that target detection results often contain multiple security targets, while each first violation identification result image contains only one type of first violation identification behavior, to improve image merging efficiency, the violation type label of each first violation behavior can be annotated to the corresponding position in the first target image based on the coordinate information of the first violation behavior in each first violation identification result image, thus obtaining the first target image.
[0067] Optionally, such as Figure 4The target detection result image output by the target detection model, as well as the merged first target image, can be stored in the database. Specifically, the fields stored in the database mainly include: the image identifier (Identity document, ID) of the video surveillance image, the violation type label, the name of the security target, and the first target image. The first target image is saved in base64 format.
[0068] Step S240: Obtain the preset safety regulations corresponding to each preset violation type, thus obtaining multiple preset safety regulations.
[0069] Understandably, multiple violation recognition models can be used to identify violations of various preset violation types. Similarly, the safety regulations knowledge base also contains multiple preset safety regulations, each corresponding to at least one preset violation type. Likewise, the corresponding preset safety regulations fields and their contents retrieved from the safety regulations knowledge base will be added to the image ID data of the corresponding video surveillance image in the database.
[0070] Step S250: Based on multiple preset safety regulations and the target detection results, and using a pre-trained multimodal large model to re-judge multiple first violation identification results, multiple second violation identification results are obtained.
[0071] Specifically, based on multiple preset security regulations and the target detection results, a target question text is generated, which is used to ask whether the security target complies with multiple preset security regulations; the target question text and the first target image are input into the multimodal large model to obtain the output target answer text; based on the target answer text, multiple first violation identification results are re-evaluated to obtain multiple second violation identification results.
[0072] Similarly, if the first violation is not included in the multiple second violations, then according to the second preset annotation method, the violation identifier in the first target image is replaced with the non-violation identifier in the violation type label of the first violation, to obtain a second target image. The second preset annotation method is different from the first preset annotation method. Based on the coordinate information of the second violations in the multiple second violation recognition result images, and according to the first preset annotation method, the violation identifiers in the violation type labels of the multiple second violations are all labeled to the corresponding positions in the second target image to obtain a third target image. The third target image and safety prompt information are then output.
[0073] The specific implementation details of step S250 can be found in the aforementioned embodiments and will not be repeated here.
[0074] For example, the second violation identification result of "inaccurate", the identification result of other violation types of "does not comply with safety regulations", and the rectification opinions can all be written into the data corresponding to the image id 1. Specifically, a single data entry for a surveillance video image in the final database can contain the fields and data values in Table 1 below:
[0075] In some implementations, at specified intervals, all monitoring video images with inaccurate re-judgment results can be retrieved from the database, and a third sample image set for each violation recognition model can be constructed based on all monitoring video images with inaccurate re-judgment results. Based on the third sample image set, each violation recognition model can be iteratively optimized, thereby improving the violation recognition accuracy of each violation recognition model.
[0076] In other implementations, the number of surveillance video images whose identification results of each violation identification model in the database are incorrect is counted. If the number of images is greater than a preset threshold, all surveillance video images whose identification results of the violation identification model are incorrect are obtained. A fourth sample image set for the corresponding violation identification model is constructed based on the obtained surveillance video images. Based on the fourth sample image set, the corresponding violation identification model is iteratively optimized. This is to specifically optimize the violation identification model with low violation identification accuracy, thereby improving the violation identification accuracy of the violation identification model.
[0077] In this embodiment, considering that small models have fewer parameters and higher recognition efficiency, but lower accuracy and a higher risk of false alarms, while multimodal large models have more parameters and lower real-time recognition, but stronger text understanding and a better overall image understanding, a collaborative approach of small and large models is adopted to identify violations. The small model first quickly captures the violations on-site, and then the large model uses the recognition results to query relevant safety regulations for a professional and accurate re-judgment, while supplementing the target detection results with additional recognition. This violation re-judgment reduces the false alarm rate of violation alarms and effectively reduces the workload of manual re-judgment; simultaneously, the large model's supplementary image recognition reduces the false negative rate of violations to some extent. The multimodal large model can also provide corresponding rectification suggestions for violations, thus completing a closed loop in the business chain to a certain extent.
[0078] Please refer to Figure 5 The diagram shows a structural block diagram of a traffic violation identification device 300 according to an embodiment of this application. The device 300 may include: an image acquisition module 310, a traffic violation identification module 320, a target detection module 330, a traffic regulation acquisition module 340, and a traffic violation review module 350.
[0079] The image acquisition module 310 is used to acquire monitoring video images of the power operation site.
[0080] The violation recognition module 320 is used to identify violations in the surveillance video image using a pre-trained violation recognition model to obtain a first violation recognition result. The violation recognition model is used to identify violations of a preset violation type.
[0081] The target detection module 330 is used to perform security target detection on the surveillance video image using a pre-trained target detection model, and obtain the target detection result.
[0082] The regulation acquisition module 340 is used to acquire the preset safety regulations corresponding to the preset violation type.
[0083] The violation review module 350 is used to review the first violation identification result based on the preset safety regulations and the target detection result, and by using a pre-trained multimodal large model, to obtain a second violation identification result.
[0084] Optionally, the first violation identification result includes a first violation identification result image, in which a violation type label of the first violation behavior is marked, and the coordinate information of the first violation behavior in the surveillance video image; the target detection result includes a target detection result image, in which the target name of the identified safety target is marked, and the coordinate information of the safety target in the surveillance video image. The violation identification device 300 may further include: an identifier adding module, which can be used to mark the violation type label of the first violation behavior to the corresponding position in the target detection image according to the coordinate information of the first violation behavior in the first violation identification result image before the second violation identification result is obtained by re-judging the first violation identification result based on the preset safety regulations and the target detection result and using a pre-trained multimodal large model, thereby obtaining a first target image.
[0085] Optionally, the violation review module 350 may include a question generation unit, an answer generation unit, and a violation review unit. The question generation unit can generate target question text based on the preset safety regulations and the target detection result. The target question text is used to inquire whether the safety target complies with the preset safety regulations. The answer generation unit can input the target question text and the first target image into the multimodal large model to obtain the output target answer text. The violation review unit can review the first violation identification result based on the target answer text to obtain the second violation identification result.
[0086] Optionally, the labeling module can be specifically used to label the violation label in the violation type label of the first violation behavior to the corresponding position in the target result detection image according to the first preset labeling method, so as to obtain the first target image.
[0087] Optionally, the second violation recognition result includes a second violation recognition result image, in which a violation type label of the identified second violation is marked, and the coordinate information of the second violation in the surveillance video image. The violation recognition device 300 may further include: a label replacement module, a label addition module, and a safety prompt module. The label replacement module can be used to, after marking the violation type label of the first violation to the corresponding position in the target result detection image according to a first preset labeling method to obtain a first target image, if the second violation does not include the first violation, then according to a second preset labeling method, replace the violation label in the first target image with a non-violation label from the violation type label of the first violation, to obtain a second target image. The second preset labeling method is different from the first preset labeling method. The label addition module can be used to, based on the coordinate information of the second violation in the second violation recognition result image and according to the first preset labeling method, mark the violation label from the violation type label of the second violation to the corresponding position in the second target image, to obtain a third target image. The safety alert module is used to output the third target image and safety alert information, the safety alert information being used to alert the safety target that the second violation has occurred.
[0088] Optionally, there are multiple violation recognition models, and different violation recognition models are used to identify different preset violation types. There are multiple first violation recognition results and multiple preset safety regulations. The violation review module 350 can be specifically used to: review multiple first violation recognition results based on multiple preset safety regulations and the target detection results, and use a pre-trained multimodal large model to obtain multiple second violation recognition results.
[0089] Optionally, each first violation identification result includes a first violation identification result image, each first violation identification result image is labeled with the identified first violation behavior, and the coordinate information of the first violation behavior in the surveillance video image; the target detection result includes a target detection result image, the target detection result image is labeled with the target name of the identified security target, and the coordinate information of the security target in the surveillance video image. The violation identification device 300 may further include: an identifier adding module. The identifier adding module can be used to, before performing a re-judgment on multiple first violation identification results based on the preset safety regulations and the target detection results, and using a pre-trained multimodal large model to obtain multiple second violation identification results, label the violation type of each first violation behavior in each first violation identification result image to the corresponding position in the first target image, based on the coordinate information of the first violation behavior in each first violation identification result image, to obtain a first target image.
[0090] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0091] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.
[0092] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0093] In summary, considering that small models have fewer parameters and higher recognition efficiency but lower accuracy and are prone to false alarms, while multimodal large models have more parameters and lower real-time recognition, but have stronger text understanding capabilities and a greater advantage in understanding the overall image, a collaborative approach of small and large models is adopted for identifying violations. The small model first quickly captures the violations on-site, and then the large model uses the recognition results to query relevant safety regulations for a professional and accurate re-judgment, while supplementing the recognition with the target detection results. Specifically, the violation re-judgment reduces the false alarm rate of violation alarms and effectively reduces the workload of manual re-judgment; at the same time, the use of the large model for supplementary image recognition reduces the false negative rate of violations to some extent. The multimodal large model can also provide corresponding rectification suggestions for violations, thus completing a closed loop in the business chain to a certain extent.
[0094] Please see Figure 6 , Figure 6This is a block diagram illustrating an electronic device 400 according to an exemplary embodiment. Figure 6 As shown, the electronic device 400 may include a processor 401 and a memory 402. The electronic device 400 may also include one or more of a multimedia component 403, an input / output (I / O) interface 404, and a communication component 405.
[0095] The processor 401 controls the overall operation of the electronic device 400 to complete all or part of the steps in the aforementioned violation identification method. The memory 402 stores various types of data to support the operation of the electronic device 400. This data may include, for example, instructions for any application or method operating on the electronic device 400, and application-related data such as contact data, sent and received messages, pictures, audio, video, etc. The memory 402 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 403 may include a screen and audio components. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in memory 402 or transmitted via communication component 405. The audio component also includes at least one speaker for outputting audio signals. I / O interface 404 provides an interface between processor 401 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons may be virtual or physical buttons. Communication component 405 is used for wired or wireless communication between the electronic device 400 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IoT, eMTC, or other 5G technologies, or combinations thereof, is not limited here. Therefore, the corresponding communication component 405 may include: a Wi-Fi module, a Bluetooth module, an NFC module, etc.
[0096] In an exemplary embodiment, the electronic device 400 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described violation identification method.
[0097] Please refer to Figure 7 This diagram illustrates a structural block diagram of a computer-readable storage medium provided in an embodiment of this application. The computer-readable medium 500 stores program code that can be called by a processor to execute the methods described in the above method embodiments.
[0098] The computer-readable storage medium 500 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 500 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 500 has storage space for program code 510 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 510 may be compressed, for example, in a suitable form.
[0099] In some embodiments, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the steps in the above-described method embodiments.
[0100] The preferred embodiments of this disclosure have been described in detail above with reference to the accompanying drawings. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, various simple modifications can be made to the technical solutions of this disclosure, and these simple modifications all fall within the protection scope of this disclosure.
[0101] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.
[0102] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.
Claims
1. A method for identifying traffic violations, characterized in that, The method includes: Acquire monitoring video images of the power operation site; Using a pre-trained violation recognition model, the surveillance video image is used to identify violations, resulting in a first violation recognition result. The violation recognition model is used to identify violations of a preset violation type. Using a pre-trained target detection model, security target detection is performed on the surveillance video image to obtain the target detection result; Obtain the preset safety regulations corresponding to the preset violation type; Based on the preset safety regulations and the target detection results, and using a pre-trained multimodal large model to re-evaluate the first violation identification result, a second violation identification result is obtained.
2. The method according to claim 1, characterized in that, The first violation identification result includes a first violation identification result image, in which the violation type label of the first violation behavior is marked, and the coordinate information of the first violation behavior in the surveillance video image; The target detection result includes a target detection result image, in which the target detection result image is labeled with the target name of the identified security target and the coordinate information of the security target in the surveillance video image; Before obtaining the second violation identification result by re-judging the first violation identification result based on the preset safety regulations and the target detection result, and using a pre-trained multimodal large model, the method further includes: Based on the coordinate information of the first violation in the first violation recognition result image, the violation type label of the first violation is marked at the corresponding position in the target result detection image to obtain the first target image.
3. The method according to claim 2, characterized in that, The second violation identification result is obtained by re-evaluating the first violation identification result based on the preset safety regulations and the target detection result, and using a pre-trained multimodal large model. This includes: Based on the preset security regulations and the target detection results, a target question text is generated, which is used to ask whether the security target complies with the preset security regulations. The target question text and the first target image are input into the multimodal large model to obtain the output target answer text; The first violation identification result is re-evaluated based on the target answer text to obtain the second violation identification result.
4. The method according to claim 2, characterized in that, The step of labeling the violation type tag of the first violation behavior to the corresponding position in the target result detection image to obtain the first target image includes: According to the first preset annotation method, the violation identifier in the violation type label of the first violation behavior is annotated to the corresponding position in the target result detection image to obtain the first target image.
5. The method according to claim 4, characterized in that, The second violation identification result includes a second violation identification result image, in which the violation type label of the identified second violation behavior is marked, and the coordinate information of the second violation behavior in the surveillance video image; After labeling the violation type tag of the first violation behavior to the corresponding position in the target result detection image according to the first preset labeling method to obtain the first target image, the method further includes: If the second violation does not include the first violation, then according to the second preset annotation method, the violation identifier in the first target image is replaced with the non-violation identifier in the violation type label of the first violation to obtain the second target image. The second preset annotation method is different from the first preset annotation method. Based on the coordinate information of the second violation in the second violation recognition result image, and according to the first preset annotation method, the violation identifier in the violation type label of the second violation is annotated to the corresponding position in the second target image to obtain the third target image. The third target image and security alert information are output, wherein the security alert information is used to indicate that the security target has committed the second violation.
6. The method according to any one of claims 1-5, characterized in that, The number of violation recognition models is multiple, and different violation recognition models are used to identify different preset violation types of violations. The number of first violation recognition results is multiple, and the number of preset safety regulations is multiple. The second violation identification result is obtained by re-evaluating the first violation identification result based on the preset safety regulations and the target detection result, and using a pre-trained multimodal large model. This includes: Based on multiple preset safety regulations and target detection results, and using a pre-trained multimodal large model to re-judge multiple first violation identification results, multiple second violation identification results are obtained.
7. The method according to claim 6, characterized in that, Each first violation identification result includes a first violation identification result image, each first violation identification result image is marked with the identified first violation behavior, and the coordinate information of the first violation behavior in the surveillance video image; the target detection result includes a target detection result image, the target detection result image is marked with the target name of the identified security target and the coordinate information of the security target in the surveillance video image; Before obtaining multiple second violation identification results by re-judging multiple first violation identification results based on the preset safety regulations and the target detection results, and using a pre-trained multimodal large model, the method further includes: Based on the coordinate information of the first violation in each of the first violation recognition result images, the violation type label of each of the first violations is marked at the corresponding position in the first target image to obtain the first target image.
8. A traffic violation identification device, characterized in that, The device includes: The image acquisition module is used to acquire monitoring video images of the power operation site; The violation identification module is used to identify violations in the surveillance video image using a pre-trained violation identification model to obtain a first violation identification result. The violation identification model is used to identify violations of a preset violation type. The target detection module is used to perform security target detection on the surveillance video image using a pre-trained target detection model, and obtain the target detection result; The regulation acquisition module is used to acquire the preset safety regulations corresponding to the preset violation type; The violation review module is used to review the first violation identification result based on the preset safety regulations and the target detection result, and by using a pre-trained multimodal large model, to obtain a second violation identification result.
9. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 7.