Detection alarm method and device based on multiple video frames

By detecting multiple video frames and setting the alarm frame threshold, the problem of low accuracy of single frame detection results in the prior art is solved, and the accuracy and reliability of factory safety alarms are improved.

CN120047871APending Publication Date: 2025-05-27PIPE NETWORK MANAGEMENT BRANCH OF BEIJING WATERWORKS GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510127909.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the prior art, it is judged whether an alarm is triggered based on the detection results of a single video frame. The accuracy is low and it is easy to trigger the alarm by mistake, affecting the normal operation of the factory.

Method used

By inputting the captured factory video data into the trained factory safety detection algorithm, the detection results of multiple video frames are obtained, and the factory safety alarm is determined based on the alarm frame threshold.

Benefits of technology

Improve the accuracy of triggering factory safety alarms, reduce false alarms caused by single frame mischeck or accidental fluctuations, and ensure the reliability and effectiveness of alarms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047871A_ABST
    Figure CN120047871A_ABST
Patent Text Reader

Abstract

The invention provides a detection alarm method and device based on multiple video frames, and the method comprises the steps: inputting shot plant video data into a trained plant safety detection algorithm corresponding to each detection content, and obtaining the detection results of multiple video frames; obtaining an alarm frame number threshold value of each detection content; and if the number of the alarm video frames of which the detection result is the alarm result is greater than the alarm frame threshold value of the corresponding detection content, triggering the plant safety alarm of the corresponding detection content. According to the invention, the accuracy of triggering plant safety alarm can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video frame detection and alarm, and specifically relates to a detection and alarm method and device based on multiple video frames. Background Art

[0002] With the development of technology, the technology of detecting and alarming based on surveillance videos captured by cameras is more and more widely used. Specifically, in the prior art, as long as the detection result of one video frame is an alarm result, the alarm will be triggered. However, since the video data captured by the camera includes multiple video frames in one second, the time corresponding to each video frame is very short. That is, the accuracy of judging whether to trigger the alarm based on the detection result of one video frame is very low, and it is easy to trigger the alarm by mistake. For the safety monitoring of the factory building, frequent false alarms will seriously affect the normal operation of the factory building. Summary of the Invention

[0003] The purpose of this application is to provide a detection and alarm method and device based on multiple video frames, which can overcome the disadvantages and deficiencies in the prior art.

[0004] The first aspect of the embodiments of this application provides a detection and alarm method based on multiple video frames, including:

[0005] Input the captured factory building video data into the factory building safety detection algorithms corresponding to each detection content that have been trained to obtain the detection results of multiple video frames;

[0006] Obtain the alarm frame number threshold for each of the detection contents;

[0007] If the number of alarm video frames with the detection result being an alarm result is greater than the alarm frame threshold corresponding to the detection content, trigger the factory building safety alarm corresponding to the detection content.

[0008] Further, the step of triggering the factory building safety alarm corresponding to the detection content when the number of alarm video frames with the detection result being an alarm result is greater than the alarm frame threshold corresponding to the detection content includes:

[0009] Generate a result queue according to the queue length threshold corresponding to the detection content;

[0010] Add each of the detection results to the result queue in sequence. When the detection result is an alarm result and is added to the result queue, the cumulative number of alarm video frames is incremented by one;

[0011] When the cumulative number of alarm video frames in the result queue is greater than the alarm frame threshold corresponding to the detection content, trigger the factory building safety alarm corresponding to the detection content.

[0012] Further, the step of sequentially adding each of the detection results to the result queue includes:

[0013] When adding the current detection result to the result queue, if the number of detection results existing in the result queue is equal to the queue length threshold, remove the detection result at the front end of the result queue, and add the current detection result to the end of the result queue.

[0014] Further, the step of removing the detection result at the front end of the result queue includes:

[0015] If the removed detection result is an alarm result, the cumulative number of alarm video frames is decreased by one.

[0016] Further, the factory building safety detection algorithm is obtained through the following steps:

[0017] Obtain a plurality of training samples;

[0018] According to the label types of each training sample, group the plurality of training samples to obtain a plurality of sample groups corresponding to training different detection algorithms; each of the sample groups includes a plurality of training samples of at least one type;

[0019] Train each of the initial detection algorithms of the laboratory server cluster with each sample group to obtain a plurality of different factory building safety detection algorithms.

[0020] Further, the factory building safety detection algorithm includes a backbone module, an enhancement module, a neck module, and a head module;

[0021] The step of inputting the captured factory building video data into the factory building safety detection algorithm corresponding to each detection content that has been trained to obtain detection results of a plurality of video frames includes:

[0022] Input the factory building video data into the backbone module of the current factory building safety detection algorithm for feature extraction processing to obtain first features of each video frame;

[0023] Input the first features of each video frame into the enhancement module of the current factory building safety detection algorithm for enhanced convolution processing to obtain second features of each video frame;

[0024] Input the second features of each video frame into the neck module of the current factory building safety detection algorithm for bidirectional feature convolution processing to obtain third features of each video frame;

[0025] Input the third features of each video frame into the head module of the current factory building safety detection algorithm for feature prediction processing to obtain detection results of a plurality of the video frames of the current factory building safety detection algorithm.

[0026] Furthermore, the backbone module includes a feature extraction layer and a plurality of cascaded combination convolutional layers in sequence; wherein, the combination convolutional layer includes a first convolutional network and a second convolutional network;

[0027] The step of inputting the factory building video data into the backbone module of the current factory building safety detection algorithm for feature extraction processing to obtain the first feature of each video frame includes:

[0028] Input the video data into the feature extraction layer for feature extraction processing to obtain the fourth feature of each video frame;

[0029] Input the fourth feature of each video frame into the first convolutional network of the first combination convolutional layer for convolutional processing to obtain the fifth feature of each video frame;

[0030] Input the fifth feature of each video frame into the second convolutional network of the same combination convolutional layer for convolutional processing to obtain the sixth feature of each video frame;

[0031] Input the sixth feature of each video frame into the next-level combination convolutional layer for convolutional processing, and determine the sixth feature of each video frame output by the last-level combination convolutional layer as the first feature.

[0032] Furthermore, the second convolutional network includes a plurality of cascaded phantom convolutional units and convolutional fusion units in sequence;

[0033] The step of inputting the fifth feature of each video frame into the second convolutional network of the same combination convolutional layer for convolutional processing to obtain the sixth feature of each video frame includes:

[0034] Input the fifth feature into a plurality of cascaded phantom convolutional units for convolutional processing to obtain the seventh feature output by each phantom convolutional unit;

[0035] Input the fifth feature and the seventh feature output by each phantom convolutional unit into the convolutional fusion unit for convolutional fusion to obtain the sixth feature.

[0036] Furthermore, the step of inputting the fifth feature of each video frame into the second convolutional network of the same combination convolutional layer for convolutional processing to obtain the sixth feature of each video frame includes:

[0037] Obtain the sixth feature through the following formula:

[0038]

[0039] Y is the sixth feature, x is the fifth feature, cat(·) is the data concatenation function, y iIt represents the seventh feature output after i cascaded phantom convolution units, and n is the number of cascaded phantom convolution units.

[0040] The second aspect of the embodiments of the present application provides a detection and alarm device based on multiple video frames, including:

[0041] A detection result acquisition module, configured to input the captured factory building video data into the factory building safety detection algorithms corresponding to each detection content that have been trained, and obtain the detection results of multiple video frames;

[0042] An alarm frame number threshold acquisition module, which acquires the alarm frame number thresholds for each of the detection contents;

[0043] A result generation module, configured to trigger the factory building safety alarm corresponding to the detection content if the number of alarm video frames whose detection result is an alarm result is greater than the alarm frame threshold corresponding to the detection content.

[0044] Compared with the prior art, the present application inputs the captured factory building video data into the factory building safety detection algorithms corresponding to each detection content that have been trained, and obtains the detection results of multiple video frames; acquires the alarm frame number thresholds for each detection content; if the number of alarm video frames whose detection result is an alarm result is greater than the alarm frame threshold corresponding to the detection content, triggers the factory building safety alarm corresponding to the detection content, achieving the technical effect of judging whether to trigger the factory building safety alarm according to the detection results of multiple video frames. It will not trigger an alarm due to misdetection or accidental fluctuations in a single frame, but requires continuous observation of problems within a certain period of time to consider that the behavior is abnormal, improving the accuracy of triggering the factory building safety alarm. Moreover, the present application performs feature extraction processing on each video frame through the first convolutional network and the second convolutional network including multiple sequentially cascaded phantom convolution units and convolution fusion units, which can retain the original details of the factory building video data to the greatest extent and enhance the accuracy of the detection results output by the factory building safety detection algorithm.

[0045] In order to understand the present application more clearly, the following will describe the specific implementation manners of the present application in conjunction with the accompanying drawings. Description of the Drawings

[0046] Figure 1 It is a flowchart of the detection and alarm method based on multiple video frames according to an embodiment of the present application.

[0047] Figure 2 It is a schematic diagram of the algorithm model of the factory building safety detection algorithm according to an embodiment of the present application.

[0048] Figure 3 It is a schematic diagram of the second convolutional network of the backbone module of the factory building safety detection algorithm according to an embodiment of the present application.

[0049] Figure 4 Schematic diagram of the target enhancement layer of the enhancement module of the factory building safety detection algorithm according to an embodiment of the present application.

[0050] Figure 5 Schematic diagram of the module connection of the factory building safety detection device according to an embodiment of the present application.

[0051] 1. Detection result acquisition module; 2. Alarm frame number threshold acquisition module; 3. Result generation module. Detailed implementation manners

[0052] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0053] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the embodiments of the present application.

[0054] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and do not have to be used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances. The singular forms of "a", "the" and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. The word "if" / "when" used herein can be interpreted as "when...", "when...", or "in response to a determination".

[0055] In addition, in the description of the present application, unless otherwise stated, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0056] As shown in the appendix of the specification Figure 1 , which is a flowchart of the detection and alarm method based on multiple video frames according to an embodiment of the present application, including:

[0057] S1: Input the captured factory building video data into the factory building safety detection algorithm trained for each detection content to obtain the detection results of multiple video frames.

[0058] Among them, the factory building safety detection algorithms corresponding to each detection content are shown in the following table:

[0059]

[0060]

[0061] S2: Obtain the alarm frame number threshold for each of the detection contents.

[0062] Among them, different detection contents correspond to different alarm frame number thresholds. For example, the alarm frame number threshold for helmet detection (detecting whether a helmet is worn) is 3 frames, and the alarm frame threshold for reflective vest detection (detecting whether a reflective vest is worn) is 8 frames.

[0063] S3: If the number of alarm video frames in the detection result is the alarm result, and is greater than the alarm frame threshold corresponding to the detection content, trigger the factory building safety alarm corresponding to the detection content.

[0064] In a feasible embodiment, S3: If the number of alarm video frames in the detection result is the alarm result, and is greater than the alarm frame threshold corresponding to the detection content, the step of triggering the factory building safety alarm corresponding to the detection content includes:

[0065] S31: Generate a result queue according to the queue length threshold corresponding to the detection content;

[0066] S32: Add each of the detection results to the result queue in sequence. When the detection result is the alarm result and is added to the result queue, the cumulative number of alarm video frames is incremented by one;

[0067] S33: When the cumulative number of alarm video frames in the result queue is greater than the alarm frame threshold corresponding to the detection content, trigger the factory building safety alarm corresponding to the detection content.

[0068] In a feasible embodiment, the step of S32: adding each of the detection results to the result queue in sequence includes:

[0069] When adding the current detection result to the result queue, if the number of detection results existing in the result queue is equal to the queue length threshold, remove the detection result at the front end of the result queue, and add the current detection result to the end of the result queue.

[0070] In a feasible embodiment, the step of removing the detection result at the front end of the result queue includes:

[0071] If the removed detection result is the alarm result, the cumulative number of alarm video frames is decremented by one.

[0072] Compared with the prior art, the present application inputs the captured video data of the factory building into the factory building safety detection algorithms corresponding to each detection content that have been trained, and obtains the detection results of multiple video frames; obtains the alarm frame number thresholds for each detection content; if the number of alarm video frames with the detection result being an alarm result is greater than the alarm frame threshold for the corresponding detection content, triggers the factory building safety alarm corresponding to the detection content, achieving the technical effect of determining whether to trigger the factory building safety alarm based on the detection results of multiple video frames, and will not trigger an alarm due to misdetection or accidental fluctuations of a single frame, but requires observing problems continuously within a certain time period to consider that the behavior is abnormal, improving the accuracy of triggering the factory building safety alarm.

[0073] In a feasible embodiment, the factory building safety detection algorithm is trained through the following steps:

[0074] Obtain multiple training samples;

[0075] According to the label types of each training sample, group the multiple training samples to obtain multiple sample groups corresponding to training different detection algorithms; each of the sample groups includes multiple training samples of at least one type;

[0076] Train each of the initial detection algorithms of the laboratory server cluster with each sample group to obtain multiple different factory building safety detection algorithms.

[0077] As shown in the appendix of the specification Figure 2 , in a feasible embodiment, the factory building safety detection algorithm includes a backbone module, an enhancement module, a neck module, and a head module;

[0078] The step S1 of inputting the captured video data of the factory building into the factory building safety detection algorithms corresponding to each detection content that have been trained and obtaining the detection results of multiple video frames includes:

[0079] S11: Input the factory building video data into the backbone module of the current factory building safety detection algorithm for feature extraction processing to obtain the first features of each video frame;

[0080] S12: Input the first features of each video frame into the enhancement module of the current factory building safety detection algorithm for enhanced convolution processing to obtain the second features of each video frame;

[0081] S13: Input the second features of each video frame into the neck module of the current factory building safety detection algorithm for two-way feature convolution processing to obtain the third features of each video frame;

[0082] S14: Input the third feature of each video frame into the head module of the current factory building safety detection algorithm for feature prediction processing to obtain the detection results of multiple video frames of the current factory building safety detection algorithm.

[0083] In a feasible embodiment, the backbone module includes a feature extraction layer and multiple cascaded combination convolutional layers in sequence; wherein, the combination convolutional layer includes a first convolutional network and a second convolutional network.

[0084] The step S11 of inputting the factory building video data into the backbone module of the current factory building safety detection algorithm for feature extraction processing to obtain the first feature of each video frame includes:

[0085] S111: Input the video data into the feature extraction layer for feature extraction processing to obtain the fourth feature of each video frame.

[0086] S112: Input the fourth feature of each video frame into the first convolutional network of the first combination convolutional layer for convolutional processing to obtain the fifth feature of each video frame.

[0087] S113: Input the fifth feature of each video frame into the second convolutional network of the same combination convolutional layer for convolutional processing to obtain the sixth feature of each video frame.

[0088] S114: Input the sixth feature of each video frame into the next-level combination convolutional layer for convolutional processing, and determine the sixth feature of each video frame output by the last-level combination convolutional layer as the first feature.

[0089] As shown in the appendix of the specification Figure 3 , in a feasible embodiment, the second convolutional network includes multiple cascaded phantom convolutional units and convolutional fusion units.

[0090] The step S113 of inputting the fifth feature of each video frame into the second convolutional network of the same combination convolutional layer for convolutional processing to obtain the sixth feature of each video frame includes:

[0091] S1131: Input the fifth feature into multiple cascaded phantom convolutional units for convolutional processing to obtain the seventh feature output by each phantom convolutional unit.

[0092] S1132: Input the fifth feature and the seventh feature output by each phantom convolutional unit into the convolutional fusion unit for convolutional fusion to obtain the sixth feature.

[0093] In a feasible embodiment, the step S113: inputting the fifth feature of each video frame into the second convolutional network of the same combined convolutional layer for convolutional processing to obtain the sixth feature of each video frame, can be expressed by the following formula:

[0094] The sixth feature is obtained through the following formula:

[0095]

[0096] Y is the sixth feature, x is the fifth feature, cat(·) is the data concatenation function, and y i represents the seventh feature output after i cascaded phantom convolutional units, and n is the number of cascaded phantom convolutional units.

[0097] In a feasible embodiment, the backbone module includes a plurality of sequentially cascaded combined convolutional layers; the enhancement module includes a first convolutional layer, a second convolutional layer, and a target enhancement layer;

[0098] The step S12: inputting the first feature of each video frame into the enhancement module for enhanced convolutional processing to obtain the second feature of each video frame, includes:

[0099] S121: Inputting the first features output by the penultimate and antepenultimate combined convolutional layers of the backbone module into the first convolutional layer and the second convolutional layer respectively for convolutional processing to obtain a first enhanced feature and a second enhanced feature.

[0100] S122: Inputting the first feature output by the last combined convolutional layer of the backbone module into the target enhancement layer for enhanced convolutional processing to obtain a third enhanced feature.

[0101] As shown in the appendix of the specification Figure 4 , the target enhancement layer first smoothly extracts a concise feature map through 1×1 convolution and 3×3 convolution, and then performs a dilated residual attention operation on the feature map to filter branch features through convolutions with different dilation rates. Each branch in the dilated residual attention operation has a unique receptive field, which can form a comprehensive feature representation. Among them, due to the relatively small target characteristics of the antenna interference source, in this embodiment, the dilation rate 5 used in the convolution is replaced with the dilation rate 3 to minimize the redundancy in the receptive field. Finally, we use convolutional channel compression to perform convolutional compression on the output results of multiple branches in the dilated residual attention operation to obtain a comprehensive third enhanced feature.

[0102] In the above embodiments, by using the first convolutional network and the second convolutional network including a plurality of sequentially cascaded phantom convolutional units and convolutional fusion units to perform feature extraction processing on each video frame, the original details of the plant video data can be retained to the greatest extent, the accuracy of the detection results output by the plant safety detection algorithm can be enhanced, and by enhancing and fusing features of different dimensions, the plant video data can be detected more comprehensively in combination with multiple dimensions to improve the accuracy of the detection results output by the plant safety detection algorithm.

[0103] As shown in the appendix of the specification Figure 5 , the second aspect of the embodiments of the present application provides a detection and alarm device based on multiple video frames, including:

[0104] A detection result acquisition module 1, configured to input the captured plant video data into a trained plant safety detection algorithm corresponding to each detection content to obtain detection results of multiple video frames;

[0105] An alarm frame number threshold acquisition module 2, which acquires the alarm frame number threshold for each of the detection contents;

[0106] A result generation module 3, configured to trigger a plant safety alarm for the corresponding detection content if the number of alarm video frames whose detection results are alarm results is greater than the alarm frame threshold for the corresponding detection content.

[0107] It should be noted that when the detection and alarm device based on multiple video frames provided in the second embodiment of the present application executes the detection and alarm method based on multiple video frames, only the above division of each functional network is used for illustration. In actual application, the above functions can be allocated to different functional networks according to needs, that is, the internal structure of the device is divided into different functional networks to complete all or part of the functions described above. In addition, the detection and alarm device based on multiple video frames provided in the second embodiment of the present application and the detection and alarm method based on multiple video frames in the first embodiment of the present application belong to the same concept. The implementation process is shown in the method embodiment and will not be elaborated here.

[0108] The device embodiments described above are only illustrative. The components described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the networks can be selected according to actual needs to achieve the purpose of the solution of the present application. Those of ordinary skill in the art can understand and implement it without creative work.

[0109] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0110] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the selected functions in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the selected functions in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.

[0111] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the selected functions in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.

[0112] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0113] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.

[0114] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be a network of computer-readable instructions, data structures, programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0115] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0116] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A detection alarm method based on multiple video frames, characterized in that: include: The captured factory video data is input into the trained factory safety detection algorithm corresponding to each detection content to obtain the detection results of multiple video frames; Obtaining the alarm frame number threshold of each detection content; If the number of alarm video frames of the detection result being an alarm result is greater than the alarm frame threshold of the corresponding detection content, the factory safety alarm of the corresponding detection content is triggered.

2. The detection alarm method based on multiple video frames according to claim 1, characterized in that: If the number of alarm video frames of the detection result being an alarm result is greater than the alarm frame threshold of the corresponding detection content, the step of triggering the factory building safety alarm of the corresponding detection content includes: Generate a result queue according to a queue length threshold corresponding to the detection content; Add each of the detection results to the result queue in sequence, and when the detection result is an alarm result and is added to the result queue, the accumulated number of alarm video frames is increased by one; When the accumulated number of alarm video frames in the result queue is greater than the alarm frame threshold of the corresponding detection content, the factory safety alarm of the corresponding detection content is triggered.

3. The detection alarm method based on multiple video frames according to claim 2 is characterized in that: The step of sequentially adding each of the detection results to the result queue comprises: When adding the current detection result to the result queue, if the number of detection results existing in the result queue is equal to the queue length threshold, remove the frontmost detection result of the result queue and add the current detection result to the end of the result queue.

4. The detection alarm method based on multiple video frames according to claim 3 is characterized in that: The step of removing the frontmost test result of the result queue comprises: If the removed detection result is an alarm result, the accumulated number of alarm video frames is reduced by one.

5. The detection alarm method based on multiple video frames according to claim 1, characterized in that: The plant safety detection algorithm is trained by the following steps: Obtain multiple training samples; According to the label type of each training sample, multiple training samples are grouped to obtain multiple sample groups corresponding to training different detection algorithms; each of the sample groups includes multiple training samples of at least one type; Each sample group is used to train each of the initial detection algorithms of the laboratory server cluster to obtain a plurality of different plant safety detection algorithms.

6. The detection alarm method based on multiple video frames according to claim 1, characterized in that: The plant safety detection algorithm includes a trunk module, an enhancement module, a neck module and a head module; The step of inputting the captured factory video data into the trained factory safety detection algorithm corresponding to each detection content to obtain the detection results of multiple video frames includes: Inputting the factory video data into the backbone module of the current factory safety detection algorithm for feature extraction processing to obtain the first feature of each video frame; Inputting the first feature of each video frame into the enhancement module of the current factory building safety detection algorithm for enhanced convolution processing to obtain the second feature of each video frame; Inputting the second feature of each video frame into the neck module of the current factory building safety detection algorithm for bidirectional feature convolution processing to obtain the third feature of each video frame; The third feature of each video frame is input into the head module of the current factory building safety detection algorithm for feature prediction processing to obtain the detection results of the multiple video frames of the current factory building safety detection algorithm.

7. The detection alarm method based on multiple video frames according to claim 6, characterized in that: The backbone module includes a feature extraction layer and a plurality of sequentially cascaded combined convolutional layers; wherein the combined convolutional layer includes a first convolutional network and a second convolutional network; The step of inputting the factory video data into the backbone module of the current factory safety detection algorithm for feature extraction processing to obtain the first feature of each video frame includes: Inputting the factory video data into the feature extraction layer for feature extraction processing to obtain the fourth feature of each video frame; Inputting the fourth feature of each video frame into the first convolutional network of the first combined convolutional layer for convolution processing to obtain the fifth feature of each video frame; Inputting the fifth feature of each video frame into the second convolution network of the same combined convolution layer for convolution processing to obtain the sixth feature of each video frame; The sixth feature of each video frame is input into the next-level combined convolution layer for convolution processing, and the sixth feature of each video frame output by the last-level combined convolution layer is determined as the first feature.

8. The detection alarm method based on multiple video frames according to claim 7, characterized in that: The second convolutional network includes a plurality of phantom convolution units and convolution fusion units cascaded in sequence; The step of inputting the fifth feature of each video frame into the second convolution network of the same combined convolution layer for convolution processing to obtain the sixth feature of each video frame includes: Inputting the fifth feature into a plurality of phantom convolution units cascaded in sequence for convolution processing, to obtain a seventh feature output by each phantom convolution unit; The fifth feature and the seventh feature output by each phantom convolution unit are input into the convolution fusion unit for convolution fusion to obtain the sixth feature.

9. The detection alarm method based on multiple video frames according to claim 6, characterized in that: The step of inputting the fifth feature of each video frame into the second convolution network of the same combined convolution layer for convolution processing to obtain the sixth feature of each video frame includes: The sixth characteristic is obtained by the following formula: Y is the sixth feature, x is the fifth feature, cat(·) is the data concatenation function, y i represents the seventh feature output by i cascaded phantom convolution units, and n is the number of cascaded phantom convolution units.

10. A detection alarm device based on multiple video frames, characterized in that: include: The detection result acquisition module is used to input the captured factory video data into the trained factory safety detection algorithm corresponding to each detection content to obtain the detection results of multiple video frames; An alarm frame number threshold acquisition module is used to acquire the alarm frame number threshold of each detection content; The result generating module is used to trigger the factory safety alarm of the corresponding detection content if the number of alarm video frames of the detection result as the alarm result is greater than the alarm frame threshold of the corresponding detection content.