A video monitoring anomaly detection method, device and equipment
By using an improved YOLOv5 model and CA attention mechanism, combined with the LeakyRlue activation function, efficient flame fault detection in liquid rocket engine video monitoring was achieved, solving the problem of low detection efficiency in existing technologies and enabling timely identification of flame faults and issuance of alarms.
Patent Information
- Application Number
- CN202210612915.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2042-05-31
AI Technical Summary
In existing technologies, video monitoring for anomaly detection in liquid rocket engines is inefficient and slow, and relying on manual operation makes it impossible to identify faults in a timely manner.
An improved YOLOv5 model, combined with the CA attention mechanism and the LeakyRlue activation function, is used to identify flames in video images of liquid rocket engines and generate flame fault alarm information.
It improves detection accuracy and speed, and can issue instantaneous alarm signals in the event of a fire fault, reducing manual intervention and improving the accuracy and efficiency of fault identification.
Smart Images

Figure CN114898273B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of model identification, and in particular to a video monitoring anomaly detection method, device and equipment. BACKGROUND
[0002] Liquid rocket engines (LREs) are the power source of liquid rockets and the core component of rocket systems. Generally, the working conditions of LREs are very harsh, i.e., high temperature, high pressure, high speed, and unstable working requirements. Due to the extreme working conditions, various faults may occur, such as flame, leakage, and sensor falling off. Although the redline cutoff system and many other methods have been applied to LREs and have achieved good performance in monitoring the running process of LREs, inevitable faults may still occur and cannot be detected. The current method for handling these faults relies on manual operation, which means that special personnel are needed to check the video recording frame by frame and determine whether a fault occurs and the fault location. However, in actual implementation, manual identification of fault models usually needs to interrupt the operation. Moreover, the identification efficiency and accuracy are low. Therefore, proper video monitoring of LREs can save a lot of resources. The prosperity of deep learning makes it possible. In recent years, deep learning has been applied to target detection and computer vision. Ren et al. proposed a faster region-based convolutional neural network (R-CNN), which combines R-CNN and a region proposal network, and makes great progress in detection accuracy and speed by reducing the computational cost. Vaswani et al. proposed the concept of attention mechanism. They believe that in the detection of objects, the location where the target is located should be the obvious part, and Chen et al. created the MEGA algorithm, which uses spatiotemporal random sampling frames around the key frames in a video to help detect targets in the key frames. In the field of target detection based on deep learning, after so much effort, the required requirements have been met.
[0003] You only look once (YOLO) was first proposed by Joseph Redmon in 2016
[10] . YOLO uses the idea of regarding target detection as a regression problem. It is based on an independent end-to-end network that gradually grids the entire image and outputs the center coordinates, the height and width of the surrounding box, and the confidence of a certain target falling in a grid at the same time, thereby completing the positioning and classification of the target. In this way, positioning and classification are completed in one step, the detection speed is greatly improved, but the detection accuracy is reduced.
[0004] Therefore, there is an urgent need to provide a more reliable video monitoring anomaly detection scheme. SUMMARY
[0005] The application aims to provide a video monitoring anomaly detection method, device and equipment, and aims to solve the problems of low detection efficiency and low detection progress in the prior art.
[0006] In order to achieve the above-mentioned purpose, the application provides the following technical scheme.
[0007] In a first aspect, the application provides a video monitoring anomaly detection method, which comprises the following steps:
[0008] Video image data in a ground thermal test process of a liquid rocket engine is acquired;
[0009] The video image data is input into a trained YOLOv5 model, and whether the video image data contains a flame image is identified to obtain an identification result; the trained YOLOv5 model uses a CA attention mechanism, and an activation function used by the CA attention mechanism is a LeakyRlue function;
[0010] When the identification result indicates that the video image data contains a flame image, position information of the flame is determined; there is a connection in a spatial dimension between flame pixels in the flame image;
[0011] Based on the position information, flame fault alarm information is generated.
[0012] In a second aspect, the application provides a video monitoring anomaly detection device, which comprises the following modules:
[0013] A video image data acquisition module is configured to acquire video image data in a ground thermal test process of a liquid rocket engine;
[0014] An identification module is configured to input the video image data into a trained YOLOv5 model, identify whether the video image data contains a flame image, and obtain an identification result; the trained YOLOv5 model uses a CA attention mechanism, and an activation function used by the CA attention mechanism is a LeakyRlue function;
[0015] A flame position information determination module is configured to determine position information of a flame when the identification result indicates that the video image data contains a flame image; there is a connection in a spatial dimension between flame pixels in the flame image;
[0016] A flame fault alarm information generation module is configured to generate flame fault alarm information based on the position information.
[0017] In a third aspect, the application provides a video monitoring anomaly detection device, which comprises the following modules:
[0018] A communication unit / communication interface is configured to acquire video image data in a ground thermal test process of a liquid rocket engine.
[0019] A processing unit / processor is configured to input the video image data into a trained YOLOv5 model, identify whether the video image data contains a flame image, and obtain an identification result.
[0020] When the identification result indicates that the video image data contains a flame image, position information of the flame is determined.
[0021] Based on the position information, flame failure alarm information is generated.
[0022] The application also provides a computer storage medium, which stores instructions, and when the instructions are executed, the above-mentioned video monitoring abnormality detection method is realized.
[0023] Compared with the prior art, the application provides a video monitoring abnormality detection scheme. BRIEF DESCRIPTION OF DRAWINGS
[0024] The accompanying drawings, which are included to provide a further understanding of the application, form a part of the application and, along with the specification, serve to explain the application. The illustrative embodiments of the application and their description serve to explain the application. They do not limit the application. In the drawings:
[0025] Figure 1 A flowchart of a video monitoring abnormality detection method provided by the application is shown.
[0026] Figure 2 A schematic diagram of the trained YOLOv5 model in the video monitoring anomaly detection method provided by the present solution is shown in the figure below:
[0027] Figure 3 A schematic diagram of the detection result of the original YOLOv5 is shown in the figure below:
[0028] Figure 4 A schematic diagram of the detection result of the combination of YOLOv5 and the original CA module is shown in the figure below:
[0029] Figure 5 A schematic diagram of the video monitoring anomaly detection result provided by the present solution is shown in the figure below:
[0030] Figure 6 A schematic diagram of the structure of the video monitoring anomaly detection device provided by the present solution is shown in the figure below:
[0031] Figure 7 A schematic diagram of the structure of the video monitoring anomaly detection device provided by the present solution is shown in the figure below. DETAILED DESCRIPTION
[0032] In order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, the terms "first", "second", etc. are used to distinguish the same or similar items with basically the same function and effect. For example, the first threshold and the second threshold are only used to distinguish different thresholds, and do not limit the order. Those skilled in the art can understand that the terms "first", "second", etc. do not limit the number and execution order, and the terms "first", "second", etc. also do not necessarily mean different.
[0033] It should be noted that in the present application, the words "exemplary" or "for example" are used to indicate an example, illustration or description. Any embodiment or design scheme described as "exemplary" or "for example" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the use of the words "exemplary" or "for example" is intended to present the relevant concept in a specific manner.
[0034] In the present application, "at least one" means one or more, and "multiple" means two or more. The association relationship between the associated objects is described by "and / or", which means that there can be three kinds of relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can represent a, b, c, the combination of a and b, the combination of a and c, the combination of b and c, or the combination of a, b and c, where a, b and c can be single or multiple.
[0035] During the ground hot test of a liquid rocket engine, various faults may occur. Although the existing redline cutoff system plays an important role in LRE fault monitoring, it cannot guarantee that all faults can be timely alarmed. Therefore, various monitoring methods can still play a role as an auxiliary means. For obvious fault patterns that can be identified by the naked eye, such as flames, leaks, and sensor shedding, video detection algorithms configured with cameras can also be used for detection. Therefore, if video monitoring cameras and redline cutoff systems are used simultaneously, better fault diagnosis results can be obtained. However, current video detection still mainly relies on manual operation. In recent years, target detection methods based on deep learning have made great progress, making it possible for automatic detection algorithms to replace manual detection. The CA attention mechanism is proposed for natural language processing. Later, it is also applied to the field of computer vision. Then various attention modules are created, such as squeeze and excitation network (SE), convolutional block attention module (CBAM), and coordinate attention (CA). Regardless of the principle, all modules can improve detection accuracy without sacrificing detection speed, which is what YOLO lacks, and can allocate higher weights to target regions of interest.
[0036] The present scheme aims to distinguish whether there is a flame in the video, and improves the YOLO architecture by adding an improved CA block, which obtains higher detection accuracy and faster detection speed on real data sets compared with the prior art.
[0037] Next, the scheme provided by the embodiments of the present application will be described in conjunction with the accompanying drawings:
[0038] Figure 1A flowchart of a video monitoring anomaly detection method provided by the application. From the perspective of the program, the execution subject of the flowchart can be a server corresponding to a video monitoring security management platform. The security management platform can refer to a platform for controlling the safety performance of a liquid rocket engine. The platform can run on a fixed terminal or a mobile terminal.
[0039] As shown in Figure 1 , the flowchart can include the following steps:
[0040] Step 110: Obtain video image data in the ground thermal test process of the liquid rocket engine.
[0041] In the specific implementation process of the scheme, the liquid rocket engine ground thermal test process is monitored, so that the video image data in the liquid rocket engine ground thermal test process is obtained through the camera. The video is continuous multi-frame image data.
[0042] Step 120: input the video image data into the trained YOLOv5 model, identify whether the video image data contains a flame image, and obtain an identification result; the trained YOLOv5 model uses a CA attention mechanism, and the activation function used by the CA attention mechanism is a LeakyRlue function.
[0043] The YOLOv5 model input end, Backbone, Neck, and Prediction four parts, wherein Backbone: a convolutional neural network that aggregates and forms image features at different image granularities. Neck: a series of network layers that mix and combine image features, and pass the image features to the prediction layer. (Generally FPN or PANET). Head: predict the image features to generate a bounding box and predict the category.
[0044] Taking the YOLOv5s structure as an example, in the first Focus structure, the number of convolution kernels is 32 during the last convolution operation, so the size of the feature map becomes 304*304*32 after the Focus structure. Of course, the more the number of convolution kernels, the thicker the feature map, i.e. the wider the width, and the stronger the learning ability of the network to extract features. For YOLOv5, whether it is V5s, V5m, V5l or V5x, its Backbone, Neck and Head are consistent. The only difference is the depth and width settings of the model. Only by modifying these two parameters can the network structure of the model be adjusted.
[0045] CA (Coordinate attention) attention mechanism, in order to obtain attention in the width and height of the image and encode the precise position information, the input feature map is first divided into width and height directions for global average pooling, respectively obtaining the feature maps in the width and height directions. Unlike the channel attention that converts the feature tensor into a single feature vector through 2-dimensional global pooling, the coordinate attention decomposes the channel attention into two 1-dimensional feature encoding processes, respectively aggregating the features along the two spatial directions. In this way, long-range dependencies can be captured along one spatial direction, while precise position information can be preserved along the other spatial direction. Then the generated feature maps are respectively encoded into a pair of direction-aware and position-sensitive attention maps, which can be applied to the input feature map complementarily to enhance the representation of the attention object.
[0046] Activation functions can be divided into two categories: saturated activation functions and non-saturated activation functions. Among them, ReLU and its variants belong to non-saturated activation functions, and Leaky Relu differs from Relu in that a very small constant leak is reserved on the negative axis, so that when the input information is less than 0, the information is not completely lost, and a corresponding reservation is made, that is, ReLU has no gradient when the value is less than zero, and LeakyReLU gives a very small gradient when the value is less than 0.
[0047] In the improved YOLOv5 model, not only is the CA attention mechanism added, but also the activation function used by the CA attention mechanism is the LeakyRlue function, and a residual layer is added to compensate for global information, which can improve the accuracy and detection efficiency of video monitoring anomaly recognition.
[0048] Step 130: When the identification result indicates that the video image data contains a flame image, the position information of the flame is determined; there is a spatial dimension relationship between the flame pixels in the flame image.
[0049] When a flame is identified in the video, the position information of the flame can be further located. Since the shape of the flame is not fixed and is always changing in actual application, it cannot be detected based on points alone, but global information in the image needs to be considered.
[0050] Step 140: Based on the position information, flame failure alarm information is generated.
[0051] The flame failure alarm information can alarm the failure that occurs during the ground hot test of the liquid rocket engine. When prompting, in addition to prompting the existence of flame failure, the specific flame position information can also be prompted to prompt relevant personnel to handle the failure in a timely manner.
[0052] Figure 1 The method in the method, by acquiring video image data in the ground thermal test process of the liquid rocket engine; the video image data is input into the trained YOLOv5 model, whether the video image data contains a flame image is identified, and an identification result is obtained; the trained YOLOv5 model uses a CA attention mechanism, and the activation function used by the CA attention mechanism is a LeakyRlue function; when the identification result indicates that the video image data contains a flame image, the position information of the flame is determined; there is a connection in the spatial dimension between the flame pixels in the flame image; based on the position information, flame fault alarm information is generated. By adding an improved CA attention mechanism to improve the YOLOv5 architecture, and using the LeakyRlue function as the activation function, the flame target in a specific frame of the video is detected, which has a faster convergence speed, and the overall scheme obtains higher detection accuracy and faster detection speed on the actual data set, and a detection result with higher confidence can be obtained, and in the ground thermal test of the liquid rocket engine, an instantaneous alarm signal can be sent when a fire fault occurs.
[0053] The method based on Figure 1 The method, and some specific embodiments of the method are also provided in the embodiments of the present specification, which are described below.
[0054] Optionally, the video image data is input into the trained YOLOv5 model to identify whether the video image data contains a flame image, and an identification result is obtained, which can specifically include:
[0055] The video image data is preprocessed based on the CA attention mechanism to obtain a feature vector;
[0056] The feature vector is split into one-dimensional feature vectors in the width direction and the height direction.
[0057] Optionally, the YOLOv5 model can at least include an input layer, a residual layer, a convolution layer, a fully connected layer and an output layer; wherein the input layer receives the video image data; the convolution layer is used for extracting a feature vector from the video image data; the convolution kernel size of the convolution layer is 7;
[0058] The weights in the fully connected layer are updated to obtain a flame feature vector corresponding to the video in the ground thermal test process of the liquid rocket engine;
[0059] The output layer is used for outputting a flame detection result according to the flame feature vector.
[0060] The average pooling and global maximum pooling are used in the residual layer to compensate for the global spatial information of the CA attention mechanism. When the average pooling is performed, one-dimensional feature vectors in the width direction and one-dimensional feature vectors in the height direction are pooled.
[0061] In actual application scenarios, the YOLOv5 model in the present scheme needs to be trained before application. The training process can be implemented based on the following steps:
[0062] A training sample set and a verification sample set are obtained. Each image in the training sample set and the verification sample set includes at least one flame target.
[0063] The training sample set is input into an initial YOLOv5 model to obtain a preliminary training result.
[0064] The preliminary training result is compared with the verification sample set to obtain a comparison result.
[0065] Based on the comparison result, the training parameters in the initial YOLOv5 model are adjusted until the comparison result meets the preset requirement, thereby obtaining a trained YOLOv5 model.
[0066] Optionally, based on the position information, flame failure alarm information is generated, which can specifically include:
[0067] Based on the position information, the flame position is determined. Based on the flame position and the flame size, the failure level is determined. Based on the failure level, the flame failure alarm information is generated. The flame failure alarm information at least contains the failure level and the flame position information. Specifically, the alarm information can be one or more of voice information, text information, and image information, that is, different information prompt modes can be selected according to actual application scenarios. The present scheme does not make specific limitations on this.
[0068] The existing CA model focuses on selecting an input channel R, G or B, especially the channel that contributes most in the detection process, which is called channel attention. However, channel attention only focuses on the channel dimension and ignores the information contained in the spatial dimension. In order to solve this problem, as shown in formulas (1), (2) and (3), the full-channel attention is decomposed into two one-dimensional feature encoding processes along the width and height dimensions, respectively, to aggregate features along the two spatial directions.
[0069]
[0070]
[0071]
[0072] wherein Z represents a feature extraction result of a certain layer of convolution; Z H represents a feature extraction result component in the H direction; Z W represents a feature extraction result component in the W direction; i represents a feature in the H direction; j represents a feature in the W direction; h represents a one-dimensional component in the H direction, so the H-dimensional feature value of the feature point is h; w represents a one-dimensional component in the W direction, so the W-dimensional feature value of the feature point is w; x(·) represents a convolution feature extraction result.
[0073] Compared with the existing attention block that simply focuses on channel attention, the CA model achieves better detection accuracy and does not cause greater computational load, thereby reducing the detection speed. However, the trick of aggregating by direction ignores the two-dimensional relationship information. For a certain type of target, especially those with obvious common similarities, such as the flame in the study, there must be an inherent relationship between all the pixels in the picture. Usually, all the pixels in the flame are red or similar to red. Therefore, they are related in the wide spatial dimension. In order to compensate for the lack of global spatial information of the original CA block, a global average pooling layer and a global maximum pooling are added before the feature is split into two one-dimensional vectors. In addition, a large convolution kernel convolution layer is used for sampling, and the kernel size is 7.
[0074] On the other hand, the activation function used in the existing CA block is the Sigmod function. Sigmod performs well in increasing nonlinearity, but for detecting the existence of fire in a specific frame of video, the LeakyRlue function has a faster convergence speed and has a certain linearity. Therefore, the LeakyRlue function is used in the CA attention mechanism.
[0075] Figure 2 The schematic diagram of the YOLOv5 model trained in the video monitoring anomaly detection method provided by the present scheme is shown in FIG. 1. Figure 2 As shown in FIG. 1, the input data is processed by the residual layer, and the residual layer is added to the model to process the global information. The convolution layer is a large convolution kernel convolution layer, and the pooling uses average pooling and global maximum pooling. When average pooling is used, one-dimensional feature vectors in two dimensions of the width direction and the height direction are processed, and the LeakyRlue function replaces the Sigmoid function in the CA attention mechanism.
[0076] To improve the calculation speed, the block is connected to the YOLOv5 backbone next to the focus layer. The attention block can be used as a pre-processing step for the feature map. In order to carry out flame detection, a dataset is established, which contains 1089 pictures in the training set and 287 pictures in the validation set. All pictures in the dataset contain at least one flame target, and all targets have good labels. A test set is formed, which consists of 120 pictures about fire for inference. In order to verify the effect of the LeakyRule activation function, all Sigmod in the CA block is changed to LeakyRule without adding an additional pooling layer, which is symbolized as LeakyRule CA module in Table 1.
[0077] The original YOLOv5, YOLOv5 with the original CA module, LeakyRule CA module and the improved CA module in the present scheme are trained and interfered, and the detection accuracy and detection speed of the three methods are compared. The results are shown in Table 1.
[0078] Table 1: Comparison of detection results of different methods
[0079]
[0080] From the results in Table 1, it can be seen that the present scheme indeed improves the detection accuracy by about 3% compared with the original YOLOv5. However, the detection result is actually more accurate than the accuracy value shows, and the detection speed is improved by 60%, which is about 16% higher than the original CA block.
[0081] The detection effect of the original YOLOv5, YOLOv5 with the original CA module and the improved CA module in the present scheme can also be explained in combination with Figures 3-5 , the detection result of the original YOLOv5 is shown in Figure 3 , the detection result of YOLOv5 combined with the original CA module is shown in Figure 4 , and the video monitoring abnormal detection result provided by the present scheme is shown in Figure 5 . As Figures 3-5 shown, for the same picture, the scheme applied in the present scheme has a higher confidence value and better detection performance, greatly improving the detection speed. In the present scheme, the feature mapping will be easier to perform the convolution step, and replacing the LeakyRule function will result in faster convergence speed.
[0082] The performance of the present scheme is better in small target recognition, and can give a higher confidence detection result. Experiments in the ground thermal test of liquid rocket engine show that the improved CA block algorithm can send an instantaneous alarm signal when a fire fault occurs.
[0083] Based on the same idea, the embodiment of the present specification also provides a video monitoring abnormality detection device. Figure 6 The structure diagram of the video monitoring abnormality detection device provided by the present application is shown in Figure 1. Figure 6 As shown, it can include:
[0084] The video image data acquisition module 610 is configured to acquire video image data in the ground heat test of the liquid rocket engine.
[0085] The recognition module 620 is configured to input the video image data into the trained YOLOv5 model to identify whether the video image data contains a flame image, and obtain an identification result. The trained YOLOv5 model uses a CA attention mechanism, and the activation function used by the CA attention mechanism is a LeakyRlue function.
[0086] The flame position information determination module 630 is configured to determine the position information of the flame when the identification result indicates that the video image data contains a flame image. The flame pixels in the flame image are related in the spatial dimension.
[0087] The flame fault alarm information generation module 640 is configured to generate flame fault alarm information based on the position information.
[0088] Based on the device in Figure 6 , there are some specific implementation modules, which will be described below:
[0089] Optionally, the recognition module 620 can specifically include:
[0090] The preprocessing unit is configured to pre-process the video image data based on the CA attention mechanism to obtain a feature vector.
[0091] The feature vector is split into one-dimensional feature vectors in the width direction and the height direction.
[0092] Optionally, the YOLOv5 model at least includes:
[0093] An input layer, a residual layer, a convolution layer, a fully connected layer, and an output layer.
[0094] The input layer receives the video image data.
[0095] The convolution layer is configured to extract a feature vector from the video image data. The convolution kernel size of the convolution layer is 7.
[0096] The fully connected layer updates the weight to obtain a flame feature vector corresponding to the video in the ground heat test of the liquid rocket engine.
[0097] The output layer is configured to output a flame detection result according to the flame feature vector.
[0098] Optionally, the device can further include:
[0099] The global information compensation module is configured to compensate global spatial information of the CA attention mechanism by using average pooling and global maximum pooling in the residual layer.
[0100] Optionally, when the residual layer performs average pooling, one-dimensional feature vectors in the width direction and one-dimensional feature vectors in the height direction are pooled.
[0101] Optionally, the device can further include:
[0102] The sample acquisition module is configured to acquire a training sample set and a verification sample set; each image in the training sample set and the verification sample set includes at least one flame target;
[0103] The preliminary training module is configured to input the training sample set into an initial YOLOv5 model to obtain a preliminary training result.
[0104] The result comparison module is configured to compare the preliminary training result with the verification sample set to obtain a comparison result.
[0105] The parameter adjustment module is configured to adjust training parameters in the initial YOLOv5 model based on the comparison result until the comparison result meets a preset requirement, so as to obtain a trained YOLOv5 model.
[0106] Based on the same idea, the embodiments of the present specification also provide a video monitoring abnormality detection device. Figure 7 The video monitoring abnormality detection device provided by an invention can include:
[0107] The communication unit / communication interface is configured to acquire video image data in a ground thermal test process of a liquid rocket engine;
[0108] The processing unit / processor is configured to input the video image data into the trained YOLOv5 model to identify whether the video image data includes a flame image, so as to obtain an identification result; the trained YOLOv5 model uses a CA attention mechanism, and an activation function used by the CA attention mechanism is a LeakyRlue function;
[0109] When the identification result indicates that the video image data includes a flame image, position information of the flame is determined; there is a spatial dimension relationship between flame pixels in the flame image.
[0110] Based on the position information, flame failure alarm information is generated.
[0111] As shown in Figure 7 , the terminal device can further include a communication line. The communication line can include a path for transmitting information between the components.
[0112] Optionally, as shown in Figure 7 , the terminal device can further include a memory. The memory is used to store computer execution instructions for executing the scheme of the present application, and is controlled by the processor to execute. The processor is used to execute the computer execution instructions stored in the memory, so as to realize the method provided by the embodiment of the present application.
[0113] As shown in Figure 7 , the memory can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, an optical disk storage (including a compact disk, a laser disk, an optical disk, a digital versatile disk, a Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited to this. The memory can exist independently and be connected to the processor through the communication line. The memory can also be integrated with the processor.
[0114] Optionally, the computer execution instructions in the embodiment of the present application can also be referred to as application program codes, which are not specifically limited by the embodiment of the present application.
[0115] In a specific implementation, as an embodiment, as shown in Figure 7 , the processor can include one or more CPUs, such as CPU0 and CPU1 in Figure 7 .
[0116] In a specific implementation, as an embodiment, as shown in Figure 7 , the terminal device can include a plurality of processors, such as the processors in Figure 7 . Each of the processors can be a single-core processor or a multi-core processor.
[0117] Based on the same idea, the embodiments of the present specification also provide a computer storage medium corresponding to the above-mentioned embodiments, and the computer storage medium stores instructions, and when the instructions are executed, the following are achieved:
[0118] Acquiring video image data in a ground thermal test process of a liquid rocket engine;
[0119] Inputting the video image data into a trained YOLOv5 model, identifying whether the video image data contains a flame image, and obtaining an identification result; the trained YOLOv5 model uses a CA attention mechanism, and an activation function used by the CA attention mechanism is a LeakyRlue function;
[0120] When the identification result indicates that the video image data contains a flame image, determining position information of the flame; the flame pixels in the flame image have a connection in a spatial dimension;
[0121] Based on the position information, generating flame fault alarm information.
[0122] The above mainly introduces the scheme provided by the embodiments of the present application from the perspective of interaction between various modules. It can be understood that, in order to achieve the above functions, each module contains a corresponding hardware structure and / or software unit for executing each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of the examples described in the embodiments disclosed in the present text, the present application can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed by hardware or computer software driven hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0123] The embodiments of the present application can divide the functional modules according to the above-mentioned method examples, for example, each functional module can be divided according to each function, or two or more functions can be integrated in one processing module. The above-mentioned integrated module can be realized in the form of hardware or software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, and is only a logical functional division. When actually implemented, there can be another division method.
[0124] The processor in the present specification can also have the function of a memory. The memory is used to store computer execution instructions for executing the scheme of the present application, and is controlled by the processor to execute. The processor is used to execute the computer execution instructions stored in the memory, so as to realize the method provided by the embodiments of the present application.
[0125] The memory can be read-only memory (ROM) or other type of static storage devices that can store static information and instructions, random access memory (RAM) or other type of dynamic storage device that can store information and instructions, electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disk storage, optical disk storage (including compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), Blu-ray discs, etc.), magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited to this. The memory can exist independently and be connected to the processor through a communication line. The memory can also be integrated with the processor.
[0126] Optionally, the computer-executed instructions in the embodiments of the present application can also be referred to as application codes, and the embodiments of the present application do not make specific limitations thereto.
[0127] The method disclosed in the embodiments of the present application can be applied to a processor or implemented by the processor. The processor can be an integrated circuit chip with a signal processing capability. In the implementation process, the steps of the above method can be completed by the integrated logic electric circuit or the instruction of software form in the processor. The processor mentioned above can be a general processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The disclosed methods, steps and logic block diagrams in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register or other mature storage medium in the art. The storage medium is located in the memory, and the processor reads the information in the memory and combines the hardware to complete the steps of the above method.
[0128] In a possible implementation, a computer readable storage medium is provided, and the computer readable storage medium stores instructions. When the instructions are executed, the instructions are used to implement the logic operation control method and / or the logic operation reading method in the above embodiments.
[0129] In the above embodiments, the implementation can be achieved by software, hardware, firmware or any combination thereof, entirely or partially. When implemented by software, the implementation can be in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer programs or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are performed entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a terminal, a user equipment or other programmable apparatus. The computer programs or instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another computer readable storage medium, for example, the computer programs or instructions can be transferred from one website site, computer, server or data center to another website site, computer, server or data center through wired or wireless manner. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center and the like integrated with one or more available media. The available media can be a magnetic medium, such as a floppy disk, a hard disk, a magnetic tape; an optical medium, such as a digital video disc (DVD); and a semiconductor medium, such as a solid state disk (SSD).
[0130] Although the present application is described herein in conjunction with various embodiments, it is understood that other variations of the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed application, from an inspection of the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude a plurality. A single processor or other unit can fulfill the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to an advantage.
[0131] Although the present application has been described in connection with the preferred embodiments thereof with reference to the specific content thereof, it will be apparent to those skilled in the art that various modifications and changes can be made thereto without departing from the spirit and scope of the application. Accordingly, it is intended that the present application cover all such modifications and changes as fall within the scope of the application, along with all equivalents thereof. It will be understood by those within the art that, in general, terms used herein, and especially to the immediately preceding description and claims attached hereto, are intended to be given their broadest interpretation consistent with the specification and the patent statutes.
Claims
1. A video surveillance anomaly detection method, characterized by, The method is applied to ground thermal test of a liquid rocket engine, and the method comprises the following steps: Obtaining video image data in the process of ground thermal test of the liquid rocket engine; Inputting the video image data into a trained YOLOv5 model to identify whether the video image data contains a flame image, and obtaining an identification result; the trained YOLOv5 model uses a CA attention mechanism, and a Sigmoid function in the CA attention mechanism is replaced with a LeakyRlue function to detect a flame target in a specific frame of the video image data; the CA attention module is connected to a backbone of the YOLOv5 model next to a focal layer; When the identification result indicates that the video image data contains a flame image, determining position information of the flame; the flame pixels in the flame image are related in the spatial dimension; Based on the position information, generating flame fault alarm information; the step of inputting the video image data into the trained YOLOv5 model to identify whether the video image data contains a flame image and obtaining an identification result specifically comprises the following steps: Based on the CA attention mechanism, pre-processing the video image data to obtain a feature vector; Splitting the feature vector into one-dimensional feature vectors in the width direction and the height direction; average pooling and global maximum pooling are used in a residual layer of the YOLOv5 model to compensate for global spatial information of the CA attention mechanism, and the residual layer performs pooling on the one-dimensional feature vectors in the width direction and the one-dimensional feature vectors in the height direction when performing average pooling; The CA attention mechanism is decomposed into two one-dimensional feature encoding processes along the width and height dimensions to respectively gather features along the two spatial directions: wherein, represents the feature extraction result of a certain layer convolution; represents the feature extraction result component in the direction; represents the feature extraction result component in the direction; represents the feature in the direction; represents the feature in the direction; represents the feature value of all feature points in the one-dimensional component in the direction; is ; represents the feature value of all feature points in the one-dimensional component in the direction; is ; represents the convolution feature extraction result.
2. The method of claim 1, wherein, The YOLOv5 model at least comprises: An input layer, a residual layer, a convolution layer, a fully connected layer and an output layer; The input layer receives the video image data; The convolution layer is used for extracting a feature vector from the video image data; the convolution kernel size of the convolution layer is 7; The fully connected layer updates weights to obtain a flame feature vector corresponding to the video in the process of ground thermal test of the liquid rocket engine; The output layer is used for outputting a flame detection result according to the flame feature vector.
3. The method of claim 1, wherein, Before the step of inputting the video image data into the trained YOLOv5 model, the following steps are further included: Obtaining a training sample set and a verification sample set; each image in the training sample set and the verification sample set at least includes one flame target; Inputting the training sample set into an initial YOLOv5 model to obtain a preliminary training result; Comparing the preliminary training result with the verification sample set to obtain a comparison result; Based on the comparison result, adjusting training parameters in the initial YOLOv5 model until the comparison result meets a preset requirement, and obtaining the trained YOLOv5 model.
4. The method of claim 1, wherein, Based on the position information, determining a flame position; determine a fault level based on the flame position and the flame size; generate flame fault alarm information based on the fault level, the flame fault alarm information at least containing the fault level and the flame position information.
5. A video monitoring abnormality detection apparatus characterized by comprising: The device is applied to liquid rocket engine ground thermal test, and the device comprises: a video image data acquisition module configured to acquire video image data in a liquid rocket engine ground thermal test process; an identification module configured to input the video image data into a trained YOLOv5 model, identify whether the video image data contains a flame image, and obtain an identification result; the trained YOLOv5 model uses a CA attention mechanism, and a Sigmoid function in the CA attention mechanism is replaced with a LeakyRlue function to detect a flame target in a specific frame of the video image data; the CA attention module is connected to a backbone of the YOLOv5 model next to a focal layer; a flame position information determination module configured to determine flame position information when the identification result indicates that the video image data contains a flame image; the flame pixels in the flame image are related in a spatial dimension; a flame fault alarm information generation module configured to generate flame fault alarm information based on the position information; the identification module specifically comprises: a preprocessing unit configured to preprocess the video image data based on the CA attention mechanism to obtain a feature vector; the feature vector is split into one-dimensional feature vectors in a width direction and a height direction; average pooling and global maximum pooling are used in a residual layer of the YOLOv5 model to compensate for global spatial information of the CA attention mechanism, and the residual layer performs pooling on the one-dimensional feature vectors in the width direction and the one-dimensional feature vectors in the height direction when performing average pooling; the CA attention mechanism is decomposed into two one-dimensional feature encoding processes along the width and height dimensions to respectively gather features along the two spatial directions: wherein, represents the feature extraction result of a certain layer convolution; represents the feature extraction result component in the direction; represents the feature extraction result component in the direction; represents the feature in the direction; represents the feature in the direction; represents the feature value of all feature points in the one-dimensional component in the direction; ; represents the feature value of all feature points in the one-dimensional component in the direction; ; represents the convolution feature extraction result. 6. A video surveillance anomaly detection device, characterized by, The device is applied to liquid rocket engine ground thermal test, and the device comprises: a communication unit / communication interface configured to acquire video image data in a liquid rocket engine ground thermal test process; a processing unit / processor configured to input the video image data into a trained YOLOv5 model, identify whether the video image data contains a flame image, and obtain an identification result; the trained YOLOv5 model uses a CA attention mechanism, and a Sigmoid function in the CA attention mechanism is replaced with a LeakyRlue function to detect a flame target in a specific frame of the video image data; the CA attention module is connected to a backbone of the YOLOv5 model next to a focal layer; when the identification result indicates that the video image data contains a flame image, determine flame position information; the flame pixels in the flame image are related in a spatial dimension; generate a flame fault alarm information based on the position information; the video image data is input into the trained YOLOv5 model to identify whether the video image data contains a flame image, and an identification result is obtained, specifically comprising: the video image data is preprocessed based on the CA attention mechanism to obtain a feature vector; the feature vector is split into one-dimensional feature vectors in the width direction and the height direction; the residual layer of the YOLOv5 model adopts average pooling and global maximum pooling to compensate for the global spatial information of the CA attention mechanism, and when performing average pooling, the residual layer pools the one-dimensional feature vectors in the width direction and the one-dimensional feature vectors in the height direction; the CA attention mechanism is decomposed into two one-dimensional feature encoding processes along the width and height dimensions to respectively gather features along the two spatial directions: wherein, represents the feature extraction result of a certain layer convolution; represents the feature extraction result component in the direction; represents the feature extraction result component in the direction; represents the feature in the direction; represents the feature in the direction; represents the feature value of all feature points in the one-dimensional component in the direction; is ; represents the feature value of all feature points in the one-dimensional component in the direction; is ; represents the convolution feature extraction result.
7. A computer storage medium, characterized in that the computer storage medium stores instructions, and when the instructions are executed, the video monitoring anomaly detection method of any one of claims 1-4 is implemented.
Citation Information
Patent Citations
Air conditioner outdoor unit image intelligent detection method and device based on Attention + YOLOv3 and medium
CN113139945A
Real-time flame detection method and device based on improved CenterNet
CN113627284A
Fire-fighting fire source detection method and device oriented to small sample condition and storage medium
CN114140732A
Unmanned target detection method, device, equipment and medium
CN114359851A