Alarm video synthesis method

Through the time-divided and phased alarm video synthesis method, the CLIP model and event tracking model extract keyframes are used to solve the problem of insufficient computing resources and network bandwidth of the video surveillance system in intelligent management of water conservancy engineering, and efficient alarm video processing and transmission are achieved, improving the real-time and reliability of the system.

CN120281992AActive Publication Date: 2025-07-08NANJING NARI WATER RESOURCES & HYDROPOWER TECH CO LTD +1
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510756896.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

In the intelligent management of water conservancy projects, when the video surveillance system processes and transmits a large number of alarm videos, insufficient computing resources and network bandwidth leads to system performance degradation, network congestion and delays, especially in remote areas, county and township water conservancy station resources, resulting in limited decision-making delays.

Method used

The alarm video synthesis method is adopted, and the video data is processed in time segments and stages, and preliminary alarm videos are quickly synthesized during the day, and the video is improved by using idle resources at night. The CLIP model and event tracking model are used to extract keyframes, and efficient alarm videos are synthesized and transmitted.

Benefits of technology

It improves the practicality and integrity of alarm videos, ensures efficient operation of the system, reduces network bandwidth requirements, reduces packet loss rate and transmission delay, and improves the efficiency and reliability of the video surveillance system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281992A_ABST
    Figure CN120281992A_ABST
Patent Text Reader

Abstract

The invention provides an alarm video synthesis method, which relates to the technical field of computers, and comprises the following steps: S1, collecting original video data; s2, in response to the alarm instruction, starting a synthesis process of an alarm video; s3, judging that the alarm instruction is in the daytime, and synthesizing a preliminary alarm video; obtaining a key period picture, selecting a key frame from a video picture per second of the key period picture, performing event tracking processing on an image of the key frame to obtain an identified synthetic picture, and synthesizing a preliminary alarm video; s4, during the idle period at night, the preliminary alarm video is perfected; and supplementing frames, performing re-identification processing and integration on the images in the key period, and synthesizing an alarm video. According to the method, efficient operation of the system is guaranteed through a time-phased and staged video synthesis strategy, the practicability and the integrity of the alarm video are improved, and the efficiency and the reliability of the system are high. Network congestion is avoided, the transmission speed is high, the real-time performance is high, and jamming and collapse are avoided; and the system performance and the final alarm video quality are considered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method for synthesizing alarm videos. Background Art

[0002] In the field of intelligent management of water conservancy projects, the combination of video surveillance and artificial intelligence (AI) analysis technology has become a key means to improve project safety and management efficiency. Especially in scenarios such as water diversion projects, management of sluice pump stations, and slope monitoring. The sub-stations of the AI algorithm platform are deployed on the edge side of county and township-level water conservancy facilities (such as pump stations, sluice gates, and slope monitoring points), undertaking the core tasks of real-time monitoring and risk warning. However, in the actual deployment and application process, especially when dealing with complex and changeable natural environments and sudden events (such as heavy rain and floods), it is extremely easy to cause multiple cameras (such as 10 cameras) to issue alarms concurrently, instantaneously generating a large amount of video data to be processed and transmitted. Usually, these video data need to be processed into alarm videos of 10 seconds before and after the alarm event and transmitted back to the centralized control center of the main station through the network channel for review and decision-making. At this time, the following technical problems exist: 1. When generating alarm videos, the system needs to process a large amount of video data and alarm events. Video encoding and processing consume a large amount of computing resources, resulting in a decline in system performance, and even freezing or crashing.

[0003] 2. The transmission of video data occupies a large amount of network bandwidth, resulting in network congestion, increasing the delay of video transmission, and affecting the real-time nature of alarms.

[0004] 3. County and township water conservancy stations are often deployed in remote or areas with weak network infrastructure. The hardware resources of the deployed sub-stations and the transmission bandwidth with the main station and other resources are limited. The instantaneous bandwidth demand for the backhaul of alarm videos generated by multiple concurrent alarms far exceeds the available bandwidth, resulting in a soaring packet loss rate, further exacerbating the transmission delay, forming a vicious cycle, and increasing the risk of decision-making delays. Therefore, solving the problems of concurrent processing and transmission timeliness caused by insufficient resources is a technical problem that needs to be solved urgently. Summary of the Invention

[0005] The purpose of the present invention is to solve the deficiencies existing in the prior art, and a method for synthesizing alarm videos is proposed.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions: A method for synthesizing alarm videos includes the following steps: S1: Acquisition of original video data; During daily operation, video stream data is collected through cameras, and the collected video stream data is stored in a database to form original video data; S2: Respond to the alarm instruction and start the synthesis process of the alarm video; Set the duration of the alarm video. Respond to the alarm instruction and obtain the key period images from the original video data. The key period images cover the video images of half the duration of the alarm video before and after the alarm occurs; The alarm instruction includes the alarm instruction time and the alarm source. The alarm instruction time is the time when the alarm instruction is triggered, and the alarm source is the original video data collected by the camera that triggers the alarm; S3: Judge that the alarm instruction is during the day and synthesize the preliminary alarm video; Judge whether the alarm instruction is during the day or at night according to the alarm instruction time. Identify that the alarm instruction time is during the day. In each second of the video images in the key period images, only select one representative image. Through these key frames, quickly splice the video segments to synthesize the preliminary alarm video; S31: Judge whether the alarm instruction is during the day; Set the time periods for day and night. Compare the alarm instruction time with the time periods for day and night. Judge whether the alarm instruction is within the day time period. If it is within the day time period, it is determined that the alarm instruction is during the day, and then execute step S32; otherwise, the alarm instruction is at night. If the alarm instruction is at night, directly proceed to step S4; S32: Select key frames in each second of the video images in the key period images; S33: Process the key frames and quickly splice the key frames in chronological order into a video segment with the duration of the alarm video; Pre-train an event tracking model. Use the training data to train the model to obtain the event tracking model. Preprocess the images of the key frames and input them into the corresponding event tracking model to output the recognition result. The recognition result is a set of coordinates. Color the coordinates of the recognition result to obtain the synthesized picture after recognition. Splice the synthesized pictures in chronological order of each frame into a video with the duration of the alarm video; S4: Improve the preliminary alarm video during the idle period at night; During the idle period at night, complete the remaining video frames within the alarm time period. Through reprocessing and integration of the key period images, synthesize a complete alarm video with normal frame rate and smooth picture.

[0007] Further, step S1 includes the following steps: Plan the areas. Set cameras in each area as needed. There are multiple cameras, which monitor different events respectively. Each camera has an identification code, and the collected video stream data is stored as the original video data according to the identification code of the camera respectively; Build a correspondence table between cameras and monitoring events, including camera identification codes, monitoring events, and the areas where the cameras are located; Monitoring events include water accumulation, floating objects, personnel intrusion, and water level.

[0008] Further, step S32 includes: S321: Split the key-period video into multiple second-long videos by the second; S322: The CLIP model encodes the video frames to generate feature vectors; First, preprocess each frame in the second-long video, including resizing, normalizing, and converting to the Tensor format; Next, load the pre-trained CLIP model, using the transformers library or the CLIP image encoder of OpenAI; input the preprocessed second-long video in the Tensor format into the CLIP image encoder to generate feature vectors of a fixed dimension; Finally, traverse all the frames in the processed second-long video, and store the feature vectors of all the frames in the second-long video in the same list; S323: Calculate the similarity and output the key frames; Initialize an empty list as the key-frame list for storing key frames, take the first frame of the second-long video as the initial key frame, and store the feature vector of the first frame in the key-frame list. Traverse the feature vectors of the subsequent frames, calculate the cosine similarity between the feature vector of the current frame and the feature vector of the previous key frame, store the calculated cosine similarity, compare the magnitudes of all the cosine similarities, add the feature vector of the current frame with the smallest cosine similarity to the key-frame list, output the current frame as the key frame, and store it in the image format; record the frame number and timestamp of the key frame; S324: Traverse all the second-long videos, repeat steps S322 and S322 to output the key frames of each second-long video, and store them in the image format.

[0009] Further, step S33 includes: S331: Train the event tracking model; Use the deep neural network algorithm, YOLO algorithm, or OCR algorithm to train the event tracking model; install the model library, use the training data to train the model to obtain the bounding box coordinates of the recognized targets; the training data includes videos or images of various events, and the events include water accumulation, floating objects, personnel intrusion, and water level; train different event tracking models according to different events and store them in the server; S332: Preprocess the images of the key frames; S333: Input the preprocessed images into the corresponding event tracking models to obtain the recognition results; According to the alarm source, the identification code of the surveillance camera is obtained. According to the corresponding table of surveillance events, the surveillance event corresponding to the camera identification code is obtained, and the event tracking model that is the same as the surveillance event is called. The preprocessed image is input into the event tracking model to obtain the recognition result, which is the bounding box coordinates of the recognition target. S334: Coloring the recognition result to obtain a composite image; Use OpenCV to draw a bounding box on the key frame image according to the bounding box coordinates and fill the bounding box area with color; finally, store the image as a composite image; S335: splicing the synthesized images into a preliminary warning video according to the time sequence of each frame; According to the synthetic images stored in the preliminary synthetic list in the order of timestamps, create a video writer, set the display time of each synthetic image, traverse all the synthetic images in the synthetic list, write the synthetic images into the video frame by frame in the order of time to form a preliminary alarm video, and finally transmit the preliminary alarm video to the display for display.

[0010] Further, step S4 includes: S41: Video frame completion: reacquire all video frames of the key period, and remove the key frames with recorded frame numbers; S42: Video frame processing: pre-process the image of the video frame and input it into the event tracking model, and after recognition processing, color the recognition result to generate a composite image of the video frame; traverse and process all video frames to obtain the composite image of the video frame, and store it in the composite video list; S43: video encoding and processing: obtaining a composite picture from the preliminary composite list, inserting it into the composite video list in the order of timestamps, and completing the video frames in the composite video list; re-encoding and processing the completed video frames using an encoder; S44: Video synthesis: synthesizing the processed video frames in time sequence to generate a final alarm video; S45: Quality check and optimization: Perform quality check on the synthesized alarm video to ensure that the video's clarity, smoothness and other indicators meet the requirements; if there are any problems with the video quality, further optimization processing will be performed; S46: Video storage and transmission: The synthesized final alarm video is transmitted and stored in a database.

[0011] Furthermore, step S31 further includes: detecting server resources, and judging whether the server is idle according to the occupancy of each resource, and directly executing step S4 if the server is idle; otherwise, executing step S32; Server resources include bandwidth, memory, and graphics cards; the occupancy of resources refers to the usage of various resources on the server, expressed as a utilization rate; set the determination thresholds for each resource. The determination threshold for bandwidth is 60%, the determination threshold for memory is set to 70%, and the determination threshold for the graphics card is 60%. Compare the utilization rate of each resource with the corresponding determination threshold. If the utilization rate of the resource is less than the determination threshold, it is determined to be idle; otherwise, it is determined to be under high load, and step S32 is executed.

[0012] Further, in step S2, the duration of the alarm video is set to 10 seconds. According to the alarm instruction time, 5 seconds before the alarm instruction time and 5 seconds after the alarm instruction time are intercepted from the original video data of the alarm source.

[0013] Further, the filled colors include green and red; the transparency of the color is adjusted to 50%-60%.

[0014] Further, the optimization process includes removing redundant frames, setting a suitable frame rate, video encoding optimization, image quality enhancement, and noise reduction processing.

[0015] Through the above methods and steps, the system of the present invention can make full use of computing resources and network bandwidth during the idle period at night to improve the preliminary alarm video, thereby enhancing the practicability and integrity of the alarm video.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) Through the video synthesis strategy of dividing time periods and stages, the present invention not only ensures the efficient operation of the system but also maximally improves the practicability and integrity of the alarm video, which helps to enhance the efficiency and reliability of the entire video surveillance and alarm system. (2) The alarm video synthesis method of the present invention takes into account both the system performance during the peak business period during the day and ensures the quality of the final alarm video. (3) Although the frame rate of the initially synthesized video is low, it is sufficient to present the general situation of the alarm event in the first place and meet the preliminary alarm requirements. The number of video frames processed is small, which reduces the events processed by the video, ensures the real-time nature of the preliminary alarm. The low frame rate of the preliminary alarm video ensures the size of the video, avoids occupying a large amount of network bandwidth during transmission, avoids network congestion, improves the transmission speed, has low bandwidth requirements, reduces the packet loss probability, and ensures the real-time nature of the alarm video. (4) By processing in time periods and stages, the present invention ensures that when multiple tasks are processed concurrently and the server resources are highly occupied, the preliminary alarm video is obtained by processing a small number of key frames, with less computing resources occupied by the processing, ensuring the operation performance, avoiding overloading of resource occupation, and preventing situations of freezing and crashing. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flowchart of the steps of a method for synthesizing an alarm video according to the present invention. Detailed implementation mode

[0018] To further understand the purpose, structure, features and functions of the present invention, the following is a detailed description in conjunction with embodiments.

[0019] A method for synthesizing alarm videos, comprising the following steps: S1: Acquisition of original video data During daily operation, video stream data is collected through a camera and the collected video stream data is stored in a database to form original video data for retrieval and use.

[0020] When deploying, areas will be planned, and cameras will be set in each area as needed. Multiple cameras can be set, each monitoring different events; each camera has an identification code, and the identification code of the camera is unique and serves as the identity identifier of the camera. The collected video stream data is stored as original video data according to the identification codes of the cameras respectively.

[0021] Construct a correspondence table between the camera and the monitored event, including the camera identification code, the monitored event, the area where the camera is located, etc.; The monitored events include water accumulation, floating objects, personnel intrusion, water level, etc.

[0022] S2: Respond to the alarm instruction and start the synthesis process of the alarm video The duration of the alarm video is set to 10 seconds. Responding to the alarm instruction, obtain the key period images from the original video data; the key period images cover the video images of 5 seconds before the alarm occurs and 5 seconds after the alarm occurs. Covering the process of the event, through the clipped key period images, the video processing volume is reduced, the processing time is saved, and the real-time performance of the alarm video synthesis is ensured.

[0023] The alarm instruction includes the alarm instruction time, the alarm source, etc. The alarm instruction time is the time when the alarm instruction is triggered.

[0024] The alarm source is the original video data collected by the camera that triggers the alarm. According to the alarm instruction time, 5 seconds before the alarm instruction time and 5 seconds after the alarm instruction time are intercepted from the original video data of the alarm source.

[0025] S3: Judge that the alarm instruction is in the daytime and synthesize a preliminary alarm video Judge whether the alarm instruction is in the daytime or at night according to the alarm instruction time. Identify that the alarm instruction time is in the daytime. In each second of the video images in the key period images, only select one representative image, and quickly splice these key frames into a 10-second video clip to preliminarily synthesize the video; Including the following steps: S31: Determine whether the alarm instruction is during the day; Set the day as the time period from 9:00 to 18:00, the night as the time periods from 18:00 to 24:00 and from 0:00 to 9:00; compare the time of the alarm instruction with the time periods of the day and the night, and determine whether the alarm instruction is within the time period of the day. If it is within the time period of the day, it is determined that the alarm instruction is during the day, and then step S32 is executed; otherwise, the alarm instruction is at night. If the alarm instruction is at night, directly proceed to step S4.

[0026] Furthermore, detect the server resources, and based on the occupancy of each resource, determine whether the server is idle. If it is idle, directly execute step S4; otherwise, execute step S32; Server resources include bandwidth, memory, graphics cards, etc.; the occupancy of each resource is the usage of various resources on the server, expressed by the utilization rate; set the determination thresholds for each resource. For example, the determination threshold for bandwidth is 60%, the determination threshold for memory is set to 70%, and the determination threshold for the graphics card is 60%. Compare the utilization rate of each resource with the corresponding determination threshold. If the utilization rate of the resource is less than the determination threshold, it is determined to be idle; otherwise, it is determined to be high load, and step S32 is executed. Reasonably plan the computing resources, optimize the utilization rate of computing resources, ensure the stability of operation, and avoid freezing and crashing.

[0027] S32: Select key frames in each second of the video frames in the critical period; S321: Split the video frames in the critical period into multiple one - second videos by seconds; S322: The CLIP model encodes the video frames to generate feature vectors.

[0028] First, pre - process each frame in the one - second video, including resizing, normalizing, and converting to the Tensor format; Next, load the pre - trained CLIP model, which can use the transformers library or the CLIP image encoder of OpenAI; input the one - second video in the pre - processed Tensor format into the CLIP image encoder to generate feature vectors with a fixed dimension, which can be stored as 512 - dimensional. The CLIP model can capture the semantic information of the image, generate high - dimensional feature vectors, and the generated feature vectors have semantic consistency, which is suitable for cross - frame comparison.

[0029] Finally, traverse all the frames in the processed one - second video, and store the feature vectors of all the frames of the one - second video in the same list; this is convenient for comparing video frames.

[0030] S323: Calculate the similarity and output the key frames; Initialize an empty list as the key frame list to store key frames. Take the first frame of the second video as the initial key frame, and store the feature vector of the first frame in the key frame list. Traverse the feature vectors of subsequent frames, calculate the cosine similarity between the feature vector of the current frame and the feature vector of the previous key frame, according to the formula: Cosine similarity = ; where A and B are the feature vectors of the current frame and the previous key frame respectively, and are the norms of the feature vectors A and B. Store the calculated cosine similarity, compare the magnitudes of all cosine similarities, add the feature vector of the current frame with the minimum cosine similarity to the key frame list, output the current frame as a key frame, and store it in image format; record the frame number and timestamp of the key frame. Using cosine similarity can effectively measure the similarity of high-dimensional vectors, with simple and efficient calculation, meeting the requirements of large-scale frame processing and fast processing speed. By extracting key frames, the differences between key frames are ensured, redundancy is avoided, important changes in the video are retained, and the extracted key frames are more representative, ensuring the general situation of the alarm time during actual use.

[0031] The frame number of a video frame refers to the unique identifier of each frame in the video sequence, used to identify each frame in the video stream during video encoding. The timestamp is the time mark corresponding to each frame of video data.

[0032] S324: Traverse all second videos, repeat steps S322 and S322 to output the key frames of each second video, and store them in image format.

[0033] S33: Process the key frames and quickly splice the key frames in chronological order into a 10-second video clip; Pre-train an event tracking model, train the model using training data to obtain the event tracking model, preprocess the images of the key frames and input them into the corresponding event tracking model to output the recognition result, where the recognition result is a set of coordinates; color the coordinates of the recognition result to obtain the recognized composite picture, and splice the composite pictures in chronological order of each frame into a 10-second video.

[0034] Although the frame rate of this preliminary synthesized video is low, it is sufficient to present the general situation of the alarm event in the first time, meeting the preliminary alarm requirements. The number of video frames processed is small, which is conducive to the events processed by the video, ensuring the real-time nature of the preliminary alarm. The low frame rate of the preliminary alarm video ensures the size of the video, avoids occupying a large amount of network bandwidth during transmission, avoids network congestion, improves the transmission speed, and ensures the real-time nature of the alarm video.

[0035] S331: Train the event tracking model; Train an event tracking model using deep neural network algorithms, YOLO algorithms, or OCR algorithms; install a model library, and use training data to train the model to obtain the bounding box coordinates of the recognition target. The training data includes videos or images of various events, and the events are the same as the monitored events, including waterlogging, floating objects, personnel intrusion, water level, etc. According to the different events, different event tracking models are trained respectively and stored in the server. Ensure that images, texts, etc. can be tracked and recognized. The training of the tracking model is the prior art content in this field and is not the creative solution of this application, so it will not be elaborated here.

[0036] S332: Preprocessing of the key-frame image; Including processing such as resizing, normalizing, binarizing, and denoising the key-frame image; S333: Input the preprocessed image into the corresponding event tracking model to obtain the recognition result; According to the alarm source, obtain the identification code of the monitoring camera. According to the corresponding table of the monitored events, obtain the monitored events corresponding to the camera identification code, call the event tracking model identical to the monitored event, input the preprocessed image into the event tracking model, and obtain the recognition result. The recognition result is the bounding box coordinates of the recognition target. Through the alarm source, it is possible to determine the use of the event tracking model. Through multiple models, different events can be processed and recognized separately during use, achieving concurrent processing, realizing precise event processing, optimizing the utilization of resources, reducing resource waste, and through training different recognition models, the time recognition accuracy is high, avoiding misrecognition or missed recognition.

[0037] S334: Color the recognition result to obtain a composite picture; Use OpenCV to draw a bounding box on the key-frame image according to the bounding box coordinates, and fill the color (such as green, red, etc.) within the bounding box area, adjust the transparency of the color to 50%-60%, and finally store the picture as a composite picture. Using color coloring can ensure that the process of the triggered event can be intuitively displayed when the video is played, with good display effect, facilitating analysis, and strong usability.

[0038] S335: Stitch the composite pictures into a 10-second video in the chronological order of each frame; According to the composite pictures, store them in the preliminary composite list in the chronological order of the timestamps. Create a video writer, set the display time of each composite picture (set each picture to be displayed for 1 second), traverse all the composite pictures in the composite list, and write the composite pictures into the video frame by frame in the chronological order to form a 10-second preliminary alarm video. Finally, transmit the preliminary alarm video to the display for display.

[0039] S4: Improve the preliminary alarm video during the idle period at night; During the night or idle periods, the remaining video frames within the alarm time period are filled in. By reprocessing and integrating the key-period images, a complete 10-second alarm video with normal frame rate and smooth images is synthesized.

[0040] S41: Video frame filling: All video frames of the key-period images are retrieved again, and the key frames with recorded frame numbers are excluded. The remaining video frames are those not selected during the preliminary synthesis. The purpose of this step is to ensure the integrity and smoothness of the alarm video.

[0041] S42: Video frame processing: The processing method is the same as S33. The images of the video frames are preprocessed and then input into the event tracking model. After recognition processing, the recognition results are colored to generate composite pictures of the video frames. All video frames are processed iteratively to obtain the composite pictures of the video frames, which are then stored in the composite video list. S43: Video encoding and processing: The composite pictures are obtained from the preliminary synthesis list and inserted into the composite video list in timestamp order to fill in the video frames in the composite video list. An encoder is used to re-encode and process the filled video frames to ensure that the quality and frame rate of the video meet the standards.

[0042] S44: Video synthesis: The processed video frames are synthesized in chronological order to generate a complete 10-second alarm video with normal frame rate and smooth images.

[0043] S45: Quality inspection and optimization: The synthesized alarm video is subjected to quality inspection to ensure that indicators such as clarity and smoothness meet the requirements. If problems are found in the video quality, further optimization processing can be carried out. Optimization processing includes removing redundant frames, setting an appropriate frame rate, video encoding optimization, image quality enhancement, noise reduction processing, etc. Ensure that the final alarm video has high quality and ensure the smoothness and clarity of the alarm video.

[0044] S46: Video storage and transmission: The synthesized 10-second alarm video is transmitted and stored in the database.

[0045] Through the above methods and steps, the system of the present invention can make full use of computing resources and network bandwidth during the night idle period to improve the preliminary alarm video, thereby enhancing the practicality and integrity of the alarm video.

[0046] The present invention has been described by the above related embodiments. However, the above embodiments are only examples for implementing the present invention. It must be pointed out that the disclosed embodiments do not limit the scope of the present invention. On the contrary, modifications and refinements made without departing from the spirit and scope of the present invention fall within the scope of patent protection of the present invention.

Claims

1. A method for synthesizing alarm videos, characterized in that: It includes the following steps: S1: Acquisition of original video data; During daily operation, video stream data is collected through a camera, and the collected video stream data is stored in a database to form original video data; S2: Respond to the alarm instruction and start the synthesis process of the alarm video; Set the duration of the alarm video, respond to the alarm instruction, and obtain the key period pictures from the original video data; The key period pictures cover the video pictures of half the duration of the alarm video before and after the alarm occurs; The alarm instruction includes the alarm instruction time and the alarm source; The alarm instruction time is the time when the alarm instruction is triggered, and the alarm source is the original video data collected by the camera that triggers the alarm; S3: Judge that the alarm instruction is during the day and synthesize the preliminary alarm video; Judge whether the alarm instruction is during the day or at night according to the alarm instruction time. Identify that the alarm instruction time is during the day. In each second of the video pictures of the key period pictures, only select one representative picture. Through these key frames, quickly splice the video segments to synthesize the preliminary alarm video; S31: Judge whether the alarm instruction is during the day; Set the daytime period and the nighttime period, compare the alarm instruction time with the daytime period and the nighttime period, judge whether the alarm instruction is within the daytime period. If it is within the daytime period, it is determined that the alarm instruction is during the day, and then execute step S32; otherwise, the alarm instruction is at night. If the alarm instruction is at night, directly proceed to step S4; S32: Select key frames in each second of the video pictures in the key period pictures; S33: Process the key frames and quickly splice the key frames in chronological order into a video segment of the duration of an alarm video; Pre-train an event tracking model, train the model using training data to obtain the event tracking model, preprocess the images of the key frames and input them into the corresponding event tracking model, and output the recognition result, where the recognition result is a set of coordinates; color the coordinates of the recognition result to obtain the synthesized picture after recognition, and splice the synthesized pictures in chronological order of each frame into a video of the duration of an alarm video; S4: Improve the preliminary alarm video during the nighttime idle period; During the nighttime idle period, complete the remaining video frames within the alarm time period, and through reprocessing and integration of the key period pictures, synthesize a complete alarm video with normal frame rate and smooth pictures.

2. The method for synthesizing an alarm video according to claim 1, wherein: Step S1 includes the following steps: Plan the areas, set cameras in each area as needed. There are multiple cameras, each monitoring different events; each camera has an identification code, and the collected video stream data is stored as original video data according to the identification code of the camera; Construct a correspondence table between the cameras and the monitored events, including the camera identification code, the monitored event, and the area where the camera is located; The monitored events include water accumulation, floating objects, personnel intrusion, and water level.

3. The method for synthesizing an alarm video according to claim 2, characterized in that: Step S32 includes: S321: Split the key period pictures into multiple one-second videos by seconds; S322: The CLIP model encodes the video frames to generate feature vectors; First, preprocess each frame in the one-second video, including resizing, normalizing, and converting to the Tensor format; Next, load the pre-trained CLIP model and use the transformers library or OpenAI's CLIP image encoder. Input the pre-processed Tensor format second video into the CLIP image encoder to generate a fixed-dimensional feature vector. Finally, all frames in the second video are traversed and processed, and the feature vectors of all frames in the second video are stored in the same list; S323: Calculate similarity and output key frames; Initialize an empty list as a key frame list for storing key frames, take the first frame of the second video as the initial key frame, and store the feature vector of the first frame in the key frame list, traverse the feature vectors of subsequent frames, calculate the cosine similarity between the feature vectors of the current frame and the previous key frame, store the calculated cosine similarity, compare the sizes of all cosine similarities, obtain the feature vector of the current frame with the smallest cosine similarity, add it to the key frame list, output the current frame as a key frame, and store it in image format; record the frame number and timestamp of the key frame; S324: traverse all the second videos, repeat step S322 and step S322 to output the key frame of each second video, and store it in image format.

4. The method for synthesizing an alarm video according to claim 3, wherein: Step S33 includes: S331: training event tracking model; Use deep neural network algorithm, YOLO algorithm or OCR algorithm to train event tracking model; install model library, use training data to train model, and obtain bounding box coordinates of identified targets; training data includes videos or images of various events, including water accumulation, floating objects, human intrusion, and water level; train different event tracking models according to different events and store them in the server; S332: preprocessing of key frame images; S333: inputting the preprocessed image into the corresponding event tracking model to obtain a recognition result; According to the alarm source, the identification code of the surveillance camera is obtained. According to the corresponding table of surveillance events, the surveillance event corresponding to the camera identification code is obtained, and the event tracking model that is the same as the surveillance event is called. The preprocessed image is input into the event tracking model to obtain the recognition result, which is the bounding box coordinates of the recognition target. S334: Coloring the recognition result to obtain a composite image; Use OpenCV to draw a bounding box on the key frame image according to the bounding box coordinates and fill the bounding box area with color; finally, store the image as a composite image; S335: splicing the synthesized images into a preliminary warning video according to the time sequence of each frame; According to the synthetic images stored in the preliminary synthetic list in the order of timestamps, create a video writer, set the display time of each synthetic image, traverse all the synthetic images in the synthetic list, write the synthetic images into the video frame by frame in the order of time to form a preliminary alarm video, and finally transmit the preliminary alarm video to the display for display.

5. The method for synthesizing an alarm video according to claim 4, characterized in that: Step S4 includes: S41: Video frame completion: reacquire all video frames of the key period, and remove the key frames with recorded frame numbers; S42: Video frame processing: pre-process the image of the video frame and input it into the event tracking model, and after recognition processing, color the recognition result to generate a composite image of the video frame; traverse and process all video frames to obtain the composite image of the video frame, and store it in the composite video list; S43: video encoding and processing: obtaining a composite picture from the preliminary composite list, inserting it into the composite video list in the order of timestamps, and completing the video frames in the composite video list; re-encoding and processing the completed video frames using an encoder; S44: Video synthesis: synthesizing the processed video frames in time sequence to generate a final alarm video; S45: Quality check and optimization: Perform quality check on the synthesized alarm video to ensure that the video's clarity and smoothness meet the requirements; if there are any problems with the video quality, further optimization will be performed; S46: Video storage and transmission: The synthesized final alarm video is transmitted and stored in a database.

6. The method for synthesizing an alarm video according to claim 1, wherein: Step S31 also includes: detecting server resources, judging whether the server is idle according to the occupancy of each resource, and directly executing step S4 if the server is idle; otherwise, executing step S32; Server resources include bandwidth, memory, and graphics card; resource occupancy refers to the usage of various resources on the server, which is expressed as usage rate; set the judgment threshold of each resource, the judgment threshold of bandwidth is 60%, the judgment threshold of memory is set to 70%, and the judgment threshold of graphics card is 60%, compare the usage rate of each resource with the corresponding judgment threshold, if the resource usage rate is less than the judgment threshold, it is judged as idle, otherwise it is judged as high load, and execute step S32.

7. The method for synthesizing an alarm video according to claim 1, wherein: In step S2, the duration of the alarm video is set to 10 seconds, and according to the alarm instruction time, 5 seconds of video before the alarm instruction time and 5 seconds of video after the alarm instruction time are intercepted from the original video data of the alarm source.

8. The method for synthesizing an alarm video according to claim 4, wherein: The filling colors include green and red; adjust the color transparency to 50%-60%.

9. The method for synthesizing an alarm video according to claim 5, wherein: The optimization processing includes removing redundant frames, setting a suitable frame rate, optimizing video encoding, enhancing image quality, and reducing noise.

Citation Information

Patent Citations

  • Method and device for flame detection based on video image analysis

    CN104598895A

  • Video information acquisition method and device

    CN112419639A

  • Video data transmission method and system and computer readable storage medium

    CN116017038A

  • Distributed PLC edge computing system and data processing method thereof

    CN116540619A

  • Time-lapse video recording method and device based on AI

    CN119135817A