An ultra-high-definition video detection method and system

By employing frame-by-frame image processing, dynamic object recognition, and resolution optimization, the problems of resource waste and low detection accuracy in ultra-high-definition video detection have been solved, enabling efficient and accurate video surveillance and analysis.

CN120164143BActive Publication Date: 2025-11-18B&M MODERN MEDIA INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510200610.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-11-18
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

Traditional video detection methods suffer from resource waste and low detection accuracy when processing ultra-high-definition videos, especially in complex scenes and at different resolutions.

Method used

High-quality video data is generated through frame-by-frame image segmentation, effective video re-stitching, keyframe background segmentation, dynamic object recognition, scene event detection, and resolution adjustment, enabling abnormal behavior detection and resolution optimization.

Benefits of technology

It improves the accuracy and efficiency of video detection, reduces resource consumption, ensures the timely detection and handling of abnormal behavior, and enhances the reliability and adaptability of video surveillance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164143B_ABST
    Figure CN120164143B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image processing, in particular to an ultra-high-definition video detection method and system. The method comprises the following steps: acquiring a source video; performing frame-by-frame image cutting on the source video to obtain frame-by-frame images of the source video; performing effective video re-splicing on the frame-by-frame images of the source video to generate effective videos; performing key frame image background segmentation on the effective videos to generate key frame foreground object images; performing dynamic object identification on the key frame foreground object images to generate dynamic object identification data; performing dynamic object marking on the effective videos according to the dynamic object identification data to generate dynamic object labels; and performing scene event detection on the effective videos by using the dynamic object labels to generate a scene event image set. Through frame-by-frame image processing, accurate dynamic object identification, abnormal behavior detection and resolution optimization, the application improves the accuracy and detection performance of detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an ultra-high-definition video detection method and system. Background Technology

[0002] Early video technologies primarily focused on standard definition (SD) and high definition (HD), with SD typically offering 480p resolution and HD offering 720p and 1080p resolutions. As demand for image quality increased, ultra-high definition (UHD) video technologies with 4K (3840x2160 pixels) and 8K (7680x4320 pixels) resolutions gradually became mainstream. High-definition television (HDTV) technology was gradually adopted, but its testing methods remained focused on basic frame rate, resolution, and color gamut. With the widespread adoption of 4K televisions in the mid-2010s, video testing methods became more complex, encompassing metrics such as color accuracy, dynamic range (HDR), and frame rate. To ensure video quality, industry standards such as the Rec.2020 color space and HDR10 were introduced, leading to more refined testing methods. In recent years, the introduction of artificial intelligence and machine learning technologies has significantly improved the accuracy and efficiency of video testing. Deep learning models, especially convolutional neural networks (CNNs), are used in fields such as video frame quality assessment, noise detection and removal, and edge sharpening. However, traditional video detection often relies on preset rules to identify abnormal behavior and lacks the ability to dynamically detect events in complex scenes. At the same time, when processing videos of different resolutions, there are often problems of wasted resources or poor processing results, which leads to low detection accuracy and performance. Summary of the Invention

[0003] Therefore, it is necessary to provide an ultra-high-definition video detection method and system to solve at least one of the above-mentioned technical problems.

[0004] To achieve the above objectives, an ultra-high-definition video detection method is provided, the method comprising the following steps:

[0005] Step S1: Obtain the source video; perform frame-by-frame image segmentation on the source video to obtain frame-by-frame images of the source video; perform effective video re-segmentation on the frame-by-frame images of the source video to generate an effective video;

[0006] Step S2: Perform background segmentation on keyframe images of the valid video to generate keyframe foreground object images; perform dynamic object recognition on the keyframe foreground object images to generate dynamic object recognition data; mark dynamic objects in the valid video based on the dynamic object recognition data to generate dynamic object labels.

[0007] Step S3: Use dynamic object tags to perform scene event detection on valid videos to generate a scene event image set; perform abnormal behavior detection on the scene event image set to generate abnormal behavior detection data; compare the abnormal behavior detection data with the preset standard abnormal behavior detection threshold to generate an abnormal scene image set and a normal scene image set.

[0008] Step S4: Obtain resolution resource data; adjust the resolution performance of the normal scene image set and the abnormal scene image set according to the resolution resource data to generate a normal scene low-resolution image set and an abnormal scene high-resolution image set; reconstruct the video from the normal scene low-resolution image set and the abnormal scene high-resolution image set to generate a resolution-optimized video for performing ultra-high-definition video detection.

[0009] This invention allows for independent processing of each frame through frame-by-frame image segmentation, reducing errors in overall video processing and improving detection accuracy. Effective video re-stitching removes irrelevant or invalid parts from the source video, generating higher-quality, effective video and providing more reliable input data for subsequent analysis. Background segmentation extracts foreground object images from keyframes, making subsequent object recognition more accurate and efficient. Dynamic object recognition accurately identifies and labels dynamic elements in the video, improving the accuracy of object tracking and scene analysis. Scene event detection identifies important events or scene changes in the video, enhancing the monitoring capability of specific activities. Abnormal behavior detection identifies and labels abnormal behavior, helping to quickly discover and respond to potential problems or security risks, ensuring the effectiveness of video surveillance. By adjusting the resolution of normal and abnormal scenes, resource usage is optimized, reducing the processing burden on normal scenes while improving the processing clarity of abnormal scenes. Resolution optimization ensures detail preservation in abnormal scenes at high resolution, enhancing the detectability of abnormal behavior and improving the overall accuracy and efficiency of video detection. Therefore, this invention improves the accuracy and performance of detection through frame-by-frame image processing, accurate dynamic object recognition, abnormal behavior detection, and resolution optimization.

[0010] Preferably, step S1 includes the following steps:

[0011] Step S11: Acquire the source video using a recording device;

[0012] Step S12: Perform frame-by-frame image segmentation on the source video based on the preset timestamp to obtain frame-by-frame images of the source video;

[0013] Step S13: Overlay adjacent frames of the source video frame by frame to generate adjacent frame overlap data;

[0014] Step S14: Effective video re-segmentation is performed on each frame of the source video using overlapping data from adjacent frames to generate an effective video.

[0015] This invention utilizes preset timestamps to segment source video frame by frame, ensuring accurate segmentation of each frame and providing high-quality image data for subsequent processing. The adjacent frame overlap step ensures continuity and consistency between frames, facilitating subsequent video re-stitching and resulting in a smoother, more natural re-stitched video. By processing the overlapping data of adjacent frames, the source video can be effectively re-stitched frame by frame, removing redundant and invalid frames to generate a more concise and useful video. These processing steps significantly improve the overall video quality, resulting in a final video with high clarity and coherence, while also reducing jitter and other common visual imperfections. The concise video generated by effective video re-stitching not only improves quality but also optimizes file size, facilitating storage and transmission and reducing storage space and bandwidth consumption.

[0016] Preferably, step S14 includes the following steps:

[0017] Step S141: Perform grayscale conversion on each frame of the source video to generate grayscale images of each frame of the source video;

[0018] Step S142: Perform feature point detection on each frame of the grayscale image of the source video to obtain frame image feature points; perform feature description on the frame image feature points to generate frame image feature descriptors;

[0019] Step S143: Perform image feature point matching on each frame of grayscale image of the source video using frame image feature descriptors to generate frame-by-frame image feature point matching data; perform image deviation analysis on each frame of grayscale image of the source video using overlapping data of adjacent frame images to generate frame-by-frame image deviation data.

[0020] Step S144: Use frame-by-frame image feature point matching data and frame-by-frame image deviation data to remove invalid frame images from the source video frame-by-frame images to obtain valid video frame images; perform time-series frame-by-frame stitching on the valid video frame images to generate a valid video.

[0021] This invention simplifies image data and reduces computational complexity by performing grayscale conversion on each frame of the source video, while retaining important visual information, providing a more efficient data foundation for subsequent feature point detection. Utilizing feature point detection and feature descriptor generation techniques, key features in each frame can be accurately captured, generating rich frame image feature descriptors and providing a reliable basis for feature point matching. Image feature point matching using frame image feature descriptors ensures accurate matching between adjacent frames, improving matching accuracy and generating reliable frame-by-frame image feature point matching data. Image deviation analysis of each frame's grayscale image using overlapping data from adjacent frames effectively identifies deviation information between images, generating frame-by-frame image deviation data and providing a precise basis for invalid frame removal. Combining image feature point matching data and image deviation data, invalid frames in the source video can be removed, retaining valid frames. A coherent and valid video is then generated by sequentially stitching each frame, improving the overall video quality. These processing steps significantly improve the video's coherence and stability, reducing visual jitter and frame skipping, resulting in a smoother and more natural final video. The resulting video file after effective video stitching is more concise, occupies less storage space, and is easier to store and transmit, reducing bandwidth and storage resource consumption. The entire video processing is automated, employing efficient algorithms from grayscale conversion, feature point detection, and feature point matching to invalid frame removal and temporal stitching, significantly improving the efficiency and accuracy of video processing.

[0022] Preferably, step S2 includes the following steps:

[0023] Step S21: Extract keyframe images from the valid video to obtain keyframe images;

[0024] Step S22: Perform background segmentation on the keyframe image to generate the foreground object image of the keyframe;

[0025] Step S23: Perform object localization and detection on the foreground object image of the keyframe to generate object localization data; perform dynamic object recognition on the effective video using the object localization data to generate dynamic object recognition data.

[0026] Step S24: Mark dynamic objects in the valid video based on the dynamic object recognition data and generate dynamic object tags.

[0027] This invention extracts keyframe images from effective videos, extracting representative and important frames that contain the main dynamic information, providing simplified and efficient foundational data for subsequent processing. Background segmentation of the keyframe images effectively separates foreground objects from background information, making foreground objects stand out and providing clear targets for object localization and detection. Object localization detection of foreground objects in keyframe images generates object localization data that accurately identifies the position and boundaries of objects in the video, laying the foundation for subsequent dynamic object recognition. Using object localization data for dynamic object recognition in effective videos allows tracking of dynamic objects, generating detailed dynamic object recognition data, and improving the accuracy and reliability of video analysis. Dynamic object tagging of effective videos based on dynamic object recognition data identifies all detected dynamic objects, generating intuitive dynamic object labels for easy observation and analysis. Through automated keyframe extraction, background segmentation, object localization and detection, and dynamic object recognition, the efficiency of video analysis can be significantly improved, reducing manual intervention and quickly obtaining high-quality analysis results. Dynamic object tagging adds more descriptive information to the video, making the video content richer and more intuitive, facilitating understanding and application.

[0028] Preferably, step S24 includes the following steps:

[0029] Step S241: Perform inter-frame difference analysis on the effective video based on the dynamic object recognition data to obtain dynamic object motion region data; perform region segmentation on the effective video based on the dynamic object motion region data to obtain dynamic object motion region image.

[0030] Step S242: Use dynamic object recognition data to perform edge feature detection on the image of the moving area of ​​the dynamic object, and generate edge feature data, wherein the edge feature detection includes shape detection, texture analysis and color extraction;

[0031] Step S243: Perform optical flow tracing on the motion region image of the dynamic object based on edge feature data to generate dynamic object motion characteristic data; divide the effective video into dynamic object images according to the dynamic object motion characteristic data to generate live object images and non-live object images.

[0032] Step S244: Dynamically label the effective video using live and non-live images to generate dynamic object tags.

[0033] This invention utilizes dynamic object recognition data to perform inter-frame differencing on valid video, accurately locating the motion region of dynamic objects. This significantly improves the accuracy of region segmentation, resulting in clear images of the motion regions of dynamic objects. Edge feature detection is then performed on these motion region images. Through shape detection, texture analysis, and color extraction, comprehensive edge feature data is generated. This data more accurately describes the shape, texture, and color of objects, improving object recognition accuracy. Optical flow tracing based on the edge feature data effectively captures the motion characteristics of dynamic objects, generating detailed motion characteristic data that reflects the object's trajectory and speed changes, providing deeper motion analysis. Using this motion characteristic data to segment valid video images distinguishes between living and non-living objects. This segmentation facilitates further analysis and processing of different types of dynamic objects, improving the accuracy and detail of video analysis. Labeling valid video with both living and non-living images generates accurate tags for each dynamic object in the video. These dynamic object tags include not only location and boundary information but also motion characteristics and type information, making the labeling more comprehensive and detailed. Automated inter-frame differencing, region segmentation, edge feature detection, and optical flow tracing significantly improve video processing efficiency, reduce manual intervention, and enable fast and accurate video analysis. Dynamic object tagging contains rich feature information and motion data, making video content display more intuitive and detailed, facilitating user understanding and application.

[0034] Preferably, step S3 includes the following steps:

[0035] Step S31: Use dynamic object tags to perform scene event detection on valid videos, thereby generating a scene event image set;

[0036] Step S32: Use scene event detection data to perform image behavior detection on the scene event image set and generate behavior detection data;

[0037] Step S33: Construct an anomaly detector; use the anomaly detector to detect abnormal behavior in the behavior detection data and generate abnormal behavior detection data;

[0038] Step S34: Compare the abnormal behavior detection data with the preset standard abnormal behavior detection threshold. When the abnormal behavior detection data is greater than or equal to the preset standard abnormal behavior detection threshold, mark the scene event image set as abnormal scene events based on the abnormal behavior detection data to generate an abnormal scene image set. When the abnormal behavior detection data is less than the preset standard abnormal behavior detection threshold, mark the scene event image set as normal scene events based on the abnormal behavior detection data to generate a normal scene image set.

[0039] This invention utilizes dynamic object tags to perform scene event detection on valid videos. It extracts scenes with specific events from the video, generating a scene event image set. This step effectively filters out key event scenes in the video, providing foundational data for subsequent behavior detection. By performing image behavior detection on the scene event image set, various behaviors in the images can be identified and analyzed, generating detailed behavior detection data. This behavior detection data reflects dynamic changes and behavioral characteristics within the scene, providing deeper behavior analysis. An anomaly detector is constructed to detect abnormal behaviors based on the behavior detection data, identifying potential abnormal behaviors. This process efficiently and accurately detects abnormal behaviors in the video, generating abnormal behavior detection data. Comparing the abnormal behavior detection data with a preset standard abnormal behavior detection threshold distinguishes between normal and abnormal behaviors. When the detection data is greater than or equal to the preset threshold, the scene event image set is labeled with abnormal scene events based on the abnormal behavior detection data, generating an abnormal scene image set; otherwise, a normal scene image set is generated. This step accurately labels abnormal and normal scene events in the video, improving the accuracy of abnormal behavior recognition. Automated scene event detection, behavior detection, and abnormal behavior detection steps can significantly improve video processing efficiency, reduce manual intervention, and achieve fast and accurate video analysis and labeling. Precise abnormal behavior detection and labeling can significantly enhance the security and reliability of video surveillance systems, enabling timely detection and handling of abnormal behavior and preventing potential security risks. The image sets of both abnormal and normal scenes contain rich behavioral and event information, making video analysis more comprehensive and detailed, facilitating user understanding and application.

[0040] Preferably, step S31 includes the following steps:

[0041] Step S311: Use dynamic object tags to classify the effective video into object regions and generate effective object region videos, which include biological region videos and non-biological region videos.

[0042] Step S312: Perform behavior recognition on the biological region video to generate biological behavior recognition data; perform pose tracking on the biological region video using the biological behavior recognition data to generate biological behavior motion trajectory images; perform motion pattern analysis on the biological behavior motion trajectory images to generate biological region motion pattern data.

[0043] Step S313: Perform temporal localization on the video of the abiotic region to generate temporal localization data of the abiotic region; use the temporal localization data of the abiotic region to perform motion trajectory analysis on the video of the abiotic region to generate abiotic motion trajectory image; perform motion path analysis on the abiotic motion trajectory image to generate abiotic motion path data.

[0044] Step S314: Separate scene event images from the video of the effective object area using biological region motion pattern data and non-biological region movement path data to generate a scene event image set.

[0045] This invention utilizes dynamic object tags to classify object regions in valid videos, accurately separating biological and non-biological regions to generate valid object region videos. This step significantly improves the accuracy and efficiency of subsequent analysis. Behavior recognition is performed on biological region videos to generate biological behavior recognition data, allowing for detailed recording and analysis of various biological behaviors. Further, posture tracking generates biological behavior trajectory images, and motion pattern analysis produces biological region motion pattern data. This data provides in-depth analysis of biological behavior, offering strong support for biological behavior research and applications. Temporal localization is performed on non-biological region videos to generate non-biological region temporal localization data. Movement trajectory analysis and movement path analysis generate non-biological movement trajectory images and non-biological region movement path data. These analyses help understand changes and movement patterns in non-biological regions, providing strong evidence for the monitoring and prediction of non-biological behavior. Scene event image separation is performed on valid object region videos using biological region motion pattern data and non-biological region movement path data, generating detailed scene event image sets. This step separates complex scene event images, laying the foundation for subsequent scene event detection and analysis. Automated object region classification, behavior recognition, temporal localization, and movement path analysis significantly improve video processing efficiency, reduce manual intervention, and enable rapid and accurate video analysis. Scene event image sets contain rich behavioral and movement information of both biological and non-biological regions, making video analysis more comprehensive and detailed, facilitating user understanding and application. Precise biological behavior recognition and non-biological movement path analysis can significantly enhance the security and reliability of video surveillance systems, enabling timely detection and handling of abnormal behaviors and events.

[0046] Preferably, step S33 includes the following steps:

[0047] Step S331: Divide the behavior detection data into a dataset to obtain a model training set and a model test set;

[0048] Step S332: Train the model on the training set using the Isolation Forest algorithm to generate an anomaly detection detector; use the test set to iterate and optimize the anomaly detection detector to generate an anomaly detector.

[0049] Step S333: Import the behavior detection data into the anomaly detector to detect abnormal behavior and generate abnormal behavior detection data.

[0050] This invention divides the behavior detection data into a training set and a test set, ensuring a reasonable allocation of training and testing data and improving the efficiency of model training and the accuracy of testing. Using the Isolation Forest algorithm to train the model on the training set enables efficient detection of anomalous behavior in the data. The Isolation Forest algorithm exhibits high performance when processing high-dimensional data and can accurately identify anomalies. Optimizing and iterating the anomaly detector using the test set continuously improves its performance, making it more accurate and reliable when processing real-world data. This optimization and iteration process effectively enhances the model's detection capability and robustness. Importing the behavior detection data into the anomaly detector generates accurate anomaly detection data. The trained anomaly detector can efficiently identify anomalous behavior in videos, supporting timely handling of abnormal events. Through the Isolation Forest algorithm and the optimization and iteration process, the accuracy of anomaly detection can be significantly improved, reducing the possibility of false positives and false negatives, and ensuring the reliability of detection results. Automated anomaly detection steps can significantly improve the efficiency of video surveillance, reduce manual intervention, and achieve rapid and accurate anomaly behavior identification and processing. Accurate anomaly detection can significantly improve the security and reliability of video surveillance systems, enabling timely identification and handling of potential security risks and ensuring environmental safety. Anomaly detection data contains detailed information on behavioral anomalies, making video analysis more comprehensive and detailed, and easier for users to understand and apply.

[0051] Preferably, step S4 includes the following steps:

[0052] Step S41: Obtain resolution resource data;

[0053] Step S42: Perform a first resolution performance adjustment on the normal scene image set based on the resolution resource data to generate a normal scene low-resolution image set; perform a second resolution performance adjustment on the abnormal scene image set based on the resolution resource data to generate an abnormal scene high-resolution image set.

[0054] Step S43: Perform video reconstruction on the low-resolution image set of normal scenes and the high-resolution image set of abnormal scenes to generate a resolution-optimized video for ultra-high-definition video detection.

[0055] This invention, by acquiring resolution resource data and adjusting the resolution of the image set according to actual needs, can effectively utilize system resources, avoid unnecessary energy consumption, and improve video processing efficiency. Adjusting the resolution of normal scene image sets to low resolution and abnormal scene image sets to high resolution based on the resolution resource data helps save resources in normal scenes while providing higher image detail in abnormal scenes, ensuring the accuracy of anomaly detection. Video reconstruction using the low-resolution image set of normal scenes and the high-resolution image set of abnormal scenes generates a resolution-optimized video. This approach balances the needs of normal monitoring and anomaly detection, improving overall video quality and detection effectiveness. Using low-resolution images in normal scenes helps reduce the computational load and energy consumption of video processing, extending equipment lifespan and improving the overall performance and efficiency of the system. Using high-resolution images in abnormal scenes provides more image detail, improving the accuracy and reliability of anomaly detection, ensuring timely detection and handling of abnormal events. The generated resolution-optimized video can be used for ultra-high-definition video inspection operations, meeting the needs of high-quality video monitoring and analysis, and providing higher resolution and clearer image quality for video inspection systems. By adjusting the resolution of normal scene image sets to low resolution, significant storage space can be saved. Simultaneously, storing abnormal scene images at high resolution ensures the integrity and availability of critical data. Dynamic adjustment of resolution resource data allows for flexible adjustment of image resolution according to different scenario requirements, enhancing the system's adaptability and flexibility to meet the needs of various application scenarios. Through reasonable resolution adjustment and video reconstruction, a high-quality video can be generated, providing a reliable foundation for subsequent monitoring, analysis, and detection.

[0056] This specification provides an ultra-high-definition video detection system for performing the above-described ultra-high-definition video detection method. The ultra-high-definition video detection system includes:

[0057] The effective video filtering module is used to acquire source video; to segment the source video frame by frame to obtain source video frame by frame images; and to re-sew the source video frame by frame images into effective video to generate effective video.

[0058] The dynamic object recognition module is used to perform background segmentation of keyframe images in valid videos to generate keyframe foreground object images; to perform dynamic object recognition on keyframe foreground object images to generate dynamic object recognition data; and to mark dynamic objects in valid videos based on dynamic object recognition data to generate dynamic object labels.

[0059] The event analysis module is used to detect scene events in valid videos using dynamic object tags, thereby generating a scene event image set; to detect abnormal behavior in the scene event image set, generating abnormal behavior detection data; and to compare the abnormal behavior detection data with a preset standard abnormal behavior detection threshold, generating an abnormal scene image set and a normal scene image set.

[0060] The resolution adjustment module is used to acquire resolution resource data; adjust the resolution performance of normal scene image sets and abnormal scene image sets according to the resolution resource data to generate normal scene low-resolution image sets and abnormal scene high-resolution image sets; reconstruct the video from the normal scene low-resolution image sets and abnormal scene high-resolution image sets to generate resolution-optimized video for performing ultra-high-definition video detection.

[0061] The beneficial effects of this invention lie in ensuring that only clear and effective video segments are selected through frame-by-frame image segmentation and effective video re-stitching, eliminating useless or redundant data. This improves the efficiency and accuracy of subsequent processing and ensures the reliability of the analysis results. Through keyframe background segmentation and dynamic object recognition, dynamic objects in the video can be accurately detected and labeled. The generated dynamic object tags provide a clear basis for subsequent event analysis, improving the accuracy and efficiency of dynamic object recognition. Scene event detection using dynamic object tags generates a scene event image set containing key event information. Abnormal behavior detection further refines the distinction between abnormal and normal scenes, ensuring the system can promptly detect and handle abnormal situations, enhancing the system's anomaly detection capabilities. Resolution performance is adjusted for normal and abnormal scenes based on resolution resource data. Normal scenes use low-resolution image sets to reduce resource consumption, while abnormal scenes use high-resolution image sets to provide more detailed information, ensuring high-quality video monitoring and analysis. Video reconstruction using low-resolution image sets for normal scenes and high-resolution image sets for abnormal scenes generates resolution-optimized videos that simultaneously meet the needs of resource optimization and high-definition detection. This optimization process provides clear video quality for ultra-high-definition video detection operations, improving the accuracy and effectiveness of detection. By processing normal scenes at low resolution, computational and storage resources are saved, while processing abnormal scenes at high resolution ensures the quality of critical data. This layered resolution adjustment effectively balances resource usage and data quality, reducing system load. The system comprehensively utilizes video filtering, dynamic object recognition, event analysis, and resolution adjustment technologies to enhance overall video monitoring and analysis capabilities, ensuring reliable analysis results across various scenarios. This system is applicable to multiple fields such as security monitoring, behavior analysis, and abnormal event detection. It can adjust video processing strategies according to different needs and scenarios, improving the intelligence and accuracy of various applications. Through efficient video processing and optimization, users can obtain clearer and more accurate video data, improving their understanding and application of monitored content, and enhancing user experience. The system can flexibly adjust processing strategies according to different video scenes and resolution requirements, adapting to changes in different environments and application scenarios, improving system adaptability and flexibility. Therefore, this invention improves detection accuracy and performance through frame-by-frame image processing, precise dynamic object recognition, abnormal behavior detection, and resolution optimization. Attached Figure Description

[0062] Figure 1 This is a flowchart illustrating the steps of an ultra-high-definition video detection method.

[0063] Figure 2 for Figure 1 A detailed flowchart illustrating the implementation steps of step S2.

[0064] Figure 3for Figure 1 A detailed flowchart illustrating the implementation steps of step S3.

[0065] Figure 4 for Figure 1 A detailed flowchart illustrating the implementation steps of step S4.

[0066] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0067] The technical method of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this invention.

[0068] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.

[0069] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0070] To achieve the above objectives, please refer to Figures 1 to 4 A method for detecting ultra-high-definition video, the method comprising the following steps:

[0071] Step S1: Obtain the source video; perform frame-by-frame image segmentation on the source video to obtain frame-by-frame images of the source video; perform effective video re-segmentation on the frame-by-frame images of the source video to generate an effective video;

[0072] Step S2: Perform background segmentation on keyframe images of the valid video to generate keyframe foreground object images; perform dynamic object recognition on the keyframe foreground object images to generate dynamic object recognition data; mark dynamic objects in the valid video based on the dynamic object recognition data to generate dynamic object labels.

[0073] Step S3: Use dynamic object tags to perform scene event detection on valid videos to generate a scene event image set; perform abnormal behavior detection on the scene event image set to generate abnormal behavior detection data; compare the abnormal behavior detection data with the preset standard abnormal behavior detection threshold to generate an abnormal scene image set and a normal scene image set.

[0074] Step S4: Obtain resolution resource data; adjust the resolution performance of the normal scene image set and the abnormal scene image set according to the resolution resource data to generate a normal scene low-resolution image set and an abnormal scene high-resolution image set; reconstruct the video from the normal scene low-resolution image set and the abnormal scene high-resolution image set to generate a resolution-optimized video for performing ultra-high-definition video detection.

[0075] This invention allows for independent processing of each frame through frame-by-frame image segmentation, reducing errors in overall video processing and improving detection accuracy. Effective video re-stitching removes irrelevant or invalid parts from the source video, generating higher-quality, effective video and providing more reliable input data for subsequent analysis. Background segmentation extracts foreground object images from keyframes, making subsequent object recognition more accurate and efficient. Dynamic object recognition accurately identifies and labels dynamic elements in the video, improving the accuracy of object tracking and scene analysis. Scene event detection identifies important events or scene changes in the video, enhancing the monitoring capability of specific activities. Abnormal behavior detection identifies and labels abnormal behavior, helping to quickly discover and respond to potential problems or security risks, ensuring the effectiveness of video surveillance. By adjusting the resolution of normal and abnormal scenes, resource usage is optimized, reducing the processing burden on normal scenes while improving the processing clarity of abnormal scenes. Resolution optimization ensures detail preservation in abnormal scenes at high resolution, enhancing the detectability of abnormal behavior and improving the overall accuracy and efficiency of video detection. Therefore, this invention improves the accuracy and performance of detection through frame-by-frame image processing, accurate dynamic object recognition, abnormal behavior detection, and resolution optimization.

[0076] In this embodiment of the invention, reference Figure 1 The above is a schematic diagram of the steps of an ultra-high-definition video detection method according to the present invention. In this example, the ultra-high-definition video detection method includes the following steps:

[0077] Step S1: Obtain the source video; perform frame-by-frame image segmentation on the source video to obtain frame-by-frame images of the source video; perform effective video re-segmentation on the frame-by-frame images of the source video to generate an effective video;

[0078] In this embodiment of the invention, the acquisition method of the source video is determined. Specifically, the source video is recorded by a camera device, extracted from a database, or downloaded from the Internet. The format and quality of the source video are ensured to meet processing requirements. Video processing software (such as Adobe Premiere Pro, Final Cut Pro, etc.) or programming tools (such as Python's OpenCV library) are used to parse the source video frame by frame. Frame-by-frame parsing refers to extracting each frame from the video and saving it as a separate image file. The steps of frame-by-frame image segmentation include: opening the source video file in the video processing software, setting parsing parameters, such as the frame rate (how many frames to extract per second), and specifically implementing frame-by-frame parsing through manual settings or automated tools. Each parsed frame image is saved to a designated folder, with filenames numbered sequentially to ensure the frame order is not disrupted. The frame-by-frame images are then filtered and processed, removing invalid or poor-quality frames and retaining only valid frames. The criteria for determining valid frames specifically include image sharpness, brightness, contrast, and consistency of frame content. The steps for selecting valid frames include: inspecting each frame one by one, removing blurry, overly dark, overly bright, or other interfering images, and performing necessary image processing on the valid frames, such as noise reduction, contrast enhancement, and brightness adjustment. The steps for re-stitching the video include: importing all selected valid frame images into the video processing software, setting the frame rate as needed (usually using the original video's frame rate), stitching all valid frames sequentially into a new video file, and selecting a suitable video format and encoding method when exporting.

[0079] Step S2: Perform background segmentation on keyframe images of the valid video to generate keyframe foreground object images; perform dynamic object recognition on the keyframe foreground object images to generate dynamic object recognition data; mark dynamic objects in the valid video based on the dynamic object recognition data to generate dynamic object labels.

[0080] In this embodiment of the invention, keyframes are extracted from effective videos using video analysis technology. A keyframe refers to a representative and important frame in the video. The method for selecting keyframes is specifically based on scene changes, time intervals, or content analysis. Image processing techniques are used to perform background segmentation on the keyframes. Background segmentation refers to separating foreground objects from background regions to generate an image containing only the foreground objects. Common methods include: color-based segmentation: distinguishing foreground objects from the background using color features; motion-based segmentation: extracting moving objects as foreground using inter-frame motion information; and deep learning methods: using pre-trained convolutional neural networks (such as Mask R-CNN) for foreground and background segmentation. The segmented background image is saved as a keyframe foreground object image, with the filename corresponding to the original keyframe to ensure the order of the foreground object images and consistency with the video frames. Dynamic object recognition algorithms are used to analyze the keyframe foreground object images to identify dynamic objects in the image. Common methods include edge detection and feature point matching for identifying dynamic objects. The recognition result for each foreground object is recorded, including the object's category, location (boundary box coordinates), and other information. These recognition results are saved as dynamic object recognition data, specifically in formats such as JSON and XML, containing object information for each frame. The generated dynamic object recognition data is imported into video processing software or programming tools to prepare for labeling valid videos. Based on the recognition data, each identified dynamic object in the video is labeled. The labeling method involves drawing bounding boxes of the recognized objects in the video frames, adding object category labels next to the bounding boxes, and, for consecutive frames, drawing the motion trajectories of the objects. The labeled valid videos are then exported, generating video files containing dynamic object labels. This labeled video can be used for further analysis, demonstration, or training of machine learning models.

[0081] Step S3: Use dynamic object tags to perform scene event detection on valid videos to generate a scene event image set; perform abnormal behavior detection on the scene event image set to generate abnormal behavior detection data; compare the abnormal behavior detection data with the preset standard abnormal behavior detection threshold to generate an abnormal scene image set and a normal scene image set.

[0082] In this embodiment of the invention, valid videos containing dynamic object tags are imported into video processing software or programming tools. Based on the dynamic object tags in the video, events occurring in each video frame or frame sequence are analyzed. The scene event detection method includes: pre-defining certain scene events (such as an object entering a specific area, an object stopping movement, etc.), and detecting the occurrence of these events through programming logic. A model is trained to recognize specific scene events, and event detection is performed by inputting dynamic object tags and video frame data into the model. The video frames corresponding to the detected scene events are saved as image files, forming a scene event image set. Image filenames are ensured to be numbered according to event time or sequence for easy subsequent processing. A model suitable for abnormal behavior detection is selected or trained, specifically based on traditional methods (such as statistical analysis, pattern recognition) or deep learning methods (such as convolutional neural networks, recurrent neural networks). The generated scene event image set is input into the abnormal behavior detection model for image analysis. Based on the model's output, whether the behavior in each image is abnormal is recorded. The results of abnormal behavior detection should include information such as image filename, abnormal score, or label. The detection results are saved as abnormal behavior detection data, specifically in JSON, XML, or CSV format, recording the detection results for each image in detail. Read the abnormal behavior detection data and the preset standard abnormal behavior detection threshold. The threshold is a fixed value, or it can be dynamically adjusted based on historical data. Compare the anomaly score or label of each image with the threshold. If the score is higher than the threshold, the image is considered to contain abnormal behavior; otherwise, the image is considered normal. Based on the comparison results, save the images in the scene event image set to the abnormal scene image set and the normal scene image set respectively. Ensure that the image file name and the detection result are consistent for subsequent processing and analysis.

[0083] Step S4: Obtain resolution resource data; adjust the resolution performance of the normal scene image set and the abnormal scene image set according to the resolution resource data to generate a normal scene low-resolution image set and an abnormal scene high-resolution image set; reconstruct the video from the normal scene low-resolution image set and the abnormal scene high-resolution image set to generate a resolution-optimized video for performing ultra-high-definition video detection.

[0084] In this embodiment of the invention, resolution resource data specifically includes image resolution specifications (e.g., 720p, 1080p, 4K, etc.) and resource parameters required for video reconstruction (e.g., computing power, storage space, etc.). Resolution resource data is obtained by querying a resource management system or a preset configuration file. Images in the normal scene image set are adjusted to a lower resolution, specifically by using image processing tools (e.g., OpenCV) to scale the images and reduce their resolution to a specified low resolution specification. Images in the abnormal scene image set are adjusted to a higher resolution. Common methods include: using super-resolution algorithms (e.g., SRCNN, ESRGAN) to reconstruct images at high resolution. Low-resolution normal scene images and high-resolution abnormal scene images are recombine in chronological order to generate a new video. Low-resolution normal scene videos and high-resolution abnormal scene videos are merged to generate a resolution-optimized video file. Video merging is performed using video editing tools (e.g., FFmpeg). Video reconstruction is performed using video processing tools (e.g., FFmpeg).

[0085] Preferably, step S1 includes the following steps:

[0086] Step S11: Acquire the source video using a recording device;

[0087] Step S12: Perform frame-by-frame image segmentation on the source video based on the preset timestamp to obtain frame-by-frame images of the source video;

[0088] Step S13: Overlay adjacent frames of the source video frame by frame to generate adjacent frame overlap data;

[0089] Step S14: Effective video re-segmentation is performed on each frame of the source video using overlapping data from adjacent frames to generate an effective video.

[0090] In this embodiment of the invention, the device (such as a camera or mobile phone) for recording is determined and prepared, ensuring that the device is stable and can clearly record the required video content. The source video is recorded as needed, ensuring that the video contains all the content that needs to be analyzed and processed. The timestamp interval required for video processing is set (e.g., frames per second). Video processing tools (such as FFmpeg or OpenCV) are used to cut the source video into frame-by-frame images, ensuring that each frame has timestamp information. Image processing algorithms (such as SIFT, SURF, or ORB) are used to perform feature point matching on adjacent frames, calculating the overlapping portions of adjacent frames. The overlapping data for each pair of adjacent frames is recorded, including the location information of the overlapping area and the matching feature point information. Based on the overlapping data, the overlapping portions of adjacent frames are stitched together to generate a continuous and valid video. Video compositing tools (such as FFmpeg or OpenCV) are used to reassemble the stitched image sequence into a complete video file.

[0091] Preferably, step S14 includes the following steps:

[0092] Step S141: Perform grayscale conversion on each frame of the source video to generate grayscale images of each frame of the source video;

[0093] Step S142: Perform feature point detection on each frame of the grayscale image of the source video to obtain frame image feature points; perform feature description on the frame image feature points to generate frame image feature descriptors;

[0094] Step S143: Perform image feature point matching on each frame of grayscale image of the source video using frame image feature descriptors to generate frame-by-frame image feature point matching data; perform image deviation analysis on each frame of grayscale image of the source video using overlapping data of adjacent frame images to generate frame-by-frame image deviation data.

[0095] Step S144: Use frame-by-frame image feature point matching data and frame-by-frame image deviation data to remove invalid frame images from the source video frame-by-frame images to obtain valid video frame images; perform time-series frame-by-frame stitching on the valid video frame images to generate a valid video.

[0096] In this embodiment of the invention, source video images are read frame by frame. Each frame is converted into a grayscale image to reduce computational complexity and improve feature detection accuracy. Feature detection algorithms (such as SIFT, SURF, or ORB) are used to detect feature points in each grayscale image. Feature descriptors are generated for the detected feature points to facilitate subsequent feature matching. Feature matching algorithms (such as BFMatcher) are used to match the feature descriptors of adjacent frames to generate feature point matching data. Based on the feature point matching data of adjacent frames, the deviation between images is analyzed, and the deviation information between each pair of adjacent frames is recorded. According to the feature point matching data and image deviation data, invalid frames (such as duplicate frames, blurred frames, or irrelevant frames) are removed, and only valid frames are retained. The retained valid frames are stitched together frame by frame in chronological order to generate a continuous and valid video.

[0097] As an example of the present invention, reference is made to Figure 2 As shown, in this example, step S2 includes:

[0098] Step S21: Extract keyframe images from the valid video to obtain keyframe images;

[0099] Step S22: Perform background segmentation on the keyframe image to generate the foreground object image of the keyframe;

[0100] Step S23: Perform object localization and detection on the foreground object image of the keyframe to generate object localization data; perform dynamic object recognition on the effective video using the object localization data to generate dynamic object recognition data.

[0101] Step S24: Mark dynamic objects in the valid video based on the dynamic object recognition data and generate dynamic object tags.

[0102] In this embodiment of the invention, representative frames are selected as keyframes by analyzing the timestamps and motion information of the video. Keyframes are typically frames that change significantly in the video, helping to reduce redundant information and highlight important events. Keyframes are extracted from the effective video using video processing tools or libraries (such as OpenCV), specifically based on inter-frame differences or specific algorithms (such as K-means clustering) to select representative frames. Background modeling techniques (such as Gaussian mixture models, background subtraction algorithms, etc.) are used to separate the background and foreground of the video frames. A background subtraction algorithm is applied to each keyframe to extract foreground object images; this can be implemented using background segmentation functions in OpenCV or other computer vision libraries. Object detection algorithms (such as YOLO, SSD, Faster R-CNN, etc.) are used to process the foreground object images, detecting and locating objects. A trained deep learning model can be used for object localization to obtain information such as the object's position and category. The object localization data is applied to the entire effective video to identify dynamically changing objects. By comparing the object localization data of the keyframes with the foreground images of other frames, changes in dynamic objects are identified and tracked. Based on the dynamic object recognition data, the identified dynamic objects are marked in each frame of the video. Draw markers (such as borders, labels, etc.) in video frames to mark the identified objects, generate a label file or layer containing dynamic object information, record the position information, category and motion trajectory of each object, and save the marker data as additional information for the video file or generate a label file (such as XML or JSON format).

[0103] Preferably, step S24 includes the following steps:

[0104] Step S241: Perform inter-frame difference analysis on the effective video based on the dynamic object recognition data to obtain dynamic object motion region data; perform region segmentation on the effective video based on the dynamic object motion region data to obtain dynamic object motion region image.

[0105] Step S242: Use dynamic object recognition data to perform edge feature detection on the image of the moving area of ​​the dynamic object, and generate edge feature data, wherein the edge feature detection includes shape detection, texture analysis and color extraction;

[0106] Step S243: Perform optical flow tracing on the motion region image of the dynamic object based on edge feature data to generate dynamic object motion characteristic data; divide the effective video into dynamic object images according to the dynamic object motion characteristic data to generate live object images and non-live object images.

[0107] Step S244: Dynamically label the effective video using live and non-live images to generate dynamic object tags.

[0108] In this embodiment of the invention, consecutive frames are extracted from valid video. Differential operations are performed on adjacent frames to calculate the inter-frame variation regions. Dynamic object motion region data is generated, typically a binary image or mask of the variation regions. The dynamic object motion region data is processed using segmentation algorithms (such as thresholding, region growing, or image segmentation models). A dynamic object motion region image is obtained, in which the regions of the dynamic objects are marked. An edge detection algorithm (such as Canny edge detection) is applied to the dynamic object motion region image. Shape features, texture features, and color features are extracted. An optical flow algorithm (such as the Lucas-Kanade optical flow algorithm) is used to track the motion of the object in the video frames. Dynamic object motion characteristic data, including motion direction, speed, and other information, is generated. Based on the optical flow data, the dynamic object motion region image is divided into live and non-live images. In each frame of the valid video, labeling data for live and non-live images is applied. Borders, labels, or other markers are drawn to highlight the dynamic objects. The labeling information is saved as a label file or video supplementary information. The labels specifically include detailed information such as the category, location, and motion characteristics of the dynamic object.

[0109] As an example of the present invention, reference is made to Figure 3 As shown, step S3 in this example includes:

[0110] Step S31: Use dynamic object tags to perform scene event detection on valid videos, thereby generating a scene event image set;

[0111] Step S32: Use scene event detection data to perform image behavior detection on the scene event image set and generate behavior detection data;

[0112] Step S33: Construct an anomaly detector; use the anomaly detector to detect abnormal behavior in the behavior detection data and generate abnormal behavior detection data;

[0113] Step S34: Compare the abnormal behavior detection data with the preset standard abnormal behavior detection threshold. When the abnormal behavior detection data is greater than or equal to the preset standard abnormal behavior detection threshold, mark the scene event image set as abnormal scene events based on the abnormal behavior detection data to generate an abnormal scene image set. When the abnormal behavior detection data is less than the preset standard abnormal behavior detection threshold, mark the scene event image set as normal scene events based on the abnormal behavior detection data to generate a normal scene image set.

[0114] In this embodiment of the invention, a video frame dataset is constructed by extracting frames from valid videos. Object detection algorithms (such as YOLO and SSD) are applied to detect dynamic objects in each frame. Labels are generated for the detected objects. Scene understanding models (such as CNN and RNN) are used to identify events in the video. Scene events are detected by analyzing object labels and time-series data using event detection algorithms. A scene event image set containing the detected event frames is generated. A behavior detection model (such as 3D-CNN or a temporal model) is selected or trained to analyze the scene event image set. A behavior detection algorithm is applied to each image to extract behavior features. The behavior features are converted into behavior detection data, and the behavior information for each event is recorded. An appropriate anomaly detection method (such as Isolation Forest, One-Class SVM, or Autoencoder) is selected. An anomaly detection model is trained using normal behavior data. The model needs to be able to distinguish between normal and abnormal behavior. The trained anomaly detector is applied to analyze the behavior detection data. Abnormal behavior detection data is generated, and abnormal behaviors are marked. A preset standard abnormal behavior detection threshold is determined. The threshold is specifically set through statistical analysis of normal behavior data or based on the opinions of domain experts. The abnormal behavior detection data is compared with a standard threshold. If the abnormal behavior detection data is greater than or equal to the threshold: the scene event image set is marked as an abnormal scene event. An abnormal scene image set is generated, and abnormal frames are marked. If the abnormal behavior detection data is less than the threshold: the scene event image set is marked as a normal scene event. A normal scene image set is generated, and normal frames are marked.

[0115] Preferably, step S31 includes the following steps:

[0116] Step S311: Use dynamic object tags to classify the effective video into object regions and generate effective object region videos, which include biological region videos and non-biological region videos.

[0117] Step S312: Perform behavior recognition on the biological region video to generate biological behavior recognition data; perform pose tracking on the biological region video using the biological behavior recognition data to generate biological behavior motion trajectory images; perform motion pattern analysis on the biological behavior motion trajectory images to generate biological region motion pattern data.

[0118] Step S313: Perform temporal localization on the video of the abiotic region to generate temporal localization data of the abiotic region; use the temporal localization data of the abiotic region to perform motion trajectory analysis on the video of the abiotic region to generate abiotic motion trajectory image; perform motion path analysis on the abiotic motion trajectory image to generate abiotic motion path data.

[0119] Step S314: Separate scene event images from the video of the effective object area using biological region motion pattern data and non-biological region movement path data to generate a scene event image set.

[0120] In this embodiment of the invention, object detection models (such as YOLO and SSD) are used to detect objects in each frame of a valid video, generating object labels. Objects are categorized into biotic and non-biotic regions. Biobiotic regions include humans and animals, while non-biotic regions include vehicles and buildings. Classification algorithms (such as CNN classifiers) are used to further confirm object categories. The object regions in the video are labeled as biotic region videos and non-biotic region videos. Behavior recognition models (such as 3D-CNN and temporal models) are used to perform behavior analysis on the biotic region videos. Biological behavior recognition data is generated, recording the behavior category and features of each biotic region. Using the biological behavior recognition data, pose tracking (such as pose estimation models like OpenPose) is performed on the organisms in the biotic region videos. Biological behavior trajectory images are generated, displaying the organism's movement trajectory in the video. Motion pattern analysis is performed on the biological behavior trajectory images to extract motion features (such as direction, speed, and frequency). Biological region motion pattern data is generated, describing the organism's movement behavior patterns. Temporal localization is performed on the non-biotic region videos, and temporal analysis models (such as LSTM and GRU) are used for time series data analysis. Generate temporal localization data for abiotic regions, recording the temporal changes of objects in the video. Use the temporal localization data to analyze the movement trajectories of objects in the video within the abiotic regions. Generate abiotic movement trajectory images, displaying the object's movement path. Perform path analysis on the abiotic movement trajectory images, extracting path features (such as the shape and direction of the movement path). Generate abiotic region movement path data, describing the movement paths of abiotic objects. Integrate biotic region motion pattern data and abiotic region movement path data to analyze scene events in the video. Separate scene event images based on the motion patterns of biotic regions and the movement paths of abiotic regions. Identify scene events in the video frames and generate a scene event image set. The scene event image set contains image frames with significant event features.

[0121] Preferably, step S33 includes the following steps:

[0122] Step S331: Divide the behavior detection data into a dataset to obtain a model training set and a model test set;

[0123] Step S332: Train the model on the training set using the Isolation Forest algorithm to generate an anomaly detection detector; use the test set to iterate and optimize the anomaly detection detector to generate an anomaly detector.

[0124] Step S333: Import the behavior detection data into the anomaly detector to detect abnormal behavior and generate abnormal behavior detection data.

[0125] In this embodiment of the invention, behavior detection data is collected and organized. This data should include behavioral features extracted from videos, such as action type, speed, and frequency. The behavior detection dataset is divided into a model training set and a model test set. The training set is used to train the Isolation Forest algorithm model and typically accounts for 70%-80% of the data. The test set is used to verify the model's performance and for optimization and typically accounts for 20%-30% of the data. The partitioning method specifically uses random partitioning or cross-validation to ensure that the training and test sets are representative. The Isolation Forest algorithm is used to train the model on the training set. By constructing multiple decision trees, data points are "isolated" to identify anomalies. The algorithm's advantage is that it performs well on large-scale datasets and is suitable for high-dimensional data. The training process includes setting the model's hyperparameters, such as the number of trees and the maximum tree depth. These parameters need to be adjusted empirically or through cross-validation. The model test set is imported into the trained Isolation Forest model for anomaly detection to generate prediction results. Based on the detection results of the test set, the model's hyperparameters are optimized: the parameters of the Isolation Forest algorithm (such as the number of trees, the number of segmentation features, etc.) are adjusted to improve the model's accuracy. The model undergoes iterative optimization to ensure its effectiveness in detecting anomalous behavior, generating a final anomaly detector. Anomaly detection is then performed using the optimized model. Behavior detection data is imported into the optimized anomaly detector for testing. Necessary preprocessing, such as standardization and feature scaling, is applied to the behavior detection data to adapt it to the model input. The anomaly detector analyzes the behavior detection data, generating anomalous behavior detection data, including anomaly scores or labels for each data point, indicating whether it represents anomalous behavior. Output results are generated, such as the specific type of anomalous behavior, anomaly scores, and a list of anomalous data points. Detected anomalous behaviors can be visualized to aid in further analysis and understanding of the anomalies.

[0126] As an example of the present invention, reference is made to Figure 4 As shown, step S4 in this example includes:

[0127] Step S41: Obtain resolution resource data;

[0128] Step S42: Perform a first resolution performance adjustment on the normal scene image set based on the resolution resource data to generate a normal scene low-resolution image set; perform a second resolution performance adjustment on the abnormal scene image set based on the resolution resource data to generate an abnormal scene high-resolution image set.

[0129] Step S43: Perform video reconstruction on the low-resolution image set of normal scenes and the high-resolution image set of abnormal scenes to generate a resolution-optimized video for ultra-high-definition video detection.

[0130] In this embodiment of the invention, the source of resolution resource data is determined, including device capability data, network bandwidth, storage capacity, etc. Relevant data is collected, specifically obtained from system monitoring, device specifications, or network analysis tools. The system's resolution resources are quantified, including the highest supported resolution, currently available bandwidth, storage space, etc. This data is recorded and organized for subsequent resolution performance adjustments. A normal scene image set, generated in step S31, is used. Based on the resolution resource data, image downsampling techniques are used to adjust the normal scene image set to a low resolution. Specifically, image downsampling algorithms (such as bilinear interpolation, bicubic interpolation) are used to reduce the image size to a specified low resolution, resulting in a normal scene low-resolution image set. An abnormal scene image set, generated in step S34, is used. Based on the resolution resource data, image super-resolution techniques are used to adjust the abnormal scene image set to a high resolution. Specifically, image super-resolution algorithms (such as convolutional neural network (CNN) super-resolution, generative adversarial network (GAN)) are used to improve the image resolution, resulting in an abnormal scene high-resolution image set. The process involves several steps: First, a collection of low-resolution images from normal scenes is compiled into a video stream, which is then generated using video encoding techniques such as H.264 and H.265. Second, a collection of high-resolution images from abnormal scenes is compiled into a video stream, which is then generated using high-resolution video encoding techniques such as HEVC and VP9. Finally, the low-resolution videos from normal scenes and the high-resolution videos from abnormal scenes are combined into a single video stream using video editing tools or programming libraries such as FFmpeg. Based on the required resolution resource data, video encoding settings are adjusted to optimize video quality and smoothness, ensuring that the generated video meets the requirements for ultra-high-definition video detection. Ultra-high-definition video detection tools (such as deep learning-based video quality assessment models) are used to detect the resolution-optimized video, assess whether the video quality and resolution meet the standards, detect anomalies in the video, generate a detection report, and provide video quality assessment results and anomaly markers.

[0131] This specification provides an ultra-high-definition video detection system for performing the above-described ultra-high-definition video detection method. The ultra-high-definition video detection system includes:

[0132] The effective video filtering module is used to acquire source video; to segment the source video frame by frame to obtain source video frame by frame images; and to re-sew the source video frame by frame images into effective video to generate effective video.

[0133] The dynamic object recognition module is used to perform background segmentation of keyframe images in valid videos to generate keyframe foreground object images; to perform dynamic object recognition on keyframe foreground object images to generate dynamic object recognition data; and to mark dynamic objects in valid videos based on dynamic object recognition data to generate dynamic object labels.

[0134] The event analysis module is used to detect scene events in valid videos using dynamic object tags, thereby generating a scene event image set; to detect abnormal behavior in the scene event image set, generating abnormal behavior detection data; and to compare the abnormal behavior detection data with a preset standard abnormal behavior detection threshold, generating an abnormal scene image set and a normal scene image set.

[0135] The resolution adjustment module is used to acquire resolution resource data; adjust the resolution performance of normal scene image sets and abnormal scene image sets according to the resolution resource data to generate normal scene low-resolution image sets and abnormal scene high-resolution image sets; reconstruct the video from the normal scene low-resolution image sets and abnormal scene high-resolution image sets to generate resolution-optimized video for performing ultra-high-definition video detection.

[0136] The beneficial effects of this invention lie in ensuring that only clear and effective video segments are selected through frame-by-frame image segmentation and effective video re-stitching, eliminating useless or redundant data. This improves the efficiency and accuracy of subsequent processing and ensures the reliability of the analysis results. Through keyframe background segmentation and dynamic object recognition, dynamic objects in the video can be accurately detected and labeled. The generated dynamic object tags provide a clear basis for subsequent event analysis, improving the accuracy and efficiency of dynamic object recognition. Scene event detection using dynamic object tags generates a scene event image set containing key event information. Abnormal behavior detection further refines the distinction between abnormal and normal scenes, ensuring the system can promptly detect and handle abnormal situations, enhancing the system's anomaly detection capabilities. Resolution performance is adjusted for normal and abnormal scenes based on resolution resource data. Normal scenes use low-resolution image sets to reduce resource consumption, while abnormal scenes use high-resolution image sets to provide more detailed information, ensuring high-quality video monitoring and analysis. Video reconstruction using low-resolution image sets for normal scenes and high-resolution image sets for abnormal scenes generates resolution-optimized videos that simultaneously meet the needs of resource optimization and high-definition detection. This optimization process provides clear video quality for ultra-high-definition video detection operations, improving the accuracy and effectiveness of detection. By processing normal scenes at low resolution, computational and storage resources are saved, while processing abnormal scenes at high resolution ensures the quality of critical data. This layered resolution adjustment effectively balances resource usage and data quality, reducing system load. The system comprehensively utilizes video filtering, dynamic object recognition, event analysis, and resolution adjustment technologies to enhance overall video monitoring and analysis capabilities, ensuring reliable analysis results across various scenarios. This system is applicable to multiple fields such as security monitoring, behavior analysis, and abnormal event detection. It can adjust video processing strategies according to different needs and scenarios, improving the intelligence and accuracy of various applications. Through efficient video processing and optimization, users can obtain clearer and more accurate video data, improving their understanding and application of monitored content, and enhancing user experience. The system can flexibly adjust processing strategies according to different video scenes and resolution requirements, adapting to changes in different environments and application scenarios, improving system adaptability and flexibility. Therefore, this invention improves detection accuracy and performance through frame-by-frame image processing, precise dynamic object recognition, abnormal behavior detection, and resolution optimization.

[0137] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.

[0138] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A method for detecting ultra-high-definition video, characterized in that, Includes the following steps: Step S1: Obtain the source video; The source video is segmented frame by frame to obtain the source video frame by frame images; Effective video re-stitching is performed on each frame of the source video to generate an effective video; Step S2: Perform background segmentation on keyframe images of the effective video to generate keyframe foreground object images; perform dynamic object recognition on the keyframe foreground object images to generate dynamic object recognition data. Based on the dynamic object recognition data, dynamic objects are marked in valid videos to generate dynamic object tags; Step S3: Use dynamic object tags to detect scene events in valid videos, thereby generating a scene event image set; Perform abnormal behavior detection on the scene event image set to generate abnormal behavior detection data; compare the abnormal behavior detection data with the preset standard abnormal behavior detection threshold to generate abnormal scene image set and normal scene image set; Step S4: Obtain resolution resource data; Based on the resolution resource data, the resolution performance of the normal scene image set and the abnormal scene image set is adjusted to generate a low-resolution image set of normal scene and a high-resolution image set of abnormal scene. The video is reconstructed from a set of low-resolution images of normal scenes and a set of high-resolution images of abnormal scenes to generate a resolution-optimized video for ultra-high-definition video inspection.

2. The ultra-high-definition video detection method according to claim 1, characterized in that, Step S1 includes the following steps: Step S11: Acquire the source video using a recording device; Step S12: Perform frame-by-frame image segmentation on the source video based on the preset timestamp to obtain frame-by-frame images of the source video; Step S13: Overlay adjacent frames of the source video frame by frame to generate adjacent frame overlap data; Step S14: Effective video re-segmentation is performed on each frame of the source video using overlapping data from adjacent frames to generate an effective video.

3. The ultra-high-definition video detection method according to claim 2, characterized in that, Step S14 includes the following steps: Step S141: Perform grayscale conversion on each frame of the source video to generate grayscale images of each frame of the source video; Step S142: Perform feature point detection on each frame of the grayscale image of the source video to obtain frame image feature points; perform feature description on the frame image feature points to generate frame image feature descriptors; Step S143: Perform image feature point matching on each frame of grayscale image of the source video using frame image feature descriptors to generate frame-by-frame image feature point matching data; perform image deviation analysis on each frame of grayscale image of the source video using overlapping data of adjacent frame images to generate frame-by-frame image deviation data. Step S144: Use frame-by-frame image feature point matching data and frame-by-frame image deviation data to remove invalid frame images from the source video frame-by-frame images to obtain valid video frame images; perform time-series frame-by-frame stitching on the valid video frame images to generate a valid video.

4. The ultra-high-definition video detection method according to claim 1, characterized in that, Step S2 includes the following steps: Step S21: Extract keyframe images from the valid video to obtain keyframe images; Step S22: Perform background segmentation on the keyframe image to generate the foreground object image of the keyframe; Step S23: Perform object localization and detection on the foreground object image of the keyframe to generate object localization data; perform dynamic object recognition on the effective video using the object localization data to generate dynamic object recognition data. Step S24: Mark dynamic objects in the valid video based on the dynamic object recognition data and generate dynamic object tags.

5. The ultra-high-definition video detection method according to claim 4, characterized in that, Step S24 includes the following steps: Step S241: Perform inter-frame difference analysis on the effective video based on the dynamic object recognition data to obtain dynamic object motion region data; perform region segmentation on the effective video based on the dynamic object motion region data to obtain dynamic object motion region image. Step S242: Use dynamic object recognition data to perform edge feature detection on the image of the moving area of ​​the dynamic object, and generate edge feature data, wherein the edge feature detection includes shape detection, texture analysis and color extraction; Step S243: Perform optical flow tracing on the motion region image of the dynamic object based on edge feature data to generate dynamic object motion characteristic data; divide the effective video into dynamic object images according to the dynamic object motion characteristic data to generate live object images and non-live object images. Step S244: Dynamically label the effective video using live and non-live images to generate dynamic object tags.

6. The ultra-high-definition video detection method according to claim 1, characterized in that, Step S3 includes the following steps: Step S31: Use dynamic object tags to perform scene event detection on valid videos, thereby generating a scene event image set; Step S32: Use scene event detection data to perform image behavior detection on the scene event image set and generate behavior detection data; Step S33: Construct an anomaly detector; use the anomaly detector to detect abnormal behavior in the behavior detection data and generate abnormal behavior detection data; Step S34: Compare the abnormal behavior detection data with the preset standard abnormal behavior detection threshold. When the abnormal behavior detection data is greater than or equal to the preset standard abnormal behavior detection threshold, mark the scene event image set as abnormal scene events based on the abnormal behavior detection data to generate an abnormal scene image set. When the abnormal behavior detection data is less than the preset standard abnormal behavior detection threshold, mark the scene event image set as normal scene events based on the abnormal behavior detection data to generate a normal scene image set.

7. The ultra-high-definition video detection method according to claim 6, characterized in that, Step S31 includes the following steps: Step S311: Use dynamic object tags to classify the effective video into object regions and generate effective object region videos, which include biological region videos and non-biological region videos. Step S312: Perform behavior recognition on the biological region video to generate biological behavior recognition data; perform pose tracking on the biological region video using the biological behavior recognition data to generate biological behavior motion trajectory images; perform motion pattern analysis on the biological behavior motion trajectory images to generate biological region motion pattern data. Step S313: Perform temporal localization on the video of the abiotic region to generate temporal localization data of the abiotic region; use the temporal localization data of the abiotic region to perform motion trajectory analysis on the video of the abiotic region to generate abiotic motion trajectory image; perform motion path analysis on the abiotic motion trajectory image to generate abiotic motion path data. Step S314: Separate scene event images from the video of the effective object area using biological region motion pattern data and non-biological region movement path data to generate a scene event image set.

8. The ultra-high-definition video detection method according to claim 6, characterized in that, Step S33 includes the following steps: Step S331: Divide the behavior detection data into a dataset to obtain a model training set and a model test set; Step S332: Train the model on the training set using the Isolation Forest algorithm to generate an anomaly detection detector; use the test set to iterate and optimize the anomaly detection detector to generate an anomaly detector. Step S333: Import the behavior detection data into the anomaly detector to detect abnormal behavior and generate abnormal behavior detection data.

9. The ultra-high-definition video detection method according to claim 1, characterized in that, Step S4 includes the following steps: Step S41: Obtain resolution resource data; Step S42: Perform a first resolution performance adjustment on the normal scene image set based on the resolution resource data to generate a normal scene low-resolution image set; perform a second resolution performance adjustment on the abnormal scene image set based on the resolution resource data to generate an abnormal scene high-resolution image set. Step S43: Perform video reconstruction on the low-resolution image set of normal scenes and the high-resolution image set of abnormal scenes to generate a resolution-optimized video for ultra-high-definition video detection.

10. An ultra-high-definition video inspection system, characterized in that, For performing the ultra-high-definition video detection method as described in claim 1, the ultra-high-definition video detection system comprises: The effective video filtering module is used to acquire source video; to segment the source video frame by frame to obtain source video frame by frame images; and to re-sew the source video frame by frame images into effective video to generate effective video. The dynamic object recognition module is used to perform background segmentation of keyframe images in valid videos to generate keyframe foreground object images; to perform dynamic object recognition on keyframe foreground object images to generate dynamic object recognition data; and to mark dynamic objects in valid videos based on dynamic object recognition data to generate dynamic object labels. The event analysis module is used to detect scene events in valid videos using dynamic object tags, thereby generating a scene event image set; to detect abnormal behavior in the scene event image set, generating abnormal behavior detection data; and to compare the abnormal behavior detection data with a preset standard abnormal behavior detection threshold, generating an abnormal scene image set and a normal scene image set. The resolution adjustment module is used to acquire resolution resource data; adjust the resolution performance of normal scene image sets and abnormal scene image sets according to the resolution resource data to generate normal scene low-resolution image sets and abnormal scene high-resolution image sets; reconstruct the video from the normal scene low-resolution image sets and abnormal scene high-resolution image sets to generate resolution-optimized video for performing ultra-high-definition video detection.

Citation Information

Patent Citations

  • Pedestrian identification system based on multi-ball multi-gun camera array

    CN105979210A

  • Video conference exception reconstruction method, computer device and storage medium

    CN118118620A