Video classification method based on AI

By deploying a closed-loop mechanism that links multi-source classification detection with event triggering in video classification, the problem of interference from bright saturated spots on video classification is solved. This enables accurate detection and unified feature extraction of bright saturated spots, significantly reduces the false positive rate, and improves the accuracy and consistency of video classification.

CN121236668AActive Publication Date: 2025-12-30XIAMEN XINZHUN TECHNOLOGY CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511513018.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2025-12-30
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

In existing technologies for video classification, the abnormal amplification of feature maps caused by bright saturation spots affects the image recognition results. Furthermore, single detection methods are easily affected by data noise, and the lack of a unified standard for data feature extraction and modeling leads to misjudgments.

Method used

A closed-loop mechanism of multi-source classification detection and event triggering is adopted. Brightness and chromaticity threshold, temporal consistency, temporal spectrum and morphological geometry template acquisition modules are deployed at the input, temporal sequence, spectrum and morphological layers respectively. Through the unification and dynamic adjustment of frame acquisition frequency and pixel spatial granularity, multi-dimensional detection and judgment of high brightness saturation spots are performed.

Benefits of technology

It significantly reduces the false positive rate, improves the accuracy and consistency of video classification, avoids asynchronous and multi-scale distortion, and enhances the temporal consistency and interpretability of video-level judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121236668A_ABST
    Figure CN121236668A_ABST
Patent Text Reader

Abstract

The invention discloses an AI-based video classification method, which relates to the technical field of image recognition, and comprises the following steps: S0, in a video frame sequence which is disassembled frame by frame, deploying a brightness and chrominance threshold acquisition module in an image input layer and a preprocessing channel for extracting pixel-level brightness and chrominance characteristics, and acquiring a brightness and chrominance threshold value; a time sequence consistency acquisition module is embedded into a middle layer of a frame sequence analysis channel and is used for capturing time consistency of brightness and texture changes between adjacent frames, and a time domain frequency spectrum acquisition module is deployed in a video signal frequency analysis channel after feature extraction and is used for detecting inter-frame frequency oscillation and optical flow energy distribution. A morphological geometric template acquisition module is arranged in a high-level semantic modeling layer and a deconstruction layer and is used for identifying geometric boundaries, contours and morphological structure features in video frames, S1, each acquisition module works and detects a highlight saturated spot phenomenon, and the method has the characteristic of targeted adjustment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of image recognition, and particularly relates to an AI-based video classification method. BACKGROUND

[0002] In AI-based video classification of image recognition, a mainstream method is to divide a video into frames, extract features of each frame by using a 2DCNN model, capture time sequence dynamics in frame sequences by using an RNN model, and then classify the video after time sequence fusion. When raindrops / snowflakes / lens stains and the like suddenly enter a field of view in a video image, various types of highlight saturation spots are generated under light sources. Convolution kernels and edge operators in an image processing algorithm are most sensitive to the high-frequency changes, which may cause abnormal amplification of feature maps and affect the result of image recognition.

[0003] The highlight saturation spots are divided into scene optical interaction types, lens optical system types, decoding post-processing artifact types and time sequence motion-induced types. The highlight saturation spot detection methods that can be used include brightness chroma threshold judgment, time sequence consistency judgment, time domain spectrum judgment and morphological geometric template judgment. Each detection method has a plurality of video image acquisition modules and analysis modules that work in coordination to improve analysis efficiency and analysis accuracy.

[0004] In the prior art, one highlight saturation spot detection method is often used to correspond to one or multiple types of highlight saturation spots. Single basis is susceptible to data noise interference, leading to misjudgment. If information synthesis is directly used, the frame acquisition frequency and pixel space granularity suitable for different types of highlight saturation spot detection are quite different, and a unified data feature extraction and modeling standard is not established. After normalization processing, data consistency is damaged. If not targetedly adjusted, the normal feature extraction requirement is deviated. Therefore, it is necessary to design a targeted AI-based video classification method. SUMMARY

[0005] The application aims to provide an AI-based video classification method to solve the problems in the background art.

[0006] To solve the above technical problems, the application provides the following technical scheme: an AI-based video classification method, comprising the following steps: S0, in the frame-by-frame disassembled video frame sequence, the luminance chrominance threshold acquisition module is arranged in the image input layer and the preprocessing channel, used for extracting the pixel-level luminance and chrominance features, the time sequence consistency acquisition module is embedded in the middle layer of the frame sequence analysis channel, used for capturing the time consistency of the luminance and texture changes between adjacent frames, the time domain spectrum acquisition module is arranged in the frequency analysis channel after the feature extraction, used for detecting the inter-frame frequency oscillation and the light flow energy distribution, and the morphological geometric template acquisition module is arranged in the high-level semantic modeling layer and the deconstruction layer, used for identifying the geometric boundary, contour and morphological structure features in the video frame; S1, each acquisition module works and detects the highlight saturation spot phenomenon, each acquisition module outputs the detection result in the initial frame acquisition frequency and pixel space granularity, and detects whether the highlight saturation spot phenomenon occurs through the corresponding analysis module; S2, when a certain acquisition module detects that a certain type of highlight saturation spot occurs, the other type of acquisition module most related to the feature domain of the acquisition module is adjusted in the frame acquisition frequency and the pixel space granularity, so that the acquisition module to be adjusted and the current acquisition module are unified in the frame acquisition frequency and the pixel space granularity; S3, with the continuation of the frequency, the frame acquisition frequency and the pixel space granularity of the acquisition module to be adjusted are dynamically adjusted according to whether the same highlight saturation spot phenomenon is continuously detected; S4, when the frame acquisition frequency and the pixel space granularity of each acquisition module are not unified, the detection data is processed by using the data normalization module, and the frame acquisition frequency and the pixel space granularity are normalized; S5, based on the frame-by-frame structure analysis result, the spatial morphological features of the video frame are extracted and judged by using the morphological and geometric template method.

[0007] According to the technical scheme, in S0, it is clear that which detection method corresponds to each type of highlight saturation spot, which is specifically: S0-1, scene optical interaction type highlight saturation spot, that is, the luminance chrominance threshold detection is used as the dominant judgment method, and the time sequence consistency detection and the time domain spectrum detection are used as the auxiliary judgment method, caused by the reflection, refraction and direct reflection of the object surface and the external light source; S0-2, lens optical system type highlight saturation spot, that is, the time sequence consistency detection is used as the dominant judgment method, and the luminance chrominance threshold detection and the morphological geometric template detection are used as the auxiliary judgment method, caused by the lens stain, oil film, halo, ghosting and flare; S0-3, decoding post-processing artifact type highlight saturation spot, that is, the time domain spectrum detection is used as the dominant judgment method, and the luminance chrominance threshold detection is used as the auxiliary judgment method, caused by the compression ringing, block effect and over-sharpening; S0-4. Temporal motion-induced high-brightness saturation spots, i.e. instantaneous high-brightness areas caused by raindrops, snowdrops, stroboscopic flashes, and rapid movement, are judged primarily by temporal consistency detection, with temporal spectrum detection and brightness and chromaticity threshold detection as auxiliary judgment methods.

[0008] According to the above technical solution, in step S2, the unification of frame acquisition frequency and pixel spatial granularity specifically involves: S2-1. Each acquisition module, by default, detects the type of bright saturated spot corresponding to its primary judgment method, and uses its default frame acquisition frequency. and pixel space granularity Output the detection results, where This refers to the number of data acquisition modules; S2-2, Order No. The acquisition module detected a bright saturation spot phenomenon, and its default frame acquisition frequency was [missing information]. The default pixel space granularity is The initial probability of this type of bright saturation spot phenomenon occurring is: At this point, it is necessary to... The frame acquisition frequency and pixel spatial granularity of each acquisition module were adjusted. The default frame acquisition frequency and pixel spatial granularity before the adjustment were as follows: and This makes its adjusted frame acquisition frequency Adjusted pixel space granularity .

[0009] According to the above technical solution, the dynamic adjustment in S3 specifically includes: S3-1, Order No. Each acquisition module maintains a frame acquisition frequency. and pixel space granularity continued During the time period, if Within the time period When the acquisition module detects the same type of bright saturation spot phenomenon again, observe the first... If each acquisition module simultaneously detects the same type of bright saturation spot phenomenon, and if so, the probability of this type of bright saturation spot phenomenon occurring is increased, with the adjusted probability being... ,in For the first When the first acquisition module is used as an auxiliary judgment method, it is related to the first... The probability increment caused by the detection of the same type of bright saturation spot phenomenon by each acquisition module, if the first acquisition module detects the same type of bright saturation spot phenomenon, If none of the acquisition modules simultaneously detected the same type of bright saturation spot phenomenon, then... ; S3-2, If in Within the time period When the first acquisition module does not detect the same type of bright saturation spot phenomenon, the second... The frame acquisition frequency and pixel spatial granularity of each acquisition module are determined by... and To its initial value and Gradually recovering, making the current distance The elapsed time of the end of the time period is Then the frame acquisition frequency at this time and pixel space granularity The calculation formulas are as follows: when hour, ,when hour, ,when hour, ,when hour, ,in , This is the frequency conversion factor.

[0010] According to the above technical solution, in step S4, the normalization processing of frame acquisition frequency and pixel spatial granularity specifically involves: S4-1, Firstly in When the time period ends, the first The and the first The first acquisition frequency of each acquisition module is aligned, at which point both acquisition modules acquire data simultaneously. The next acquisition will only collect data that meets the specified frequency. and The detection data is collected only at the least common multiple of the time; other data are not collected. S4-2, the first The and the first The maximum and minimum values ​​of the detection data from each acquisition module are mapped to... Within the specified range, the dimensional differences in pixel spatial granularity between different acquisition modules are eliminated.

[0011] According to the above technical solution, in S5, the determination of the type of bright saturated spot is specifically as follows: before each detection of a bright saturated spot phenomenon by each acquisition module, the geometric feature distribution of this type of bright spot detected by the morphological geometric template acquisition module is respectively... ,in The geometric feature distribution refers to the number of geometric morphological types involved in the high-brightness saturation spots. It indicates the intensity of the template's response in ring-shaped, strip-shaped, dot-shaped, and radial morphologies. When a certain type of high-brightness saturation spot phenomenon is detected, if this type of high-brightness saturation spot leads to... Changes have occurred, and Actual detection The probability of occurrence of this highlight saturation spot type increases when the increment change is generated , wherein is the geometric feature increment influence coefficient, and if no detection is made The increment change is generated , combined with S3-1 to Calculate the final probability, when , is the probability judgment threshold value, and the highlight saturation spot type is determined to be established.

[0012] An AI-based video classification system, comprising an image recognition module, a variety of information synthesis module, and a highlight saturation spot judgment module, the image recognition module is used to collect luminance chrominance threshold, time sequence consistency, time domain spectrum, and combine morphological geometric templates for multi-dimensional detection, the variety of information synthesis module is used to unify the frame acquisition frequency and pixel space granularity for frequency reference and granularity normalization processing, dynamically adjust two output parameter characteristics according to subsequent detection results, and perform multi-source information synthesis processing, and the highlight saturation spot judgment module is used to comprehensively determine the highlight saturation spot type according to the processed data.

[0013] According to the above technical scheme, the image recognition module comprises a luminance chrominance threshold acquisition module, a time sequence consistency acquisition module, a time domain spectrum acquisition module, a morphological geometric template acquisition module, a luminance chrominance threshold analysis module, a time sequence consistency analysis module, a time domain spectrum analysis module, and a morphological geometric template analysis module, the luminance chrominance threshold acquisition module, the time sequence consistency acquisition module, and the time domain spectrum acquisition module are respectively used to collect luminance chrominance threshold, time sequence consistency, and time domain spectrum, the morphological geometric template acquisition module is used to detect the morphological geometric template in the video image, the luminance chrominance threshold analysis module, the time sequence consistency analysis module, and the time domain spectrum analysis module are respectively used to analyze the detection results of the luminance chrominance threshold, the time sequence consistency, and the time domain spectrum, and the morphological geometric template analysis module is used to analyze various geometric templates; The variety of information synthesis module comprises a frame acquisition frequency adjustment module, a pixel space granularity adjustment module, a data normalization module, a highlight saturation spot type corresponding module, and a time recording module, the time recording module is used to count the time of the occurrence of a certain highlight saturation spot form again, the frame acquisition frequency adjustment module and the pixel space granularity adjustment module are respectively used to adjust the frame acquisition frequency and the pixel space granularity of each acquisition module, and the data normalization module is used to normalize the frame acquisition frequency and the pixel space granularity of each acquisition module; The highlight saturation spot judgment module comprises a geometric shape auxiliary acquisition module and a highlight saturation spot type judgment module, the geometric shape auxiliary acquisition module is used for analyzing the morphological geometric template change of the frame-by-frame disassembled video image to assist in judging the highlight saturation spot type, and the highlight saturation spot type judgment module is used for judging the highlight saturation spot type with the maximum probability.

[0014] Compared with the prior art, the present application has the following beneficial effects: the present application proposes a closed-loop mechanism of multi-source type detection and event triggering linkage aiming at the interference of highlight saturation spots in video on classification accuracy: luminance chroma threshold, time sequence consistency, time domain spectrum and morphological geometric template acquisition are respectively deployed in the input, time sequence, frequency spectrum and shape layers, and joint determination is performed according to the spot type dominant + auxiliary strategy; once a suspicious event is detected in any channel, the related modules are linked to cooperatively observe in a short window.

[0015] By unifying heterogeneous methods to two types of controllable parameters: frame acquisition frequency and pixel space granularity, and performing adaptive alignment, dynamic adjustment and alignment normalization, asynchronous, distortion of different scales and statistical deviation caused by direct fusion are avoided, so that the misjudgment rate is significantly reduced, and the time sequence consistency and explainability of video-level determination are improved. BRIEF DESCRIPTION OF DRAWINGS

[0016] The accompanying drawings are used to provide a further understanding of the present application, and constitute a part of the specification, together with the embodiments of the present application, to explain the present application, and do not constitute a limitation on the present application. In the drawings: Figure 1 It is the overall module structure schematic diagram of the present application. DETAILED DESCRIPTION

[0017] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0018] Please refer to Figure 1 The present application provides a technical solution: an AI-based video classification method, comprising the following steps: S0, in the frame-by-frame disassembled video frame sequence, the luminance chrominance threshold acquisition module is arranged in the image input layer and the preprocessing channel, used for extracting pixel-level luminance and chrominance features, the time sequence consistency acquisition module is embedded in the middle layer of the frame sequence analysis channel, used for capturing the time consistency of luminance and texture changes between adjacent frames, the time domain frequency spectrum acquisition module is arranged in the frequency analysis channel of the video signal after feature extraction, used for detecting inter-frame frequency oscillation and optical flow energy distribution, the morphological geometric template acquisition module is arranged in the high-level semantic modeling layer and the deconstruction layer, used for identifying geometric boundaries, contours and morphological structure features in the video frame; S1, each acquisition module works and detects the high-light saturation spot phenomenon, each acquisition module outputs the detection result with the initial frame acquisition frequency and pixel space granularity, and detects whether the high-light saturation spot phenomenon occurs through the corresponding analysis module; S2, when a certain acquisition module detects that a certain type of high-light saturation spot occurs, the other type of acquisition module most related to the feature domain of the acquisition module is adjusted in frame acquisition frequency and pixel space granularity, so that the acquisition module to be adjusted and the current acquisition module are unified in frame acquisition frequency and pixel space granularity; S3, with the continuation of the frequency, the frame acquisition frequency and the pixel space granularity of the acquisition module to be adjusted are dynamically adjusted according to whether the same high-light saturation spot phenomenon is continuously detected; S4, when the frame acquisition frequency and the pixel space granularity of each acquisition module are not unified, the detection data is processed by using the data normalization module to perform normalization processing on the frame acquisition frequency and the pixel space granularity; S5, based on the frame-by-frame structure analysis result, the spatial morphological features of the video frame are extracted and judged by using the morphological and geometric template method; In view of the problem that the optimal observation granularity of different detection methods is inconsistent, the frame acquisition frequency and the pixel space granularity are used as unified parameter bundles, the parameters of the module most related to the current event are unified (S2) at the triggering time, the dynamic adjustment is made (S3) in the continuous observation, and finally the cross-module alignment is realized through data normalization (S4). The process avoids the normalization distortion and statistical deviation caused by directly mixing asynchronous and different scale data, and ensures the consistency and comparability of the fusion judgment.

[0019] The dominant detection method and the auxiliary method of different spot types (such as the frequency flash artifact mainly in the time domain spectrum, and the lens system mainly in the time sequence consistency), when fused, preferentially use the evidence channel most matched with the physical / cause, and then use other channels for verification, reduce the deviation caused by strong related false images, and improve the classification accuracy.

[0020] In S0, it is clear that various high-light saturation spot types correspond to which detection method, which is specifically: S0-1, highlight saturation spot of scene optical interaction type, which is caused by object surface reflection, refraction and direct light from external light source, and the detection of luminance chrominance threshold is used as the main judgment method, and the time sequence consistency detection and time domain spectrum detection are used as auxiliary judgment methods; S0-2, highlight saturation spot of lens optical system type, which is formed by lens stains, oil film, halo, ghost and flare, and the time sequence consistency detection is used as the main judgment method, and the luminance chrominance threshold detection and morphological geometric template detection are used as auxiliary judgment methods; S0-3, highlight saturation spot of decoding post-processing artifact type, which is caused by compression ringing, blocking effect and over-sharpening, and the time domain spectrum detection is used as the main judgment method, and the luminance chrominance threshold detection is used as the auxiliary judgment method; S0-4, highlight saturation spot of time sequence motion induced type, which is caused by raindrops, snowdrops, stroboscopic light and fast movement, and the time sequence consistency detection is used as the main judgment method, and the time domain spectrum detection and luminance chrominance threshold detection are used as auxiliary judgment methods; The present application arranges luminance chrominance threshold, time sequence consistency, time domain spectrum and morphological geometric template four types of collection and analysis modules (see step S0, S0-1-S0-4 of right 2) in input layer, time sequence analysis layer, spectrum analysis layer and semantic structure layer, and performs typing detection on highlight saturation spots of scene optical interaction type, lens optical system type, decoding post-processing artifact type and time sequence motion induced type. Compared with a single method, the cross-domain feature mutual evidence effectively suppresses false positives caused by white background, over-sharpening, stroboscopic and compression artifacts, and significantly reduces the false positive rate.

[0021] In S2, the frame acquisition frequency and the pixel space granularity are unified, specifically: S2-1, each collection module detects the highlight saturation spot type corresponding to the main judgment method by default, and the detection result is output by using the default frame acquisition frequency and the pixel space granularity , wherein is the number of collection modules; S2-2, the first collection module detects the highlight saturation spot phenomenon, the default frame acquisition frequency is , and the default pixel space granularity is , then the initial probability of the occurrence of the highlight saturation spot phenomenon of this type is , at this time, the frame acquisition frequency and the pixel space granularity of the first collection module need to be adjusted, the default frame acquisition frequency and the pixel space granularity before adjustment are and , so that the adjusted frame acquisition frequency , and the adjusted pixel space granularity ; In S3, the dynamic adjustment is specifically as follows: S3-1, Order No. Each acquisition module maintains a frame acquisition frequency. and pixel space granularity continued During the time period, if Within the time period When the acquisition module detects the same type of bright saturation spot phenomenon again, observe the first... If each acquisition module simultaneously detects the same type of bright saturation spot phenomenon, and if so, the probability of this type of bright saturation spot phenomenon occurring is increased, with the adjusted probability being... ,in For the first When the first acquisition module is used as an auxiliary judgment method, it is related to the first... The probability increment caused by the detection of the same type of bright saturation spot phenomenon by each acquisition module, if the first acquisition module detects the same type of bright saturation spot phenomenon, If none of the acquisition modules simultaneously detected the same type of bright saturation spot phenomenon, then... ; S3-2, If in Within the time period When the first acquisition module does not detect the same type of bright saturation spot phenomenon, the second... The frame acquisition frequency and pixel spatial granularity of each acquisition module are determined by... and To its initial value and Gradually recovering, making the current distance The elapsed time of the end of the time period is Then the frame acquisition frequency at this time and pixel space granularity The calculation formulas are as follows: when hour, ,when hour, ,when hour, ,when hour, ,in , For frequency conversion factors; In S4, the normalization processing of frame acquisition frequency and pixel spatial granularity is specifically performed as follows: S4-1, Firstly in When the time period ends, the first The and the first The first acquisition frequency of each acquisition module is aligned, at which point both acquisition modules acquire data simultaneously. The next acquisition will only collect data that meets the specified frequency. and the detection data of the least common multiple moment, and other data is not collected; S4-2, the maximum and minimum of the detection data of the first and the second acquisition module are mapped into the interval of , eliminating the dimensional difference of pixel space granularity between different acquisition modules; In S5, the type of the highlight saturation spot is determined. Before each detection of the highlight saturation spot, the geometric feature distribution amount of the highlight saturation spot type detected by the morphological geometric template acquisition module is respectively , wherein is the number of geometric morphological types involved in the highlight saturation spot, and the geometric feature distribution amount refers to the response intensity of the template in the annular, strip, point, and radial morphologies. When a certain type of highlight saturation spot phenomenon is detected, if the type of the highlight saturation spot causes to change, and , the actual detection produces an incremental change, then the probability of the occurrence of the type of the highlight saturation spot is , wherein is the geometric feature incremental influence coefficient. If no incremental change of is detected, then , the final probability of is calculated in combination with S3-1, when , is the probability judgment threshold, and the type of the highlight saturation spot is determined to be correct; The occurrence probability of different types of highlight saturation spots is calculated by the change of the geometric feature distribution amount of the morphological geometric template, realizing quantitative and interpretable type determination. This method does not depend on single-frame brightness abnormalities, but captures the incremental response of geometric features such as annular, strip, point, and radial, and can effectively distinguish different spot types caused by optical reflection, lens stains, or motion artifacts. Compared with traditional threshold detection, this mechanism significantly reduces the false positives caused by local highlights, and maintains temporal consistency and determination stability in multi-frame fusion.

[0022] An AI-based video classification system includes an image recognition module, a multi-information synthesis module, and a highlight saturation spot judgment module. The image recognition module is used to collect brightness and chroma threshold values, temporal consistency, and time domain spectrum, and performs multi-dimensional detection in combination with a morphological geometric template. The multi-information synthesis module is used to perform unified frequency reference and granularity normalization processing on frame acquisition frequency and pixel space granularity, dynamically adjusts two output parameter characteristics according to subsequent detection results, and performs multi-source information synthesis processing. The highlight saturation spot judgment module is used to comprehensively determine the type of the highlight saturation spot based on the processed data. The image recognition module comprises a brightness chroma threshold acquisition module, a time sequence consistency acquisition module, a time domain spectrum acquisition module, a morphological geometric template acquisition module, a brightness chroma threshold analysis module, a time sequence consistency analysis module, a time domain spectrum analysis module, and a morphological geometric template analysis module. The brightness chroma threshold acquisition module, the time sequence consistency acquisition module, and the time domain spectrum acquisition module are respectively used for acquiring brightness chroma threshold, time sequence consistency, and time domain spectrum. The morphological geometric template acquisition module is used for detecting a disassembled morphological geometric template in a video image. The brightness chroma threshold analysis module, the time sequence consistency analysis module, and the time domain spectrum analysis module are respectively used for analyzing detection results of brightness chroma threshold, time sequence consistency, and time domain spectrum. The morphological geometric template analysis module is used for analyzing templates of various geometric shapes. The multiple information comprehensive module comprises a frame acquisition frequency adjustment module, a pixel space granularity adjustment module, a data normalization module, a highlight saturation spot type corresponding module, and a time recording module. The time recording module is used for counting time of reoccurrence of a certain highlight saturation spot form. The frame acquisition frequency adjustment module and the pixel space granularity adjustment module are respectively used for adjusting frame acquisition frequency and pixel space granularity of each acquisition module. The data normalization module is used for normalizing frame acquisition frequency and pixel space granularity of each acquisition module. The highlight saturation spot judgment module comprises a geometric shape auxiliary acquisition module and a highlight saturation spot type judgment module. The geometric shape auxiliary acquisition module is used for analyzing changes of disassembled morphological geometric templates of a video image frame by frame to assist in judging a highlight saturation spot type. The highlight saturation spot type judgment module is used for judging a maximum probability highlight saturation spot type.

[0023] It should be noted that the relational terms herein such as first and second and the like are used only to differentiate one entity or operation from another, and do not necessarily require or imply that any such entity or operation exists in any actual relationship or order. Moreover, the terms "include", "contain", and any other variant thereof are intended to cover non-exclusive inclusion, so that processes, methods, articles, and devices including a series of elements not only include those elements, but also include other elements not explicitly listed, or inherent to such processes, methods, articles, and devices.

[0024] Finally, it should be noted that the above only describes the preferred embodiments of the present application, and is not intended to limit the present application. Although the foregoing embodiments of the present application have been described in detail, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or equivalently replace some technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. An AI-based video classification method, characterized in that: Comprise the following steps: S0, in the frame-by-frame disassembly video frame sequence, luminance chroma threshold acquisition module is arranged in the image input layer and the pretreatment channel, for extracting the pixel level luminance and chroma characteristics, the time sequence consistency acquisition module is embedded in the middle layer of the frame sequence analysis channel, for capturing the time consistency of the luminance and texture changes between adjacent frames, the time domain spectrum acquisition module is arranged in the frequency analysis channel of the video signal after feature extraction, for detecting the frequency oscillation and the light flow energy distribution between frames, the morphological geometric template acquisition module is set in the high-level semantic modeling layer and the deconstruction layer, for identifying the geometric boundary, contour and morphological structure characteristics in the video frame; S1, each acquisition module works and detects the highlight saturation spot phenomenon, each acquisition module outputs the detection result with the initial frame acquisition frequency and pixel space granularity, and detects whether the highlight saturation spot phenomenon occurs through the corresponding analysis module; S2, when a certain acquisition module detects that a certain type of highlight saturation spot occurs, the other type of acquisition module most related to the feature domain of the acquisition module is adjusted in frame acquisition frequency and pixel space granularity, so that the acquisition module to be adjusted and the current acquisition module are unified in frame acquisition frequency and pixel space granularity; S3, with the continuation of the frequency, according to whether the same highlight saturation spot phenomenon is continuously detected, the frame acquisition frequency and the pixel space granularity of the acquisition module to be adjusted are dynamically adjusted; S4, when the frame acquisition frequency and the pixel space granularity of each acquisition module are not unified, the detection data is processed by using the data normalization module, and the frame acquisition frequency and the pixel space granularity are normalized; S5, based on the frame-by-frame structure analysis result, the spatial morphological characteristics of the video frame are extracted and judged by using the morphological and geometric template method.

2. The AI-based video classification method of claim 1, wherein: In the S0, it is clear that which detection method corresponds to each type of highlight saturation spot, which is specifically: S0-1, scene optical interaction type highlight saturation spot, that is, the highlight saturation spot caused by object surface reflection, refraction and external light source direct incidence takes the luminance chroma threshold detection as the dominant judgment method, and the time sequence consistency detection and the time domain spectrum detection as the auxiliary judgment method; S0-2, lens optical system type highlight saturation spot, that is, the highlight saturation spot formed by lens stains, oil film, halo, ghosting and flare takes the time sequence consistency detection as the dominant judgment method, and the luminance chroma threshold detection and the morphological geometric template detection as the auxiliary judgment method; S0-3, decoding post-processing artifact type highlight saturation spot, that is, the pseudo highlight generated by compression ringing, block effect and over-sharpening takes the time domain spectrum detection as the dominant judgment method, and the luminance chroma threshold detection as the auxiliary judgment method; S0-4, time sequence motion induced type highlight saturation spot, that is, the transient highlight area caused by raindrops, snowflakes, stroboscopic light and rapid movement takes the time sequence consistency detection as the dominant judgment method, and the time domain spectrum detection and the luminance chroma threshold detection as the auxiliary judgment method.

3. The AI-based video classification method of claim 2, wherein: In the S2, the unification in the frame acquisition frequency and the pixel space granularity is specifically: S2-1, each acquisition module detects the highlight saturated spot type corresponding to the leading judgment mode by default, and acquires frames at its default frame acquisition frequency and pixel spatial granularity output the detection result, wherein is the number of acquisition modules; S2-2, Order No. The acquisition module detected a bright saturation spot phenomenon, and its default frame acquisition frequency was [missing information]. The default pixel space granularity is The initial probability of this type of bright saturation spot phenomenon occurring is: At this point, it is necessary to... The frame acquisition frequency and pixel spatial granularity of each acquisition module were adjusted. The default frame acquisition frequency and pixel spatial granularity before the adjustment were as follows: and This makes its adjusted frame acquisition frequency Adjusted pixel space granularity .

4. The AI-based video classification method of claim 3, wherein: In the S3, the dynamic adjustment is specifically: S3-1, Order No. Each acquisition module maintains a frame acquisition frequency. and pixel space granularity continued During the time period, if Within the time period When the acquisition module detects the same type of bright saturation spot phenomenon again, observe the first... If each acquisition module simultaneously detects the same type of bright saturation spot phenomenon, and if so, the probability of this type of bright saturation spot phenomenon occurring is increased, with the adjusted probability being... ,in For the first When the first acquisition module is used as an auxiliary judgment method, it is related to the first... The probability increment caused by the detection of the same type of bright saturation spot phenomenon by each acquisition module, if the first acquisition module detects the same type of bright saturation spot phenomenon, If none of the acquisition modules simultaneously detected the same type of bright saturation spot phenomenon, then... ; S3-2、if in the time period, the first acquisition module does not detect the same type of highlight saturation phenomenon, the frame acquisition frequency and the pixel spatial granularity of the first acquisition module are gradually restored to the initial values of the first acquisition module and the second acquisition module, respectively, and the distance between the first acquisition module and the second acquisition module is gradually reduced to the initial distance between the first acquisition module and the second acquisition module. wherein f1 and f2 are frequency conversion coefficients.​​​​​​​​​​​​​​​​​​​​ 5. The AI-based video classification method of claim 4, wherein: In the S4, the normalization processing of the frame acquisition frequency and the pixel space granularity is specifically: S4-1, Firstly in When the time period ends, the first The and the first The first acquisition frequency of each acquisition module is aligned, at which point both acquisition modules acquire data simultaneously. The next acquisition will only collect data that meets the specified frequency. and The detection data is collected only at the least common multiple of the time; other data are not collected. S4-2, the maximum and minimum of the detection data of the first and second acquisition modules are mapped into the interval of [0, 1], eliminating the dimensional difference of pixel space granularity between different acquisition modules. ​​​ 6. The AI-based video classification method of claim 5, wherein: The type of the highlight saturation spot is determined in S5, specifically, before each detection of the highlight saturation spot, the morphological geometric template acquisition module detects the geometric feature distribution of the type of the highlight saturation spot , wherein is the number of geometric morphologies involved in the highlight saturation spot, and the geometric feature distribution refers to the response intensity of the template on the annular, strip, point and radial morphologies. When a certain type of highlight saturation spot is detected, if the type of the highlight saturation spot causes to change, and actually detects to produce an incremental change, then the probability of the type of the highlight saturation spot occurring is , wherein is a geometric feature incremental influence coefficient, if no incremental change is detected , the final probability of is calculated in combination with S3-1, when , is a probability judgment threshold, and the type of the highlight saturation spot is determined to be correct.

7. An AI-based video classification system, characterized by: The image recognition module is used for collecting luminance chroma threshold, time sequence consistency and time domain spectrum, and combining morphological geometric templates for multi-dimensional detection.

8. The AI-based video classification system of claim 7, wherein: The image recognition module includes luminance chroma threshold collection module, time sequence consistency collection module, time domain spectrum collection module, morphological geometric template collection module, luminance chroma threshold analysis module, time sequence consistency analysis module, time domain spectrum analysis module and morphological geometric template analysis module. The multi-information comprehensive module includes frame collection frequency adjustment module, pixel space granularity adjustment module, data normalization module, high-light saturated spot type corresponding module and time recording module. The high-light saturated spot judgment module includes geometric morphology auxiliary collection module and high-light saturated spot type judgment module.

Citation Information

Patent Citations

  • Video behavior recognition system and method based on space-time sequence model

    CN114743144A

  • Laser positioning marking method, device and equipment for helmet video acquisition

    CN119741378A

  • Video coding method and system based on multi-channel concurrent software and hardware mixing

    CN120455690A

  • Gynecological tumor image processing method and system based on AI multi-modal image analysis

    CN120747029A

  • Video processing for storage or transmission

    GB9517436D0