An image recognition tool control system for the field of aviation

Through the illumination posture recognition, contour coherence screening and time response extraction modules, the recognition error problem of traditional aerial image recognition systems under dynamic illumination and posture changes is solved, the stability and accuracy are improved, and image recognition results with temporal continuity and structural focus are formed.

CN120635488BActive Publication Date: 2025-10-21SHANGHAI KEZHI ELECTRIC AUTOMATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511129041.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-10-21
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

Traditional aerial image recognition tool control systems have difficulty handling dynamic lighting and posture changes when processing remote sensing image sequences, leading to recognition errors. The lack of a multi-frame target screening mechanism causes target misjudgment and recognition area drift.

Method used

The illumination posture recognition module, contour coherence screening module, time response extraction module and confidence fluctuation labeling module are used to process the posture changes, edge continuity, target frequency and confidence of the image frames respectively, eliminate unstable and error targets, and form a stable image target list.

Benefits of technology

The accuracy and stability of image recognition are improved. Through multi-frame target repeated response and confidence screening, the reliability and consistency of recognition results and image clarity are ensured, and the temporal continuity and structural focus features of recognition are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635488B_ABST
    Figure CN120635488B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image recognition, in particular to an image recognition tool control system for the aviation field, which comprises an illumination posture recognition module, a contour coherence screening module, a time response extraction module, a confidence fluctuation marking module and a collection frame reservation module.In the present application, the stable area is screened according to the posture and illumination change of the image frame, the clear image content is extracted in combination with the edge gray scale aggregation feature, the coherent closed area is reserved by eliminating the structure fracture and the twisted path, the image structure expression integrity is improved, the key entity appearing continuously is further identified based on the multiple frame target repeated response frequency, the scattered interference content is excluded, the fluctuation target is screened out by using the confidence section distribution stability, the recognition result is ensured to be reliable and consistent, finally the frame image recognition value is judged in combination with the image definition and the target concentration, the image sequence with time sequence continuity and structure focusing feature is formed, and the recognition accuracy and output stability are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and in particular to an image recognition tool control system for the aviation field. Background Art

[0002] Image recognition technology lies at the intersection of information processing and intelligent analysis. It primarily studies how computer systems can automatically detect, classify, and identify objects, scenes, and features in images. Core areas include image acquisition, image preprocessing, feature extraction, image understanding, and classification and recognition. This technology is widely used in a variety of high-tech applications, including security surveillance, facial recognition, autonomous driving, medical image analysis, industrial inspection, and aerospace. With the advancement of artificial intelligence and deep learning, image recognition technology has gradually evolved from traditional rule-based or shallow feature-based approaches to intelligent recognition methods based on large-scale data and deep neural networks, boasting enhanced autonomous learning capabilities and recognition accuracy. Traditional aviation-specific image recognition tool control systems are systems used to perform image data recognition and subsequent command control in aviation applications. These systems primarily focus on image data analysis and processing tasks in flight monitoring, navigation assistance, and target identification. These systems typically employ image recognition methods based on fixed algorithmic models, such as edge detection-based contour extraction, template matching-based image alignment, and principal component analysis-based feature dimension compression. These systems implement predetermined operational controls based on the image recognition results through embedded control units.

[0003] During the processing process, traditional systems only perform analysis tasks based on static image information and lack the ability to dynamically distinguish posture changes and lighting offsets between frames. When faced with scenes with strong environmental disturbances in remote sensing image sequences, it is difficult to filter out distorted frame content, causing subsequent recognition logic to be falsely triggered. The extraction of image areas relies on static edge calculation and fixed template matching, and is unable to identify the continuity of edge structures, resulting in structural fragmentation or local deformation areas being mixed into the target range. There is a lack of a screening mechanism for the frequency of occurrence of image targets between multiple frames, resulting in single-frame interference elements being mislabeled as stable targets, and no continuous analysis of the recognition confidence distribution is performed. When confidence drift occurs frequently, there is a lack of a elimination strategy, which can easily cause problems such as target misjudgment, feature overlap, and recognition area drift in image recognition applications. Summary of the Invention

[0004] The purpose of the present invention is to solve the shortcomings of the prior art and to propose an image recognition tool control system for the aviation field.

[0005] In order to achieve the above objectives, the present invention adopts the following technical solutions: An image recognition tool control system for the aviation field includes:

[0006] The illumination attitude recognition module obtains the attitude angle and illumination direction of the current image frame of the high-altitude fixed-wing remote sensing aircraft, determines whether the attitude change between adjacent frames exceeds the stability threshold, analyzes the grayscale distribution of the image edge blocks, and screens the image frames to generate a recognizable image frame label set;

[0007] The contour continuity screening module extracts the target edge path in the image based on the identifiable image frame marker set, evaluates the edge continuity and closure degree, eliminates structural breaks and distorted blocks, and screens structurally coherent and boundary-complete areas to form a structurally complete image target group;

[0008] The temporal response extraction module counts the number of consecutive appearances of the target in the frame sequence from the structure-complete image target group, removes the targets that appear in a single frame or scattered frames, marks the target areas that appear continuously in multiple frames, and constructs a multi-frame response image target set;

[0009] The confidence fluctuation marking module analyzes the confidence changes of the multi-frame response image target set, filters out targets with confidence fluctuations, retains only areas that maintain high confidence for a long time, and outputs a list of identified stable image targets.

[0010] As a further solution of the present invention, the recognizable image frame tag set includes posture angle data, lighting direction parameters, edge grayscale distribution characteristics, clarity discrimination tags, and image frame validity identifiers; the structural integrity image target group includes edge connectivity information, path consistency index, closure score, structural integrity tag, and boundary integrity area identifier; the multi-frame response image target set includes intra-frame continuous recognition count, frequent response target index, and multi-frame matching consistency tag; the identified stable image target list includes confidence value stability interval, confidence trend concentration, and confidence fluctuation range tags; the image recognition tool control results include stable target quantity index, edge clarity score, and multi-target aggregation evaluation;

[0011] The stability threshold refers to a critical value used to determine whether the posture change between adjacent image frames is too large. If the value exceeds this value, it is considered unstable;

[0012] The high-confidence region refers to an image target region that appears continuously in multiple frames of images and whose recognition confidence remains at a high level for a long time.

[0013] As a further solution of the present invention, the lighting gesture recognition module includes:

[0014] The attitude change determination submodule collects the aircraft attitude angle and illumination direction data based on the current image frame, calculates the pitch, roll, and yaw angle differences between adjacent image frames, calls the set attitude change stability limit threshold, compares and judges the attitude angle differences, selects image frames that do not exceed the stability limit, and obtains a stable image frame marker set;

[0015] The edge grayscale analysis submodule calls the image frame corresponding to the stable image frame marker set, extracts the edge area of ​​the image, calculates the grayscale value distribution concentration and the pixel gradient change amplitude of the edge area, generates the edge structure change value and the grayscale aggregation value, and performs a joint analysis to generate the edge aggregation structure strength;

[0016] The image clarity screening submodule sets edge clarity and light-dark structure thresholds based on the edge aggregation structure strength, judges the structure strength value and the threshold, screens image frames that meet the clarity and light-dark structure recognition conditions, calculates and obtains edge clarity matching values, and classifies and marks image frames with matching values ​​greater than the judgment threshold to obtain a set of identifiable image frame labels;

[0017] The attitude change stability limit threshold is a preset difference value for judging whether the pitch angle, roll angle and yaw angle changes between adjacent image frames are within a stable range;

[0018] The judgment threshold is a numerical limit for determining whether the image edge clarity matching value reaches the recognizable standard.

[0019] As a further solution of the present invention, the profile consistency screening module includes:

[0020] The edge path extraction submodule extracts the target edge path in the image based on the identifiable image frame marker set, counts the edge pixel sequence, connects adjacent pixels according to the index to form a contour segment, and obtains the edge path connection metric;

[0021] The direction continuity judgment submodule calls the edge path connection metric and calculates the direction angle difference and the weighted result of the path segment length based on the path segment direction sequence. It also combines the displacement vector amplitude and the breakpoint area density to calculate the path direction continuity improvement value, and screens it in combination with the set continuity reference interval to obtain the path direction continuity result.

[0022] The closed area screening submodule identifies whether the edge path forms a closed structure based on the path direction continuity result, detects the connectivity status of the closed area boundary and the number of distortion positions, eliminates the fractured and structurally abnormal areas, and obtains a structurally complete image target group;

[0023] The displacement vector amplitude is the vector distance of the intensity of spatial position change between adjacent pixels on the edge path, which measures the smoothness of the path continuity;

[0024] The continuity reference interval is the numerical range of the angle difference and path characteristic index for determining whether the direction change of the edge path is smooth and reasonable.

[0025] As a further solution of the present invention, the time response extraction module includes:

[0026] The target counting submodule traverses the target area in the image sequence frame by frame based on the structured complete image target group, and counts the number of times the target is continuously recognized in the frame sequence according to the target number to obtain the target continuous recognition frequency;

[0027] The frequency elimination and screening submodule calls the target continuous recognition frequency, screens the targets whose continuous recognition times reach the set threshold, eliminates the scattered targets that appear in a single frame and frame interval, and obtains the target response stability result;

[0028] The response target labeling submodule extracts and labels the target areas that remain in the continuous multiple frames according to the target response stability result, unifies the numbering and establishes the frame sequence index, and obtains the multi-frame response image target set;

[0029] The target whose number of consecutive recognitions reaches the set threshold value refers to a target area whose number of consecutive recognitions in the image sequence is not less than the preset number limit and has stable temporal characteristics;

[0030] The frame sequence index is a set of numbers that identifies the position and time sequence of the target in the image frame sequence, and tracks the spatiotemporal distribution of the multi-frame response target.

[0031] As a further solution of the present invention, the confidence fluctuation marking module includes:

[0032] The confidence sequence extraction submodule extracts the recognition confidence records of the target in the continuous frames based on the multi-frame response image target set, establishes a confidence value sequence according to the target number, and obtains the target confidence change trajectory;

[0033] The confidence concentration screening submodule calls the target confidence change trajectory, compares the distribution characteristics of the target confidence value in the frame sequence, screens the targets whose confidence remains within the concentrated segment interval, eliminates the areas with confidence fluctuation range, and obtains the confidence stable distribution interval;

[0034] The stable target marking submodule marks the target area with a continuously stable confidence according to the confidence stability distribution interval, records the number and the frame segment index, and obtains a list of identified stable image targets;

[0035] The confidence value sequence refers to a set of recognition confidence values ​​recorded in chronological order for the same target in consecutive image frames, reflecting the changing trend of recognition stability.

[0036] As a further solution of the present invention, the system further includes a collection frame retention module:

[0037] The acquisition frame retention module filters image frames from the identified stable image target list, determines the number of targets in the frame and the clarity of the contours, eliminates imaging blur and target sparse frames, retains image frames with structures and multi-target aggregation, and outputs image recognition tool control results.

[0038] As a further solution of the present invention, the acquisition frame retention module includes:

[0039] The target frame extraction submodule extracts the image frame number information corresponding to the stable target based on the identified stable image target list, locates and extracts the corresponding frame image from the original frame sequence, and obtains the frame index result where the target is located;

[0040] The image clarity judgment submodule calls the frame index result where the target is located, detects the edge clarity index of the target area in the image and the grayscale change of the overall image structure, determines the blur degree and structural boundary integrity, and obtains the image imaging clarity;

[0041] The frame sequence screening submodule determines whether the recognition area is clear and the target number threshold is met based on the image imaging clarity and the target number statistics, excludes sparse and blurred image frames, and obtains the image recognition tool control result;

[0042] The blur degree and structural boundary integrity are comprehensive indicators that measure whether the edge clarity of the target area in the image and the overall structural outline are continuous and complete, and evaluate the imaging quality of the image;

[0043] The target aggregation number threshold is a limit for determining whether the number of recognized targets in an image frame reaches a recognition density standard.

[0044] Compared with the prior art, the advantages and positive effects of the present invention are:

[0045] In the present invention, by eliminating structural breaks and distorted paths, retaining coherent closed areas, and improving the integrity of image structure expression, the continuously appearing key entities are further identified based on the frequency of repeated responses of multi-frame targets, and scattered interference content is eliminated. The fluctuating targets are screened out using the stability of the confidence segment distribution to ensure that the recognition results are reliable and consistent. Finally, the frame image recognition value is judged in combination with image clarity and target concentration, forming an image sequence with temporal continuity and structural focus characteristics, thereby enhancing recognition accuracy and output stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is a system flow chart of the present invention;

[0047] Figure 2 This is a flow chart of the lighting posture recognition module of the present invention;

[0048] Figure 3This is a flow chart of the profile coherence screening module of the present invention;

[0049] Figure 4 This is a flow chart of the time response extraction module of the present invention;

[0050] Figure 5 This is a flow chart of the confidence fluctuation marking module of the present invention;

[0051] Figure 6 This is a flow chart of the acquisition frame retention module of the present invention. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0053] In the description of the present invention, it should be understood that the terms "length," "width," "up," "down," "front," "back," "left," "right," "vertical," "horizontal," "top," "bottom," "inside," "outside," and the like, indicating positions or relationships, are based on the positions or relationships shown in the accompanying drawings and are intended only to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limiting the present invention. Furthermore, in the description of the present invention, "plurality" means two or more, unless otherwise expressly and specifically defined.

[0054] See also Figure 1 , an image recognition tool control system for the aviation field includes:

[0055] The illumination attitude recognition module obtains the attitude angle and illumination direction data of the current image frame of the high-altitude fixed-wing remote sensing aircraft, determines whether the attitude change amplitude between adjacent frames exceeds the stability limit, extracts the grayscale cluster distribution of the edge blocks in the image, judges the edge clarity and light-dark structure, selects image frames that meet the recognition clarity conditions, and obtains a set of recognizable image frame markers;

[0056] The contour continuity screening module extracts the target edge path in the image based on the recognizable image frame marker set, makes a continuity judgment on the direction, identifies the degree of edge closure, and removes image blocks with breaks and structural distortions. It selects target areas with coherent structures and closed boundaries as the retained objects, thus obtaining a structurally complete image target group.

[0057] The temporal response extraction module extracts the number of times each target is continuously recognized in the frame sequence from the structurally complete image target group, records the frequently appearing image entities, removes targets that appear in single frames or scattered frames, and marks the target areas that respond continuously in multiple consecutive frames to generate a multi-frame response image target set;

[0058] The confidence fluctuation marking module, based on the continuous confidence records of the targets in the multi-frame response image target set, screens whether the recognition value remains in the set segment for a long time, marks the targets with stable confidence performance, and removes the target areas with wide fluctuations in confidence values ​​to generate a list of stable recognition image targets;

[0059] The acquisition frame retention module extracts the frame image of the target from the list of identified stable image targets, determines the number of stable targets in the image and the clarity of the outline, excludes frames with sparse content and blurred imaging, retains image frames with clear identification areas and multi-target aggregation features, and outputs the image recognition tool control results.

[0060] The recognizable image frame label set includes posture angle data, lighting direction parameters, edge grayscale distribution characteristics, clarity discrimination labels, and image frame validity identification. The structural integrity image target group includes edge connectivity information, path consistency index, closure score, structural integrity label, and boundary integrity area identification. The multi-frame response image target set includes intra-frame continuous recognition count, frequent response target index, and multi-frame matching consistency mark. The recognition stable image target list includes confidence value stability interval, confidence trend concentration, and confidence fluctuation range label. The image recognition tool control results include stable target number index, edge clarity score, and multi-target aggregation evaluation.

[0061] See also Figure 2 , the illumination gesture recognition module includes:

[0062] The attitude change determination submodule collects the aircraft attitude angle and illumination direction data based on the current image frame, calculates the pitch, roll, and yaw angle differences between adjacent image frames, calls the set attitude change stability limit threshold, compares and judges the attitude angle differences, selects image frames that do not exceed the stability limit, and obtains a stable image frame marker set;

[0063] The attitude change judgment submodule aligns the three-axis attitude angle data of the onboard IMU with the real-time acquired image frames frame by frame through the synchronous clock during system initialization, and then records a unique number and millisecond timestamp for each frame. The module first reads the pitch angle, roll angle, and yaw angle values ​​of the current frame and stores them in the cache array, and then immediately reads the corresponding three-axis angles in the cache of the previous frame, and obtains the attitude angle difference vector between the current frame and the previous frame through three simple subtraction operations. The absolute value function is then called on the difference vector to unify the positive and negative directions, and then enters the stability threshold comparison link; the threshold is determined according to the aircraft model and control sensitivity during the system deployment stage, where the pitch angle stability threshold is defined as "the pitch angle change in a single frame shall not exceed half of the maximum pitch rate of a flight". In this embodiment The roll angle stability threshold is defined as "the roll angle change of a single frame shall not exceed the horizontal hold control margin" and is set to 1.0°. The yaw angle stability threshold is defined as "the yaw angle change of a single frame shall not be greater than the reference heading hold error" and is set to 1.0°. The module compares the three differences with the corresponding thresholds one by one. If all three are less than or equal to the threshold, the current frame index is written into the stable frame list. Otherwise, it is directly discarded and the next frame data is read. To ensure real-time performance, the module sets a maximum execution time of 5ms for each difference calculation and threshold comparison. When the sampling frequency is 50Hz, it can meet the online screening requirements. The final stable image frame marker set only contains the frame numbers and timestamps whose three-axis attitude angle changes are in the stable range, which is used as the input of the subsequent image processing link.

[0064] The edge grayscale analysis submodule calls the image frame corresponding to the stable image frame marker set, extracts the edge area of ​​the image, calculates the grayscale value distribution concentration and the pixel gradient change amplitude of the edge area, generates the edge structure change value and grayscale aggregation value, and performs a joint analysis to generate the edge aggregation structure strength;

[0065] After receiving the stable image frame mark, the edge grayscale analysis submodule first reads the corresponding original image in the list order. When performing the edge clipping operation on each image, a fixed 80px ring band is cropped from the four sides of the image to the center as the edge area. After the cropping is completed, the pixel values ​​in the area are immediately converted to grayscale and mapped to the 0-1 floating point interval. Then, a 5×5 sliding window is used to traverse the edge area with a window step of 1px. The standard deviation of the grayscale value of each pixel in the window is calculated and the average is accumulated to obtain the edge grayscale concentration of the frame; the grayscale concentration threshold is defined as "the minimum standard deviation mean that indicates whether the grayscale change in the edge area is sufficient to reflect the texture details". In the current embodiment, it is 0.15. When the concentration is lower than the threshold, it means that the overall grayscale difference of the image edge is too small and it is easy to produce aliasing. Blur; then, a differential operation is performed on the submodule in the horizontal and vertical directions respectively to obtain the average gradient amplitude in the two directions, and the square root of the square of the sum of the two is obtained to obtain the edge structure change value. The structure change threshold is defined as "the minimum gradient amplitude sufficient to reveal the clarity of the object contour", which is 0.18 here. If the structure change value is lower than the threshold, it means that the edge contour lacks sharpness; the processing flow writes the grayscale concentration and the structure change value into the frame feature queue and enters the next module. If both indicators are not lower than the corresponding threshold, the logical flag "edge feature meets the conditions" is marked in the frame header information, otherwise it is marked as "edge feature is insufficient" for rapid filtering by subsequent submodules. The entire analysis process is strictly controlled to be completed within 10ms to ensure real-time edge quality assessment of high frame rate image streams.

[0066] The image clarity screening submodule sets the edge clarity and light-dark structure thresholds based on the edge aggregation structure strength, judges the structure strength value and the threshold, and screens image frames that meet the clarity and light-dark structure recognition conditions using the formula:

[0067] ;

[0068] Obtain edge definition matching values ​​through calculation, classify and mark image frames whose matching values ​​are greater than a judgment threshold, and obtain a set of identifiable image frame marks;

[0069] in, represents the edge definition matching value, Representative Frame edge grayscale aggregation value, Representative Frame edge area structure change value, Representative The difference value of the brightness and darkness distribution of the frame image, Representative The frame definition reference value, Represents the total number of image frames, Indicates the summation process of all frames;

[0070] The image clarity screening submodule performs image screening based on the edge aggregation structure strength obtained above. Its core is to calculate the edge clarity matching value. , the formula is:

[0071] ;

[0072] Assume that this group of image frames has 3 frames in total (i.e. ), the parameters of each frame are as follows (described in text form):

[0073] Frame 1: Grayscale aggregation value , structural change value , the difference between bright and dark distribution , clarity benchmark value ;

[0074] Frame 2: Grayscale aggregation value , structural change value , the difference between bright and dark distribution , clarity benchmark value ;

[0075] Frame 3: Grayscale aggregation value , structural change value , the difference between bright and dark distribution , clarity benchmark value ;

[0076] Substitute the formulas in turn to calculate:

[0077] Calculation process of the first frame:

[0078] ;

[0079] ;

[0080] Calculation process of the second frame:

[0081] ;

[0082] ;

[0083] Calculation process of the third frame:

[0084] ;

[0085] ;

[0086] Summarize the overall clarity matching value:

[0087] ;

[0088] If the image clarity judgment threshold set by the system is 0.012, then due to , indicating that the overall clarity match of the image frames meets the recognition requirements, and all three frames are included in the set of recognizable image frame labels. This result indicates that the match value exceeds the clarity judgment benchmark, and the image edge information can be used for subsequent detection and analysis tasks. The formula is beneficial because by introducing the structural change intensity term and the combined influence of the brightness and darkness distribution differences, it eliminates the influence of misjudgment caused by uneven image contrast or blurred edges, thereby ensuring accurate screening of clear images in the overall system.

[0089] Dimensional normalization method of each parameter:

[0090] Normalization of grayscale aggregation values:

[0091] The grayscale aggregation value indicates the concentration of pixel grayscale distribution in the edge region of an image. Its original unit is grayscale, typically ranging from 0 to 255. To ensure uniformity in the calculation results and avoid interference caused by inconsistent units, the grayscale values ​​must be normalized. Specifically, the grayscale value of each pixel is divided by the maximum grayscale value of 255 to convert it to a decimal between 0 and 1. This normalized value is then statistically analyzed using a local sliding window to calculate the overall aggregation level, which is used as the weight term in subsequent product operations.

[0092] Normalization of structural change values:

[0093] The structural change value measures the intensity of image contour or gradient changes within edge regions, and its original unit is grayscale difference per pixel. Because the theoretical maximum value of gradients in an image can reach 255 grayscale levels per pixel or even higher, normalization is performed based on the maximum theoretical value. Normalization is performed by dividing the actual measured structural change amplitude by the set maximum gradient amplitude. For example, if the maximum reference value is 360, the value is normalized to a range between 0 and 1, making it comparable across different images or devices.

[0094] Normalization of the difference between bright and dark distribution values:

[0095] The brightness / darkness distribution difference value reflects the grayscale difference between high-light and low-light areas within the edge of the image and is a grayscale unit. Normalization also involves dividing by 255 to convert it to a dimensionless value, allowing it to be calculated uniformly with the structural change value, avoiding numerical imbalances caused by different scales.

[0096] Normalization of clarity benchmark value:

[0097] The clarity benchmark defines the minimum standard for an image to be considered "clear." Its original units are consistent with the structural change value and are normalized to the maximum reference value to maintain the same dimensionality as structural change and brightness differences. The normalized benchmark value is also a dimensionless value, ensuring consistency and comparability in subsequent calculations.

[0098] Definition of edge sharpness matching value:

[0099] The edge sharpness matching value is a unified evaluation metric for the clarity of edge structure strength across multiple frames. Its purpose is to screen images with complete edge structures and concentrated grayscale variations within a sequence of multiple frames. It simultaneously examines three characteristics—edge grayscale concentration, structural intensity variations, and brightness differences—and combines them with a clarity benchmark to produce a numerical value representing the overall image edge quality. A higher value indicates a more pronounced image, while a lower value indicates an image that is less recognizable.

[0100] The calculation principle of edge clarity matching value:

[0101] The algorithm first extracts pixel data from the edge of the image, normalizes its grayscale values, and evaluates the grayscale standard deviation of these pixels within the local area. This determines whether the grayscale distribution is concentrated, resulting in a value reflecting edge aggregation. The algorithm then measures the degree of change in the structural contours of the image—the magnitude of the grayscale gradient at the edge—and analyzes the overall grayscale distribution difference between bright and dark areas. Together, these two metrics form a composite indicator of image structural clarity.

[0102] Next, the difference between this structural clarity and a preset clarity benchmark is compared to determine whether the image performs better than the pre-set standard. This difference is then multiplied by the previously calculated grayscale concentration value to amplify the impact of grayscale-concentrated frames on the overall clarity score. After completing these steps for all frames, the average of the corresponding scores for each frame is taken to determine the overall clarity match value.

[0103] Finally, this matching value is compared with the screening threshold set by the system. If it is greater than the threshold, it means that the image clarity meets the recognition requirements at an average level and can be marked as a valid image; if it is lower than the threshold, the image edge is considered unclear and will not enter the recognition frame sequence.

[0104] See also Figure 3 , the profile coherence screening module includes:

[0105] The edge path extraction submodule extracts the target edge path in the image based on the identifiable image frame marker set, counts the edge pixel sequence, connects adjacent pixels according to the index to form contour segments, and obtains the edge path connection metric;

[0106] The edge path extraction submodule is based on a recognizable image frame marker set, loads image data frame by frame, and performs edge feature extraction operations. First, the grayscale image corresponding to the image frame is called and gradient enhancement processing is performed. After the image edge details are enhanced by the filter, a fixed threshold edge extraction operation is used. For example, the threshold pair is set to 80 and 160 to obtain an edge binary image. The obtained edge image is scanned in rows and columns to extract a set of edge points with a pixel value of 255. Each edge pixel records its two-dimensional coordinate position such as (45, 118). Then, the pixel connection process is executed to determine whether there is a grayscale connection between each edge point and its 8 neighboring pixels, that is, whether there is a continuous edge response. If so, an index connection is established between the two. For example, if the current pixel is (45, 118), if its adjacent position (46, 119) is also an edge point, the two points are connected as an edge line segment and then connected. The starting and ending points of the path segment are recorded in the data structure. After traversing all edge points, a set of path segments is generated. The number of pixels in each path segment is counted as the pixel length of the path segment. For example, if a path segment contains 25 pixels, its path length is 25. Then, sequential connection measurement is performed on the path segment to determine whether there is a break jump. A jump is defined as the difference in pixel coordinates between two consecutive pixel points in the path segment is greater than 2 pixel units. For example, if point A in a path segment is (50, 100) and point B is (53, 103), the difference is 3, which is a break point. The number of break points in all paths is counted, and the proportion of break points in the path length in each path segment is recorded as the basis for break density. Finally, the starting point, end point, pixel sequence, connectivity information and break position index of all path segments are output to form edge path connection measurement data for the next module call.

[0107] The direction continuity judgment submodule calls the edge path connection metric and calculates the direction angle difference and the weighted result of the path segment length based on the path segment direction sequence. At the same time, it combines the displacement vector amplitude and the breakpoint area density and uses the formula:

[0108] ;

[0109] The path direction continuity improvement value is obtained by calculation, and the set continuity reference interval is combined for screening to obtain the path direction continuity result;

[0110] in, is the continuity improvement value of the path direction, For the The direction angle of the segment path, For the The direction angles of adjacent path segments, is the angular fluctuation entropy, which measures the local directional stability. For path Segment length, is the modulus of the direction vector, indicating the magnitude of the velocity, For the The Euclidean distance between the starting and ending points of the segment, For the The local fracture density of the segment represents the number of fracture points per unit length. For the The average angle between the segment and the target direction, : is an adjustable weight coefficient used to control the weight of angle difference, fluctuation entropy, density and target deviation in the score.

[0111] Parameter setting basis and experimental range description:

[0112] 、 : represents the weight of the direction angle difference and the local direction fluctuation entropy;

[0113] : Experimentally, the value range is [1.5, 3.5], and the middle value is taken;

[0114] : Indicates the degree of control of directional deviation, and can still make distinctions under conditions of small angle deviation;

[0115] Path direction angle range: [0°, 180°];

[0116] The fracture density usually does not exceed 0.5 / pixel and is set in the interval [0, 0.5];

[0117] The measured range of directional entropy is [0.05, 0.3].

[0118] Example Setting: Number of Path Segments , set the following parameters:

[0119] Path segment 1:

[0120] ;

[0121] Path segment 2:

[0122] ;

[0123] Path segment 3:

[0124] ;

[0125] Molecular calculation (sum of directional weighted terms):

[0126] Path segment 1:

[0127] ;

[0128] Path segment 2:

[0129] ;

[0130] Path segment 3:

[0131] ;

[0132] The sum of the numerators is:

[0133] ;

[0134] Denominator calculation (calculated and summed independently for each path segment):

[0135] Path segment 1:

[0136] ;

[0137] Path segment 2:

[0138] ;

[0139] Path segment 3:

[0140] ;

[0141] The sum of the denominators is:

[0142] ;

[0143] Final calculated value:

[0144] ;

[0145] Result analysis:

[0146] According to the preset direction coherence reference interval: : non-coherent path; : moderately coherent; : Highly coherent;

[0147] Therefore, the directional continuity value calculated in this example is approximately 14.10, which is "moderately coherent" and can be retained for further screening.

[0148] Formula innovation description:

[0149] The benefit of this formula is that by weighting the directional angle difference and the directional entropy value together and incorporating them into the calculation, and coupling them with the length weight, it enhances the comprehensive judgment ability of local directional stability and directional consistency. At the same time, by introducing the cubic root of the fracture density to regulate the growth amplitude of the denominator, the error amplification problem caused by short path fractures is avoided, thereby providing a more stable and adjustable path direction continuity judgment indicator in the overall system.

[0150] Dimensional normalization method of each parameter:

[0151] In the original calculation formula, each parameter has a different physical dimension and numerical range. They need to be dimensionally normalized to make them comparable and able to participate in the calculation of the same formula. Normalization methods include linear normalization, logarithmic normalization, or power conversion, as follows:

[0152] Direction angle difference ;

[0153] The unit of this item is angle (°), the normalization method is linear normalization, the maximum direction difference is set to 180°, and the normalization formula is: angle difference divided by 180;

[0154] That is, the normalized range is [0, 1], which represents the degree of difference between the direction angles of the two path segments.

[0155] Directional Fluctuation Entropy ;

[0156] The entropy value is a dimensionless statistic, and its original range is generally between 0.05 and 0.3. The maximum and minimum normalization method is used, and the maximum value is set to 0.3 and the minimum value is 0. After normalization, it is in the range of [0, 1].

[0157] If the direction of the local path changes dramatically, the entropy value is high; if the direction is stable, the entropy value approaches 0.

[0158] Path length ;

[0159] The original unit is the number of pixels. The maximum path length is set to the longest edge path observed in the system, for example, 100 pixels. The normalization method is to divide the path length by 100;

[0160] If you need to retain the contribution of the path to the total score, you can use it as a weighting factor in the calculation after normalization.

[0161] Direction vector modulus ;

[0162] It is essentially the length of the direction vector, with units of pixels / frame (or displacement per unit time). In static image analysis, it is usually normalized to a unit vector (with a maximum value of 1), so no special normalization is required. If there is a speed greater than 1, it can be normalized according to the maximum vector length.

[0163] Euclidean distance between the start and end points ;

[0164] The unit is pixel. The maximum point distance within a reasonable path segment is set to 10 pixels. Normalized to the current point distance divided by 10, the range is [0, 1].

[0165] Local fracture density ;

[0166] The unit is the number of breakpoints divided by the path length (i.e., number of breaks / pixel). If the maximum density is set to 0.5, the normalization process is the current density divided by 0.5, resulting in a range of [0, 1].

[0167] Since this term is calculated in the form of one-third power, it remains positive after normalization to avoid the risk of a near-zero denominator due to small density.

[0168] Directional deviation angle ;

[0169] The unit is degree. The normalization method is to divide the deviation angle by the maximum possible deviation 180. The result range is [0, 1]. The smaller the deviation angle, the more consistent the path direction is with the target direction.

[0170] Weight coefficient

[0171] These are dimensionless coefficients and do not need to be normalized. Their function is to adjust the weight ratio of each item in the calculation. They can be set through experiments to a combination of values ​​that meets the expected output quality.

[0172] Definition of path direction continuity improvement value

[0173] The Path Directional Coherence Continuity Improvement value is a comprehensive measure of the directional consistency and continuity of a group of edge path segments in image space. This value is essentially a proportional metric, constructed by taking the weighted sum of the "directional similarity" and "directional stability" of multiple path segments as the numerator, and combining the "fracture strength," "directional deviation," and "displacement mutation" of each path segment as the denominator, as factors that hinder directional discontinuity.

[0174] By definition, the higher the value, the stronger the directional coherence and stability of the path as a whole; if the value is lower, it means that there are large angle mutations, dense fractures or directional drift in the path.

[0175] This indicator has good discriminability and is particularly suitable for the path screening stage in the image edge extraction process to eliminate edge path segments with unstable structures and obvious direction jumps.

[0176] The operation process follows the following basic principles:

[0177] Contribution of molecular term construction direction consistency

[0178] The degree of deviation from directional consistency for a single path segment is calculated by combining the absolute value of the directional difference between each segment and its adjacent segments with the entropy of the directional angle fluctuation (a statistical indicator of directional variation). This factor, multiplied by the path length, represents the path's contribution to overall directional consistency. The directional consistency contributions of all path segments are summed to form the overall directional consistency factor.

[0179] The denominator constructs the direction discontinuity penalty factor

[0180] The denominator is based on the path segment unit and constructs the interference term of directional continuity from three dimensions:

[0181] The distance between the starting and ending points of the path itself and the modulus of the unit direction vector are combined to determine the "linearity of the path";

[0182] The density of fracture points, treated to one-third power, reflects “local structural integrity”;

[0183] The average angle with the expected target direction constitutes the "direction deviation influence term"; the above three terms are combined and summed, and the denominator is used to represent the "direction continuity obstruction factor" of this path segment.

[0184] Construct the final index in the form of ratio

[0185] By dividing the directional consistency contribution by the barrier factor, we obtain a directional coherence score for each path segment, which is then used to determine whether to retain the path. This score is highly discriminatory, and in actual implementation, we set a coherence judgment interval to perform classification processing.

[0186] In summary, the design idea of ​​this value is to "reward the good and punish the bad": the scores of paths with small directional angle differences, low fluctuation entropy, stable paths, few breaks and close to the target direction will be significantly increased; otherwise, the scores will decrease until they are eliminated, completing the path direction continuity screening task.

[0187] The closed area screening submodule identifies whether the edge path forms a closed structure based on the path direction continuity results, detects the connectivity status of the closed area boundary and the number of distortion positions, eliminates fractures and structural abnormalities, and obtains a structurally complete image target group;

[0188] The closed area screening submodule calls the aforementioned path direction consistency result and performs closed structure recognition operation on all edge path segments that have passed the direction consistency judgment. First, the coordinate difference judgment is performed on the starting pixel and the ending pixel in the path segment. If the Euclidean distance between the two points is less than or equal to 2 pixel units, it is considered to constitute a closed path segment. For example, if the starting point is (120, 240) and the end point is (121, 239), the sum of the squares of the coordinate differences is 2, and the square root is approximately 1.41, which meets the closure condition. After confirming that the path segment forms a closed area, the boundary connectivity judgment is continued. By traversing all adjacent pixel points in the path, it is confirmed whether the image coordinate adjacency relationship is met. If there is a jump of more than 2 pixel units between any two points, it is marked as a distortion point, such as (85 , 67) is followed by (88, 71), and the difference is 5, which is discontinuous. The number of distortion points is counted and their positions are recorded in the path segment. Then, the internal fracture of each closed path is counted. The fracture point is determined by the interruption position of the previous path segment. The fracture density index is obtained by accumulating the total number of fracture points and dividing it by the total length of the path segment. For example, if the path segment length is 100 pixels and there are 6 fracture points, the fracture density is 6%. If the fracture density threshold is set to 8%, the path is still a valid area. If it exceeds the threshold, the path is eliminated. Repeat the judgment process to process all closed path segments, and finally select a set of closed areas with satisfactory boundary connectivity, complete structure, and distortion points below the upper limit as a structurally complete image target group. This structure is used for subsequent image matching and shape contour analysis.

[0189] See also Figure 4 , the time response extraction module includes:

[0190] The target counting submodule is based on the target group of the complete image. It traverses the target area in the image sequence frame by frame, and counts the number of times it is continuously recognized in the frame sequence according to the target number to obtain the target continuous recognition frequency.

[0191] First, the input continuous image frame sequence is decoded frame by frame, and the target detection operation is performed on each frame image. Here, the target detection adopts the region extraction method based on convolution features. All visible target areas are identified in each frame and assigned unique numbers. For example, three target areas are detected in the first frame, which are recorded as ID_01, ID_02, and ID_03 respectively. Next, the inter-frame matching process is entered. By traversing the frame by frame, the target numbers are compared and matched with the previous frame starting from the second frame. Three key parameters are extracted for each pair of target areas: the Euclidean distance of the center point, the area ratio, and the grayscale histogram correlation coefficient. In each group of target area comparisons, if the set threshold conditions are met, the two target areas are considered to be the same target continuous recognition results. For example, when the Euclidean distance of the center point is set to ≤5 pixels, the area ratio is between 0.85 and 1.15, and the grayscale histogram correlation coefficient is ≥0.75, they are considered to be the same target. The setting basis of the above thresholds is the target movement speed, camera frame rate and image resolution during the actual shooting process, through the following The process is determined. First, pedestrian targets with a moving speed of less than 0.5 m / s are collected in a sequence with a resolution of 1920×1080. At 30 frames / second, the maximum pixel position offset of the same target in consecutive frames does not exceed 5 pixels. Area changes are mainly caused by angle or slight occlusion. In the experiment, it was observed that an area ratio of 0.85 to 1.15 can cover most targets. Therefore, the area ratio threshold range is set based on this. The lower limit of 0.75 for the grayscale histogram correlation is based on statistical analysis of the impact of image illumination changes. Combined with 100 sets of target samples, it was found that more than 75% of the same targets in consecutive frames have a grayscale correlation coefficient above 0.75, so this is set as the judgment basis. During the execution process, all target numbers are accumulated with the number of times they meet the above matching conditions in the sequence. For example, if ID_12 meets the matching conditions in frames 2 to 7, it is counted as 6 times. The final structure data is {ID_12: 6, ID_08: 3, ID_03: 1, ...} for subsequent screening processing.

[0192] The frequency elimination and screening submodule calls the target continuous recognition frequency, screens the targets whose continuous recognition times reach the set threshold, eliminates scattered targets that appear in a single frame and frame interval, and obtains a stable target response result;

[0193] First, a mapping structure between the target number and its continuous recognition frequency is constructed. Then, the frequency value of each target is read in turn and compared with the preset threshold. The setting process of the preset threshold is configured according to different target response scenarios. For example, in public security monitoring, it is usually required that the target appear in at least 3 consecutive frames to be considered as a valid response target. The source of this threshold is based on the typical video acquisition frame rate of 25 frames / second. Considering factors such as rapid movement or short-term occlusion of the target, statistical analysis shows that under the condition of 3-frame continuous recognition, the target recognition accuracy reaches more than 93%. Therefore, the minimum response frequency threshold is set to 3. For frequencies below this threshold, the target recognition accuracy is 93%. During the execution, the frequency value of each number in the structure is traversed, and the targets corresponding to numbers less than 3 are directly removed from the data structure. For example, in {ID_12: 6, ID_08: 3, ID_03: 1}, ID_03 will be eliminated, and the remaining targets constitute a stable target response set. This operation avoids interfering targets caused by single-frame false detection or intermittent recognition from entering the final analysis stage. The threshold can be adjusted in different practical applications. For example, in industrial visual inspection, the image acquisition frequency is as high as 60 frames / second, and the target motion is stable, then the threshold can be increased to above 5 to increase the judgment accuracy.

[0194] The response target annotation submodule extracts and annotates the target areas that remain in the continuous multi-frame state according to the target response stability results, unifies the numbering and establishes the frame sequence index, and obtains the multi-frame response image target set;

[0195] First, extract the target number set that has been screened and retained, and trace back its frame position index in the sequence. For each target number, scan the image sequence frame by frame, record the frame number that meets the continuous recognition condition, and build a number-frame sequence mapping table. For example, the frame sequence corresponding to ID_12 is [2, 3, 4, 5, 6, 7]. Then, extract the bounding box coordinates of the corresponding target in each frame image, and mark the target area in the image frame as a rectangular box. At the same time, superimpose the number information in the upper left corner of the target, such as "ID_12". The marking color uses a unique identification color for easy distinction. Then unify the numbering rules for encoding to ensure that the consistency of the number across frames is not lost. For example, if ID_12 is in the image in the second frame, For pedestrians at the center position, although the regional position in subsequent frames has a slight offset but still meets the similarity judgment conditions, its number in each frame remains ID_12. Finally, all image frames with stable target numbers and their annotation data in all frames are organized into a target response sequence set. For example, the images with ID_12 annotation in frames [2, 3, 4, 5, 6, 7] are output as a set of data sets for subsequent analysis or visualization. The frame sequence index is also saved in a structured table, which contains three items: number, start frame, end frame, and cumulative number of frames. For example, {number: ID_12, start frame: 2, end frame: 7, cumulative number of frames: 6}, to ensure that subsequent target behavior analysis has index basis and consistency.

[0196] See also Figure 5 , the confidence fluctuation marking module includes:

[0197] The confidence sequence extraction submodule extracts the recognition confidence records of the target in consecutive frames based on the multi-frame response image target set, establishes a confidence value sequence according to the target number, and obtains the target confidence change trajectory;

[0198] First, a frame sequence table is constructed for each target number. All frames in the image sequence containing the target are traversed one by one. The confidence score returned by the recognition module is read for each frame. This score reflects the system's confidence in identifying the target as the correct category in the current frame and is usually expressed as a decimal. During the traversal process, the confidence scores of each frame are arranged in frame order to form a confidence change vector. Missing frames are marked. When handling missing frames, placeholder values ​​can be set to maintain sequence consistency. After confidence extraction, data is uniformly normalized for the confidence vectors of each target number to eliminate the scale inconsistency caused by score fluctuations between different frames, making it easier to perform subsequent trend analysis and change assessment. During this process, the relative position of the highest and lowest values ​​in the sequence is calculated, and the confidence data is linearly mapped to a unified interval to ensure that the confidence sequences of different targets have a consistent comparison basis. After completion, the system saves each target number and its corresponding confidence change sequence as the basic data structure for subsequent confidence feature analysis. The entire process achieves dynamic expression and abstract modeling of target confidence changes without changing the original image content.

[0199] The confidence concentration screening submodule calls the target confidence change trajectory, compares the distribution characteristics of the target confidence value in the frame sequence, screens the targets whose confidence remains within the concentrated segment interval, eliminates the areas with confidence fluctuation range, and obtains the confidence stable distribution interval;

[0200] The confidence concentration screening submodule calls the target confidence change trajectory, performs statistical analysis on the confidence sequence corresponding to each target number, and identifies the target with relatively stable confidence value changes. The execution process checks the fluctuation range and overall distribution of the confidence value in each sequence to determine whether its confidence performance has stable characteristics. First, the confidence change interval is identified, and the fluctuation amplitude is determined by calculating the difference between the highest and lowest confidence values ​​in the sequence. Then, the degree of dispersion within the sequence is examined to evaluate whether the confidence value is concentrated in the entire frame segment. If a target has similar confidence scores in multiple consecutive frames, changes slowly, and there is no obvious mutation or short-term fluctuation, it is considered to have a concentrated confidence distribution. In terms of the setting, we referred to the target recognition results in a large number of actual shooting videos, and selected the situations where the confidence values ​​were continuously distributed in multiple frame segments with small fluctuations as representative samples of stable recognition. We compared the differences in the confidence changes of different samples, and determined the reasonable range of fluctuation amplitude and concentration. During screening, the system compared the confidence sequence characteristics of the target numbers one by one, and only retained the target numbers whose confidence value changes were within the set range and without obvious abnormalities. Other targets with drastic confidence jumps or low average confidence values ​​in a short period of time will be eliminated, and finally a set of target numbers with stable confidence distribution will be formed, and their specific position intervals in the frame sequence will be attached for subsequent labeling processing.

[0201] The stable target annotation submodule annotates the target area with a continuously stable confidence level according to the confidence stability distribution interval, records the number and the frame index, and obtains a list of identified stable image targets.

[0202] The stable target annotation submodule annotates the corresponding targets in the image frame sequence according to the confidence stability distribution interval. It first reads the target number and its stable frame range retained in the previous stage. For each number, it extracts the target bounding box coordinates and confidence score information in the frame where it appears. The target annotation box is drawn on the original image and the target number and the confidence score of the corresponding frame are annotated within the box or at the edge. To improve visual recognition efficiency, a unified annotation style and color are used for all confident stable targets, making it easier for users to quickly identify such targets in the image. During the annotation process, if the target position in a frame shifts slightly but the number remains the same and the confidence score does not change suddenly, the original number is retained for annotation. The annotation remains consistent throughout the entire frame segment. The system also records auxiliary information such as the start frame, end frame, and confidence score change characteristics of each target to form a structured target index table, which provides an annotation basis for subsequent image retrieval or target tracking. The final output is an image frame sequence annotated with the target number and confidence score. The supporting structure data table clearly records the content, so that stable recognition targets can maintain consistent identification throughout the sequence.

[0203] See also Figure 6 , the acquisition frame retention module includes:

[0204] The target frame extraction submodule extracts the image frame number information corresponding to the stable target based on the identification of the stable image target list, locates and extracts the corresponding frame image from the original frame sequence, and obtains the frame index result where the target is located;

[0205] The target frame extraction submodule is based on identifying a list of stable image targets. It first reads the target number of each item in the list and its corresponding start and end frame information. It then indexes all frames in the frame sequence where this number appears. After indexing is completed, the corresponding frame is located in the original image sequence based on the frame index information, and the image file paths of these frames are stored in correspondence with the image index numbers. During the extraction process, it is necessary to ensure that the image reading order is consistent with the number correspondence to avoid image frame misalignment or number offset. Frames that do not contain stable target numbers are skipped during this process. For example, if the record number ID_03 in the stable image target list appears in frames 12 to 16, the submodule reads the 12th, 13th, 14th, 15th, and 16th frames from the image sequence and records the file index. At the same time, this number is written into the image frame label table, forming a bidirectional mapping relationship between the target number and its corresponding frame image. This mapping facilitates the subsequent rapid retrieval of the corresponding frame image based on the target number. After the extraction is completed, a complete frame image index data structure is formed. The storage format includes the target number, frame sequence number, and image path, providing image source information for the next stage image clarity assessment module.

[0206] The image clarity judgment submodule calls the target frame index result, detects the edge clarity index of the target area in the image and the grayscale change of the overall image structure, determines the blur degree and structural boundary integrity, and obtains the image clarity;

[0207] The image clarity judgment submodule calls the target frame index result, and performs edge structure recognition processing on the area containing the target number in each frame image. First, the target area boundary box area corresponding to the target number in the image is extracted, and the edge clarity analysis is performed after the image area is intercepted to determine whether the image has obvious edge blur. The clarity assessment mainly focuses on the gradient change of the target area contour edge and the continuity of the grayscale transition. After the edge area extraction is completed, the degree of change of the image grayscale distribution is counted and compared with the preset reference grayscale range. If the image has an obvious grayscale slow-changing area at the target boundary, it means that the edge blur is high. If the target boundary presents a sudden grayscale jump, it means that the edge clarity is high. On this basis, the image is analyzed. The overall grayscale structure change level of the image is calculated, and structural texture judgment is performed on the entire image frame. If the number of overall texture details of the image is small and the contrast of the local area is poor, it means that the image has an imaging blur problem. The image clarity judgment threshold is introduced in the clarity judgment process. The threshold is set according to the quality level of the manually annotated image samples in the experimental evaluation. By observing a large number of manually annotated image samples with clear structural boundaries and sufficient texture details as a reference, the edge grayscale change amplitude range and image grayscale contrast distribution characteristics are counted, and the minimum edge clarity value and the overall structural grayscale contrast range are set. The clarity judgment finally outputs the imaging clarity label corresponding to each frame of the image to form a frame image clarity evaluation result, which provides a basis for subsequent image screening.

[0208] The frame sequence screening submodule determines whether the recognition area is clear and the target number threshold is met based on the image clarity and the target number statistics, excludes sparse and blurred image frames, and obtains the image recognition tool control results;

[0209] The frame sequence screening submodule judges all frames to be processed frame by frame based on the image clarity and the target number statistics, and screens out image frames with high image blur and small number of targets. First, it reads the image clarity label to determine whether it is lower than the preset clarity standard. The standard is based on the dual indicators of edge gradient value and image grayscale contrast. If the image edge structure is not obvious or the overall grayscale distribution is flat, it is recorded as a low-definition frame. At the same time, the number of stable target numbers identified in each frame is read, the number is quantitatively evaluated, and compared with the set target aggregation threshold. The threshold is based on The average number of targets in the recognition image in a typical application scenario is set. In the monitoring scenario, three to five targets are a reasonable range. If the number of targets in a frame is less than two categories or the total number is too small, it is judged as a sparse content image frame. During the execution process, the dual conditions of low-definition frames and frames with insufficient target numbers are combined for judgment. If an image frame meets both conditions at the same time, it is directly removed from the frame sequence. If only one of them is met, the priority is determined to retain some frames for subsequent redundant analysis. Finally, a set of image frames with complete image structure and reasonable number of targets is obtained. This set can be called by downstream image recognition tools and execute subsequent analysis processes.

[0210] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still falls within the scope of protection of the technical solution of the present invention.

Claims

1. An image recognition tool control system for the aviation field, characterized in that: The system comprises: The illumination attitude recognition module obtains the attitude angle and illumination direction of the current image frame of the high-altitude fixed-wing remote sensing aircraft, determines whether the attitude change between adjacent frames exceeds the stability threshold, analyzes the grayscale distribution of the image edge blocks, and screens the image frames to generate a recognizable image frame label set; The contour continuity screening module extracts the target edge path in the image based on the identifiable image frame marker set, evaluates the edge continuity and closure degree, eliminates structural breaks and distorted blocks, and screens structurally coherent and boundary-complete areas to form a structurally complete image target group; The temporal response extraction module counts the number of consecutive appearances of the target in the frame sequence from the structure-complete image target group, removes the targets that appear in a single frame or scattered frames, marks the target areas that appear continuously in multiple frames, and constructs a multi-frame response image target set; The confidence fluctuation marking module analyzes the confidence changes of the multi-frame response image target set, filters out targets with confidence fluctuations, retains only areas that maintain high confidence for a long time, and outputs a list of identified stable image targets.

2. The image recognition tool control system for the aviation field according to claim 1, characterized in that: The recognizable image frame tag set includes posture angle data, lighting direction parameters, edge grayscale distribution characteristics, clarity discrimination tags, and image frame validity identifiers. The structural integrity image target group includes edge connectivity information, path consistency index, closure score, structural integrity tag, and boundary integrity area identifier. The multi-frame response image target set includes intra-frame continuous recognition count, frequent response target index, and multi-frame matching consistency tag. The identified stable image target list includes confidence value stability interval, confidence trend concentration, and confidence fluctuation range tags. The image recognition tool control results include stable target quantity index, edge clarity score, and multi-target aggregation evaluation. The stability threshold refers to a critical value used to determine whether the posture change between adjacent image frames is too large. If the value exceeds this value, it is considered unstable; The high-confidence region refers to an image target region that appears continuously in multiple frames of images and whose recognition confidence remains at a high level for a long time.

3. The image recognition tool control system for the aviation field according to claim 2, characterized in that: The illumination gesture recognition module includes: The attitude change determination submodule collects the aircraft attitude angle and illumination direction data based on the current image frame, calculates the pitch, roll, and yaw angle differences between adjacent image frames, calls the set attitude change stability limit threshold, compares and judges the attitude angle differences, selects image frames that do not exceed the stability limit, and obtains a stable image frame marker set; The edge grayscale analysis submodule calls the image frame corresponding to the stable image frame marker set, extracts the edge area of ​​the image, calculates the grayscale value distribution concentration and the pixel gradient change amplitude of the edge area, generates the edge structure change value and the grayscale aggregation value, and performs a joint analysis to generate the edge aggregation structure strength; The image clarity screening submodule sets edge clarity and light-dark structure thresholds based on the edge aggregation structure strength, judges the structure strength value and the threshold, screens image frames that meet the clarity and light-dark structure recognition conditions, calculates and obtains edge clarity matching values, and classifies and marks image frames with matching values ​​greater than the judgment threshold to obtain a set of identifiable image frame labels; The attitude change stability limit threshold is a preset difference value for judging whether the pitch angle, roll angle and yaw angle changes between adjacent image frames are within a stable range; The judgment threshold is a numerical limit for determining whether the image edge clarity matching value reaches the recognizable standard.

4. The image recognition tool control system for the aviation field according to claim 3, characterized in that: The profile coherence screening module includes: The edge path extraction submodule extracts the target edge path in the image based on the identifiable image frame marker set, counts the edge pixel sequence, connects adjacent pixels according to the index to form a contour segment, and obtains the edge path connection metric; The direction continuity judgment submodule calls the edge path connection metric and calculates the direction angle difference and the weighted result of the path segment length based on the path segment direction sequence. It also combines the displacement vector amplitude and the breakpoint area density to calculate the path direction continuity improvement value, and screens it in combination with the set continuity reference interval to obtain the path direction continuity result. The closed area screening submodule identifies whether the edge path forms a closed structure based on the path direction continuity result, detects the connectivity status of the closed area boundary and the number of distortion positions, eliminates the fractured and structurally abnormal areas, and obtains a structurally complete image target group; The displacement vector amplitude is the vector distance of the intensity of spatial position change between adjacent pixels on the edge path, which measures the smoothness of the path continuity; The continuity reference interval is the numerical range of the angle difference and path characteristic index for determining whether the direction change of the edge path is smooth and reasonable.

5. The image recognition tool control system for the aviation field according to claim 4, characterized in that: The time response extraction module includes: The target counting submodule traverses the target area in the image sequence frame by frame based on the structured complete image target group, and counts the number of times the target is continuously recognized in the frame sequence according to the target number to obtain the target continuous recognition frequency; The frequency elimination and screening submodule calls the target continuous recognition frequency, screens the targets whose continuous recognition times reach the set threshold, eliminates the scattered targets that appear in a single frame and frame interval, and obtains the target response stability result; The response target labeling submodule extracts and labels the target areas that remain in the continuous multiple frames according to the target response stability result, unifies the numbering and establishes the frame sequence index, and obtains the multi-frame response image target set; The target whose number of consecutive recognitions reaches the set threshold value refers to a target area whose number of consecutive recognitions in the image sequence is not less than the preset number limit and has stable temporal characteristics; The frame sequence index is a set of numbers that identifies the position and time sequence of the target in the image frame sequence, and tracks the spatiotemporal distribution of the multi-frame response target.

6. The image recognition tool control system for the aviation field according to claim 5, characterized in that: The confidence fluctuation marking module includes: The confidence sequence extraction submodule extracts the recognition confidence records of the target in the continuous frames based on the multi-frame response image target set, establishes a confidence value sequence according to the target number, and obtains the target confidence change trajectory; The confidence concentration screening submodule calls the target confidence change trajectory, compares the distribution characteristics of the target confidence value in the frame sequence, screens the targets whose confidence remains within the concentrated segment interval, eliminates the areas with confidence fluctuation range, and obtains the confidence stable distribution interval; The stable target marking submodule marks the target area with a continuously stable confidence according to the confidence stability distribution interval, records the number and the frame segment index, and obtains a list of identified stable image targets; The confidence value sequence refers to a set of recognition confidence values ​​recorded in chronological order for the same target in consecutive image frames, reflecting the changing trend of recognition stability.

7. The image recognition tool control system for the aviation field according to claim 6, characterized in that: The system also includes an acquisition frame retention module: The acquisition frame retention module filters image frames from the identified stable image target list, determines the number of targets in the frame and the clarity of the contours, eliminates imaging blur and target sparse frames, retains image frames with structures and multi-target aggregation, and outputs image recognition tool control results.

8. The image recognition tool control system for the aviation field according to claim 7, characterized in that: The acquisition frame retention module includes: The target frame extraction submodule extracts the image frame number information corresponding to the stable target based on the identified stable image target list, locates and extracts the corresponding frame image from the original frame sequence, and obtains the frame index result where the target is located; The image clarity judgment submodule calls the frame index result where the target is located, detects the edge clarity index of the target area in the image and the grayscale change of the overall image structure, determines the blur degree and structural boundary integrity, and obtains the image imaging clarity; The frame sequence screening submodule determines whether the recognition area is clear and the target number threshold is met based on the image imaging clarity and the target number statistics, excludes sparse and blurred image frames, and obtains the image recognition tool control result; The blur degree and structural boundary integrity are comprehensive indicators that measure whether the edge clarity of the target area in the image and the overall structural outline are continuous and complete, and evaluate the imaging quality of the image; The target aggregation number threshold is a limit for determining whether the number of recognized targets in an image frame reaches a recognition density standard.

Citation Information

Patent Citations

  • Rapid three-dimensional reconstruction method for unmanned aerial vehicle mine inspection scene

    CN120339534A

  • Image segmentation method and apparatus, and device and storage medium

    WO2022133627A1