Intelligent image recognition method for multimedia ROI region
Patent Information
- Application Number
- CN202610899514.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-09-25
AI Technical Summary
若识别系统无法根据ROI置信度图对候选区域进行边界自适应修正和局部增强,也无法依据连续时刻的识别置信度计算异常累积置信度,则难以区分瞬时干扰与真实异常发展过程,进而影响预警的准确性和及时性
本发明区别于现有整图识别或固定框截取方式,核心在于以可见光帧、红外热成像帧和压缩关键帧中的瞬态变化边界进行时间同步,并结合目标部件布局模板形成包含候选框和背景校正环带的初始感兴趣区域候选区域。该手段使不同媒体源中同一微小异常位置在时间和空间上对应,背景校正环带又能扣除反光、热扩散和压缩失真带来的环境干扰,从而解决微小感兴趣区域被整图背景淹没、不同媒体源错位导致误检的问题。
Smart Images

Figure CN122821085A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimedia image recognition technology, and more specifically to an intelligent image recognition method for multimedia ROI regions. Background Technology
[0002] Multimedia image recognition technology is widely used in scenarios such as industrial inspection, equipment monitoring, security early warning, and abnormal status identification. Existing methods typically acquire visible light images, infrared thermal images, or compressed video keyframes to perform target detection or image classification on the scene to be identified, thereby determining whether there are abnormal targets in the scene. However, in complex scenarios such as energy storage compartments, power cabinets, underground pipe corridors, and enclosed machine rooms, abnormal targets often do not appear in the salient areas of the entire image, but are concentrated in small ROI areas near specific components, such as interface gaps, heat dissipation holes, pressure relief vents, connection terminals, or partially obscured edges. These types of ROI areas are small, have weak boundaries, and change rapidly. If only the whole image is recognized or a single image source is used for judgment, the effective features are easily overwhelmed by background information.
[0003] In practical applications, visible light frames are easily affected by light flicker, reflections, shadows, and low-light noise; infrared thermal imaging frames are easily affected by heat diffusion, thermal ghosting, and low resolution; and compressed keyframes are prone to block artifacts, edge blurring, and loss of detail. Due to differences in acquisition time, imaging scale, and viewing angle between different media sources, the same anomalous ROI does not behave completely consistently in different images. Existing ROI identification methods mostly use fixed template selection or single-frame candidate box detection, which makes it difficult to simultaneously utilize media quality features, spatial texture features, cross-media consistency features, and inter-frame micro-change features. This leads to the misjudgment of reflections, condensation, noise, or compression artifacts as anomalous ROIs when there is slight shift in the target area, local occlusion, or image compression distortion. It also easily misses small, real, but still in the early stages of anomalous regions.
[0004] Especially in unattended or high-risk equipment monitoring scenarios, abnormal ROIs typically exhibit temporal characteristics such as gradual diffusion across consecutive frames, slow accumulation of temperature or brightness, and directional boundary shifts. If the recognition system cannot adaptively correct and locally enhance candidate regions based on the ROI confidence map, nor can it calculate the cumulative confidence of anomalies based on the recognition confidence at consecutive time points, it becomes difficult to distinguish between transient interference and the actual development process of anomalies, thus affecting the accuracy and timeliness of early warnings. Therefore, accurately determining the target ROI region in a multimedia image sequence and combining multi-source features, local enhancement, and continuous temporal verification to achieve intelligent recognition of the ROI region has become an urgent technical problem to be solved. Summary of the Invention
[0005] The purpose of this invention is to provide an intelligent image recognition method for multimedia ROI regions to address the shortcomings of the prior art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: an intelligent image recognition method for multimedia ROI regions, comprising: The system acquires visible light frames, infrared thermal imaging frames, and compressed keyframes of the scene to be identified, and performs time synchronization based on transient change boundaries to obtain a multimedia image sequence. Based on the target component layout template, the multimedia image sequence is coarsely localized to obtain the initial region of interest candidate region containing candidate boxes and background correction rings. Media quality features, spatial texture features, cross-media consistency features and inter-frame micro-variation features are extracted to form candidate feature groups. A confidence map of the region of interest is formed based on the candidate feature groups, and the target region of interest is determined. Boundary correction and local enhancement are performed based on the confidence gradient and inter-frame micro-change features to obtain the image patch of the target region of interest. Based on the target region of interest image patch and candidate feature group, the region of interest category, location boundary and recognition confidence are obtained, and the final region of interest recognition result is output according to the recognition confidence at consecutive time steps.
[0007] Preferably, time synchronization based on transient change boundaries includes: The original acquisition order of visible light frames, infrared thermal imaging frames, and compressed keyframes within the same acquisition cycle is preserved and unified to the same image area; Extract the visible light variation boundary, infrared variation boundary, and compressed keyframe variation boundary formed by brightness and temperature changes in adjacent frames, respectively. Using the visible light change boundary as a reference, the infrared thermal imaging frames and compressed key frames are shifted and calibrated. Frame groups that meet the preset conditions in terms of boundary center distance, boundary overlap ratio, and change direction are used as synchronous frame groups and arranged in the acquisition order to obtain a multimedia image sequence.
[0008] Preferably, coarse positioning is performed based on the target component layout template, including: Extract the edge of the target component structure within the same frame area of the synchronized frame group, and determine the positioning reference line based on the distance between the edges of adjacent target component structures; The reference position of the target component is determined based on the direction, position, and spacing of the positioning baseline and its fit with the layout template of the target component. Candidate boxes are captured centered on the reference position of the target component, and a background correction ring is formed outside the candidate boxes. The candidate boxes and the background correction ring are used together as the initial region of interest candidate regions.
[0009] Preferably, the extraction of media quality features includes: Calculate the average edge intensity of the candidate box and the background correction ring separately, and use the difference between the two as the sharpness difference; The number of block breaks is counted in the compressed keyframes, the reflective saturation area is counted in the visible light frames, and the occlusion ratio is calculated based on the edge length that should appear in the target component layout template and the actual effective edge length. Media quality characteristics are composed of poor clarity, number of blocky breaks, reflective saturation area, and proportion of obscured edges.
[0010] Preferably, the extraction of spatial texture features and cross-media consistency features includes: Within the candidate box, the length of edge discontinuity, the consistency of fine line direction, and the distribution of gray-level abrupt changes are statistically analyzed to form spatial texture features; Unify the candidate bounding boxes in visible light frames, infrared thermal imaging frames, and compressed keyframes to the same pixel coordinates; The overlap values at boundary locations, brightness change locations, and temperature change locations are calculated separately to form cross-media consistency features.
[0011] Preferably, the extraction of inter-frame micro-variation features includes: Calculate the displacement direction of the candidate box center point at the current time relative to the candidate box center point at the previous time, and the displacement direction of the candidate box center point at the previous time relative to the candidate box center point at the previous two time points; The outer contour expansion range is obtained by the difference between the effective edge enclosing area of the candidate box at the current moment and the average effective edge enclosing area of the previous two moments. Based on the difference between the average brightness and average temperature of the candidate boxes at the current moment and the corresponding average values at the previous two moments, the local brightness and temperature increment is obtained, and together with the media quality features, spatial texture features, and cross-media consistency features, a candidate feature group is formed.
[0012] Preferably, forming a confidence map of the region of interest and determining the target region of interest includes: A map is constructed using the pixel coordinates of the initial candidate region of interest, and quality correction values are assigned to the locations of block breaks, reflective saturation, occluded edges, and continuous effective edges based on media quality characteristics. The location of local anomalies is determined based on spatial texture features, and the location corresponding to the same gray-level abrupt change in the background correction ring is subtracted. Based on cross-media consistency characteristics and inter-frame micro-change characteristics, continuous enhancement positions are determined. The quality correction value, local anomaly indication position, and continuous enhancement position are written into the same map to obtain the region of interest confidence map. Connected regions that meet the preset conditions in terms of both confidence and connectivity area are determined as the target region of interest.
[0013] Preferably, boundary correction includes: Pixel-by-pixel sampling is performed along the initial boundary of the target region of interest, both inward and outward, and the stable boundary position is determined based on the continuous decrease in confidence between adjacent sampling positions. Based on the candidate box center offset direction, outer contour expansion range, and local brightness temperature increment, determine the outward expansion correction side and the inward contraction correction side; The contour of the target region of interest is reconstructed according to the stable position of the boundary. The outward correction side is included in the target region of interest and the inward correction side is removed, resulting in a corrected region of interest whose boundary is limited to the candidate region of the initial region of interest.
[0014] Preferred, localized enhancement includes: Calculate the brightness difference, temperature difference, and edge intensity difference between the corrected region of interest and the background correction ring, respectively. Based on the average brightness and average temperature of the background correction ring, local contrast enhancement and thermal difference enhancement are performed on the correction region of interest; The discontinuous edges within the region of interest with an interval not exceeding a preset number of pixels and a confidence level meeting a preset condition are continuously enhanced, and then cropped to obtain the target region of interest image block.
[0015] Preferably, the final region of interest identification result is output, including: Extract the concentrated brightness locations, concentrated temperature locations, and continuous edge locations from the target region of interest image block, and match their locations with candidate feature groups to form a recognition feature sequence; Based on the correspondence between reflective saturation, blocky fracture, occlusion edge break, unidirectional gray-scale abrupt change, synchronous brightness and temperature change and continuous boundary expansion in the identification feature sequence, calculate the category confidence of normal area, reflective interference area, compression artifact area and abnormal warning area, and determine the category and location boundary of region of interest; The cumulative confidence of anomalies is calculated based on the recognition confidence, the degree of overlap of location boundaries, and the continuity of local brightness temperature increments at the current time and the two previous time points. When the cumulative confidence of anomalies reaches the preset output value, the final region of interest recognition result is output.
[0016] The technical effects and advantages provided by the present invention in the above technical solution are as follows: This invention differs from existing whole-image recognition or fixed-frame extraction methods. Its core lies in temporal synchronization using transient change boundaries in visible light frames, infrared thermal imaging frames, and compressed keyframes, combined with a target component layout template to form an initial region of interest (ROI) candidate region containing candidate boxes and a background correction ring. This method ensures that the same minute anomaly location in different media sources corresponds in both time and space. Furthermore, the background correction ring can eliminate environmental interference caused by reflections, thermal diffusion, and compression distortion, thereby solving the problems of minute ROIs being submerged by the overall image background and false detections due to misalignment between different media sources.
[0017] The invention further differs in that it does not directly output results based on single-frame brightness or temperature anomalies. Instead, it incorporates media quality features, spatial texture features, cross-media consistency features, and inter-frame micro-variation features into the region of interest confidence map. It then utilizes the confidence gradient and brightness-temperature changes over consecutive time intervals to expand, shrink, and locally enhance the boundaries. This method can separate compression artifacts, reflective saturation, occlusion edges from the truly continuously expanding anomaly regions, making the final location boundary more closely match the actual anomaly edge and reducing false alarms caused by transient noise. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0019] Figure 1 This is a flowchart of an intelligent image recognition method for multimedia ROI regions according to the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] Example 1, please refer to Figure 1 As shown in this embodiment, an intelligent image recognition method for multimedia ROI regions includes: The visible light frame, infrared thermal imaging frame, and compressed key frame of the scene to be identified are acquired and time-synchronized to obtain a multimedia image sequence.
[0022] In one embodiment, when acquiring visible light frames, infrared thermal imaging frames, and compressed keyframes of the scene to be identified, the original acquisition order of each type of frame is preserved within the same acquisition cycle. Visible light frames are converted into visible light grayscale images based on their brightness channels, infrared thermal imaging frames are converted into infrared grayscale images based on their temperature values, and compressed keyframes are decoded to form keyframe grayscale images based on their brightness channels. Visible light grayscale images, infrared grayscale images, and keyframe grayscale images acquired sequentially within the same acquisition cycle are grouped into a frame group to be calibrated. To avoid incomparability of boundary positions due to different imaging scales, the infrared grayscale images and keyframe grayscale images in the frame group to be calibrated are scaled proportionally to the pixel size of the visible light grayscale images, and the overlapping area at the center of the image, which is included in all three types of grayscale images, is extracted as the same image area.
[0023] When extracting transient change boundaries in a frame group to be calibrated, the pixel change between two adjacent frames is first calculated for each type of grayscale image within the same image area. The pixel change is the absolute value of the difference between the grayscale value of a pixel in the current frame and the grayscale value of the pixel at the same position in the previous frame. Then, all pixel changes within the same image area are arranged in ascending order, and the value in the middle is taken as the median change. The absolute value of the difference between each pixel change and the median change is then calculated, and all absolute values are averaged to obtain the average discrete change. The median change is added to twice the average discrete change to obtain the change judgment value. Pixels whose pixel change reaches the change judgment value are recorded as changed pixels, and pixels that do not reach the change judgment value are recorded as stable pixels. Pixels among the changed pixels that have at least one adjacent stable pixel are recorded as boundary pixels, and the adjacency relationship uses eight directions: top, bottom, left, right, and four diagonal directions. Boundaries with eight consecutive boundary pixels are used as transient change boundaries, thereby obtaining visible light change boundaries, infrared change boundaries, and compressed keyframe change boundaries.
[0024] During time synchronization, the infrared thermal imaging frames and compressed keyframes are shifted and calibrated along the original acquisition sequence, using the visible light change boundary as a reference. The shift range is limited to two frames forward, the current frame, and two frames backward. After each shift, the matching deviation value between the infrared change boundary or compressed keyframe change boundary and the visible light change boundary is calculated. The matching deviation value is obtained by adding three terms: the first term is the pixel distance between the center points of the two change boundaries; the second term is the proportion of overlapping pixels of the two change boundaries to the number of pixels of the reference change boundary, and subtracting this proportion from 1 gives the boundary overlap deviation; the third term is the change direction deviation, which is determined by the movement direction of the current boundary center point relative to the corresponding boundary center point of the previous frame, and the angle between the two movement directions is divided by 180 degrees to obtain the change direction deviation. The shift position with the lowest matching deviation value is taken as the calibration position.
[0025] Once the infrared thermal imaging frame and the compressed keyframe have both achieved their calibration positions, the visible light frame, the calibrated infrared thermal imaging frame, and the calibrated compressed keyframe are combined into a synchronization frame group. Within the synchronization frame group, if the center point distance between the visible light change boundary, the infrared change boundary, and the compressed keyframe change boundary meets the following conditions: no more than 3 pixels between their centers, no less than 70% overlap, and an angle between their change directions not exceeding 22.5 degrees, the synchronization frame group is retained; otherwise, it is discarded. Finally, according to the acquisition sequence corresponding to the synchronization frame group, each synchronization frame group is arranged sequentially to obtain a multimedia image sequence. By utilizing the transient change boundary formed by brightness and heat changes to achieve inter-frame alignment, even in the absence of a unified hardware trigger signal or when there is a delay in the compressed keyframe, the temporal position of the same anomalous ROI remains consistent across the visible light frame, the infrared thermal imaging frame, and the compressed keyframe, providing stable input for subsequent initial ROI candidate region determination.
[0026] Based on the target component layout template, images in the multimedia image sequence are coarsely located to obtain initial ROI candidate regions. Media quality features, spatial texture features, cross-media consistency features, and inter-frame micro-variation features are extracted to form ROI candidate feature groups.
[0027] In one embodiment, when coarsely locating images in a multimedia image sequence based on a target component layout template, visible light frames, infrared thermal imaging frames, and compressed keyframes at the same time are first selected, and edge extraction is performed within the same image area of the three types of frames. Visible light frames and compressed keyframes use brightness values for calculation, while infrared thermal imaging frames use grayscale values converted from temperature values. For each pixel, the absolute values of the differences between its left and right adjacent pixels and the absolute values of the differences between its upper and lower adjacent pixels are calculated, and the two absolute values are added together to obtain the edge intensity of that pixel. The edge intensities in the three types of frames are arranged from high to low, and the top 15% of pixels are selected as effective edge pixels. Then, pixel segments with a continuous length of 12 pixels and a directional deviation of no more than 5 degrees among the effective edge pixels are determined as the edge of the target component structure. The absolute value of the difference between the vertical distance between two adjacent target component structural edges and the distance between adjacent components recorded in the target component layout template is taken. If the absolute value does not exceed 8% of the distance between adjacent components in the target component layout template, the two target component structural edges are determined as positioning reference lines.
[0028] The reference position of the target component is determined based on the fit between the positioning baseline and the target component layout template. The fit calculation includes three terms: the first is the angle between the direction of the positioning baseline and the direction of the corresponding edge in the target component layout template, divided by 90 degrees; the second is the pixel distance between the midpoint of the positioning baseline and the midpoint of the corresponding edge of the template, divided by the diagonal length of the current image; the third is the absolute value of the difference between the distance between adjacent positioning baselines and the distance between adjacent components in the template, divided by the distance between adjacent components in the template. These three terms are added together to obtain the fit value, and the target component layout template position with the lowest fit value is taken as the matching position at the current moment. Based on the component center order at this matching position, the component centers recorded in the target component layout template are converted to the current image coordinates to obtain the target component reference position.
[0029] After obtaining the reference position of the target component, candidate bounding boxes are extracted centered on this reference position. The width of the candidate bounding box is equal to the pixel width of the component outline recorded in the target component layout template, converted to the current image, plus twice the outward expansion width recorded in the target component layout template; the height of the candidate bounding box is obtained in the same way. A ring-shaped region is formed by extending 6 pixels outward from the outer edge of the candidate bounding box, and the overlapping portions with adjacent candidate bounding boxes are removed to obtain a background correction ring. The candidate bounding box and the background correction ring together constitute the initial region of interest candidate region. The background correction ring participates in subsequent feature calculations to subtract the background effects caused by illumination, thermal diffusion, and compression distortion in the vicinity of the same location.
[0030] Media quality characteristics consist of sharpness difference, number of block breaks, reflective saturation area, and occlusion edge ratio. Sharpness difference is calculated by calculating the average edge intensity of all pixels within the candidate bounding box and the background correction ring, then subtracting the average edge intensity of the background correction ring from the average edge intensity of the candidate bounding box. The number of block breaks is calculated by checking the horizontal and vertical group boundaries in groups of 8 pixels in the compressed keyframe. A block break is recorded when the brightness difference between pixels on either side of the group boundary reaches 1.5 times the average brightness difference between pixels in adjacent groups. Reflective saturation area is the number of pixels in the visible light frame with a brightness value of 250 and a temperature increase of no more than 1 degree Celsius between adjacent infrared thermal imaging frames. Occlusion edge ratio is the difference between the edge length that should appear in the target component layout template and the actual effective edge length, then divided by the expected edge length.
[0031] Spatial texture features consist of edge discontinuity length, fine line direction consistency, and gray-level abrupt change distribution. Edge discontinuity length is the sum of the lengths of consecutive pixel segments on the edge of the same target component structure within the candidate box where no effective edge pixels appear; consecutive pixel segments shorter than 3 pixels are not included. Fine line direction consistency is calculated as follows: a 3x3 pixel window is slid within the candidate box, recording the direction of highest edge intensity in each window. This direction is divided into four categories: horizontal, vertical, left-slanted, and right-slanted. The number of windows in the category with the most occurrences is divided by the total number of windows with effective edge pixels to obtain the fine line direction consistency. Gray-level abrupt change distribution is calculated as follows: the candidate box is divided into four quadrants, and the number of effective edge pixels in each quadrant is counted, recording four distribution values in the order of top-left, top-right, bottom-left, and bottom-right.
[0032] Cross-media consistency features are obtained through the correspondence of candidate regions of the same initial region of interest across three types of frames. First, candidate boxes in visible light frames, infrared thermal imaging frames, and compressed keyframes are unified to the same pixel coordinates. Then, the overlap values for boundary positions, brightness change positions, and temperature change positions are calculated separately. The boundary position overlap value is the number of valid edge pixels shared by all three types of frames divided by the number of valid edge pixels in any one of the three types of frames. The brightness change position overlap value is the number of brightness change pixels that overlap in the visible light frame and the compressed keyframe divided by the total number of brightness change pixels in both. The temperature change position overlap value is the number of temperature change pixels that overlap with brightness change pixels in the infrared thermal imaging frame and the visible light frame divided by the total number of change pixels in both. These three values, in sequence, constitute the cross-media consistency features.
[0033] Inter-frame micro-variation features are obtained jointly from the initial region of interest (ROI) candidate regions at the current time and the previous two time steps. First, the lateral and longitudinal displacements of the candidate box center point at the current time relative to the candidate box center point at the previous time step are calculated. Then, the lateral and longitudinal displacements of the candidate box center point at the previous time step relative to the candidate box center points at the previous two time steps are calculated. The angle between the two sets of displacement directions is used to represent the continuity of the candidate box center offset direction. The outline expansion amplitude is the effective edge-enclosed area of the candidate box at the current time step minus the average effective edge-enclosed area of the previous two time steps. Local brightness and temperature increments include brightness increments and temperature increments. The brightness increment is the average brightness of the candidate box at the current time step minus the average brightness of the candidate boxes at the previous two time steps. The temperature increment is the average temperature of the candidate box at the current time step minus the average temperature of the candidate boxes at the previous two time steps. Media quality features, spatial texture features, cross-media consistency features, and inter-frame micro-variation features are arranged in a fixed order to form candidate feature groups.
[0034] ROI confidence maps are generated based on ROI candidate feature groups, and regions whose confidence scores meet the threshold and whose connected areas meet the preset range are identified as target ROI regions.
[0035] In one embodiment, when forming a region of interest confidence map based on candidate feature groups, a blank map with the same pixel coordinates as the initial region of interest candidate regions is first created, with an initial confidence score of 50 for each location in the map. Each location within the candidate box undergoes quality correction according to media quality features. Pixels located at blocky break locations lose 18 points, pixels at reflective saturation locations lose 30 points, and pixels at occluded edge locations lose 25 points; pixels at continuous effective edge locations gain 15 points. When a location belongs to multiple quality-deduction locations simultaneously, the corresponding deduction values are added together; the lower limit of the quality correction value is -50 points, and the upper limit is 20 points. This results in a quality correction value for each location, which weakens unrealistic abnormal responses caused by compression artifacts, specular reflections, and local occlusion during the confidence formation stage, while preserving effective changes near the true edges of the target component.
[0036] Subsequently, the locations of local anomalies are determined based on spatial texture features. The method for determining the concentrated location of edge discontinuity length is as follows: within the candidate box, find discontinuous segments with consecutive missing valid edge pixels and a length of at least 3 pixels; mark the location within a 2-pixel extension of both ends of the discontinuous segment as the first indicator location. The method for determining the continuous location of fine ridge direction is as follows: slide a 3x3 pixel window pixel by pixel within the candidate box; when the fine ridge direction of four consecutive windows belongs to the same direction category, mark the common area covered by the four windows as the second indicator location. The method for determining the concentrated location of grayscale abrupt change distribution is as follows: in the distribution values of the four quadrants, if the number of valid edge pixels in any quadrant accounts for more than 40% of the total number of valid edge pixels in the four quadrants, and the difference in the number of effective edge pixels between this quadrant and its adjacent quadrants reaches more than 12% of the total number of effective edge pixels in the four quadrants, mark the location 1 pixel beyond the effective edge pixels in this quadrant as the third indicator location.
[0037] The first, second, and third indicator positions are superimposed to form the initial local anomaly indicator positions. For each of the initial local anomaly indicator positions, a background correction ring is extended in the opposite direction to the center of the candidate box. If there is a unidirectional gray-level abrupt change in the background correction ring, and the angle between the direction of the gray-level abrupt change and the direction of the gray-level abrupt change at that position does not exceed 11.25 degrees, 15 points are deducted from the local anomaly score for that position. Among the positions that are not deducted, those belonging to all three indicator positions receive a local anomaly score of 35 points; those belonging to two indicator positions receive a local anomaly score of 22 points; and those belonging to only one indicator position receive a local anomaly score of 8 points. Positions with a local anomaly score greater than 0 points are considered local anomaly indicator positions.
[0038] The continuous enhancement positions are then determined based on cross-media consistency characteristics and inter-frame micro-change characteristics. In visible light frames, infrared thermal imaging frames, and compressed keyframes, a cross-media overlapping position is defined as one where the pixel distance between boundary positions does not exceed 2 pixels, the pixel distance between brightness change positions does not exceed 2 pixels, and the pixel distance between temperature change positions and brightness change positions does not exceed 3 pixels. Cross-media overlapping positions must also meet inter-frame continuity conditions: the angle between the current candidate box center offset direction and the previous offset direction does not exceed 22.5 degrees; the current outline expansion amplitude is positive, and the absolute value of the current outline expansion amplitude minus the average outline expansion amplitude of the previous two moments does not exceed 6% of the candidate box area; the brightness increment reaches the average brightness change of the background correction ring plus twice the average brightness discrete change; the temperature increment reaches the average temperature change of the background correction ring plus twice the average temperature discrete change. If all conditions are met, the position is determined as a continuous enhancement position, and the continuous enhancement score is 28 points; if only the cross-media overlapping condition is met but the inter-frame continuity condition is not met, the continuous enhancement score is 10 points.
[0039] The confidence score for each location in the region of interest confidence map is calculated using the formula: 50 points plus a quality correction value, plus a local anomaly score, and plus a continuous enhancement score. A score below 0 is recorded as 0, and a score above 100 is recorded as 100. The preferred preset confidence score is 72. The preset confidence score is set as follows: Take 20 sets of anomaly-free collected data, calculate the confidence score for each map surface using the above method, average the confidence scores of the highest 5% of locations in each map surface, then calculate the overall average of the 20 averages, and add twice the average dispersion value. A score below 68 is recorded as 68, and a score above 82 is recorded as 82.
[0040] It should be noted that connected regions are determined using an 8-neighborhood connection method. The lower limit of the preset pixel area range is determined by 1% of the candidate box area; if 1% of the candidate box area is less than 16 pixels, the lower limit is 16 pixels. The upper limit of the preset pixel area range is determined by 30% of the candidate box area; if 30% of the candidate box area is greater than 1800 pixels, the upper limit is 1800 pixels. Preferably, the lower limit is 16 pixels and the upper limit is 1800 pixels. Connected regions that have a confidence level reaching a preset confidence value, a connected area within the preset pixel area range, and all boundaries located within the initial region of interest candidate regions are determined as the target region of interest.
[0041] Based on the confidence gradient of the ROI confidence map and the micro-change characteristics between frames, the target ROI region is subjected to boundary adaptive correction and local enhancement to obtain the target ROI image patch.
[0042] In one embodiment, when performing boundary adaptive correction on the target region of interest (ROI), the geometric center of the ROI is first used as the inside / outside direction reference. The initial boundary of the ROI consists of the outer edge pixels in the ROI confidence map that reach a preset confidence value and form a connected region. For each boundary pixel on the initial boundary, the boundary pixel is connected to the geometric center, with the side facing the geometric center as the inside and the side away from the geometric center as the outside. Three sampling positions are taken along this direction on the inside and three sampling positions are taken on the outside, while retaining the boundary pixel itself, forming seven sampling positions. The distance between adjacent sampling positions is one pixel, and sampling positions outside the initial ROI candidate region are not included in the calculation.
[0043] The confidence gradient is calculated based on the confidence difference between two adjacent sampling positions. The confidence levels of the seven sampling positions are compared sequentially from the inside out. When the confidence level of a subsequent sampling position is lower than that of a previous sampling position, and the difference reaches 6 points, it is recorded as one effective decrease. When three consecutive effective decreases occur in the direction from the inside out, and the difference between the highest and lowest confidence levels among the seven sampling positions reaches 18 points, the position corresponding to that boundary pixel is determined as a stable boundary position. 6 points is the single decrease judgment value, and 18 points is the overall decrease judgment value. The preferred setting rule is: select 10 sets of abnormal data collection, calculate the average and average discrete values of the confidence differences between adjacent sampling positions near the initial boundary, and take the average value plus twice the average discrete value for the single decrease judgment value. If the result is lower than 4 points, take 4 points; if it is higher than 8 points, take 8 points. The overall decrease judgment value is three times the single decrease judgment value.
[0044] When determining the boundary correction direction, the candidate box center offset direction in the inter-frame micro-change features is converted into a direction line in the current image coordinates. If the angle between the outer direction of any boundary pixel on the initial boundary and this direction line does not exceed 45 degrees, the boundary pixel is located on the candidate outward expansion side; if the angle reaches 135 to 180 degrees, the boundary pixel is located on the candidate inward contraction side. For the candidate outward expansion side, the difference between the average brightness of the three outer sampling positions and the average brightness of the three inner sampling positions is calculated as the lateral brightness increment; the difference between the average temperature of the three outer sampling positions and the average temperature of the three inner sampling positions is calculated as the lateral temperature increment. If the lateral brightness increment at the current time is higher than the corresponding value at the previous time, and the corresponding value at the previous time is higher than the corresponding value at the previous two time times, and simultaneously, the lateral temperature increment at the current time is higher than the corresponding value at the previous time, and the corresponding value at the previous time is higher than the corresponding value at the previous two time times, the candidate outward expansion side is determined as the outward expansion correction side. Among the candidate inward sides, the position with a stable boundary position and a continuous descent from the inside to the outside of 3 times is determined as the inward correction side.
[0045] During contour reconstruction, all stable boundary positions are first reconnected into a closed contour using 8-neighborhood connections. When there are gaps between adjacent stable boundary positions, the gaps are filled along the shortest path between the two points. If the path length exceeds 5 pixels, the corresponding segment from the original initial boundary is used as the replacement. The outward expansion distance for the expansion correction side is calculated based on the outward expansion amplitude. The expansion amplitude is divided by the perimeter of the current target region of interest, and the result is rounded to an integer pixel. If the difference is less than 1 pixel, 1 pixel is used; if it is more than 6 pixels, 6 pixels are used. Pixels within the corresponding distance are included along the outward direction on the expansion correction side. The inward contraction distance for the contraction correction side is the pixel distance from the boundary pixel to the nearest stable boundary position. If the difference is less than 1 pixel, 1 pixel is used; if it is more than 4 pixels, 4 pixels are used. Pixels within the corresponding distance are removed along the inward direction. After completing the expansion and contraction, any portion of the boundary exceeding the initial region of interest candidate area is truncated, resulting in the corrected region of interest.
[0046] Local enhancement is based on correcting the difference between the region of interest (ROI) and the background correction ring. First, the average visible light brightness, average infrared temperature, and average edge intensity within the ROI are calculated separately. Then, the corresponding average values within the background correction ring are calculated. The brightness difference is the average brightness of the ROI minus the average brightness of the background correction ring; the temperature difference is the average temperature of the ROI minus the average temperature of the background correction ring; and the edge intensity difference is the average edge intensity of the ROI minus the average edge intensity of the background correction ring.
[0047] Local contrast enhancement is performed by centering the brightness of the background correction ring. The enhanced brightness of any pixel within the region of interest is equal to the average brightness of the background correction ring, plus the difference between the original brightness of that pixel and the average brightness of the background correction ring multiplied by the brightness enhancement factor. The brightness enhancement factor is equal to 1 plus the absolute value of the brightness difference divided by 40. A result lower than 1.1 is taken as 1.1, and a result higher than 1.6 is taken as 1.6. Enhanced brightness below 0 is recorded as 0, and brightness above 255 is recorded as 255.
[0048] The thermal enhancement is performed by centering the average temperature of the background correction ring. The enhanced temperature grayscale of any pixel within the region of interest is corrected to the grayscale corresponding to the average temperature of the background correction ring, plus the difference between the original temperature grayscale of that pixel and the grayscale corresponding to the average temperature of the background correction ring, multiplied by the thermal enhancement factor. The thermal enhancement factor is equal to 1 plus the absolute value of the temperature difference divided by 5. If the calculated result is lower than 1.05, it is taken as 1.05; if it is higher than 1.5, it is taken as 1.5.
[0049] Edge continuum enhancement is based on edge intensity difference. Within the region of interest (ROI), the interval between two valid edge pixels in the same direction does not exceed three pixels, and the confidence score of pixels within this interval in the ROI confidence map reaches 60 points. In this case, the edge intensity of the pixels within the interval is updated to the average edge intensity of the two valid edge pixels. The 60-point threshold is set by subtracting 12 points from a preset confidence value; if the score is below 55, it is set to 55; if it is above 65, it is set to 65. The corrected ROI, after local contrast enhancement, thermal difference enhancement, and edge continuum enhancement, along with the corresponding visible light segment, infrared thermal imaging segment, and compressed keyframe segment, is cropped into the target ROI image block.
[0050] Based on the target ROI image patch and its candidate ROI features, the ROI category, location boundary and recognition confidence are obtained. The cumulative confidence of anomalies is calculated based on the recognition confidence at consecutive time points, and the final ROI recognition result is output.
[0051] In one embodiment, when identifying the target region of interest (ROI) image block and its candidate feature group, the pixel coordinates of the modified ROI are first used as a common reference to perform positional mapping on the visible light segment, infrared thermal imaging segment, and compressed keyframe segment. In the visible light segment, the position where the brightness value reaches the average brightness value plus twice the average brightness dispersion value of the segment is recorded as a brightness concentration position; in the infrared thermal imaging segment, the position where the temperature value reaches the average temperature value plus twice the average temperature dispersion value of the segment is recorded as a temperature concentration position; in the compressed keyframe segment and the visible light segment, the position where the edge intensity reaches the average edge intensity value of the modified ROI plus 1.5 times the average edge intensity dispersion value, and the consecutive length in adjacent directions reaches 4 pixels, is recorded as an edge continuity position. The brightness concentration positions, temperature concentration positions, and edge continuity positions are mapped item by item to the media quality features, spatial texture features, cross-media consistency features, and inter-frame micro-variation features in the candidate feature group according to the same pixel coordinates, forming an identification feature sequence. The identification feature sequence records the existence status of reflective saturation, blocky fracture, occlusion edge break, unidirectional gray-scale abrupt change, synchronous brightness and temperature change, and continuous boundary expansion in positional order. The existence is recorded as 1, and the absence is recorded as 0.
[0052] Category confidence scores are calculated on a 100-point scale. The confidence score for normal areas is equal to 100 points, minus the percentage of reflective saturation locations multiplied by 30 points, the percentage of blocky breakage locations multiplied by 25 points, the percentage of occluded edge breaks multiplied by 25 points, the percentage of locations with synchronized brightness and temperature changes multiplied by 35 points, and the percentage of locations with continuous boundary expansion multiplied by 35 points. The confidence score for reflective interference areas is equal to the percentage of reflective saturation locations multiplied by 60 points, plus the percentage of locations with unidirectional grayscale abrupt changes multiplied by 20 points, and then minus the percentage of locations with concentrated temperature multiplied by 25 points. The confidence score for compressed artifact areas is equal to the percentage of blocky breakage locations multiplied by 55 points, plus the percentage of locations within 8-pixel group boundaries in continuous edge positions multiplied by 25 points, and then minus the percentage of locations with synchronized brightness and temperature changes multiplied by 20 points. The confidence score for an anomaly warning area is calculated as follows: (1) the percentage of locations with synchronized brightness and temperature changes multiplied by 45; (2) the percentage of locations with continuous boundary expansion multiplied by 35; (3) the percentage of locations with overlapping across media multiplied by 20; and (4) the percentage of locations with reflective saturation multiplied by 15. A confidence score below 0 is recorded as 0, and a confidence score above 100 is recorded as 100. The category with the highest confidence score is determined as the region of interest. When the confidence score difference between two categories is less than 5 points, the priority order is: anomaly warning area, compression artifact area, reflective interference area, and normal area.
[0053] The location boundary is corrected for overlap between the outer edge of the region with the highest class confidence and the boundary of the corrected region of interest. The outer edge of the region with the highest class confidence is obtained by connecting the positions that participated in the scoring in the corresponding class according to 8-neighborhood. Isolated regions with an area of less than 12 pixels are deleted. During overlap correction, positions where the distance between the outer edge of the region and the boundary of the corrected region of interest is no more than 3 pixels are retained; if the distance is more than 3 pixels but the position is both a position of synchronous brightness and temperature change and a position of continuous boundary expansion, it is retained and shrunk by 1 pixel towards the boundary of the corrected region of interest; if the distance is more than 3 pixels and the above conditions are not met, it is deleted. When there is a break after overlap, it is filled along the path with the shortest pixel distance between adjacent endpoints. If the filling length is more than 6 pixels, it is replaced by the boundary segment corresponding to the corrected region of interest, and finally a closed outer edge is obtained, which is determined as the location boundary. The highest class confidence value is used as the recognition confidence at the current time.
[0054] The cumulative confidence score for anomalies is calculated based on the current time and the two previous time points. First, a weighted average of the confidence scores over the three time points is calculated, with the current time point having a weight of 0.5, the previous time point having a weight of 0.3, and the two previous time points having a weight of 0.2. Next, the overlap of location boundaries is calculated by dividing the number of overlapping pixels between the areas enclosed by the current time point's boundary and the areas enclosed by the previous time point's boundary by the number of pixels after merging the two. The continuity of local brightness and temperature increments is set to 1 if the brightness increment at the current time point is higher than the previous time point, the previous time point's increment is higher than the two previous time points, and the temperature increment at the current time point is higher than the previous time point and the previous time point's increment is higher than the two previous time points; otherwise, it is set to 0.4. The cumulative confidence score for anomalies is equal to the weighted average multiplied by 0.6, plus the overlap of location boundaries multiplied by 100 points and then multiplied by 0.25, plus the continuity of local brightness and temperature increments multiplied by 100 points and then multiplied by 0.15.
[0055] The preset output value is preferably 78 points. The rules are as follows: Collect 20 sets of normal data and 20 sets of labeled abnormal data, and calculate the cumulative confidence score of the abnormality for each set. Take the highest cumulative confidence score of the abnormality in the normal data plus 5 points as the first value, and take the lowest cumulative confidence score of the abnormality in the labeled abnormal data minus 5 points as the second value. When the first value is not higher than the second value, the preset output value is the average of the first and second values. When the first value is higher than the second value, the preset output value is 78 points. When the cumulative confidence score of the abnormality reaches the preset output value, the final region of interest (ROI) identification result, including the ROI category, location boundary, and cumulative confidence score of the abnormality, is output. When the preset output value is not reached, the identification confidence score and location boundary at the current time are retained and used in the calculation of the next time step.
[0056] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for intelligent image recognition of multimedia ROI regions, characterized in that, include: The system acquires visible light frames, infrared thermal imaging frames, and compressed keyframes of the scene to be identified, and performs time synchronization based on transient change boundaries to obtain a multimedia image sequence. Based on the target component layout template, the multimedia image sequence is coarsely localized to obtain the initial region of interest candidate region containing candidate boxes and background correction rings. Media quality features, spatial texture features, cross-media consistency features and inter-frame micro-variation features are extracted to form candidate feature groups. A confidence map of the region of interest is formed based on the candidate feature groups, and the target region of interest is determined. Boundary correction and local enhancement are performed based on the confidence gradient and inter-frame micro-change features to obtain the image patch of the target region of interest. Based on the target region of interest image patch and candidate feature group, the region of interest category, location boundary and recognition confidence are obtained, and the final region of interest recognition result is output according to the recognition confidence at consecutive time steps.
2. The intelligent image recognition method for multimedia ROI regions according to claim 1, characterized in that, Time synchronization based on transient boundary changes includes: The original acquisition order of visible light frames, infrared thermal imaging frames, and compressed keyframes within the same acquisition cycle is preserved and unified to the same image area; Extract the visible light variation boundary, infrared variation boundary, and compressed keyframe variation boundary formed by brightness and temperature changes in adjacent frames, respectively. Using the visible light change boundary as a reference, the infrared thermal imaging frames and compressed key frames are shifted and calibrated. Frame groups that meet the preset conditions in terms of boundary center distance, boundary overlap ratio, and change direction are used as synchronous frame groups and arranged in the acquisition order to obtain a multimedia image sequence.
3. The intelligent image recognition method for a multimedia ROI region according to claim 2, characterized in that, Coarse positioning is performed based on the target component layout template, including: Extract the edge of the target component structure within the same frame area of the synchronized frame group, and determine the positioning reference line based on the distance between the edges of adjacent target component structures; The reference position of the target component is determined based on the direction, position, and spacing of the positioning baseline and its fit with the layout template of the target component. Candidate boxes are captured centered on the reference position of the target component, and a background correction ring is formed outside the candidate boxes. The candidate boxes and the background correction ring are used together as the initial region of interest candidate regions.
4. The intelligent image recognition method for multimedia ROI regions according to claim 1, characterized in that, The extraction of media quality characteristics includes: Calculate the average edge intensity of the candidate box and the background correction ring separately, and use the difference between the two as the sharpness difference; The number of block breaks is counted in the compressed keyframes, the reflective saturation area is counted in the visible light frames, and the occlusion ratio is calculated based on the edge length that should appear in the target component layout template and the actual effective edge length. Media quality characteristics are composed of poor clarity, number of blocky breaks, reflective saturation area, and proportion of obscured edges.
5. The intelligent image recognition method for a multimedia ROI region according to claim 1, characterized in that, Extraction of spatial texture features and cross-media consistency features, including: Within the candidate box, the length of edge discontinuity, the consistency of fine line direction, and the distribution of gray-level abrupt changes are statistically analyzed to form spatial texture features; Unify the candidate bounding boxes in visible light frames, infrared thermal imaging frames, and compressed keyframes to the same pixel coordinates; The overlap values at boundary locations, brightness change locations, and temperature change locations are calculated separately to form cross-media consistency features.
6. The intelligent image recognition method for multimedia ROI regions according to claim 1, characterized in that, Extraction of inter-frame micro-variation features includes: Calculate the displacement direction of the candidate box center point at the current time relative to the candidate box center point at the previous time, and the displacement direction of the candidate box center point at the previous time relative to the candidate box center point at the previous two time points; The outer contour expansion range is obtained by the difference between the effective edge enclosing area of the candidate box at the current moment and the average effective edge enclosing area of the previous two moments. Based on the difference between the average brightness and average temperature of the candidate boxes at the current moment and the corresponding average values at the previous two moments, the local brightness and temperature increment is obtained, and together with the media quality features, spatial texture features, and cross-media consistency features, a candidate feature group is formed.
7. The intelligent image recognition method for a multimedia ROI region according to claim 6, characterized in that, Generate a confidence map of the region of interest and determine the target region of interest, including: A map is constructed using the pixel coordinates of the initial candidate region of interest, and quality correction values are assigned to the locations of block breaks, reflective saturation, occluded edges, and continuous effective edges based on media quality characteristics. The location of local anomalies is determined based on spatial texture features, and the location corresponding to the same gray-level abrupt change in the background correction ring is subtracted. Based on cross-media consistency characteristics and inter-frame micro-change characteristics, continuous enhancement positions are determined. The quality correction value, local anomaly indication position, and continuous enhancement position are written into the same map to obtain the region of interest confidence map. Connected regions that meet the preset conditions in terms of both confidence and connectivity area are determined as the target region of interest.
8. The intelligent image recognition method for a multimedia ROI region according to claim 7, characterized in that, Boundary correction, including: Pixel-by-pixel sampling is performed along the initial boundary of the target region of interest, both inward and outward, and the stable boundary position is determined based on the continuous decrease in confidence between adjacent sampling positions. Based on the candidate box center offset direction, outer contour expansion range, and local brightness temperature increment, determine the outward expansion correction side and the inward contraction correction side; The contour of the target region of interest is reconstructed according to the stable position of the boundary. The outward correction side is included in the target region of interest and the inward correction side is removed, resulting in a corrected region of interest whose boundary is limited to the candidate region of the initial region of interest.
9. The intelligent image recognition method for a multimedia ROI region according to claim 8, characterized in that, Local enhancement, including: Calculate the brightness difference, temperature difference, and edge intensity difference between the corrected region of interest and the background correction ring, respectively. Based on the average brightness and average temperature of the background correction ring, local contrast enhancement and thermal difference enhancement are performed on the correction region of interest; The discontinuous edges within the region of interest with an interval not exceeding a preset number of pixels and a confidence level meeting a preset condition are continuously enhanced, and then cropped to obtain the target region of interest image block.
10. The intelligent image recognition method for a multimedia ROI region according to claim 9, characterized in that, The final region of interest identification result is output, including: Extract the concentrated brightness locations, concentrated temperature locations, and continuous edge locations from the target region of interest image block, and match their locations with candidate feature groups to form a recognition feature sequence; Based on the correspondence between reflective saturation, blocky fracture, occlusion edge break, unidirectional gray-scale abrupt change, synchronous brightness and temperature change and continuous boundary expansion in the identification feature sequence, calculate the category confidence of normal area, reflective interference area, compression artifact area and abnormal warning area, and determine the category and location boundary of region of interest; The cumulative confidence of anomalies is calculated based on the recognition confidence, the degree of overlap of location boundaries, and the continuity of local brightness temperature increments at the current time and the two previous time points. When the cumulative confidence of anomalies reaches the preset output value, the final region of interest recognition result is output.