Control method of video monitoring system
By constructing a reflection sensing vector group and analyzing the temporal trajectory consistency matrix, and dynamically updating the attitude adjustment reliability factor, the problem of camera attitude adjustment caused by reflection misjudgment in video surveillance systems is solved, realizing the accuracy and stability of camera control and ensuring the real-time effectiveness of the monitoring image.
Patent Information
- Application Number
- CN202511358662.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2025-11-28
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing video surveillance systems are prone to misinterpreting reflective shifts as target movement in environments with large areas of glass, bright surfaces, or reflective background materials. This causes the camera to continuously make incorrect posture adjustments, affecting the effectiveness of the surveillance footage and the system's continuous tracking capabilities.
By constructing a reflection sensing vector group and combining it with the target motion direction vector to calculate the angle change trend, and using the inter-frame change tensor and temporal trajectory consistency matrix to analyze image displacement anomaly patterns, a posture adjustment reliability factor evaluation function is constructed, and the evaluation function parameters are dynamically updated to achieve accurate camera posture control.
Accurately identify reflection shifts in images, avoid environmental interference from misleading camera attitude control, improve the accuracy and stability of camera control behavior, and ensure the continuous visibility of monitored targets and the real-time effectiveness of images.
Smart Images

Figure CN121037697A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video monitoring, in particular to a control method of a video monitoring system. BACKGROUND
[0002] The control of a video monitoring system refers to the orderly management and coordination of the running state, working parameters and response behavior of each component in the monitoring system through software and hardware means, so as to realize the comprehensive control and intelligent scheduling of the monitoring process. The existing control technology of a video monitoring system usually relies on a centralized or distributed management platform, which is connected with front-end camera devices, pan-tilt, storage servers, alarm modules, display terminals, etc. through wired or wireless networks, and the system is operated and configured by using a graphical interface, control protocol or preset instructions. The specific key links include the following: first, the access and control of front-end devices, such as remote control of the on-off, zoom, angle adjustment and video stream configuration of the camera; second, the recording management and strategy control, which realizes the orderly collection of recording data by setting recording plans, motion detection triggers and abnormal event recording logic; third, the transmission and network control, such as automatic switching of master-slave code streams, dynamic bandwidth adjustment and streaming media distribution strategy; in addition, there is intelligent analysis and alarm control, such as triggering real-time alarms and linking other devices through face recognition, boundary crossing detection and stay judgment; finally, the system permission and operation management control, which is used to define user permissions, log auditing and device group management operation strategies.
[0003] The existing technology has the following disadvantages: In a video monitoring system, the pose control method of the camera usually relies on the relative displacement information of the target in the image to determine whether the pitch angle of the camera needs to be adjusted. However, in a monitoring environment with a large area of glass, bright ground or reflective background material, the picture is easy to appear light offset consistent with the direction of target motion. The system will misidentify the light spot displacement caused by environmental factors as the real displacement of the target, thus triggering unnecessary pose adjustment operations. Since these reflections have continuous motion trajectories in the image, their characteristics are similar to the real motion of the target in space, which causes the system to still judge that the target has deviated from the monitoring area without actual target movement. The existing control technology of a video monitoring system cannot determine whether to perform actual pose adjustment according to the image displacement abnormal pattern in the case of light offset consistent with the direction of the target in the image during camera pose control, thus causing the camera to continuously perform incorrect pose control operations, eventually making the camera picture deviate from the real target, seriously affecting the effectiveness of the monitoring picture and the continuous tracking ability of the system, and possibly forming a monitoring blind area that cannot be detected in time.
[0004] The above information disclosed in the BACKGROUND section merely to enhance the understanding of the background of the present disclosure, and therefore it can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY
[0005] An object of the present application is to provide a control method of a video monitoring system to solve the problems in the background.
[0006] To achieve the above object, the present application provides the following technical solution: a control method of a video monitoring system, specifically comprising the following steps: S1, acquiring continuous image frames through the video monitoring system, extracting the luminance channel, edge structure, direction gradient and contrast change of the image, constructing a reflection perception vector group, and combining a target motion direction vector to calculate an angle change trend to determine whether a reflection offset consistent with the target direction appears in the picture in the camera pose control process; S2, after determining that the reflection offset consistent with the target direction appears in the picture, extracting the inter-frame change tensor of the target region and the reflection region, constructing a time sequence trajectory consistency matrix, and based on a frequency stability factor and a background similarity factor, determining an image displacement abnormal pattern in the case of the reflection offset consistent with the target direction appearing in the picture in the camera pose control process; S3, according to the determined image displacement abnormal pattern, constructing a pose adjustment credibility factor evaluation function by aggregating a disturbance stability parameter, a target trajectory confidence and a boundary offset amplitude to determine whether to perform actual pose adjustment; S4, matching the pose adjustment credibility factor with a preset evaluation threshold interval, outputting a judgment result, and performing a corresponding camera pose control behavior according to the judgment result; S5, after performing the camera pose control behavior, acquiring a controlled image state, dynamically updating the evaluation function parameters in a sliding time window, adjusting the camera pose control judgment logic according to the update result, and realizing dynamic regulation and control of the camera control behavior.
[0007] Preferably, S1 specifically comprises the following steps: S101, acquiring continuous image frames through the video monitoring system, extracting the luminance channel of each image at the pixel level, extracting the edge structure of the image based on an edge detection operator, obtaining the direction gradient information of the image through a gradient operator, and using a multi-scale window to traverse the image region to calculate the gray difference value of each region to extract the contrast change of the image; S102, using the luminance channel, edge structure, direction gradient and contrast change of the image to perform feature fusion in the same coordinate space, combining the inter-frame change trend of the image, constructing a reflection perception vector group, and the constructed reflection perception vector group represents the directional change of the brightness enhanced region in the spatial and temporal dimensions; S103, the video monitoring system identifies the motion direction vector of the target in the image, and calculates the angle change trend of the motion direction vector and each vector in the reflection light perception vector group. When the included angle in the continuous frames is maintained within the preset specified threshold and the vector trajectory is continuous, it is determined that the reflection light offset consistent with the direction of the target appears in the picture in the camera posture control process.
[0008] Preferably, S102 specifically comprises: The brightness channel, edge structure, direction gradient and contrast change extracted from the continuous image frames are respectively mapped in two-dimensional space, and the coordinate registration of each image feature map is performed through the bilinear interpolation algorithm, so that the brightness boundary, gradient direction and contrast difference have superimposability at the same pixel position; Based on the corresponding relationship of the pixel level position, the brightness channel map, edge structure map, direction gradient map and contrast change map are fused in the unified coordinate space, a four-dimensional joint feature vector is constructed for each pixel position, and the change trend of the brightness enhancement area in the continuous frames is extracted through frame difference; The reflection light perception vector group is constructed according to the four-dimensional joint feature vector of the brightness enhancement area, and the image gradient direction is encoded in the space dimension, and the brightness change direction is encoded in the time dimension, representing the directional change of the brightness enhancement area in the space and time dimensions.
[0009] Preferably, S2 specifically comprises the following steps: S201, based on the target area and the reflection area identified by the video monitoring system, the brightness channel, edge structure and direction gradient information are collected in the continuous image frames, the feature values of each region at the corresponding pixel position in each frame are subjected to difference operation in time sequence, forming the inter-frame change tensor, and the channels of the inter-frame change tensor correspond to the brightness change, the structure response change and the gradient direction change respectively; S202, linearly expand the inter-frame change tensor in the time dimension, and construct the time sequence trajectory consistency matrix according to the position continuity and feature change amplitude of each pixel in different image frames, the column of the matrix represents the image frame sequence, and the row represents the pixel trajectory in each region. The continuous change trajectory of the region as a whole in time is encoded to reveal the behavior consistency between the target and the reflection light; S203, calculate the frequency stability factor and the background similarity factor based on the time sequence trajectory consistency matrix. The frequency stability factor is obtained according to the frame proportion of the continuous frames with consistent change direction and change amplitude exceeding the threshold value, and the background similarity factor is obtained by comparing the texture gray level histogram similarity of the region. The two factors jointly constitute the image displacement abnormal pattern criterion. When the target area and the reflection area both meet the abnormal conditions in trajectory consistency and background similarity, the image displacement abnormal pattern under the condition that the reflection light offset consistent with the direction of the target appears in the picture in the camera posture control process is determined.
[0010] Preferably, S203 specifically comprises: In the time sequence trajectory consistency matrix, the feature change direction of each pixel position of the target region and the reflective region in the continuous image frames is extracted, and the number of frames in each pair of adjacent image frames in which the change direction is consistent and the change amplitude exceeds a set numerical threshold value is calculated. The frequency stability factor is calculated by the proportion of the number of frames to the total number of frames, which is used to represent the direction continuity and amplitude effectiveness of the trajectory change behavior; The gray value distribution of the target region and the reflective region in the image frames is collected, the texture distribution vector is generated by using the gray histogram, the gray distribution difference value between the regions is calculated by using the Bhattacharyya distance algorithm, and the background similarity factor is output by comparing the gray distribution difference value between the regions with a preset similarity judgment threshold value, which is used to measure the consistency degree of the image backgrounds of the two regions; The frequency stability factor and the background similarity factor are used as joint inputs, and a linear weighted discriminant function is used to construct an image displacement abnormal pattern criterion. When the frequency stability factor is greater than a first threshold value and the background similarity factor is less than a second threshold value, the image displacement abnormal pattern in the case that the camera pose control process appears a reflective shift in the picture consistent with the target direction is determined. The image displacement abnormal pattern refers to that the image region produces a linear position shift in a consistent direction in continuous image frames, the edge structure change rate in the region is lower than a preset edge stability threshold value, and the Bhattacharyya distance of the gray histogram of the region is less than a preset texture difference threshold value, which is used to represent the control misjudgment trend caused by the movement of a non-real target.
[0011] Preferably, S3 specifically comprises: On the basis of time sequence tracking of the image displacement abnormal pattern, the center coordinate offset sequence of each target region in the continuous image frames with time is extracted, and a disturbance stability parameter is constructed by using the standard deviation of the displacement amplitude in each time period in the sliding window, which is used to quantify the position stability trend of the region on the time axis; Based on the calculation result of the disturbance stability parameter, the trajectory continuity of the target region in the image frames is tracked, the target trajectory confidence is constructed by using the trajectory interruption frequency and the trajectory direction deviation amplitude, and the boundary disturbance reference index is calculated by extracting the position change range of the outer boundary contour of the target in the same frame sequence and calculating the boundary offset amplitude; The disturbance stability parameter, the target trajectory confidence and the boundary offset amplitude are standardized respectively as three input variables, which are input into a weighted fusion function to construct a pose adjustment confidence factor evaluation function, and the output pose adjustment confidence factor is compared with a pose adjustment judgment threshold value. When the pose adjustment confidence factor is greater than the judgment threshold value, it is judged that the actual pose adjustment can be executed.
[0012] Preferably, S4 specifically comprises: The preset evaluation threshold interval of the posture adjustment confidence factor is constructed, the value range of the posture adjustment confidence factor is statistically analyzed through experimental data, a plurality of continuous and non-overlapping threshold intervals are set, which correspond to different posture control levels respectively, and the evaluation threshold interval is used to divide the posture adjustment confidence factor into levels, so as to ensure that the judgment process has resolution and stability; The posture adjustment confidence factor is compared with the evaluation threshold interval one by one, and a judgment result is output according to the threshold level of the posture adjustment confidence factor, the judgment result including three states of prohibition adjustment, limited adjustment and complete adjustment, the prohibition adjustment corresponding to the posture adjustment confidence factor being lower than the first evaluation threshold, the limited adjustment corresponding to the posture adjustment confidence factor being between the first evaluation threshold and the second evaluation threshold, and the complete adjustment corresponding to the posture adjustment confidence factor being higher than the second evaluation threshold; According to the judgment result, corresponding camera posture control behaviors are executed, in the prohibition adjustment state, no change is made to the camera pitch angle, in the limited adjustment state, the camera pitch angle is corrected through a proportional fine-tuning strategy, and in the complete adjustment state, a camera posture change vector is calculated according to the target offset direction and adjusted, so as to realize the camera posture control behavior based on the judgment result.
[0013] Preferably, S5 specifically is: After the camera posture control behavior is executed, the video monitoring system is used to continuously collect the controlled image frames, the brightness channel, edge structure, direction gradient and contrast change information in the controlled image frames are extracted, the controlled image state data is formed, and the current posture adjustment corresponding image response segment is marked as an evaluation sample; The controlled image state data is input into a sliding time window with a set length, the disturbance stability parameter, target trajectory confidence and boundary offset amplitude are dynamically updated based on the continuous change trend of the image frames in the window, and the posture adjustment confidence factor is recalculated according to the updated parameters, so as to complete the dynamic update of the evaluation function parameters; According to the change relationship between the dynamically updated posture adjustment confidence factor and the historical evaluation function parameters, it is judged whether it continuously deviates from the original determination trend interval, when the continuous deviation reaches a preset fluctuation range, the judgment logic of the posture adjustment confidence factor is reset, including correcting the evaluation threshold interval or adjusting the determination level mapping rule, so as to realize the adaptive dynamic regulation and control of the camera posture control judgment logic.
[0014] In the above technical solution, the technical effects and advantages provided by the present application are: 1、The application can accurately identify whether there is a light reflection deviation consistent with the target direction in the image by constructing an information perception mechanism containing multiple features such as image brightness channel, edge structure, direction gradient and contrast change, and combining the angle change trend analysis between the reflection perception vector group and the target motion direction vector, and then analyzing the motion behavior of the target area and the reflection area by using the inter-frame change tensor and the time sequence trajectory consistency matrix, supplemented by the dual criteria of frequency stability factor and background similarity factor, to realize the effective identification of image displacement abnormal mode. This mechanism not only improves the identification accuracy of the system to the pseudo target motion, but also avoids the misdirection of the reflection interference to the camera posture control judgment, and enhances the accuracy and stability of the camera control behavior in complex monitoring environment.
[0015] 2、The application further introduces three dynamic indexes of disturbance stability parameter, target trajectory confidence and boundary offset amplitude, constructs a posture adjustment confidence factor by standardized fusion, and outputs the camera control judgment result combined with the multi-level evaluation threshold interval, to realize the grading response mechanism from "prohibition adjustment" to "complete adjustment". After the system executes the camera posture adjustment, the image state data is continuously collected by using the sliding time window, the evaluation function parameters are dynamically updated, and the judgment logic is adaptively corrected, which effectively avoids the control strategy rigidification problem caused by environmental fluctuations or reflection dynamic characteristics, improves the intelligent perception and regulation and control ability of the system, and guarantees the continuous visibility of the monitoring target and the real-time effectiveness of the picture. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description only represent some embodiments described in the present application, and other drawings can also be obtained by those skilled in the art based on these drawings.
[0017] Figure 1 The flowchart of the control method of the video monitoring system of the present application. DETAILED DESCRIPTION
[0018] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations may, however, be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these example implementations are provided so that this disclosure will be thorough and complete, and will fully convey the inventive aspects to those skilled in the art.
[0019] The present application provides a control method of a video monitoring system as shown in Figure 1 , specifically comprising the following steps: S1, acquiring continuous image frames through a video monitoring system, extracting a luminance channel, edge structure, direction gradient and contrast change of the image, constructing a reflection perception vector group, and calculating an included angle change trend in combination with a target motion direction vector to determine whether a reflection offset consistent with the target direction appears in a picture in a camera posture control process; In this embodiment, S1 specifically includes the following steps: S101, acquiring continuous image frames through a video monitoring system, performing pixel-level separation on each image frame to extract a luminance channel of the image, extracting an edge structure of the image based on an edge detection operator, obtaining a direction gradient of the image through a gradient operator, and using a multi-scale window to traverse an image region to calculate a gray difference value of each region to extract a contrast change of the image; After acquiring continuous image frames in the video monitoring system, an image channel separation operation can be performed on each image frame to convert the image from an RGB format to a YUV format, wherein the Y channel is the luminance channel and represents the gray intensity information of each pixel in the image. The extraction of the luminance channel helps to ignore color interference and only focus on the luminance structural features in the image. After the luminance channel is extracted, a Canny or Sobel edge detection operator can be used to process the luminance image to identify the edge regions with significant luminance changes in the image. Then, a gradient operator (such as a Sobel gradient) is used to calculate the gradient value and gradient direction of each pixel along the X direction and the Y direction to obtain the direction gradient map of the entire image. After obtaining the direction gradient map, a multi-scale window mechanism is introduced to traverse the entire image with different sizes of sliding windows (such as 3x3, 5x5, 7x7), and the contrast change of the region is represented by calculating the difference between the maximum and minimum values of the pixel gray scale in the window. For example, on a ceramic tile floor area, the reflection of the light spot will cause the local area gray scale to rise rapidly, and through the multi-scale window, these high-contrast change regions can be detected to assist in determining whether there is reflection interference. In the overall process, each calculation process can be realized through an image processing library such as OpenCV, and parallel acceleration in GPU can improve the processing efficiency.
[0020] Pixel-level separation refers to extracting the channel information of each pixel in the image independently, such as separating the Y channel in the YUV image for analyzing the gray scale information. Edge detection operator is a tool for identifying the boundary of gray scale mutation in the image, commonly used such as Sobel, Canny, Prewitt, etc., for discovering the object contour or texture fault in the image. Gradient operator is used to measure the direction and intensity of the change of pixel gray scale in the image, its essence is to approximate the first derivative of the image, to judge the change trend of the image in X and Y axes. Directional gradient map is the summary result of these gradient directions, used to further identify the flow direction of the texture. Multi-scale window is a local feature extraction technique of image, which analyzes the change of local area through windows of different sizes, which helps to eliminate the local misjudgment caused by single scale. For example, some diffuse reflective areas may be ignored in small size windows, but the overall gray scale mutation trend of the area can be detected in larger windows. Through the above combination, the reflective feature in the brightness change can be effectively captured, and high-quality original data support is provided for the construction of subsequent reflective perception vector.
[0021] S102, using the brightness channel, edge structure, directional gradient and contrast change of the image to fuse the features in the same coordinate space, combining the change trend between image frames, constructing a reflective perception vector group, the constructed reflective perception vector group represents the directional change of the brightness enhancement area in the spatial and temporal dimensions; S103, identifying the motion direction vector of the target in the image through the video monitoring system, and calculating the angle change trend between the motion direction vector and each vector in the reflective perception vector group, when the angle is maintained within the preset specified threshold and the vector trajectory is continuous in the continuous frames, it is determined that the reflective offset consistent with the direction of the target appears in the picture in the camera pose control process.
[0022] To realize the determination of the light reflection deviation, firstly, the video monitoring system needs to track the motion trajectory of the target in the continuous image frames, and calculate the motion direction vector based on the pixel displacement direction of the target in the image coordinate system. The formation of the direction vector is usually based on Kalman filtering combined with the optical flow method to realize robust target trajectory prediction. Then, the target motion direction vector is calculated with the included angle of each light reflection vector in the light reflection perception vector group. The included angle calculation adopts the cosine similarity principle between vectors. In a plurality of continuous frames, if the included angle between a plurality of light reflection perception vectors and the target direction vector is continuously maintained within a certain preset specified threshold, and the spatial trajectory of the light reflection vector in the image has continuity, it can be determined that the direction change of the brightness enhancement region highly simulates the actual motion direction of the target, and it is extremely likely to be light interference rather than a real target. This judgment strategy can eliminate the misleading of short-term non-directional brightness fluctuations or texture disturbances on attitude control, and improve the stability and accuracy of the judgment. For example, if the target moves to the right upper corner, the motion direction vector points to the right upper corner, and in the image, if the light reflection perception vector of a bright spot region in the continuous frames is less than ten degrees from the target direction vector, and the position shows a right upper corner continuous moving trend, it is considered that there is a light reflection deviation phenomenon.
[0023] The motion direction vector of the target is the reference basis for behavior determination, and the light reflection perception vector group is the spatial representation of the dynamic environmental interference. The dynamic calculation of the included angle between the two constitutes the core logic of the judgment mechanism. The stability of the included angle means that in a plurality of continuous frames, the change amount of the included angle is lower than the fluctuation tolerance range, thereby constructing a stability signal. The continuity of the vector trajectory is a key indicator for determining whether it has a real dynamic behavior, which is used to exclude accidental alignment caused by noise or accidental light points. The preset specified threshold is a determination standard determined after a large amount of data sample analysis, which is usually between eight degrees and fifteen degrees, representing the minimum direction difference range acceptable by the human eye, and is used to control the sensitivity of the included angle judgment. If the included angle fluctuation exceeds the threshold, it means that the direction deviation is obvious, and it can be considered as non-synchronous movement; if it continuously falls within the threshold range, there may be a pseudo-synchronous phenomenon, thereby determining it as light interference. This judgment strategy strengthens the dynamic decoupling ability between the real target and the optical interference, and is a key link in the control logic.
[0024] In this embodiment, S102 specifically is: The brightness channel, edge structure, direction gradient and contrast change extracted in the continuous image frames are respectively mapped in two-dimensional space, and the coordinate registration of each image feature map is performed through the bilinear interpolation algorithm, so that the brightness boundary, gradient direction and contrast difference have superimposability at the same pixel position; In processing continuous image frames, the extracted brightness channel, edge structure, direction gradient and contrast change need to be first converted into two-dimensional image feature maps with explicit spatial coordinates, which is called two-dimensional spatial mapping. Two-dimensional spatial mapping refers to positioning each feature value of the image to a specific pixel position in the image coordinate system, so that all feature data have spatial comparability. Due to the possible slight rotation, scaling or distortion in the image acquisition process, there is a pixel position offset between different feature maps. In order to achieve accurate pixel-level alignment, a bilinear interpolation algorithm is needed to perform coordinate registration on each image feature map. The bilinear interpolation algorithm estimates the feature value of the target position by weighted average of the four neighboring pixels around the target pixel, thereby realizing continuous alignment at non-integer coordinates, which not only ensures the position consistency between feature maps, but also avoids the sharpening or distortion of image information. In this way, the brightness boundary information, gradient direction distribution and contrast change trend have numerical superposition at the same pixel position, which can provide a consistent spatial reference for subsequent joint vector construction and improve the accuracy and robustness of feature fusion. Among them, the brightness channel provides the gray basic intensity distribution, the edge structure presents the local mutation position, the direction gradient describes the texture dominant direction, and the contrast change reflects the amplitude of brightness difference in the region. After spatial alignment, the four can form a high-dimensional integrated feature representation, which is helpful for accurately positioning the direction trajectory of the light reflection offset region.
[0025] Based on the pixel-level position correspondence, the brightness channel map, edge structure map, direction gradient map and contrast change map are fused in the unified coordinate space, a four-dimensional joint feature vector is constructed for each pixel position, and the change trend of the brightness enhancement region in the continuous frames is extracted through frame difference. After the pixel-level position correspondence is established, the luminance channel map, edge structure map, direction gradient map and contrast change map need to be fused in a unified coordinate space. The specific method is to extract the luminance value, edge intensity, gradient direction angle and local contrast amplitude of the pixel position from the four images respectively in each registered feature map based on the same pixel coordinates, and then combine the four types of numerical values into a four-dimensional joint feature vector to represent the joint attributes of the pixel in multiple feature dimensions. The four-dimensional joint feature vector can be represented as a complete local optical and structural feature description, which helps to establish a higher dimension of distinction between the glare interference area and the real target area. Next, the joint feature vector of each pixel position between consecutive image frames is calculated by frame difference to determine whether the brightness enhancement change of each pixel is continuous, significant and directional. Frame difference refers to the difference between the four-dimensional vectors of the same pixel in adjacent two frames, and then the direction and amplitude of the difference value are observed to capture the motion trajectory of the glare spot over time. This method can effectively eliminate background static areas and slight noise interference, and highlight the dynamic features of the glare offset area. In this process, the luminance channel is responsible for providing the basis for the intensity of the glare enhancement, the edge structure is used to identify the changes in the outline over time, the direction gradient is used to analyze the motion direction trend, and the contrast change is used to judge whether the brightness change has a disturbance characteristic. This fusion method greatly improves the accuracy and robustness of the glare enhancement area recognition, and is the core basis for the construction of the glare perception vector.
[0026] The four-dimensional joint feature vector of the brightness enhancement area is constructed to form a glare perception vector group, and the image gradient direction is encoded in the spatial dimension and the brightness change direction is encoded in the time dimension to represent the directional change of the brightness enhancement area in the spatial and time dimensions.
[0027] After the four-dimensional joint feature vector of each pixel position is completed, it is necessary to construct the reflection light perception vector group based on the vector set for the brightness enhancement area. The specific implementation is that first, the pixel points with significant brightness rising characteristics and significant contrast fluctuation in the continuous image frames are screened out, and the four-dimensional joint feature vectors corresponding to these pixel points are extracted as the representation basis of the reflection light candidate area. Then, the image gradient direction of each candidate pixel is spatially encoded in the two-dimensional image coordinates, specifically, the gradient direction angle is quantized into a fixed interval and mapped to a unified direction vector space, so as to reflect its spatial flow trend in the picture. Then, the brightness channel change direction of these brightness enhancement pixels in multiple continuous frames is calculated in the time dimension, and by tracking the rising path of the brightness value, it is encoded into a brightness change direction vector. Finally, the image gradient direction vector in the spatial dimension and the brightness change direction vector in the time dimension are combined to form a multi-dimensional description of each element in the reflection light perception vector group, which completely describes the spatial trend of the reflection light area in the image and the change trend with time evolution. Among them, the four-dimensional joint feature vector is the basic data structure, the image gradient direction is used to judge the propagation direction of the texture or edge, and the brightness change direction is used to identify the dynamic behavior of the light spot, and both of them determine the directionality label of the reflection light perception vector group. This encoding mechanism helps to identify the reflection light area caused by the ambient light from the perspective of dynamic evolution, and distinguish it from the target real motion, thereby providing a high-robustness directionality basis for subsequent angle trend analysis and posture control judgment.
[0028] S2, after determining that the reflection light offset consistent with the target direction appears in the picture, extracting the inter-frame change tensor of the target area and the reflection light area, constructing a time sequence trajectory consistency matrix, and determining the image displacement abnormal mode in the camera posture control process when the reflection light offset consistent with the target direction appears in the picture based on the frequency stability factor and the background similarity factor; In this embodiment, S2 specifically includes the following steps: S201, based on the target area and the reflection light area identified by the video monitoring system, collecting the brightness channel, edge structure and direction gradient information in the continuous image frames, performing difference value operation on the feature values of each region in the corresponding pixel positions in each frame in time sequence, forming an inter-frame change tensor, and the channels of the inter-frame change tensor correspond to brightness change, structure response change and gradient direction change respectively; In order to extract the change trend of the target region and the reflective region in the time dimension, the pixel coordinate range of the target region and the reflective region in each image frame can be located respectively by using the continuous image frames collected by the video monitoring system, and feature extraction is performed on the range in each frame. In the brightness channel extraction, the Y channel or the V channel is selected as the brightness representation, and the original image is separated and saved as a two-dimensional gray matrix; the edge structure information can be obtained by processing the brightness image through the Sobel, Laplacian or Canny algorithm to generate an edge response graph; the direction gradient information can be obtained by calculating the gradient components of the horizontal direction and the vertical direction through the gradient operator (such as the Scharr operator), and then the angle direction graph is obtained by using the arctangent function. After the feature graph is constructed, the feature values of the same pixel position between any two frames are selected according to the time sequence of the image frames, and the difference values are obtained, which correspond to the brightness change, the edge response intensity difference and the gradient angle difference respectively, and finally a three-channel inter-frame change tensor is constructed to represent the dynamic change of the multi-dimensional visual features of the region in the image sequence. Taking a reflective region as an example, if the brightness of the region continuously increases in a short time and the edge profile basically does not change, the change value of the brightness channel in the tensor will be significantly higher than that of other regions, the change of the structure channel is close to zero, and the change of the direction gradient channel is chaotic but locally concentrated, which reveals the abnormal properties of the region.
[0029] Pixel-level separation refers to operating on each pixel in each image frame to obtain the channel value corresponding to each pixel point from the image matrix, so as to realize fine extraction of brightness and structure; the brightness channel refers to the two-dimensional data reflecting the light intensity or gray value in the image, which usually uses the brightness component in the YUV or HSV space as the basis; the edge structure refers to the position where the gray value in the image changes abruptly, which is generated by an edge detection algorithm and represents the object outline and contrast boundary in the image; the direction gradient is the direction information of the image texture obtained by a differential operator, which is usually represented in the form of angle to represent the direction of the most severe brightness change in the image; the inter-frame change tensor is a three-dimensional array, and the three channels thereof respectively record the differences in brightness, structure and gradient caused by time evolution in continuous image frames, which is used to capture the dynamic evolution of features in the image sequence; the spatial dimension of the tensor corresponds to the image coordinates, and the channel dimension corresponds to the feature dimension, which can express the detail changes of local pixels and also can be used as input to participate in subsequent trajectory modeling and similarity analysis, thereby providing basic data basis for image behavior judgment.
[0030] S202, linearly expanding the inter-frame change tensor in the time dimension to construct a time sequence trajectory consistency matrix according to the position continuity and feature change amplitude of each pixel in different image frames, wherein the column of the matrix represents the image frame sequence, and the row represents the pixel point trajectory in each region. In order to construct the time sequence trajectory consistency matrix for identifying the consistency of the target and the reflection area behavior, the inter-frame change tensor needs to be unfolded in time sequence. For each pixel position, the brightness change amount, the edge response change amount and the direction gradient change amount of the pixel in the continuous image frames are extracted to form a three-dimensional change trajectory. The trajectory reflects the dynamic feature change of the pixel in the time dimension. In order to improve the linear readability and analysis efficiency of the trajectory, the feature trajectories can be arranged as column vectors according to the frame sequence, and each row of the matrix represents the complete feature evolution process of a pixel point in the region. For a region, the overall behavior can be described by the collective trend of all pixel trajectories. In order to quantify this trend, the position continuity calculation is introduced, that is, whether the same pixel point remains at the original position or fluctuates within the allowed offset range in the continuous frames is judged, and the stability of the feature change amplitude is combined to generate a matrix structure reflecting the consistency degree of the trajectory. By constructing such trajectory consistency matrices for the target area and the reflection area respectively and comparing them, whether the change patterns in the time dimension of the two are highly consistent can be revealed, so as to identify whether there is abnormal pseudo-target behavior.
[0031] The inter-frame change tensor is a three-dimensional structure with time as the main axis and image feature change as the channel, which can record the detailed feature change of the image region over time. Linear unfolding in the time dimension is to arrange the tensor in sequence in the time axis direction, the purpose is to maintain the time consistency of the data, so that the trajectory analysis can be carried out through matrix operation. The pixel position continuity is an important indicator for measuring whether the pixel remains stable or slides within a small range in the continuous frames, which is used to judge whether the region is truly moving. The feature change amplitude reflects the degree of visual response of the pixel point, which is a key basis for identifying dynamic regions. The time sequence trajectory consistency matrix is a two-dimensional structure, the columns correspond to the time frame sequence, the rows correspond to the pixel point trajectory, and each element in the matrix records the feature change value of a pixel at a certain time. The matrix can not only be used to describe the behavior of the pixel, but also can be analyzed by statistical method to analyze the overall change trend of the region in the time dimension, providing data support for subsequent judgment of the consistency of the region behavior and the reflection pseudo-target.
[0032] S203, calculate the frequency stability factor and the background similarity factor based on the time sequence trajectory consistency matrix, the frequency stability factor is obtained according to the frame proportion of the change direction consistency and the change amplitude exceeding the threshold in the continuous frames, the background similarity factor is obtained by comparing the texture gray level histogram similarity of the region, and the two factors jointly constitute the image displacement abnormal pattern criterion, when the target area and the reflection area both meet the abnormal conditions in the trajectory consistency and the background similarity, the image displacement abnormal pattern in the case of reflection offset in the same direction as the target direction in the camera pose control process is determined.
[0033] In this embodiment, S203 is specifically: In the time sequence trajectory consistency matrix, the feature change direction of each pixel position in the target region and the reflective region in the continuous image frames is extracted, and the number of frames in which the change direction is consistent and the change amplitude exceeds a set numerical threshold in each pair of adjacent image frames is calculated. The frequency stability factor is calculated by the proportion of the number of frames to the total number of frames, which is used to represent the direction continuity and amplitude effectiveness of the trajectory change behavior. In order to calculate the frequency stability factor, the feature change direction of each pixel in the target region and the reflective region in the continuous image frames needs to be extracted in the time sequence trajectory consistency matrix, such as the brightness channel change trend, the edge response angle change trend or the gradient vector direction change trend. By calculating the direction angle of the feature vector of each pair of adjacent frames, it can be judged whether the change direction is consistent. When a certain pixel point continuously shows the same direction change in multiple frames of images, and the change amplitude exceeds a set numerical threshold, it can be considered that the pixel has trajectory direction continuity and response significance. The number of frames meeting this condition is accumulated, and the proportion of the total number of frames is obtained. The frequency stability score of the pixel point is obtained. By statistically averaging or weightedly integrating the scores of all pixels in the region, the frequency stability factor of the entire region can be obtained. This is an important quantitative basis for measuring whether the region trajectory behavior has pseudo-consistency. For example, if multiple pixels in the reflective region maintain the same direction for eight frames out of ten frames and the brightness change is greater than a certain threshold, the frequency stability factor reaches 0.8, indicating that the region has direction stability but not real displacement behavior.
[0034] The feature change direction refers to the change trend of the same type of image feature of the pixel point in adjacent frames, such as brightness rising or falling, gradient direction clockwise rotation or counterclockwise rotation, etc. The set numerical threshold is an important parameter for controlling noise interference, which is used to filter out pixels with weak change amplitude, ensuring the stability and discrimination ability of the factor. The change direction consistency is realized by angle judgment, which is usually considered consistent when the angle is less than a certain angle. The frame number refers to the time length of the trend persistence in a time window, which is used to measure the persistence of the feature trend. The frequency stability factor is a normalized index, usually between 0 and 1, reflecting whether a region shows continuous, strong and unified change behavior in the time dimension. The direction continuity refers to the continuity of the change direction in the frame sequence, and the amplitude effectiveness measures whether the change exceeds the threshold, both of which ensure that the factor can accurately distinguish the trajectory difference between the real target and the environmental pseudo-target.
[0035] The gray value distribution of the target region and the reflective region in the image frames is collected, the texture distribution vector is generated by using the gray histogram, the gray distribution difference value between the regions is calculated by using the Bhattacharyya distance algorithm, and the background similarity factor is output by comparing the gray distribution difference value between the regions with the preset similarity judgment threshold, which is used to measure the consistency of the image background of the two regions. To determine the similarity of the target region and the reflective region in the image on the background texture, the gray value distribution of the two regions needs to be extracted in each frame of image. The gray histogram is generated by counting the number of pixels at each gray level, thereby mapping the two-dimensional image region to a one-dimensional texture vector. Based on this, the Bhattacharyya distance algorithm can be used to calculate the similarity of the two gray histograms. The algorithm measures the overlap by accumulating the product of the two distribution vectors at each gray level and taking the logarithm. The smaller the value, the closer the gray distribution of the two regions, and the higher the texture similarity. The calculation result is compared with the preset similarity judgment threshold. If the difference is lower than the threshold, it means that the texture distribution of the two regions is convergent, and they may belong to the same background entity, indicating the possibility of the reflective region "camouflaging" as the target region. For example, in an indoor hall, the reflective region may show a very similar brightness gradient to the target back due to floor reflection, resulting in an overlap in the gray histogram and leading to a false judgment of the real target extension.
[0036] The gray histogram is a basic means for describing the local brightness distribution in image processing. By dividing the pixel gray scale into a fixed number of level intervals and counting the number of pixels in each interval, the brightness texture structure of the image region can be reflected. The texture distribution vector is constructed from the histogram and is a digital representation of the local texture. The Bhattacharyya distance algorithm is an algorithm for measuring the overlap between two statistical distributions and is widely used in image matching and region recognition. Its advantage is that it can still provide stable distance measurement in the case of wide feature dimension distribution. The preset similarity judgment threshold is used to quantitatively determine whether the two regions can be considered as belonging to the same background category. This threshold is adjusted according to the false judgment sensitivity in the actual monitoring scene and is generally set through sample training or experimental experience. The selection of this threshold plays a key role in controlling the false and missed detection rates and is the core reference standard for the background similarity factor.
[0037] The frequency stability factor and the background similarity factor are used as joint inputs, and a linear weighted discriminant function is used to construct the image displacement anomaly pattern criterion. When the frequency stability factor is greater than the first threshold and the background similarity factor is less than the second threshold, the image displacement anomaly pattern is determined, which occurs when the reflective light offset in the same direction as the target direction appears in the picture during the camera pose control process. The image displacement anomaly pattern refers to the linear position displacement of the image region in the consistent direction in consecutive image frames, the region internal edge structure change rate is lower than the preset edge stability threshold, and the Bhattacharyya distance of the region gray histogram is less than the preset texture difference threshold, which is used to represent the control misjudgment trend caused by the movement of the non-real target.
[0038] To identify pseudo-motion behavior caused by reflections during camera attitude control, it is necessary to jointly model the behavioral characteristics of the target area and the reflective area. Using previously extracted frequency stability factors and background similarity factors as input signals, a linear weighted discriminant function is constructed to comprehensively evaluate whether the offset behavior is abnormal. The frequency stability factor characterizes the continuity of the trajectory direction, while the background similarity factor reflects the consistency of the regional texture. When the former exceeds a first threshold and the latter is below a second threshold, the reflective area can be considered to have a motion trend consistent with the real target, but its texture features are too similar to the background, indicating that it is not a real target movement but rather caused by environmental reflection. Further judgment is then made by combining the edge structure change rate and texture histogram differences. If the edge structure remains stable and the grayscale distribution is similar to the target, it is identified as an abnormal image displacement pattern. For example, if a ground reflection point in consecutive frames always moves along the upper right direction with high frequency stability, but its texture is similar to the floor and the edges do not produce obvious disturbances, it can be identified as a reflection artifact, and attitude adjustment should be rejected.
[0039] The first threshold refers to a reference value used to determine the frequency stability factor. Its setting reflects the consistency threshold of trajectory changes. It is usually obtained by statistically analyzing the frequency of trajectory changes during the movement of real targets to obtain an empirical range, and selecting the critical value within this range as the judgment standard. The second threshold is the critical standard for the background similarity factor. It is used to determine whether the texture of the region is excessively close to the original background. Once the similarity is lower than this threshold, it can be identified as environmental reflection rather than a real target. The preset edge stability threshold is used to control the tolerance range of the edge structure within the region, avoiding misidentification of slight edge disturbances as changes in target behavior. The calculation of the edge structure change rate is based on the change in the number of contour pixels or gradient values within the region, and its dynamic trend is analyzed by the change amplitude in continuous image frames. The combined use of these three factors constitutes the judgment standard for abnormal image displacement patterns, enabling the system to filter out false target movement caused by reflections in the control logic, effectively improving the judgment accuracy of camera attitude control.
[0040] S3. Based on the determined image displacement anomaly pattern, construct an attitude adjustment confidence factor evaluation function by aggregating perturbation stability parameters, target trajectory confidence and boundary offset amplitude, in order to determine whether to perform actual attitude adjustment; In this embodiment, S3 specifically refers to: Based on the temporal tracking of abnormal image displacement patterns, the offset sequence of the center coordinates of each target region in consecutive image frames over time is extracted. The perturbation stability parameter is constructed by the standard deviation of the displacement amplitude in each time period within the sliding window, which is used to quantify the positional stability trend of the region on the time axis. After the image displacement anomaly pattern is established, the spatial position of the target region in the continuous image frames is tracked in sequence, the center coordinates of the target region in each image are extracted, and a two-dimensional coordinate sequence formed by the change of the center coordinates over time is recorded. By setting a sliding time window of fixed length, the sequence is divided into multiple consecutive time periods, and the displacement amplitude of the target center point is calculated in each time period, and then the standard deviation of the displacement amplitude values in each time period is calculated. The standard deviation value is the disturbance stability parameter, and the smaller the value is, the more stable the position of the target in the time window is, and vice versa. This calculation process can be realized by frame-level position recognition combined with Euclidean distance calculation, and finally a stability index for judging whether the target region has a violent or abnormal jitter is formed. In actual operation, for example, in ten consecutive image frames, if the target position remains approximately unchanged in space, the standard deviation tends to zero, indicating that the disturbance is minimal, and the region can be considered to be stable and does not need to be corrected.
[0041] The center coordinates refer to the geometric center position of the target region in the image, which is usually obtained by calculating the center point of the bounding box after target segmentation. The sliding time window is a time sequence structure that slides and takes values in fixed frame numbers, which is used to capture local time-varying features in dynamic video streams. The displacement amplitude refers to the spatial distance between the center points in the current frame and the previous frame, which can be used to measure the degree of instantaneous motion. The standard deviation is a statistical description of the fluctuation degree of all displacement amplitudes in the sliding time window, and is the core calculation basis of the disturbance stability parameter. The disturbance stability parameter reflects whether there is a persistent displacement trend in the image region, and is an important basis for subsequent judgment of whether the control is triggered by mistake. This parameter does not depend on the target outline and focuses on the time stability, so it has good robustness for judging false motion caused by reflection deviation. The overall calculation method has realizability and high adaptability, and is suitable for video monitoring scenes under various backgrounds and lighting conditions.
[0042] Based on the calculation result of the disturbance stability parameter, the trajectory coherence of the target region in the image frames is tracked, the trajectory confidence of the target is constructed by the trajectory interruption frequency and the trajectory direction deviation amplitude, and the position change range of the outer boundary contour of the target is extracted in the same frame sequence, and the boundary displacement amplitude is calculated as a boundary disturbance reference index. After obtaining the disturbance stability parameters, the analysis of the motion trajectory of the target region is continued. The specific implementation is to record the position sequence of the target region in the continuous image frames, track the motion trajectory of the target by using a trajectory tracking algorithm such as Kalman filtering or SORT algorithm, and mark the occurrence frequency of the trajectory interruption event. The trajectory interruption frequency is the ratio of the number of times of unsuccessfully matching the target position in the continuous frames to the total number of frames. If the interruption is frequent, it means that the target recognition continuity is poor, and the reliability is reduced. Further, the trajectory direction deviation amplitude is calculated, that is, the angle deviation between the motion direction of each frame and the global average direction. The more intense the angle change is, the more unstable the target motion path is. The two jointly constitute the target trajectory confidence, which is used to quantify the trajectory reliability. At the same time, in the same image frame sequence, the outer boundary contour of the target region in each frame of image is extracted, and the maximum offset range of the boundary contour in space, that is, the boundary offset amplitude, is calculated. The amplitude is used as a boundary disturbance reference index to identify false boundary changes caused by external interference or reflection, and to assist in judging the abnormality of the current trajectory confidence. Through the above combined calculation method, the accuracy of trajectory confidence analysis can be improved in a visual interference environment.
[0043] Trajectory continuity refers to whether the motion path of the target in the continuous image frames is smooth and continuous, which is directly related to whether the target is correctly identified continuously. The trajectory interruption frequency is a quantitative indicator of target tracking failure, which is obtained by recording the number of times of losing the target identification in the continuous frames. The trajectory direction deviation amplitude is used to measure the difference between the local trajectory direction and the global direction, and its unit is usually angle, which can identify inconsistent movement characteristics. The target trajectory confidence is a comprehensive index calculated on the basis of the trajectory interruption frequency and the direction deviation, which is usually used to determine the reliability of the tracking result. The outer boundary contour is a spatial representation of the edge after image segmentation of the target, and the boundary offset amplitude is obtained by comparing the position difference of the bounding box of the boundary in the continuous frames, which is used to quantify the stability of the target shape on the time axis. The boundary disturbance reference index reflects the jitter behavior of the region edge, which assists in identifying whether the posture adjustment misjudgment is caused by light, reflection or identification error. The above features work together to effectively support the reliable judgment of the posture control behavior.
[0044] After the disturbance stability parameters, the target trajectory confidence and the boundary offset amplitude are standardized and treated as three input variables, they are input into a weighted fusion function to construct a posture adjustment confidence factor evaluation function, and the output posture adjustment confidence factor is compared with a posture adjustment judgment threshold. When the posture adjustment confidence factor is greater than the judgment threshold, it is judged that the actual posture adjustment can be executed.
[0045] To improve the robustness of the posture adjustment decision, it is necessary to uniformly process parameters of different dimensions and integrate them into a comprehensive evaluation index. First, the disturbance stability parameter, the target trajectory confidence and the boundary offset amplitude are normalized to avoid weight deviation caused by dimensional differences during fusion. The normalized three are input into a weighted fusion function as three input variables. The function assigns a fixed influence coefficient to each input to reflect its contribution to the final confidence factor evaluation. The weighted fusion function outputs a continuous value as the posture adjustment confidence factor, representing the overall confidence of the target tracking and boundary stability state in the current image. Then the confidence factor is compared with the preset posture adjustment decision threshold. When the output value of the confidence factor is greater than the decision threshold, the system determines that the current tracking state is stable and reliable, and the actual posture adjustment can be performed; otherwise, the current camera posture remains unchanged to avoid unnecessary adjustment operations caused by image disturbance or false recognition. This calculation method can effectively prevent false triggering of control behavior and improve the intelligent level of control decision.
[0046] The disturbance stability parameter is a time stability indicator calculated by the standard deviation of the target position change in the image frame, which measures whether the target region is disturbed continuously. The target trajectory confidence is an index to evaluate the continuity and direction consistency of target tracking, reflecting the reliability of the system in judging target motion during image analysis. The boundary offset amplitude is the maximum displacement of the image edge in consecutive frames, which is used to identify boundary abnormalities caused by reflections or false targets. Standardization is a mathematical method to map data of different scales to a unified range, including Z-score normalization or Min-Max normalization, which facilitates equal participation of different parameters in fusion. The weighted fusion function is a mathematical model that assigns weight values to different parameters based on prior experience or training results to obtain a unified evaluation result. The posture adjustment confidence factor is a continuous value after fusion, used to measure whether there is a reliable basis for posture adjustment. The posture adjustment decision threshold is a reference value set in the system, indicating that the current environment meets the adjustment conditions when the confidence factor exceeds this threshold, and the adjustment behavior can be performed. This mechanism reduces false positives through quantitative analysis, improving the accuracy and efficiency of camera control in complex environments.
[0047] S4, match the posture adjustment confidence factor with the preset evaluation threshold interval, output the judgment result, and execute the corresponding camera posture control behavior according to the judgment result; In this embodiment, S4 is specifically: The preset evaluation threshold interval of the posture adjustment credibility factor is constructed, the value range of the posture adjustment credibility factor is statistically analyzed through experimental data, a plurality of continuous and non-overlapping threshold intervals are set, which correspond to different posture control levels respectively, and the evaluation threshold interval is used for grade division of the posture adjustment credibility factor, so as to ensure that the judgment process has resolution and stability; The evaluation threshold interval of the posture adjustment credibility factor is usually completed by using a statistical analysis method based on large sample experimental data. The credibility factor samples generated in the posture adjustment process of the camera can be collected under different environmental conditions, including the credibility factor values under various situations such as real target movement, reflection interference and background change. The frequency distribution analysis is performed on these values, and the value aggregation interval of the credibility factor under different judgment conditions is extracted. According to the value concentration trend and fluctuation boundary, a plurality of continuous and non-overlapping value intervals are set, for example, a three-section partition for distinguishing between prohibited adjustment, limited adjustment and complete adjustment. This interval division method can make the evaluation system make clear judgment on posture adjustment requests with different credibility levels, and convert the fuzzy continuous variable into executable control levels, thereby improving the stability and discrimination ability of the judgment result.
[0048] The posture adjustment credibility factor is a weighted evaluation output that integrates disturbance stability parameters, target trajectory confidence and boundary offset amplitude, which is usually a continuous value and must be converted into a discrete level through an evaluation threshold interval to match the camera control logic. The evaluation threshold interval is a plurality of continuous and non-overlapping value sections, each interval representing a posture control level to avoid judgment overlap or uncertainty. Statistical analysis refers to modeling the credibility factor distribution based on standard deviation, mean, quantile or cluster center to define the interval boundary. Resolution refers to the ability of the credibility factor to fall into different level intervals generated by different inputs, and stability refers to the ability of the evaluation output to remain in a consistent level under the same environmental conditions. These technical features together ensure that the judgment logic of the credibility factor has clarity, reliability and control value.
[0049] The posture adjustment credibility factor and the evaluation threshold interval are compared one by one, and the judgment result is output according to the threshold level of the posture adjustment credibility factor, the judgment result including three states of prohibited adjustment, limited adjustment and complete adjustment, the prohibited adjustment corresponding to the posture adjustment credibility factor being lower than the first evaluation threshold, the limited adjustment corresponding to the posture adjustment credibility factor being between the first evaluation threshold and the second evaluation threshold, and the complete adjustment corresponding to the posture adjustment credibility factor being higher than the second evaluation threshold; The posture adjustment confidence factor is compared with the evaluation threshold interval one by one, which can be realized by setting the condition judgment structure. First, the posture adjustment confidence factor value calculated in the current frame is obtained, and then it is matched with the multiple threshold intervals set in advance. By determining the interval level that the value falls into, the corresponding posture control judgment result is output. For example, when the confidence factor value is less than the first evaluation threshold, the system outputs the "prohibition adjustment" state; when the value is between the first evaluation threshold and the second evaluation threshold, the system outputs the "limited adjustment" state; and when the value is greater than the second evaluation threshold, the system outputs the "full adjustment" state. This way converts continuous numerical input into three types of discrete control response through a segmented decision model, which is suitable for the hierarchical response logic of different posture confidence in the video monitoring system, and helps the system to make stable and differentiated control strategies when dealing with environmental glare interference.
[0050] The posture adjustment confidence factor is a numerical variable obtained by weighted fusion of the disturbance stability parameter, the target trajectory confidence and the boundary offset amplitude, representing the confidence strength of triggering posture adjustment in the current image. The first evaluation threshold is the preset minimum adjustment standard, which is used to divide the misjudgment area caused by interference features. The confidence factor below this value indicates that the image feature fluctuation is not enough to support the camera adjustment decision; the second evaluation threshold is the upper limit value of executing full adjustment, which is used to identify the high confidence situation that the target in the image is obviously offset and exclude glare interference. The transition interval between the first evaluation threshold and the second evaluation threshold allows the system to make limited adjustment for uncertain situations, improving the flexibility of responding to edge situations. The three judgment states constitute a complete posture control decision structure, ensuring that the system has clear action instruction output logic when facing different confidence strength inputs.
[0051] According to the judgment result, the corresponding camera posture control behavior is executed. In the prohibition adjustment state, no change is made to the camera pitch angle; in the limited adjustment state, the camera pitch angle is corrected through proportional fine-tuning strategy; in the full adjustment state, the camera posture change vector is calculated according to the target offset direction and adjusted, realizing the camera posture control behavior based on the judgment result.
[0052] When executing corresponding camera attitude control actions based on the judgment result, the system's output judgment state is first used as the trigger condition for the control command. When the judgment state is "adjustment prohibited," the control system maintains the camera's current pitch angle unchanged to prevent invalid actions due to misidentification of reflective offset. When the judgment state is "limited adjustment," the system scales the target offset by setting an adjustment ratio coefficient, generating a fine-tuning vector for the camera attitude, and driving a small adjustment of the pitch angle. This allows the system to perform safe compensation under uncertain conditions without causing drastic offsets. When the judgment state is "full adjustment," the system calculates the complete attitude change vector based on the target's actual offset direction in the image and quickly adjusts the camera's pitch angle through the servo control module to reposition the target in the center of the image. By executing attitude adjustment operations of different magnitudes based on a hierarchical control strategy, the system can effectively reduce erroneous adjustment behaviors caused by environmental interference.
[0053] Camera attitude control involves dynamic adjustment of the pitch angle and matching of execution logic. Core parameters include the judgment state, target offset direction, and attitude change vector. The "no adjustment" state filters adjustments under low-confidence conditions, ensuring system stability even with false displacement signals such as light spot interference. The proportional fine-tuning strategy corresponding to the "limited adjustment" state requires determining the adjustment coefficient based on the system's historical stability and the current offset; it typically has a small value and is used to finely compensate for potentially real but insufficiently confident target offset behavior. The attitude change vector corresponding to the "fully adjusted" state is calculated from the displacement of the target region relative to the image's center point. The vector direction determines the trend of the adjustment angle, while the vector length affects the adjustment magnitude. The system needs to perform vector analysis based on real-time tracking data before driving the camera to execute actions. This hierarchical control mechanism provides a control strategy that combines fine precision, robustness, and responsiveness for camera attitude adjustment.
[0054] S5. After executing the camera posture control behavior, the image state after control is collected, the evaluation function parameters are dynamically updated within the sliding time window, and the camera posture control judgment logic is adjusted according to the update result to realize the dynamic control of the camera control behavior.
[0055] In this embodiment, S5 specifically refers to: After executing camera posture control, the video monitoring system continuously acquires image frames after control, extracts brightness channel, edge structure, orientation gradient and contrast change information in the image frames after control, forms image state data after control, and marks the image response segment corresponding to the current posture adjustment as an evaluation sample. After executing camera attitude control, the adjusted image state needs to be continuously monitored through a video surveillance system to assess the actual impact of the adjustment on image quality and target tracking accuracy. A time frame can be set, and a series of consecutive image frames can be acquired in real time after the attitude change. The image processing module extracts features from each frame, specifically including grayscale distribution extraction of the brightness channel, contour detection of edge structures, gradient direction calculation of the orientation gradient, and grayscale difference calculation of contrast changes. These features from each frame are integrated chronologically to form the post-control image state data. The start and end times of the attitude adjustment are marked within a specific frame sequence, and image segments within that time period are extracted as evaluation samples for subsequent dynamic updates of the evaluation function and adjustments to the judgment strategy. This approach ensures that the image performance after attitude adjustment is fully captured and can be quantitatively compared with the previous state.
[0056] In the above processing, the luminance channel refers to the grayscale layer in the image color space that reflects light intensity information, generally obtained by converting an RGB image to a grayscale image; edge structure is formed by detecting image gradient changes using operators such as Sobel, Canny, or Laplacian to form edge contour information; directional gradient is the directional angle calculated based on the rate of change of luminance of each pixel in the horizontal and vertical directions, used to characterize the distribution of texture direction; contrast change reflects image sharpness and regional differences by calculating the difference in pixel grayscale values within local areas. The controlled image state data is multidimensional structured data composed of these features. The extraction of evaluation samples requires complete coverage of image frames during the pose change period to ensure the capture of image responses caused by pose changes, thereby providing a real and effective observation basis for the dynamic control mechanism.
[0057] The controlled image state data is input into a sliding time window of a set length. The disturbance stability parameters, target trajectory confidence and boundary offset amplitude are dynamically updated based on the continuous change trend of the image frames within the window. The attitude adjustment confidence factor is recalculated based on the updated parameters to complete the dynamic update of the evaluation function parameters. After the camera attitude adjustment is completed, to achieve dynamic feedback control, the acquired post-control image state data needs to be continuously input into a sliding time window of a set length. This sliding window sequentially accommodates the latest image frames in the time dimension. Whenever a new frame enters, the oldest frame is removed, ensuring that the analysis is always based on the latest image response data. Within each window period, temporal analysis is performed on the characteristics of the image frames in the window, such as brightness channel, edge structure, orientation gradient, and contrast changes, to update the perturbation stability parameters, target trajectory confidence, and boundary offset amplitude in real time. The perturbation stability parameters are obtained by calculating the standard deviation of the change in the center coordinates of the target region within the window. The target trajectory confidence is obtained based on the analysis of trajectory continuity and orientation consistency. The boundary offset amplitude is obtained by measuring the difference in the outer contour position between frames. After these parameters are updated, they are used as input variables again and fed into the weighted evaluation function to calculate the latest attitude adjustment confidence factor, realizing the adaptive update of the attitude control evaluation mechanism to cope with the dynamic changes in target behavior and image state.
[0058] In this processing, a sliding time window refers to a time series container with a fixed number of frames, often used to filter out short-term fluctuations and capture trend changes. Image state data is a structured dataset generated by the feature extraction module, mainly including brightness change curves, edge intensity maps, orientation gradient distribution maps, and local contrast change maps. The perturbation stability parameter reflects the smoothness of the target's positional changes on the time axis in the image; the smaller the value, the smaller the positional fluctuation. The target trajectory confidence measures whether the target's trajectory within the window is interrupted, reversed, or jumps, and is a quantitative representation of trajectory continuity. The boundary offset magnitude describes the degree of edge perturbation by comparing the maximum change distance of the target region boundary coordinates in adjacent frames. The updates of these three parameters can reflect the stability and target trackability after attitude adjustment in real time. The recalculation of the attitude adjustment confidence factor provides a basis for whether to continue adjusting the camera, thereby achieving refined, closed-loop dynamic control.
[0059] Based on the changing relationship between the dynamically updated attitude adjustment confidence factor and the historical evaluation function parameters, it is determined whether it continuously deviates from the original judgment trend range. When the continuous deviation reaches the preset fluctuation range, the judgment logic of the attitude adjustment confidence factor is reset, including correcting the evaluation threshold range or adjusting the judgment level mapping rule, so as to realize the adaptive dynamic control of the camera attitude control judgment logic.
[0060] During continuous adjustment of camera attitude control, it is necessary to analyze the numerical trend changes between the dynamically updated attitude adjustment confidence factor and the historical evaluation function parameters to determine whether the system output exhibits a stable deviation. If the new round of attitude adjustment confidence factor deviates from the judgment trend range formed by the historical evaluation function for several consecutive cycles, i.e., the judgment result continuously falls into an attitude control level different from the past, it is judged as a judgment logic imbalance. This imbalance may stem from increased external interference, changes in target behavior patterns, or delayed camera feedback response. To address this, an adaptive reset mechanism should be activated. Based on a preset fluctuation range, it should determine whether the deviation is significant. If it exceeds the fluctuation tolerance, the currently used evaluation threshold range should be automatically corrected, or the mapping rules between the confidence factor level and attitude control behavior should be adjusted. This dynamically updates the judgment logic, enabling the system to adapt to changing environments and improving the stability and real-time performance of the control strategy.
[0061] The evaluation function parameters mainly include disturbance stability parameters, target trajectory confidence, and boundary offset amplitude. These parameters constitute the basic data source for the attitude adjustment confidence factor. The attitude adjustment confidence factor is a control confidence value calculated in each evaluation cycle, used to guide whether the camera attitude should be adjusted. The historical evaluation function parameter sequence is a set recording all confidence factor output values within a previous time window, used to construct a trend model. The judgment trend interval is a fluctuation tolerance range defined based on historical data distribution, serving as the reference boundary for the control judgment logic. The preset fluctuation range is an acceptable offset range predefined by the system, used to determine whether the current trend deviates from the normal trajectory. When the confidence factor continuously deviates beyond this range, the system will dynamically adjust the upper and lower limits of the evaluation threshold interval based on the offset direction and amplitude, or reconstruct the level mapping relationship, making the judgment model closer to the current actual monitoring situation. This strategy enables the attitude control system to no longer rely on static rules, but to have continuous learning and adjustment capabilities.
[0062] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions according to the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means (e.g., infrared, wireless, microwave, etc.). A computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. Semiconductor media can be solid-state drives.
[0063] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0064] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0065] In the several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0066] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0067] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0068] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A control method for a video surveillance system, characterized in that, Specifically, the following steps are included: S1. Collect continuous image frames through the video surveillance system, extract the brightness channel, edge structure, orientation gradient and contrast changes of the image, construct a reflective perception vector group, and calculate the angle change trend in combination with the target motion direction vector to determine whether there is a reflective offset in the image that is consistent with the target direction during the camera attitude control process. S2. After determining that there is a reflection offset in the same direction as the target in the picture, extract the inter-frame change tensor between the target area and the reflection area, construct the temporal trajectory consistency matrix, and determine the image displacement anomaly mode when there is a reflection offset in the same direction as the target in the picture during the camera attitude control process based on the frequency stability factor and the background similarity factor. S3. Based on the determined image displacement anomaly pattern, construct an attitude adjustment confidence factor evaluation function by aggregating perturbation stability parameters, target trajectory confidence and boundary offset amplitude, in order to determine whether to perform actual attitude adjustment; S4. Match the attitude adjustment confidence factor with the preset evaluation threshold range, output the judgment result, and execute the corresponding camera attitude control behavior according to the judgment result. S5. After executing the camera posture control behavior, the image state after control is collected, the evaluation function parameters are dynamically updated within the sliding time window, and the camera posture control judgment logic is adjusted according to the update result to realize the dynamic control of the camera control behavior.
2. The control method for the video surveillance system according to claim 1, characterized in that, S1 specifically includes the following steps: S101. Acquire continuous image frames through a video surveillance system, perform pixel-level separation on each frame to extract the brightness channel of the image, extract the edge structure of the image based on the edge detection operator, obtain the orientation gradient information of the image through the gradient operator, and use a multi-scale window to traverse the image area and calculate the gray-level difference of each area to extract the contrast change of the image. S102. Feature fusion is performed using the brightness channel, edge structure, directional gradient and contrast change of the image in the same coordinate space. Combined with the inter-frame change trend of the image, a reflective perception vector group is constructed. The constructed reflective perception vector group represents the directional change of the brightness enhancement area in the spatial and temporal dimensions. S103. The motion direction vector of the target in the image is identified by the video monitoring system, and the angle change trend of the motion direction vector and each vector in the reflective sensing vector group is calculated. When there are consecutive frames where the angle is maintained within the preset specified threshold and the vector trajectory is continuous, it is determined that a reflective offset with the same direction as the target appears in the picture during the camera attitude control process.
3. The control method for the video surveillance system according to claim 2, characterized in that, S102 specifically refers to: The brightness channel, edge structure, orientation gradient and contrast change extracted from consecutive image frames are mapped in two-dimensional space, and the coordinates of each image feature map are registered by bilinear interpolation algorithm so that the brightness boundary, gradient direction and contrast difference can be superimposed at the same pixel position. Based on pixel-level position correspondence, the brightness channel map, edge structure map, orientation gradient map and contrast change map are fused in a unified coordinate space to construct a four-dimensional joint feature vector for each pixel position, and the change trend of brightness enhancement area in consecutive frames is extracted by inter-frame difference. A reflective sensing vector group is constructed based on the four-dimensional joint feature vector of the brightness enhancement region. The image gradient direction is encoded in the spatial dimension, and the brightness change direction is encoded in the temporal dimension, thus representing the directional change of the brightness enhancement region in the spatial and temporal dimensions.
4. The control method for the video surveillance system according to claim 1, characterized in that, S2 specifically includes the following steps: S201. Based on the target area and reflective area identified by the video surveillance system, acquire brightness channel, edge structure and orientation gradient information in continuous image frames, perform difference operation on the feature values of the corresponding pixel positions of each area in each frame in time order to form an inter-frame change tensor. The channels of the inter-frame change tensor correspond to the brightness change, structural response change and gradient direction change, respectively. S202. The inter-frame change tensor is linearly expanded in the time dimension. A temporal trajectory consistency matrix is constructed according to the positional continuity and feature change amplitude of each pixel in different image frames. The columns of the matrix represent the image frame order, and the rows represent the pixel trajectory in each region. The continuous change trajectory of the entire region in time is encoded to reveal the behavioral consistency between the target and the reflection. S203. Calculate the frequency stability factor and background similarity factor based on the temporal trajectory consistency matrix. The frequency stability factor is obtained based on the proportion of frames in consecutive frames with consistent change direction and change amplitude exceeding the threshold. The background similarity factor is obtained by comparing the similarity of the grayscale histograms of the regional textures. The two factors together constitute the image displacement anomaly mode criterion. When the target area and the reflective area both reach the abnormal conditions in terms of trajectory consistency and background similarity, the image displacement anomaly mode under the condition of reflective offset in the same direction as the target in the picture is determined during the camera attitude control process.
5. The control method for the video surveillance system according to claim 4, characterized in that, S203 specifically refers to: In the temporal trajectory consistency matrix, the characteristic change direction of each pixel position of the target area and the reflective area in consecutive image frames is extracted, and the number of frames with consistent change direction and change amplitude exceeding a set value threshold in each pair of adjacent image frames is calculated. The frequency stability factor is calculated by the proportion of this number of frames to the total number of frames, which is used to characterize the directional continuity and amplitude effectiveness of trajectory change behavior. The grayscale distribution of the target area and the reflective area in the image frame is collected. A texture distribution vector is generated using the grayscale histogram. The grayscale distribution difference between the regions is calculated using the Bhattacharyya distance algorithm. The grayscale distribution difference between the regions is compared with a preset similarity judgment threshold, and a background similarity factor is output to measure the degree of consistency between the image backgrounds of the two regions. The frequency stability factor and background similarity factor are used as joint inputs, and a linear weighted discriminant function is used to construct the image displacement anomaly pattern criterion. When the frequency stability factor is greater than the first threshold and the background similarity factor is less than the second threshold, the image displacement anomaly pattern is determined when there is a reflective offset in the same direction as the target in the image during the camera attitude control process. The image displacement anomaly pattern refers to the linear positional offset of the image region in the same direction in consecutive image frames, the rate of change of the edge structure inside the region is lower than the preset edge stability threshold, and the Bhattacharyya distance of the region grayscale histogram is less than the preset texture difference threshold, which is used to indicate the control misjudgment trend caused by non-real target movement.
6. The control method for the video surveillance system according to claim 1, characterized in that, S3 specifically refers to: Based on the temporal tracking of abnormal image displacement patterns, the offset sequence of the center coordinates of each target region in consecutive image frames over time is extracted. The perturbation stability parameter is constructed by the standard deviation of the displacement amplitude in each time period within the sliding window, which is used to quantify the positional stability trend of the region on the time axis. Based on the calculation results of the perturbation stability parameters, the trajectory continuity of the target region in the image frame is tracked. The confidence of the target trajectory is constructed by the trajectory interruption frequency and the trajectory direction deviation magnitude. At the same time, the range of change of the outer boundary contour position of the target is extracted in the same frame sequence, and the boundary offset magnitude is calculated as a boundary perturbation reference index. The perturbation stability parameters, target trajectory confidence, and boundary offset amplitude are standardized and used as ternary input variables. These variables are then input into a weighted fusion function to construct an attitude adjustment confidence factor evaluation function. The output attitude adjustment confidence factor is compared with the attitude adjustment judgment threshold. When the attitude adjustment confidence factor is greater than the judgment threshold, it is determined that actual attitude adjustment can be performed.
7. The control method for the video surveillance system according to claim 1, characterized in that, S4 specifically refers to: A preset evaluation threshold range for the attitude adjustment confidence factor is constructed. The range of values of the attitude adjustment confidence factor is statistically analyzed through experimental data. Multiple continuous and non-overlapping threshold ranges are set, each corresponding to a different attitude control level. The evaluation threshold range is used to classify the attitude adjustment confidence factor into levels, ensuring that the judgment process is discriminative and stable. The attitude adjustment confidence factor is compared with the evaluation threshold range one by one, and the judgment result is output according to the threshold level of the attitude adjustment confidence factor. The judgment result includes three states: no adjustment, limited adjustment, and complete adjustment. No adjustment corresponds to an attitude adjustment confidence factor lower than the first evaluation threshold, limited adjustment corresponds to an attitude adjustment confidence factor between the first evaluation threshold and the second evaluation threshold, and complete adjustment corresponds to an attitude adjustment confidence factor exceeding the second evaluation threshold. Based on the judgment result, the corresponding camera attitude control behavior is executed. In the state where adjustment is prohibited, no change is made to the camera pitch angle. In the state of limited adjustment, the camera pitch angle is corrected through a proportional fine-tuning strategy. In the state of full adjustment, the camera attitude change vector is calculated based on the target offset direction and adjusted accordingly, thereby realizing camera attitude control behavior based on the judgment result.
8. The control method for the video surveillance system according to claim 1, characterized in that, S5 specifically refers to: After executing camera posture control, the video monitoring system continuously acquires image frames after control, extracts brightness channel, edge structure, orientation gradient and contrast change information in the image frames after control, forms image state data after control, and marks the image response segment corresponding to the current posture adjustment as an evaluation sample. The controlled image state data is input into a sliding time window of a set length. The disturbance stability parameters, target trajectory confidence and boundary offset amplitude are dynamically updated based on the continuous change trend of the image frames within the window. The attitude adjustment confidence factor is recalculated based on the updated parameters to complete the dynamic update of the evaluation function parameters. Based on the changing relationship between the dynamically updated attitude adjustment confidence factor and the historical evaluation function parameters, it is determined whether it continuously deviates from the original judgment trend range. When the continuous deviation reaches the preset fluctuation range, the judgment logic of the attitude adjustment confidence factor is reset, including correcting the evaluation threshold range or adjusting the judgment level mapping rule, so as to realize the adaptive dynamic control of the camera attitude control judgment logic.
Citation Information
Cited By
Dynamic patrol-based guideboard fading detection method and system
CN121482512A
Entrance guard authorization authentication method and system based on smart campus monitoring system
CN121686333A
Real-time monitoring method, equipment and system for water pollution of sea-entering river
CN121746922A
Space sound scene self-adaptive control method with head tracking function
CN122093735A