Control method and system for monitoring camera
Through the control method of the surveillance camera, the image recognition and tracking algorithm are used to detect key target areas, calculate fuzzy feature values and out-of-focus rate, generate focus control signals, and achieve automatic focus of key targets, solving the problem of timely identification and focus in the existing technology, and improving imaging quality and intelligence level.
Patent Information
- Application Number
- CN202510787228.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-13
AI Technical Summary
When the existing surveillance camera control method handles complex scenarios where high-frequency target movements, frequent occlusions or multiple targets cross, it cannot identify and focus key targets in time, resulting in a decrease in recognition accuracy and frequent algorithm misjudgment.
The key target area is detected through image recognition and tracking algorithms, the spatial feature set is obtained, the target fuzzy feature value is calculated, the out-of-focus frame judgment mark is constructed, the target out-of-focus rate is counted, and the target out-of-focus rate is compared with the preset threshold value is generated, the focus control signal is generated, local automatic focus is achieved, and the fuzzy feature value is reacquired to evaluate the focus effect.
It effectively alleviates the problem of false triggering caused by temporary interference and occlusion, improves the ability to respond to key target imaging clarity, enhances the accuracy of image blur state determination and the degree of intelligence of focus control, and avoids resource waste and delay.
Smart Images

Figure CN120343399A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of camera control, and specifically provides a control method and system for a surveillance camera. Background Art
[0002] In the digital construction of modern cities, the security system has gradually developed into a complex interdisciplinary integration system. Among them, video surveillance technology, as a key component, is not only widely used in multiple application scenarios such as public safety management, traffic order maintenance, commercial retail management, and smart campuses, but also gradually evolving towards intelligent perception and automated response. The control method of surveillance cameras is not limited to the basic functions of shooting start / stop, angle rotation, or focal length stretching, but rather points to a high-level scheduling strategy for aspects such as the imaging quality, response rhythm, and parameter self-optimization ability of the camera. Under this strategy system, image clarity control, as a fundamental and core element, is particularly important. Especially in scenarios such as urban traffic intersections, school entrances and exits, bank lobbies, or intelligent warehouses, the recognition accuracy of key objects such as faces, license plates, gestures, and signs directly affects the overall function performance and response ability of the system. Therefore, the real-time regulation of the imaging effect of key objects has become an urgently needed direction for in-depth optimization in the control method.
[0003] In the existing camera focusing control system, most systems still adopt an autofocus mechanism triggered based on fixed interval time or static blur threshold. Although this mechanism can maintain the basic imaging quality in most cases, in complex scenarios with high-frequency target movement, frequent occlusion, or multi-target intersection, there are often problems that the target frequently enters the out-of-focus state and fails to be recognized and focused in a timely manner. This is because traditional control strategies often use blur detection as the judgment standard of "average blur of the whole image" and cannot perform region-by-region and target-by-target clarity detection for the "key object area" in the video frame. In addition, even if the frame at a certain moment is judged by the system as "clear", it may cover up the fact that key targets, such as faces or license plates, are actually out of focus in part. These defects lead to the inability of the focusing strategy to effectively interact with the target detection and target tracking modules in actual operation, making the out-of-focus problem one of the important causes of the decline in recognition accuracy and frequent algorithm misjudgment in the current video analysis system. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a control method and system for a surveillance camera, which solves the problems mentioned in the background art.
[0005] To achieve the above object, the present invention is realized through the following technical solutions: A control method for a surveillance camera includes the following steps: S1. Detect and track the key target areas in the surveillance images through image recognition and tracking algorithms, and obtain the spatial feature set FTR1 of the targets; S2. Based on the spatial feature set FTR1, extract the target blur feature values Blv of the fixed areas in each frame of the image, and obtain the fuzzy vector sequence set Ftr2; S3. Compare the preset fuzzy evaluation threshold Thf with the fuzzy vector sequence set Ftr2 traversally to construct the out-of-focus frame judgment mark set Ftr3; S4. Cumulatively count the out-of-focus frame judgment mark set Ftr3, and calculate the target out-of-focus rate Tbd within the fixed time window T; S5. Compare the obtained target out-of-focus rate Tbd with the preset out-of-focus threshold Thd, obtain the average position PosAvg and average size SizeAvg of the targets in the spatial feature set FTR1 according to the comparison result, calculate the focus control signal Foc, and apply the focus control signal Foc to perform local autofocus on the targets; S6. After performing autofocus, collect the target blur feature values Blv of the target area again, obtain the new post-execution fuzzy vector sequence set Ftr4, and then compare it with the previously collected fuzzy vector sequence set Ftr2, calculate the focus effect value Qev, and judge the improvement degree of the image quality after focusing.
[0006] Preferably, the S1 includes S11; S11. Collect the image frame Frame(t) at time t through the camera in real time, and then use the image recognition algorithm to process each image frame Frame(t), detect the key targets in the preset key target list, and extract the spatial positions and boundary information of the key targets to generate the initial target spatial features. Integrate all the initial target spatial features of the image frame Frame(t) at time t to obtain the target set Ftr11; The specific form of the target set Ftr11 is the target set Ftr11 = {Bnd(t), Pos(t), Size(t)}; where Bnd(t) represents the target bounding box at time t, specifically representing the rectangular box in the coordinates of the image frame Frame(t); Pos(t) represents the center point coordinates of the target position at time t, specifically representing the center point of the target bounding box Bnd(t) at time t; Size(t) represents the target size at time t, specifically representing the width and height of the target bounding box Bnd(t) at time t; Among them, the key target list includes faces, license plates, and people.
[0007] Preferably, the S1 includes S12; S12. Based on the obtained target set Ftr11, perform object matching between consecutive image frames Frame(t) to achieve consistent recognition of key targets, including using the DeepSORT and KCF tracking algorithms to perform consistency recognition judgment of object matching, and construct a time series tracking record of key targets, marked as the spatial feature set FTR1 of the target; The specific form of the spatial feature set FTR1 is the spatial feature set FTR1 = {Bnd(t), Pos(t), Size(t)|t ∈ T}, where T represents the time window, specifically representing the continuous recording of the target bounding box Bnd(t), the target position center point coordinates Pos(t), and the target size Size(t) of the key target within the time window T.
[0008] Preferably, the above S2 includes S21; S21. Based on the spatial feature set FTR1, use the target area defined by the target bounding box Bnd(t) as the sampling window in each image frame Frame(t). After normalization, eliminate the dimensional differences between different features, extract the image gradient variance Grd(t), Laplace response variance Lap(t), and high-frequency energy ratio Hfr(t) of the target area image at time t, and obtain the target blur feature value Blv(t) at time t through weighted calculation. After integrating all target areas, form the fuzzy vector sequence set Ftr2; The specific form of the fuzzy vector sequence set Ftr2 is the fuzzy vector sequence set Ftr2 = {Blv(t), t ∈ T}, where T represents the time window; The target blur feature value Blv(t) is obtained through the following calculation formula: ; In the formula, A1, A2, and A3 respectively represent the preset weight values of the image gradient variance Grd(t), Laplace response variance Lap(t), and high-frequency energy ratio Hfr(t) at time t, and A1 + A2 + A3 = 1, and the specific values are set by the user.
[0009] Preferably, the image gradient variance Grd(t) at time t is obtained by cropping the target area image Roi(t) at time t from each image frame Frame(t) with the target bounding box Bnd(t), and then performing the Scharr operator on the target area image Roi(t) at time t, and calculating the variance after calculating the horizontal and vertical gradients of the target area image Roi(t) at time t; The Laplace response variance Lap(t) is obtained by applying the Laplace operator for second-order derivative filtering to the target region image Roi(t) at the obtained time t, capturing the edge sharpness in the image, and calculating the variance of its response value; The high-frequency energy ratio Hfr(t) is obtained by performing Fourier transform FFT on the target region image Roi(t) at the obtained time t, transforming the image from the spatial domain to the frequency domain, and calculating the proportion of the energy of the high-frequency part in the total energy in the Fourier spectrum.
[0010] Preferably, S3 includes S31; S31. Based on the obtained set of fuzzy vector sequences Ftr2, traverse and compare with the preset fuzzy evaluation threshold Thf to determine whether the target fuzzy feature value Blv(t) at time t in each frame image Frame(t) is in an out-of-focus state, and mark it with a boolean value to obtain the out-of-focus state mark Blf(t) at time t, and construct an out-of-focus frame judgment mark set Ftr3; The specific form of the out-of-focus frame judgment mark set Ftr3 is the out-of-focus frame judgment mark set Ftr3 = {Blf(t)|t ∈ T}, where T represents the time window; Among them, the out-of-focus state mark Blf(t) at time t is obtained through the following comparison method: When the target fuzzy feature value Blv(t) at time t < the fuzzy evaluation threshold Thf, it indicates that there is an out-of-focus area in the frame image Frame(t) at time t, which means an out-of-focus frame, and mark the out-of-focus state mark Blf(t) at time t = 1; When the target fuzzy feature value Blv(t) at time t ≥ the fuzzy evaluation threshold Thf, it indicates that there is no out-of-focus area in the frame image Frame(t) at time t, which means a clear frame, and mark the out-of-focus state mark Blf(t) at time t = 0.
[0011] Preferably, S4 includes S41; S41. Based on the out-of-focus frame judgment mark set Ftr3, count the number of frames with the out-of-focus state mark Blf(t) = 1 within the time window T, and calculate the proportion of the target region in the out-of-focus frame within the time window T to obtain the out-of-focus rate numerical index Tbd, which reflects the change trend of the image quality of the target within the time window T and is used to determine whether to trigger focus control; The out-of-focus rate numerical index Tbd is obtained through the calculation formula.
[0012] Preferably, S5 includes S51; S51. Compare the obtained target defocus rate Tbd with the preset defocus threshold Thd. According to the comparison result, obtain the average position PosAvg and average size SizeAvg of the target in the spatial feature set FTR1, calculate the focus control signal Foc, and apply the focus control signal Foc to perform local autofocus on the target; The comparison result is obtained through the following comparison method: When the target defocus rate Tbd ≥ defocus threshold Thd, it indicates the result that triggers the execution of the comparison result, and perform the local autofocus operation, including obtaining the average position PosAvg and average size SizeAvg of the target in the spatial feature set FTR1, and calculating the focus control signal Foc; When the target defocus rate Tbd < defocus threshold Thd, it indicates the result that does not trigger the execution of the comparison result, and do not perform the local autofocus operation; The focus control signal Foc is obtained through the following calculation formula: ; In the formula, Foc(t) represents the focus control signal generated by the camera at time t, which is used to perform autofocus processing on the local image with an average size of SizeAvg in the area where the average position PosAvg is located. GenerateFocus represents the local autofocus function; The average position PosAvg is obtained through the following calculation formula: ; In the formula, T represents the time window, and Pos(t) represents the center point coordinates of the target position at time t; The average size SizeAvg is obtained through the following calculation formula: ; In the formula, Size(t) represents the target size at time t.
[0013] Preferably, the S6 includes S61; S61. After performing autofocus, collect the target blur feature value Blv in the target area within the time window T again, obtain the new post-execution blur vector sequence set Ftr4, and compare it frame by frame with the previously collected blur vector sequence set Ftr2. Accumulate the value difference between the post-execution blur vector sequence set Ftr4 and the blur vector sequence set Ftr2 to obtain the focus effect value Qev, and then compare it with the preset focus effect improvement evaluation threshold Qth to judge the improvement degree of the image quality after focusing.
[0014] The specific form of the post-execution blur vector sequence set Ftr4 is the post-execution blur vector sequence set Ftr4 = {Blv(t), t ∈ T}, where T represents the time window; The improvement degree of the image quality after focusing is judged as follows: When the focusing effect value Qev ≥ the focusing effect improvement evaluation threshold Qth, it indicates that the focusing behavior is effective and the image quality is improved; When the focusing effect value Qev < the focusing effect improvement evaluation threshold Qth, it indicates that the focusing behavior is ineffective and the image quality is not improved, and the local autofocus is performed again.
[0015] A control system for a surveillance camera, comprising a captured image detection module, an image feature extraction module, an image defocus judgment module, a defocus evaluation module, a focusing decision module, and an iterative optimization module; The captured image detection module detects and tracks the key target area in the surveillance image through an image recognition and tracking algorithm, and obtains the spatial feature set FTR1 of the target; The image feature extraction module extracts the target blur feature value Blv of a fixed area in each frame of image based on the spatial feature set FTR1, and obtains the blur vector sequence set Ftr2; The image defocus judgment module traverses and compares with the preset blur evaluation threshold Thf and the blur vector sequence set Ftr2 to construct a defocus frame judgment mark set Ftr3; The defocus evaluation module accumulatively statistics the defocus frame judgment mark set Ftr3, and calculates the target defocus rate Tbd within a fixed time window T; The focusing decision module compares the obtained target defocus rate Tbd with the preset defocus threshold Thd, obtains the average position PosAvg and the average size SizeAvg of the target in the spatial feature set FTR1 according to the comparison result, calculates the focusing control signal Foc, and applies the focusing control signal Foc to perform local autofocus on the target; After the iterative optimization module performs autofocus, it collects the target blur feature value Blv of the target area again, obtains the new post-execution blur vector sequence set Ftr4, then compares it with the previously collected blur vector sequence set Ftr2, calculates the focusing effect value Qev, and judges the improvement degree of the image quality after focusing.
[0016] The present invention provides a control method and system for a surveillance camera, having the following beneficial effects: (1) By calculating the target defocus rate Tbd within the time window T, the blurry state no longer depends on instantaneous image judgment but has statistical trend characteristics, effectively alleviating the false triggering problem caused by accidental factors such as temporary interference and occlusion. Further, comparing the target defocus rate Tbd with the defocus rate judgment threshold Thd to generate the focus control signal Foc, achieving automatic focusing on the key target area, avoiding the resource waste and focusing delay caused by the fixed periodic focusing of the traditional system, re-acquiring the blurry features of the focused image, constructing a new set of blurry vector sequences Ftr4, and comparing it with the set of blurry vector sequences Ftr2 to calculate the focus effect value Qev, realizing the quality closed-loop evaluation and optimization adjustment of the focusing behavior, and overcoming the problem of lack of feedback verification mechanism in the focusing control of the existing technology.
[0017] (2) Based on the obtained spatial feature set FTR1, perform fusion calculation to obtain the target blurry feature value Blv(t) at time t, and integrate the blurry feature values of all frames within the time window T to construct a set of blurry vector sequences Ftr2. This processing method not only fully characterizes the essential features of image quality through different-dimensional blurriness measurement means but also eliminates the problem of dimensional inconsistency between different blurry features through normalization and weighting processing. Compared with the existing technology that relies on a single blurriness algorithm or the average blurry index of the entire image, by synergistically fusing the image gradient variance Grd(t), Laplace response variance Lap(t), and high-frequency energy ratio Hfr(t), the expression accuracy and robustness of blurriness extraction are significantly improved, making the blurriness calculation result Blv(t) more discriminative and locally sensitive, providing more reliable and fine-grained input data support for subsequent defocus judgment and focusing decision-making.
[0018] (3) Generate the focus control signal Foc(t) to drive the camera to perform local automatic focusing on the corresponding area. After the focusing is completed, re-acquire the new blurry feature value Blv(t) within the time window T, construct the post-execution set of blurry vector sequences Ftr4, and compare it frame by frame with the set of blurry vector sequences Ftr2 in the previous stage to calculate the focus effect value Qev, which is used to quantify the improvement amplitude of image sharpness to ensure that the image quality is truly improved, constructing a focus quality closed-loop control logic with the focus effect value Qev as the final verification standard. Compared with the existing technology that only triggers automatic focusing based on blurry determination but lacks result verification feedback, this method not only realizes the evaluability of the focusing behavior but also endows the system with the adaptive iterative ability for ineffective focusing situations, effectively avoiding problems such as repeated ineffective focusing, resource waste, and unverifiable improvement of image quality, and improving the intelligent level and reliability of camera imaging quality control. Description of the Drawings
[0019] Figure 1Schematic diagram of the steps of a control method for a surveillance camera according to the present invention; Figure 2 Schematic block diagram of a control system for a surveillance camera according to the present invention; Figure 3 Schematic diagram of the change of the blurred feature value Blv(t) in the image frame before and after focusing. Specific implementation manner
[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0021] Embodiment 1 The present invention provides a control method for a surveillance camera. Please refer to Figure 1 , including the following steps: S1. Detect and track the key target area in the surveillance image through an image recognition and tracking algorithm, and obtain the spatial feature set FTR1 of the target; S2. Based on the spatial feature set FTR1, extract the target blurred feature value Blv of the fixed area in each frame of the image, and obtain the blurred vector sequence set Ftr2; S3. Traverse and compare the preset blurred evaluation threshold Thf with the blurred vector sequence set Ftr2 to construct a defocused frame judgment mark set Ftr3; S4. Cumulatively count the defocused frame judgment mark set Ftr3, and calculate the target defocus rate Tbd within a fixed time window T; S5. Compare the obtained target defocus rate Tbd with the preset defocus threshold Thd, obtain the average position PosAvg and average size SizeAvg of the target in the spatial feature set FTR1 according to the comparison result, calculate the focusing control signal Foc, and apply the focusing control signal Foc to perform local autofocus on the target; S6. After performing autofocus, collect the target blurred feature value Blv of the target area again, obtain a new blurred vector sequence set Ftr4 after execution, and then compare it with the previously collected blurred vector sequence set Ftr2 to calculate the focusing effect value Qev and judge the improvement degree of the image quality after focusing.
[0022] In this embodiment, the spatial feature set FTR1 of the key target is obtained through an image recognition and tracking algorithm, providing a stable and trackable image region basis for subsequent fuzzy feature extraction. Based on the bounding box region corresponding to the spatial feature set FTR1, the target fuzzy feature value Blv is extracted, a fuzzy vector sequence set Ftr2 is constructed, and combined with a preset fuzzy evaluation threshold Thf to determine whether the image is out of focus, forming an out-of-focus frame judgment mark set Ftr3, which overcomes the problem that the key area out-of-focus cannot be recognized in time caused by the traditional camera using the average blur of the whole image as the judgment basis. By calculating the target out-of-focus rate Tbd within the time window T, the fuzzy state no longer depends on the instantaneous image judgment, but has statistical trend characteristics, effectively alleviating the false trigger problem caused by occasional factors such as temporary interference and occlusion. Further, the target out-of-focus rate Tbd is compared with the out-of-focus rate judgment threshold Thd, and the average position PosAvg and average size SizeAvg of the target are calculated based on the spatial feature set FTR1, generating a focus control signal Foc to complete the automatic focus on the key target area, avoiding the resource waste and focus delay caused by the fixed periodic focus of the traditional system. Finally, in step S6, the fuzzy features of the focused image are re-collected, a new fuzzy vector sequence set Ftr4 is constructed, and compared with the fuzzy vector sequence set Ftr2 to calculate the focus effect value Qev, realizing the quality closed-loop evaluation and optimization adjustment of the focusing behavior, and overcoming the problem that the focus control in the prior art lacks a feedback verification mechanism. Therefore, this method can effectively improve the clarity response ability of the monitoring camera for key target imaging and enhance the determination accuracy of the system for image blur state changes and the intelligence level of focus control on the premise of keeping the hardware structure unchanged.
[0023] Embodiment 2 Specifically, S1 includes S11; S11. Real-time collect the image frame Frame(t) at time t through the camera, then use the image recognition algorithm to process each image frame Frame(t), detect the key targets in the preset key target list, and extract the spatial position and boundary information of the key targets to generate initial target spatial features. Integrate all the initial target spatial features of the image frame Frame(t) at time t to obtain the target set Ftr11; The specific form of the target set Ftr11 is the target set Ftr11 = {Bnd(t), Pos(t), Size(t)}; where Bnd(t) represents the target bounding box at time t, specifically representing the rectangular box in the coordinates in the image frame Frame(t); Pos(t) represents the coordinates of the center point of the target position at time t, specifically representing the center point of the target bounding box Bnd(t) at time t; Size(t) represents the target size at time t, specifically representing the width and height of the target bounding box Bnd(t) at time t. Among them, the list of key targets includes faces, license plates, and people.
[0024] S1 includes S12; S12. Based on the obtained target set Ftr11, perform object matching between consecutive image frames Frame(t) to achieve consistent recognition of key targets, including using DeepSORT and KCF tracking algorithms to perform consistent recognition judgment of object matching, and construct a time series tracking record of key targets, marked as the spatial feature set FTR1 of the target. The specific form of the spatial feature set FTR1 is the spatial feature set FTR1 = {Bnd(t), Pos(t), Size(t)|t ∈ T}, where T represents the time window, specifically representing the continuous record of the target bounding box Bnd(t), the coordinates of the center point of the target position Pos(t), and the target size Size(t) of the key target within the time window T.
[0025] In this embodiment, image frames Frame(t) are collected in real time through a camera, and each frame of the image is analyzed using an image recognition algorithm to accurately detect the targets in the preset list of key targets, and extract their corresponding spatial positions and boundary features to form the target set Ftr11, ensuring that the spatial information of the targets in each frame of the image is completely extracted and expressed. On this basis, using object tracking algorithms such as DeepSORT and KCF, cross-frame matching and consistent recognition of targets are achieved between consecutive image frames Frame(t), and a time series tracking record of the targets is further constructed, thereby forming the target spatial feature set FTR1, which not only maintains the coherence of the spatial structure of the key targets within the time window T, but also provides an accurate and stable regional reference for subsequent modules such as ambiguity extraction and dynamic focus control. Compared with the existing method that only performs instantaneous analysis of image frames and lacks inter-frame semantic continuity, the spatial feature set FTR1 established by this solution can significantly improve the ability to structurally understand key targets, enhance the regional positioning accuracy and response coherence of subsequent image processing algorithms from the source, and is particularly suitable for the continuous tracking and determination requirements of key targets in dynamic scenarios.
[0026] Embodiment 3 Specifically, S2 includes S21; S21. Based on the spatial feature set FTR1, in each image frame Frame(t), the target region defined by the target bounding box Bnd(t) is used as the sampling window. After normalization, the dimensional differences between different features are eliminated. The image gradient variance Grd(t), Laplace response variance Lap(t), and high-frequency energy ratio Hfr(t) of the target region image at time t are extracted. After weighted calculation, the target blur feature value Blv(t) at time t is obtained. After integrating all target regions, a fuzzy vector sequence set Ftr2 is formed; The specific form of the fuzzy vector sequence set Ftr2 is the fuzzy vector sequence set Ftr2 = {Blv(t), t ∈ T}, where T represents the time window; The target blur feature value Blv(t) is obtained through the following calculation formula: ; In the formula, A1, A2, and A3 respectively represent the preset weight values of the image gradient variance Grd(t), Laplace response variance Lap(t), and high-frequency energy ratio Hfr(t) at time t, and A1 + A2 + A3 = 1. The specific values are set by the user.
[0027] Among them, the image gradient variance Grd(t) at time t is obtained by cropping the target region image Roi(t) at time t from each image frame Frame(t) with the target bounding box Bnd(t), and then performing the Scharr operator on the target region image Roi(t) at time t, calculating the horizontal and vertical gradients of the target region image Roi(t) at time t, and then finding the variance; The Laplace response variance Lap(t) is obtained by applying the Laplace operator to the target region image Roi(t) at time t to perform second-order derivative filtering, capturing the edge sharpness in the image, and finding the variance of its response value; The high-frequency energy ratio Hfr(t) is obtained by performing the Fourier transform FFT on the target region image Roi(t) at time t, transforming the image from the spatial domain to the frequency domain, and obtaining the ratio of the energy of the high-frequency part in the Fourier spectrum to the total energy.
[0028] In this embodiment, based on the obtained spatial feature set FTR1, in each image frame Frame(t), the region defined by the target bounding box Bnd(t) is used as the sampling window to crop the target region image Roi(t). Subsequently, feature extraction processing is performed on this region image, including obtaining image gradient information using the Scharr operator and calculating its variance to form the image gradient variance Grd(t), performing second-order edge filtering using the Laplace operator and obtaining the response variance to form the Laplace response variance Lap(t), and obtaining the high-frequency part energy ratio to form the high-frequency energy ratio Hfr(t) through Fourier transform FFT. After completing the calculation of the above three sub-features, fusion calculation is performed to obtain the target blur feature value Blv(t) at time t, and the blur feature values of all frames are integrated within the time window T to construct the fuzzy vector sequence set Ftr2. This processing method not only fully characterizes the essential features of image quality through different-dimensional blur measurement means, but also eliminates the problem of dimensional inconsistency between different blur features through normalization and weighting processing. Compared with the existing technology that relies on a single blur algorithm or the average blur index of the entire image, by synergistically fusing the image gradient variance Grd(t), the Laplace response variance Lap(t), and the high-frequency energy ratio Hfr(t), the expression accuracy and robustness of blur extraction are significantly improved, making the blur calculation result Blv(t) more discriminative and sensitive to local regions, providing more reliable and fine-grained input data support for subsequent defocus judgment and focusing decision-making.
[0029] Embodiment 4 Specifically: S3 includes S31; S31. Based on the obtained fuzzy vector sequence set Ftr2, traverse and compare it with the preset fuzzy evaluation threshold Thf to determine whether the target blur feature value Blv(t) at time t in each image frame Frame(t) is in a defocus state, and mark it with a boolean value to obtain the defocus state mark Blf(t) at time t, and construct the defocus frame judgment mark set Ftr3; The specific form of the defocus frame judgment mark set Ftr3 is the defocus frame judgment mark set Ftr3 = {Blf(t)|t ∈ T}, where T represents the time window; Among them, the defocus state mark Blf(t) at time t is obtained through the following comparison method: When the target blur feature value Blv(t) at time t < the fuzzy evaluation threshold Thf, it means that there is a defocus region in the frame image Frame(t) at time t, indicating a defocus frame, and mark the defocus state mark Blf(t) at time t = 1; When the target blur eigenvalue Blv(t) of time t ≥ the blur evaluation threshold Thf, it indicates that there is no defocused area in the frame image Frame(t) at time t, representing a clear frame, and the defocus state flag Blf(t) at time t is marked as 0.
[0030] The said S4 includes S41; S41. Based on the defocus frame judgment mark set Ftr3, count the number of frames with the defocus state flag Blf(t)=1 within the time window T, and calculate the proportion of the target area in the defocus frame within the time window T to obtain the defocus rate numerical index Tbd, which reflects the change trend of the image quality of the target within the time window T and is used to judge whether to trigger the focus control; The defocus rate numerical index Tbd is obtained through the calculation formula.
[0031] In this embodiment, based on the constructed fuzzy vector sequence set Ftr2, it is compared frame by frame with the preset fuzzy evaluation threshold Thf to judge whether the target blur eigenvalue Blv(t) corresponding to time t in each frame image Frame(t) is less than the fuzzy evaluation threshold Thf, and then the defocus state flag Blf(t) is generated to construct the defocus frame judgment mark set Ftr3. On this basis, the number of frames with all defocus state flags Blf(t)=1 is cumulatively counted, and combined with the total number of frames in the time window T, the defocus rate numerical index Tbd is calculated as a fuzzy trend index within a continuous time dimension, which not only reflects the cumulative state of the target image based on single-frame blur recognition, but also overcomes the limitation that the traditional system's fuzzy judgment depends on the instantaneous frame result and is easily affected by short-time interference factors such as local occlusion and motion jitter. By using the defocus state flag Blf(t) to establish a boolean representation and performing statistical aggregation on the time window T, this method realizes the transformation of fuzzy perception from "instantaneous decision" to "trend judgment", making the triggering basis of the system's focusing control strategy more stable, temporally reasonable and anti-interference capable, and significantly improving the robustness of the overall imaging clarity judgment and the rationality of the response logic.
[0032] Embodiment 5 Please refer to Figure 1 and Figure 3 Specifically: The said S5 includes S51; S51. Compare the obtained target defocus rate Tbd with the preset defocus threshold Thd, obtain the average position PosAvg and average size SizeAvg of the target in the spatial feature set FTR1 according to the comparison result, calculate the focus control signal Foc, and apply the focus control signal Foc to perform local autofocus on the target; The comparison result is obtained through the following comparison method: When the target defocus rate Tbd ≥ defocus threshold Thd, it indicates that the comparison result triggers the execution result, and a local autofocus operation is performed, including obtaining the average position PosAvg and average size SizeAvg of the target in the spatial feature set FTR1, and calculating the focus control signal Foc; When the target defocus rate Tbd < defocus threshold Thd, it indicates that the comparison result does not trigger the execution result, and the local autofocus operation is not performed; The focus control signal Foc is obtained through the following calculation formula: ; In the formula, Foc(t) represents the focus control signal generated by the camera at time t, which is used to perform autofocus processing on the local image with an average size of SizeAvg in the area where the average position PosAvg is located. GenerateFocus represents the local autofocus function; The average position PosAvg is obtained through the following calculation formula: ; In the formula, T represents the time window, and Pos(t) represents the center point coordinates of the target position at time t; The average size SizeAvg is obtained through the following calculation formula: ; In the formula, Size(t) represents the target size at time t.
[0033] The said S6 includes S61; S61. After performing autofocus, the target blur feature value Blv in the target area within the time window T is collected again, a new post-execution blur vector sequence set Ftr4 is obtained, and then it is compared frame by frame with the previously collected blur vector sequence set Ftr2. The value difference between the post-execution blur vector sequence set Ftr4 and the blur vector sequence set Ftr2 is accumulated to obtain the focus effect value Qev, and then it is compared with the preset focus effect improvement evaluation threshold Qth to judge the improvement degree of the image quality after focusing.
[0034] The specific form of the post-execution blur vector sequence set Ftr4 is the post-execution blur vector sequence set Ftr4 = {Blv(t), t ∈ T}, where T represents the time window; The improvement degree of the image quality after focusing is judged in the following way: When the focus effect value Qev ≥ focus effect improvement evaluation threshold Qth, it indicates that the focusing behavior is effective and the image quality is improved; When the focus effect value Qev < focus effect improvement evaluation threshold Qth, it indicates that the focusing behavior is ineffective, the image quality is not improved, and the local autofocus is performed again.
[0035] Example illustration of the improvement effect of the blurred feature before and after focusing: The time window T contains 5 frames of images, numbered t = 1 to 5. The example of the target blurred feature value Blv(t) before focusing and the target blurred feature value Blv(t) after focusing for the corresponding frames is as follows: Time t: 1; Target blurred feature value Blv(t) before focusing: 0.31; Target blurred feature value Blv(t) after focusing: 0.43; Difference in time t: 0.12; Time t: 2; Target blurred feature value Blv(t) before focusing: 0.28; Target blurred feature value Blv(t) after focusing: 0.41; Difference in time t: 0.13; Time t: 3; Target blurred feature value Blv(t) before focusing: 0.30; Target blurred feature value Blv(t) after focusing: 0.42; Difference in time t: 0.12; Time t: 4; Target blurred feature value Blv(t) before focusing: 0.29; Target blurred feature value Blv(t) after focusing: 0.44; Difference in time t: 0.15; Time t: 5; Target blurred feature value Blv(t) before focusing: 0.27; Target blurred feature value Blv(t) after focusing: 0.39; Difference in time t: 0.12; According to the above differences, obtain the focusing effect value Qev: Qev = 1 / 5 * (0.12 + 0.13 + 0.12 + 0.15 + 0.12) = 0.128; The preset focusing effect improvement evaluation threshold Qth is 0.05; Obtain that the focusing effect value Qev ≥ the focusing effect improvement evaluation threshold Qth, indicating that the focusing behavior is effective and the image quality is improved.
[0036] In this embodiment, based on the target defocus rate Tbd, a logical comparison is made with the preset defocus threshold Thd. The target position center point coordinates Pos(t) and the target size Size(t) of the target within the time window T are automatically extracted from the spatial feature set FTR1. The average position PosAvg and the average size SizeAvg are calculated respectively as the basis for area control. Furthermore, a focus control signal Foc(t) is generated through the local autofocus function GenerateFocus to drive the camera to perform local autofocus operations on the corresponding area. After the focus execution is completed, a new blurred feature value Blv(t) within the time window T is collected again, and the blurred vector sequence set Ftr4 after execution is constructed. Then, a frame-by-frame comparison is made with the blurred vector sequence set Ftr2 in the previous stage, and the focus effect value Qev is calculated to quantify the improvement amplitude of the image sharpness, ensuring that the image quality is truly improved. A closed-loop control logic for focus quality with the focus effect value Qev as the final verification standard is constructed. Compared with the prior art that only triggers autofocus by fuzzy determination but lacks result verification feedback, this method not only realizes the evaluability of the focusing behavior but also endows the system with the adaptive iterative ability for ineffective focusing situations, effectively avoiding problems such as repeated ineffective focusing, resource waste, and unverifiable improvement of image quality, and improving the intelligent level and reliability of the camera imaging quality control.
[0037] Embodiment 6 A control system for a surveillance camera, please refer to Figure 2 , specifically: it includes a camera image detection module, an image feature extraction module, an image defocus judgment module, a defocus evaluation module, a focus decision-making module, and an iterative optimization module; The camera image detection module detects and tracks the key target area in the surveillance image through image recognition and tracking algorithms, and obtains the spatial feature set FTR1 of the target; The image feature extraction module extracts the target blurred feature value Blv of a fixed area in each frame of the image based on the spatial feature set FTR1, and obtains the blurred vector sequence set Ftr2; The image defocus judgment module traverses and compares with the blurred vector sequence set Ftr2 through a preset blurred evaluation threshold Thf to construct a defocus frame judgment mark set Ftr3; The defocus evaluation module accumulatively statistics the defocus frame judgment mark set Ftr3, and calculates the target defocus rate Tbd within a fixed time window T; The focus decision-making module compares the obtained target defocus rate Tbd with the preset defocus threshold Thd, obtains the average position PosAvg and the average size SizeAvg of the target in the spatial feature set FTR1 according to the comparison result, calculates the focus control signal Foc, and applies the focus control signal Foc to perform local autofocus on the target; After the iterative optimization module performs autofocus, it collects the target blur feature value Blv of the target area again, obtains a new set Ftr4 of post-execution blur vector sequences, compares it with the previously collected set Ftr2 of blur vector sequences, calculates the focus effect value Qev, and determines the improvement degree of the image quality after focusing.
[0038] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A control method for a surveillance camera, characterized in that: Including the following steps: S1. Detect and track the key target area in the surveillance image through the image recognition and tracking algorithm, and obtain the spatial feature set FTR1 of the target; S2. Based on the spatial feature set FTR1, extract the target blur feature value Blv of the fixed area in each frame of the image, and obtain the fuzzy vector sequence set Ftr2; S3. Traverse and compare with the preset fuzzy evaluation threshold Thf and the fuzzy vector sequence set Ftr2 to construct the out-of-focus frame judgment mark set Ftr3; S4. Accumulatively count the out-of-focus frame judgment mark set Ftr3, and calculate the target out-of-focus rate Tbd within the fixed time window T; S5. Compare the obtained target out-of-focus rate Tbd with the preset out-of-focus threshold Thd, obtain the average position PosAvg and average size SizeAvg of the target in the spatial feature set FTR1 according to the comparison result, calculate the focus control signal Foc, and apply the focus control signal Foc to perform local autofocus on the target; S6. After performing autofocus, collect the target blur feature value Blv of the target area again, obtain the new post-execution fuzzy vector sequence set Ftr4, and then compare with the previously collected fuzzy vector sequence set Ftr2 to calculate the focus effect value Qev and judge the improvement degree of the image quality after focusing.
2. The control method for a surveillance camera according to claim 1, characterized in that: The S1 includes S11; S11. Collect the image frame Frame(t) at time t in real time through the camera, and then use the image recognition algorithm to process each image frame Frame(t), detect the key targets in the preset key target list, and extract the spatial position and boundary information of the key targets to generate the initial target spatial features. Integrate all the initial target spatial features of the image frame Frame(t) at time t to obtain the target set Ftr11; The specific form of the target set Ftr11 is the target set Ftr11 = {Bnd(t), Pos(t), Size(t)}; where, Bnd(t) represents the target bounding box at time t, specifically representing the rectangular box in the coordinates of the image frame Frame(t); Pos(t) represents the center point coordinates of the target position at time t, specifically representing the center point of the target bounding box Bnd(t) at time t; Size(t) represents the target size at time t, specifically representing the width and height of the target bounding box Bnd(t) at time t; Among them, the key target list includes faces, license plates, and people.
3. The control method for a monitoring camera according to claim 2, characterized in that: The S1 includes S12; S12. Based on the obtained target set Ftr11, perform object matching between consecutive image frames Frame(t) to achieve consistent recognition of key targets, including using the DeepSORT and KCF tracking algorithms to perform consistent recognition judgment of object matching, and constructing the time series tracking record of key targets, marked as the spatial feature set FTR1 of the target; The specific form of the spatial feature set FTR1 is the spatial feature set FTR1 = {Bnd(t), Pos(t), Size(t) | t ∈ T}, where T represents the time window, specifically indicating the continuous recording of the target bounding box Bnd(t), the coordinates of the center point of the target position Pos(t), and the target size Size(t) of the key target within the time window T.
4. A control method for a surveillance camera according to claim 3, characterized in that: The above S2 includes S21; S21. Based on the spatial feature set FTR1, in each image frame Frame(t), the target region defined by the target bounding box Bnd(t) is used as the sampling window. After normalization, the dimensional difference between different features is eliminated, and the image gradient variance Grd(t), Laplace response variance Lap(t), and high-frequency energy ratio Hfr(t) of the target region image at time t are extracted. Through weighted calculation, the target blur feature value Blv(t) at time t is obtained, and after integrating all target regions, a fuzzy vector sequence set Ftr2 is formed; The specific form of the fuzzy vector sequence set Ftr2 is the fuzzy vector sequence set Ftr2 = {Blv(t), t ∈ T}, where T represents the time window; The target blur feature value Blv(t) is obtained through the following calculation formula: ; In the formula, A1, A2, and A3 respectively represent the preset weight values of the image gradient variance Grd(t), Laplace response variance Lap(t), and high-frequency energy ratio Hfr(t) at time t, and A1 + A2 + A3 = 1. The specific values are set by the user.
5. A control method for a surveillance camera according to claim 4, characterized in that: Among them, The image gradient variance Grd(t) at time t is obtained by cropping the target region image Roi(t) at time t from each image frame Frame(t) with the target bounding box Bnd(t), and then performing the Scharr operator on the target region image Roi(t) at time t, and calculating the variance after calculating the horizontal and vertical gradients of the target region image Roi(t) at time t; The Laplace response variance Lap(t) is obtained by applying the Laplace operator to the obtained target region image Roi(t) at time t for second-order derivative filtering, capturing the edge sharpness in the image, and calculating the variance of its response value; The high-frequency energy ratio Hfr(t) is obtained by performing the Fourier transform FFT on the obtained target region image Roi(t) at time t, transforming the image from the spatial domain to the frequency domain, and calculating the ratio of the energy of the high-frequency part in the Fourier spectrum to the total energy.
6. The control method for a surveillance camera according to claim 5, characterized in that: The above S3 includes S31; S31. Based on the obtained fuzzy vector sequence set Ftr2, it is traversed and compared with the preset fuzzy evaluation threshold Thf to determine whether the target blur feature value Blv(t) at time t in each image frame Frame(t) is in a defocus state, and it is marked with a boolean value to obtain the defocus state mark Blf(t) at time t, and a defocus frame judgment mark set Ftr3 is constructed; The specific form of the out-of-focus frame judgment mark set Ftr3 is the out-of-focus frame judgment mark set Ftr3 = {Blf(t)|t ∈ T}, where T represents the time window; Among them, the out-of-focus state mark Blf(t) at time t is obtained through the following comparison method: When the target blur feature value Blv(t) at time t < the blur evaluation threshold Thf, it means that there is an out-of-focus area in the frame image Frame(t) at time t, indicating an out-of-focus frame, and the out-of-focus state mark Blf(t) at time t is marked as 1; When the target blur feature value Blv(t) at time t ≥ the blur evaluation threshold Thf, it means that there is no out-of-focus area in the frame image Frame(t) at time t, indicating a clear frame, and the out-of-focus state mark Blf(t) at time t is marked as 0.
7. A control method for a surveillance camera according to claim 6, characterized in that: The above S4 includes S41; S41. Based on the out-of-focus frame judgment mark set Ftr3, count the number of frames with the out-of-focus state mark Blf(t) = 1 within the time window T, and calculate the proportion of the target area in the out-of-focus frame within the time window T to obtain the out-of-focus rate numerical index Tbd, which reflects the change trend of the image quality of the target within the time window T and is used to judge whether to trigger the focus control; The defocus rate numerical index Tbd is obtained by the calculation formula.
8. A control method for a surveillance camera according to claim 7, characterized in that: The above S5 includes S51; S51. Compare the obtained target out-of-focus rate Tbd with the preset out-of-focus threshold Thd, obtain the average position PosAvg and average size SizeAvg of the target in the spatial feature set FTR1 according to the comparison result, calculate the focus control signal Foc, and apply the focus control signal Foc to perform local autofocus on the target; The comparison result is obtained through the following comparison method: When the target out-of-focus rate Tbd ≥ the out-of-focus threshold Thd, it means that the result triggered by the comparison is executed, and the local autofocus operation is performed, including obtaining the average position PosAvg and average size SizeAvg of the target in the spatial feature set FTR1 and calculating the focus control signal Foc; When the target out-of-focus rate Tbd < the out-of-focus threshold Thd, it means that the result triggered by the comparison is not executed, and the local autofocus operation is not performed; The focus control signal Foc is obtained through the following calculation formula: ; In the formula, Foc(t) represents the focus control signal generated by the camera at time t, which is used to perform autofocus processing on the local image with the size of the average size SizeAvg in the area where the average position PosAvg is located, and GenerateFocus represents the local autofocus function; The average position PosAvg is obtained through the following calculation formula: ; In the formula, T represents the time window, and Pos(t) represents the central point coordinate of the target position at time t; The average size SizeAvg is obtained through the following calculation formula: ; In the formula, Size(t) represents the target size at time t.
9. A control method for a surveillance camera according to claim 8, characterized in that: The above S6 includes S61; S61. After performing autofocus, the target blur feature value Blv of the target area within the time window T is collected again to obtain a new set Ftr4 of post-execution blur vector sequences. Then, a frame-by-frame comparison is made with the previously collected set Ftr2 of blur vector sequences. The value difference between the set Ftr4 of post-execution blur vector sequences and the set Ftr2 of blur vector sequences is accumulated to obtain the focus effect value Qev. Then, a comparison is made with a preset focus effect improvement evaluation threshold Qth to determine the improvement degree of the image quality after focusing. The specific form of the set Ftr4 of post-execution blur vector sequences is the set Ftr4 = {Blv(t), t ∈ T}, where T represents the time window. The improvement degree of the image quality after focusing is determined in the following way: When the focus effect value Qev ≥ the focus effect improvement evaluation threshold Qth, it indicates that the focusing behavior is effective and the image quality is improved. When the focus effect value Qev < the focus effect improvement evaluation threshold Qth, it indicates that the focusing behavior is ineffective, the image quality is not improved, and local autofocus is performed again.
10. A control system for a surveillance camera, applied to a control method for a surveillance camera according to any one of claims 1 to 9, characterized in that: It includes a camera image detection module, an image feature extraction module, an image defocus judgment module, a defocus evaluation module, a focus decision module, and an iterative optimization module. The camera image detection module detects and tracks the key target area in the monitoring image through an image recognition and tracking algorithm to obtain the set FTR1 of target spatial features. Based on the set FTR1 of spatial features, the image feature extraction module extracts the target blur feature value Blv of a fixed area in each frame of the image to obtain the set Ftr2 of blur vector sequences. The image defocus judgment module traverses and compares with the set Ftr2 of blur vector sequences through a preset blur evaluation threshold Thf to construct a defocus frame judgment mark set Ftr3. The defocus evaluation module accumulatively statistics the defocus frame judgment mark set Ftr3 and calculates the target defocus rate Tbd within a fixed time window T. The focus decision module compares the obtained target defocus rate Tbd with a preset defocus threshold Thd, obtains the average position PosAvg and average size SizeAvg of the target in the set FTR1 of spatial features according to the comparison result, calculates the focus control signal Foc, and applies the focus control signal Foc to perform local autofocus on the target. After performing autofocus, the iterative optimization module collects the target blur feature value Blv of the target area again to obtain a new set Ftr4 of post-execution blur vector sequences. Then, it compares with the previously collected set Ftr2 of blur vector sequences, calculates the focus effect value Qev, and determines the improvement degree of the image quality after focusing.
Citation Information
Patent Citations
Automatic focusing and locating method
CN103945126A
Image focusing method and device for surveillance camera
CN104639894A
Security camera image focusing measurement method and system
CN112954315A
Photographing focusing control method based on machine vision
CN119729207A
Method, apparatus and program for image processing
US20050244077A1