Camera control method and system
By collecting and processing camera multi-dimensional perception data, detecting and repairing abnormal frames in monitoring videos, and building an intelligent early warning model and control strategy model, the problem of interruption or inconsistency of video data in the existing technology is solved, the continuity and consistency of video data is achieved, and the accuracy of analysis results is improved.
Patent Information
- Application Number
- CN202510192132.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-06
AI Technical Summary
The existing camera control methods and systems have problems that abnormal frames in monitoring videos cannot be effectively detected and repaired, resulting in interruption or inconsistency of video data, affecting the reliability of the system and the accuracy of analysis results.
By collecting multi-dimensional perceptual data of the camera, calculating the abnormality degree indicator, obtaining abnormal frames in the monitoring video data, repairing and target detection, and obtaining the video feature data set. Then, the video feature data set, camera parameter data and environment perception data were analyzed and weighted fusion to obtain the comprehensive feature data set. Build an intelligent early warning model and control strategy model, predict abnormal probability and control instructions, and track abnormal situations in real time.
The continuity and consistency of video data is achieved, the accuracy of subsequent analysis is improved, the interference caused by data loss or abnormality is reduced, and the reliability of the system and the quality of analysis results are improved.
Smart Images

Figure CN119942765A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent monitoring technology, and more specifically, to a camera control method and system. Background Art
[0002] With the acceleration of urbanization and the complication of social security, the demand for security monitoring is increasing. Whether it is public places, commercial areas, residential areas or traffic arteries, a large number of surveillance cameras need to be installed to ensure public safety. This increase in demand has promoted the continuous development and improvement of camera control methods and systems.
[0003] The patent application publication number CN113114950A discloses an IoT camera control method and control system, the method includes: obtaining a monitoring screen in a monitoring area; judging the number of moving targets in the monitoring screen; when the number of moving targets is one, controlling the first camera to rotate to monitor the moving target; when the number of moving targets is two, controlling the first camera to rotate to monitor one of the moving targets; judging whether the other moving target is in the edge area of the monitoring range of the first camera. The system includes: a monitoring screen acquisition module, a moving target number judgment module, a first control module, an edge area judgment module, a second control module, a moving direction judgment module, a third control module, a motion speed judgment module and a fourth control module. An IoT camera control method and control system of the present invention meets the effect of tracking two moving targets at the same time and reduces energy consumption.
[0004] The traditional camera control method and system have the following main problems:
[0005] Abnormal frames in surveillance videos cannot be effectively detected and repaired, resulting in interruption or inconsistency of video data; since surveillance videos are continuous, abnormal or missing frames will destroy the integrity of the video, causing blanks or errors in subsequent analysis and processing, thus affecting the reliability of the entire system; without repairing or supplementing abnormal frames, subsequent video analysis may not be able to process the complete continuous video stream, resulting in inaccurate analysis results or loss of key information; when abnormal frames are not repaired, they will be mistakenly considered as part of the normal video stream, which may cause the subsequent machine learning model to learn the wrong pattern or feature, thereby reducing the effectiveness of the model;
[0006] Different video encoding standards and display devices may lead to differences in color display. Failure to use a unified format and gamma correction will cause inconsistent display of video data on different devices, thus affecting the video fusion and analysis effects in a multi-device system. Failure to synchronize timestamps on video data may cause time misalignment of video streams from multiple sources or multiple cameras, affecting synchronous analysis across cameras. Different coded data that is not formatted in a standard way will increase the complexity of subsequent processing and storage, and may even lead to data loss or unsupported formats, affecting the stability and scalability of the entire system.
[0007] The weights of each feature data set are not reasonably allocated according to their relative importance. The fused data may not be able to fully utilize the advantages of each data set, and may even cause data distortion or information loss due to inappropriate weighting, affecting the quality of subsequent analysis results. There is no mechanism to automatically adjust the weights, and feature sets with poor data quality may still be given too high weights, affecting the accuracy of the overall data fusion and causing unstable model performance. The correlation between feature data is not fully utilized, which may cause the weighting formula to be too simple or one-sided, ignoring the dependencies and complementarities between multiple data sets.
[0008] In view of this, the present invention proposes a camera control method and system to solve the above problems. Summary of the invention
[0009] In order to overcome the above-mentioned defects of the prior art and to achieve the above-mentioned purpose, the present invention provides the following technical solution: a camera control method, comprising:
[0010] S1. Collect multi-dimensional perception data of the camera; the multi-dimensional perception data of the camera includes monitoring video data, camera parameter data and environmental perception data;
[0011] S2. Obtain abnormal frames in the surveillance video data by calculating the abnormality index; repair and detect targets in the abnormal frames in the surveillance video data to obtain a video feature data set; process the camera parameter data and the environmental perception data to obtain a parameter feature data set and an environmental feature data set;
[0012] S3, performing rank correlation analysis on the video feature dataset, the parameter feature dataset, and the environment feature dataset and fusing them to obtain a comprehensive feature dataset;
[0013] S4. Build an intelligent early warning model, which includes a generator and a discriminator. Based on the adversarial training of the generator and the discriminator, obtain a trained intelligent early warning model. Input the comprehensive feature data set into the trained intelligent early warning model to predict the abnormal probability.
[0014] S5. Compare the predicted abnormal probability with the preset abnormal probability threshold to determine whether an abnormal situation occurs; if no abnormal situation occurs, continue monitoring; if an abnormal situation occurs, trigger the intelligent early warning mechanism, automatically send an alarm and collect abnormal feature data;
[0015] S6. Build a control strategy model, input the abnormal feature data into the control strategy model, and predict the control instructions; execute the predicted control instructions through the camera to track the abnormal situation in real time.
[0016] Furthermore, the monitoring video data includes real-time video stream data and video stream metadata; the video stream metadata includes timestamp, frame rate and encoding format; the camera parameter data includes hardware parameters and configuration parameters; the hardware parameters include camera model, serial number and manufacturing date; the configuration parameters include focal length, aperture value, ISO light sensitivity, zoom parameters, rotation angle, tilt and exposure time; the environmental perception data includes light intensity, temperature parameters, humidity parameters, electromagnetic parameters and network parameters.
[0017] Furthermore, the method for acquiring the video feature data set includes:
[0018] Detect and repair abnormal frames in surveillance video data, re-encode and optimize the repaired surveillance video data, and then obtain optimized surveillance video data; perform target detection on the optimized surveillance video data and extract feature data of all targets to obtain a video feature data set; the specific steps are:
[0019] S31, converting each frame of the real-time video stream data into a grayscale image, respectively calculating the brightness and contrast of each frame; calculating the abnormality index E based on the degree of deviation of the brightness and contrast; the abnormality index E is calculated by the abnormality limit formula, and the specific mathematical formula is: Wherein, L is the brightness of each frame of real-time video stream data; is the mean brightness of the entire real-time video stream data; σ L is the standard deviation of the brightness of the entire real-time video stream data; C is the contrast of each frame of the real-time video stream data; is the mean value of the contrast of the entire real-time video stream data; σ C is the standard deviation of the contrast of the entire real-time video stream data; α is the weight coefficient of the brightness deviation; β is the weight coefficient of the contrast deviation; based on the abnormality index E, a threshold φ is set to determine the abnormal frame. If E≤φ is satisfied, the frame is determined to be a normal frame; if E>φ is satisfied, the frame is determined to be an abnormal frame;
[0020] S32, repairing the abnormal frame using a block matching method, selecting a previous frame adjacent to the abnormal frame as a reference frame; dividing the abnormal frame and the reference frame into r small blocks, each of which has a fixed size of M×N pixels; searching for the best matching block in the reference frame for each small block of the abnormal frame; obtaining the best matching block by using an absolute error and criterion, and selecting a small block with the smallest absolute error as the best matching block; calculating a motion vector from the best matching block to the corresponding small block of the abnormal frame; replacing the best matching block in the reference frame with the corresponding small block in the abnormal frame according to the motion vector, thereby repairing the abnormal frame;
[0021] S33, decoding the repaired real-time video stream data according to the encoding format in the video stream metadata, and re-encoding it into a unified standard format; optimizing the re-encoded real-time video stream data through gamma correction, and performing time series synchronization according to the timestamp in the video stream metadata, accurately matching each frame in the optimized real-time video stream data with the timestamp, and thereby obtaining optimized surveillance video data;
[0022] S34. Use the target detection algorithm to perform frame-by-frame target detection on the optimized surveillance video data, extract features of the detected targets, and obtain feature data of the targets; collect feature data of all targets, and then obtain a video feature data set.
[0023] Furthermore, the method for obtaining the comprehensive feature data set includes:
[0024] Perform rank correlation analysis on the video feature data set, parameter feature data set and environmental feature data set: sort all feature data in the video feature data set, parameter feature data set and environmental feature data set, and assign ranks; select any data from different feature data sets, and combine them into non-repeating feature data pairs; calculate the difference between the ranks of two data in the feature data pair, calculate the rank correlation coefficient based on the difference, and then obtain the rank correlation coefficient of all feature data pairs;
[0025] The video feature dataset, the parameter feature dataset and the environment feature dataset are fused by a weighted formula to obtain a comprehensive feature dataset, where the video feature dataset is denoted as P1, the parameter feature dataset is denoted as P2, and the environment feature dataset is denoted as P3;
[0026] The weighting formula is: SV = P1·λ1+P2·λ2+P3·λ3; where SV is the comprehensive feature data set; λ1 is the weight coefficient of the video feature data set; λ2 is the weight coefficient of the parameter feature data set; λ3 is the weight coefficient of the environment feature data set; the weight coefficients of the video feature data set, the parameter feature data set and the environment feature data set are obtained through the coefficient adjustment formula; the coefficient adjustment formula is: Among them, λ′ is the adjusted weight coefficient; N′ is the number of feature data pairs that appear in the data set P k The number of data in; k is the index of the feature data set; ρ is the rank correlation coefficient; P k is the kth feature data set; s k is the proportion of the kth feature data set data in all feature data pairs; (σ′) 2 is the variance of the rank correlation coefficients of all feature data pairs; G(ρ) is the sum of the absolute values of the rank correlation coefficients of all feature data pairs; K is the total number of feature data pairs; δ is the normalization factor.
[0027] Furthermore, the method for constructing an intelligent early warning model includes:
[0028] The data set is divided into a training set, a validation set, and a test set; the sample set is a subset of the data set, and each sample set includes a historical comprehensive feature data set and a corresponding abnormal probability; a generative adversarial network is used to build an intelligent early warning model, which includes a generator and a discriminator; the generator is used to generate simulated data similar to the input real data, and the discriminator is used to distinguish the simulated data from the input real data; the training set is used to train the model, initialize the parameters of the generator and the discriminator, and train the generator and the discriminator in turn and iterate continuously until the discriminator cannot distinguish the simulated data generated by the generator from the input real data and stops the iteration;
[0029] Use the validation set to evaluate the performance of the model and tune the model parameters until the model performance reaches the preset stopping condition to obtain a trained intelligent early warning model; use the test set to evaluate the performance of the model in the prediction task, input the current comprehensive feature data set into the trained intelligent early warning model to obtain the anomaly probability.
[0030] Furthermore, the method of comparing the predicted abnormal probability with a preset abnormal probability threshold to determine whether an abnormal situation occurs includes:
[0031] If the predicted abnormal probability is less than or equal to the preset abnormal rate threshold, it is determined that no abnormal situation has occurred;
[0032] If the predicted abnormal probability is greater than the preset abnormal rate threshold, an abnormal situation is determined to have occurred, triggering the intelligent early warning mechanism, automatically sending an alarm and collecting abnormal feature data.
[0033] Furthermore, the abnormal feature data includes time feature data, space feature data, type feature data, environment change data and motion trajectory data.
[0034] Furthermore, the method for constructing a control strategy model includes:
[0035] The data set is divided into a training set and a test set for model training and performance verification; the sample set is a subset of the data set, and each sample set includes historical abnormal feature data and corresponding control instructions; the control strategy model is a decision tree model, which is constructed using the CART algorithm; the input data of the model is the historical abnormal feature data, and the output label of the model is the control instruction;
[0036] The decision tree model is trained on the training data set. The decision tree divides nodes and generates regular paths based on historical abnormal feature data. Each path represents a set of feature combinations and ultimately points to a control instruction. The decision tree model is optimized through pruning methods to reduce overfitting. The accuracy of the decision tree model is verified on the test set, and the performance of the model in generating control instructions for different abnormal situations in actual scenarios is evaluated. The hyperparameters of the decision tree model are tuned according to performance feedback until the preset stopping conditions are reached, thereby obtaining a trained control strategy model. The current abnormal feature data is input into the trained control strategy model to obtain control instructions.
[0037] Furthermore, the method of tracking abnormal situations in real time by executing predicted control instructions through a camera includes: the camera executes the predicted control instructions on hardware, and adjusts the angle and focal length of the camera according to the relative position and distance changes of the abnormal situation; adjusts the abnormal feature data according to the real-time changes of the abnormal situation and tracking feedback information, and updates the control instructions of the camera in real time.
[0038] A camera control system, comprising:
[0039] The data acquisition module is used to collect multi-dimensional perception data of the camera; the multi-dimensional perception data of the camera includes monitoring video data, camera parameter data and environmental perception data;
[0040] The data processing module obtains abnormal frames in the monitoring video data by calculating the abnormality index; repairs and detects targets in the abnormal frames in the monitoring video data to obtain a video feature data set; processes the camera parameter data and environmental perception data to obtain a parameter feature data set and an environmental feature data set;
[0041] The data fusion module performs rank correlation analysis and fusion on the video feature data set, parameter feature data set and environmental feature data set to obtain a comprehensive feature data set;
[0042] The intelligent detection module is used to build an intelligent early warning model, which includes a generator and a discriminator. Based on the adversarial training of the generator and the discriminator, a trained intelligent early warning model is obtained. The comprehensive feature data set is input into the trained intelligent early warning model to predict the abnormal probability.
[0043] The risk warning module compares the predicted abnormal probability with the preset abnormal probability threshold to determine whether an abnormal situation has occurred; if no abnormal situation has occurred, monitoring will continue; if an abnormal situation has occurred, the intelligent early warning mechanism will be triggered, an alarm will be automatically sent, and abnormal feature data will be collected;
[0044] The automatic control module is used to build a control strategy model, input abnormal feature data into the control strategy model, and predict control instructions; the predicted control instructions are executed through the camera to track abnormal situations in real time.
[0045] The technical effects and advantages of a camera control method and system of the present invention are as follows:
[0046] By detecting and repairing abnormal frames in surveillance videos, the continuity and consistency of video data can be ensured, avoiding interference caused by missing or abnormal data; the repaired frames can effectively reduce noise, thereby improving the accuracy of subsequent analysis; by calculating the degree of deviation of brightness and contrast of each frame, the difference between the video frame and the normal video stream can be quantified, thereby achieving accurate abnormal frame detection; the abnormality index comprehensively considers the degree of deviation of brightness and contrast, and makes judgments by setting reasonable thresholds, which helps to reduce false positives and false negatives, avoid over-detection of normal frames as abnormal frames, or miss real abnormal frames; by adjusting the weight coefficients of brightness deviation and contrast deviation, the definition of abnormality can be flexibly optimized according to different application scenarios and video content;
[0047] When using the block matching method to repair abnormal frames, the data of adjacent frames are matched for filling, which can maintain the temporal continuity and spatial consistency of the video data and reduce the error accumulation caused by frame damage or loss; through gamma correction processing and re-encoding in standard format, the video data can be made more balanced in brightness and contrast, reducing the color deviation caused by different video encoding or equipment differences, thereby improving the consistency and applicability of the video data; through the time stamp to synchronize the time information of each frame, the time series of the video data is more accurate, avoiding analysis errors caused by time confusion or data delay;
[0048] By fusing different feature data sets through weighted formulas, different weight coefficients can be assigned according to the relative importance of different feature data sets. This weighted fusion method can ensure that in the comprehensive feature data set, data from different sources (video, camera parameters, environmental data) can be combined in the most appropriate proportion, thereby improving the effect of data fusion; the weight coefficient of each feature data set is automatically calculated and adjusted through the coefficient adjustment formula, and the weight in the fusion process can be dynamically adjusted according to the correlation and variance between different data sets; this dynamic adjustment mechanism makes feature fusion more intelligent and can automatically adapt to the characteristics of different scenarios and data sets. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 A schematic flow chart of a camera control method of the present invention;
[0050] Figure 2 The present invention is a schematic diagram of the structure of a camera control system. DETAILED DESCRIPTION
[0051] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0052] Example 1
[0053] See also Figure 1 As shown, this embodiment provides a camera control method, including:
[0054] S1. Collect multi-dimensional perception data of the camera; the multi-dimensional perception data of the camera includes monitoring video data, camera parameter data and environmental perception data;
[0055] S2. Obtain abnormal frames in the surveillance video data by calculating the abnormality index; repair and detect targets in the abnormal frames in the surveillance video data to obtain a video feature data set; process the camera parameter data and the environmental perception data to obtain a parameter feature data set and an environmental feature data set;
[0056] S3, performing rank correlation analysis on the video feature dataset, the parameter feature dataset, and the environment feature dataset and fusing them to obtain a comprehensive feature dataset;
[0057] S4. Build an intelligent early warning model, which includes a generator and a discriminator. Based on the adversarial training of the generator and the discriminator, obtain a trained intelligent early warning model. Input the comprehensive feature data set into the trained intelligent early warning model to predict the abnormal probability.
[0058] S5. Compare the predicted abnormal probability with the preset abnormal probability threshold to determine whether an abnormal situation occurs; if no abnormal situation occurs, continue monitoring; if an abnormal situation occurs, trigger the intelligent early warning mechanism, automatically send an alarm and collect abnormal feature data;
[0059] S6. Build a control strategy model, input the abnormal feature data into the control strategy model, and predict the control instructions; execute the predicted control instructions through the camera to track the abnormal situation in real time.
[0060] Surveillance video data includes real-time video stream data and video stream metadata; video stream metadata includes timestamp, frame rate and encoding format; camera parameter data includes hardware parameters and configuration parameters; hardware parameters include camera model, serial number and manufacturing date; configuration parameters include focal length, aperture value, ISO light sensitivity, zoom parameters, rotation angle, tilt and exposure time; environmental perception data includes light intensity, temperature parameters, humidity parameters, electromagnetic parameters and network parameters.
[0061] The method for obtaining the video feature dataset includes:
[0062] Detect and repair abnormal frames in surveillance video data, re-encode and optimize the repaired surveillance video data, and then obtain optimized surveillance video data; perform target detection on the optimized surveillance video data and extract feature data of all targets to obtain a video feature data set; the specific steps are:
[0063] S31, converting each frame of the real-time video stream data into a grayscale image, respectively calculating the brightness and contrast of each frame; calculating the abnormality index E based on the degree of deviation of the brightness and contrast; the abnormality index E is calculated by the abnormality limit formula, and the specific mathematical formula is: Wherein, L is the brightness of each frame of real-time video stream data; is the mean brightness of the entire real-time video stream data; σ L is the standard deviation of the brightness of the entire real-time video stream data; C is the contrast of each frame of the real-time video stream data; is the mean value of the contrast of the entire real-time video stream data; σ C is the standard deviation of the contrast of the entire real-time video stream data; α is the weight coefficient of the brightness deviation; β is the weight coefficient of the contrast deviation; based on the abnormality index E, a threshold φ is set to determine the abnormal frame. If E≤φ is satisfied, the frame is determined to be a normal frame; if E>φ is satisfied, the frame is determined to be an abnormal frame;
[0064] For example: suppose you are processing a live video stream where each frame is converted to a grayscale image. To simplify the example, consider the following hypothetical data:
[0065] The brightness of the current frame is L=50, which is the average brightness of the entire real-time video stream data. The standard deviation of brightness of the entire video stream is σ L =5, the contrast of each frame of real-time video stream data C=30, the average contrast of the entire real-time video stream data The standard deviation of the contrast of the entire real-time video stream data σ C =3, the weight coefficient of brightness deviation α = 0.7, the weight coefficient of contrast deviation β = 0.3, and the threshold of abnormality index φ = 1.5;
[0066] Based on these data, the abnormality index can be calculated as:
[0067] Compare the calculated abnormality index E with the threshold φ:
[0068] If E≤φ, the frame is determined to be a normal frame; if E>φ, the frame is determined to be an abnormal frame; in this example, E=1.201 is less than the threshold φ=1.5, so the current frame is determined to be a normal frame.
[0069] S32, repairing the abnormal frame using a block matching method, selecting a previous frame adjacent to the abnormal frame as a reference frame; dividing the abnormal frame and the reference frame into r small blocks, each of which has a fixed size of M×N pixels; searching for the best matching block in the reference frame for each small block of the abnormal frame; obtaining the best matching block by using an absolute error and criterion, and selecting a small block with the smallest absolute error as the best matching block; calculating a motion vector from the best matching block to the corresponding small block of the abnormal frame; replacing the best matching block in the reference frame with the corresponding small block in the abnormal frame according to the motion vector, thereby repairing the abnormal frame;
[0070] For example: suppose we have two consecutive frames B1 and B2, where we want to estimate the position of a 4×4 pixel block in frame B1 in frame B2 to determine the motion of the block;
[0071] A 4×4 pixel block in frame B1: Search area in frame B2: Select a 4×4 pixel block at random in frame B1, define a search area in frame B2, and calculate the sum of absolute errors between each 4×4 candidate block in the search area and the reference block; take one of the candidate blocks as Taking the calculation as an example, the absolute error sum is: |100-100|+|102-102|+...+|120-110|=39; the block with the smallest absolute error sum is selected as the best matching block;
[0072] S33, decoding the repaired real-time video stream data according to the encoding format in the video stream metadata, and re-encoding it into a unified standard format; optimizing the re-encoded real-time video stream data through gamma correction, and performing time series synchronization according to the timestamp in the video stream metadata, accurately matching each frame in the optimized real-time video stream data with the timestamp, and thereby obtaining optimized surveillance video data;
[0073] S34. Use the target detection algorithm to perform frame-by-frame target detection on the optimized surveillance video data, extract features of the detected targets, and obtain feature data of the targets; collect feature data of all targets, and then obtain a video feature data set.
[0074] The method for processing camera parameter data and environmental perception data is as follows: performing data cleaning on the acquired camera parameter data and environmental perception data, filling in missing values and removing outliers; performing standard deviation normalization on the camera parameter data and environmental perception data after data cleaning, converting them into a standard normal distribution, and obtaining a parameter feature data set and an environmental feature data set.
[0075] The method for obtaining the comprehensive feature data set includes:
[0076] Perform rank correlation analysis on the video feature data set, parameter feature data set and environmental feature data set: sort all feature data in the video feature data set, parameter feature data set and environmental feature data set, and assign ranks; select any data from different feature data sets, and combine them into non-repeating feature data pairs; calculate the difference between the ranks of two data in the feature data pair, calculate the rank correlation coefficient based on the difference, and then obtain the rank correlation coefficient of all feature data pairs;
[0077] The video feature dataset, the parameter feature dataset and the environment feature dataset are fused by a weighted formula to obtain a comprehensive feature dataset, where the video feature dataset is denoted as P1, the parameter feature dataset is denoted as P2, and the environment feature dataset is denoted as P3;
[0078] The weighting formula is: SV = P1·λ1+P2·λ2+P3·λ3; where SV is the comprehensive feature data set; λ1 is the weight coefficient of the video feature data set; λ2 is the weight coefficient of the parameter feature data set; λ3 is the weight coefficient of the environment feature data set; the weight coefficients of the video feature data set, the parameter feature data set and the environment feature data set are obtained through the coefficient adjustment formula; the coefficient adjustment formula is: Among them, λ′ is the adjusted weight coefficient; N′ is the number of feature data pairs that appear in the data set P k The number of data in; k is the index of the feature data set; ρ is the rank correlation coefficient; P k is the kth feature data set; s k is the proportion of the kth feature data set data in all feature data pairs; (σ′) 2 is the variance of the rank correlation coefficients of all feature data pairs; G(ρ) is the sum of the absolute values of the rank correlation coefficients of all feature data pairs; K is the total number of feature data pairs; δ is the normalization factor.
[0079] For example: the video feature dataset, parameter feature dataset and environment feature dataset constitute 100 feature data pairs. For each feature dataset, it is assumed that the following has been calculated: the average absolute value of the rank correlation coefficients of all feature data pairs containing video feature dataset data is 0.8; the average absolute value of the rank correlation coefficients of all feature data pairs containing parameter feature dataset data is 0.6; the average absolute value of the rank correlation coefficients of all feature data pairs containing environment feature dataset data is 0.7; the variance of the rank correlation coefficients of all feature data pairs is 0.05; the sum of the absolute values of the rank correlation coefficients of all feature data pairs is 200; it is assumed that the data of all feature datasets appear 33 times in the feature data pairs; the normalization factor is 2;
[0080] According to the coefficient adjustment formula, the weight coefficient calculated for the video feature dataset That is, the weight coefficient of the video feature dataset is 0.348.
[0081] Methods for building intelligent early warning models include:
[0082] The data set is divided into a training set, a validation set, and a test set; the sample set is a subset of the data set, and each sample set includes a historical comprehensive feature data set and a corresponding abnormal probability; a generative adversarial network is used to build an intelligent early warning model, which includes a generator and a discriminator; the generator is used to generate simulated data similar to the input real data, and the discriminator is used to distinguish the simulated data from the input real data; the training set is used to train the model, initialize the parameters of the generator and the discriminator, and train the generator and the discriminator in turn and iterate continuously until the discriminator cannot distinguish the simulated data generated by the generator from the input real data and stops the iteration;
[0083] Use the validation set to evaluate the performance of the model and tune the model parameters until the model performance reaches the preset stopping condition to obtain a trained intelligent early warning model; use the test set to evaluate the performance of the model in the prediction task, input the current comprehensive feature data set into the trained intelligent early warning model to obtain the anomaly probability.
[0084] The method of comparing the predicted abnormal probability with the preset abnormal probability threshold to determine whether an abnormal situation occurs includes:
[0085] If the predicted abnormal probability is less than or equal to the preset abnormal rate threshold, it is determined that no abnormal situation has occurred;
[0086] If the predicted abnormal probability is greater than the preset abnormal rate threshold, an abnormal situation is determined to have occurred, triggering the intelligent early warning mechanism, automatically sending an alarm and collecting abnormal feature data.
[0087] Abnormal feature data includes time feature data, space feature data, type feature data, environmental change data and motion trajectory data.
[0088] Methods for building control strategy models include:
[0089] The data set is divided into a training set and a test set for model training and performance verification; the sample set is a subset of the data set, and each sample set includes historical abnormal feature data and corresponding control instructions; the control strategy model is a decision tree model, which is constructed using the CART algorithm; the input data of the model is the historical abnormal feature data, and the output label of the model is the control instruction;
[0090] The decision tree model is trained on the training data set. The decision tree divides nodes and generates regular paths based on historical abnormal feature data. Each path represents a set of feature combinations and ultimately points to a control instruction. The decision tree model is optimized through pruning methods to reduce overfitting. The accuracy of the decision tree model is verified on the test set, and the performance of the model in generating control instructions for different abnormal situations in actual scenarios is evaluated. The hyperparameters of the decision tree model are tuned according to performance feedback until the preset stopping conditions are reached, thereby obtaining a trained control strategy model. The current abnormal feature data is input into the trained control strategy model to obtain control instructions.
[0091] The method for real-time tracking of abnormal situations by executing predicted control instructions through a camera includes: the camera executes the predicted control instructions on hardware, and adjusts the angle and focal length of the camera according to the relative position and distance changes of the abnormal situation; adjusts the abnormal feature data according to the real-time changes of the abnormal situation and tracking feedback information, and updates the control instructions of the camera in real time.
[0092] The preset abnormal probability threshold is set by the staff. Different abnormal probabilities are collected through the camera intelligent control terminal, and the average value of multiple abnormal probabilities is taken as the preset abnormal probability threshold; the abnormal degree index threshold is set in the same way.
[0093] This embodiment can ensure the continuity and consistency of video data by detecting and repairing abnormal frames in surveillance videos, and avoid interference caused by missing or abnormal data; the repaired frames can effectively reduce noise, thereby improving the accuracy of subsequent analysis; by calculating the degree of deviation of brightness and contrast of each frame, the difference between the video frame and the normal video stream can be quantified, thereby achieving accurate abnormal frame detection; the abnormality index comprehensively considers the degree of deviation of brightness and contrast, and makes a judgment by setting a reasonable threshold, which helps to reduce false positives and false negatives, avoid over-detection of normal frames as abnormal frames, or omission of real abnormal frames; by adjusting the weight coefficients of brightness deviation and contrast deviation, the definition of abnormality can be flexibly optimized according to different application scenarios and video content;
[0094] When using the block matching method to repair abnormal frames, the data of adjacent frames are matched for filling, which can maintain the temporal continuity and spatial consistency of the video data and reduce the error accumulation caused by frame damage or loss; through gamma correction processing and re-encoding in standard format, the video data can be made more balanced in brightness and contrast, reducing the color deviation caused by different video encoding or equipment differences, thereby improving the consistency and applicability of the video data; through the time stamp to synchronize the time information of each frame, the time series of the video data is more accurate, avoiding analysis errors caused by time confusion or data delay;
[0095] By fusing different feature data sets through weighted formulas, different weight coefficients can be assigned according to the relative importance of different feature data sets. This weighted fusion method can ensure that in the comprehensive feature data set, data from different sources (video, camera parameters, environmental data) can be combined in the most appropriate proportion, thereby improving the effect of data fusion; the weight coefficient of each feature data set is automatically calculated and adjusted through the coefficient adjustment formula, and the weight in the fusion process can be dynamically adjusted according to the correlation and variance between different data sets; this dynamic adjustment mechanism makes feature fusion more intelligent and can automatically adapt to the characteristics of different scenarios and data sets.
[0096] Example 2
[0097] See also Figure 2 As shown, this embodiment provides a camera control system, including:
[0098] The data acquisition module is used to collect multi-dimensional perception data of the camera; the multi-dimensional perception data of the camera includes monitoring video data, camera parameter data and environmental perception data;
[0099] The data processing module obtains abnormal frames in the monitoring video data by calculating the abnormality index; repairs and detects targets in the abnormal frames in the monitoring video data to obtain a video feature data set; processes the camera parameter data and environmental perception data to obtain a parameter feature data set and an environmental feature data set;
[0100] The data fusion module performs rank correlation analysis and fusion on the video feature data set, parameter feature data set and environmental feature data set to obtain a comprehensive feature data set;
[0101] The intelligent detection module is used to build an intelligent early warning model, which includes a generator and a discriminator. Based on the adversarial training of the generator and the discriminator, a trained intelligent early warning model is obtained. The comprehensive feature data set is input into the trained intelligent early warning model to predict the abnormal probability.
[0102] The risk warning module compares the predicted abnormal probability with the preset abnormal probability threshold to determine whether an abnormal situation has occurred; if no abnormal situation has occurred, monitoring will continue; if an abnormal situation has occurred, the intelligent early warning mechanism will be triggered, an alarm will be automatically sent, and abnormal feature data will be collected;
[0103] The automatic control module is used to build a control strategy model, input abnormal feature data into the control strategy model, and predict control instructions; the predicted control instructions are executed through the camera to track abnormal situations in real time.
[0104] Since the electronic device introduced in this embodiment is an electronic device used to implement a camera control method and system in the embodiment of this application, based on the camera control method and system introduced in the embodiment of this application, a person skilled in the art can understand the specific implementation of the electronic device of this embodiment and its various variations, so how the electronic device implements the method in the embodiment of this application is not described in detail here. As long as a person skilled in the art implements the electronic device used in a camera control method and system in the embodiment of this application, it belongs to the scope of protection of this application.
[0105] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters and thresholds in the formula are set by technicians in this field according to actual conditions.
[0106] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technical users in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. A camera control method, characterized in that: include: S1, collect multi-dimensional perception data of the camera; The camera's multi-dimensional perception data includes surveillance video data, camera parameter data, and environmental perception data; S2. Obtain abnormal frames in the surveillance video data by calculating the abnormality degree index; Repair abnormal frames and detect targets in surveillance video data to obtain video feature data sets; Processing the camera parameter data and the environment perception data to obtain a parameter feature data set and an environment feature data set; S3, performing rank correlation analysis on the video feature dataset, the parameter feature dataset, and the environment feature dataset and fusing them to obtain a comprehensive feature dataset; S4. Build an intelligent early warning model, which includes a generator and a discriminator. Based on the adversarial training of the generator and the discriminator, obtain a trained intelligent early warning model. Input the comprehensive feature data set into the trained intelligent early warning model to predict the abnormal probability. S5. Compare the predicted abnormal probability with the preset abnormal probability threshold to determine whether an abnormal situation occurs; if no abnormal situation occurs, continue monitoring; if an abnormal situation occurs, trigger the intelligent early warning mechanism, automatically send an alarm and collect abnormal feature data; S6. Build a control strategy model, input the abnormal feature data into the control strategy model, and predict the control instructions; execute the predicted control instructions through the camera to track the abnormal situation in real time.
2. The camera control method according to claim 1, characterized in that: The monitoring video data includes real-time video stream data and video stream metadata; Video stream metadata includes timestamp, frame rate and encoding format; camera parameter data includes hardware parameters and configuration parameters; hardware parameters include camera model, serial number and manufacturing date; configuration parameters include focal length, aperture value, ISO light sensitivity, zoom parameters, rotation angle, tilt and exposure time; environmental perception data includes light intensity, temperature parameters, humidity parameters, electromagnetic parameters and network parameters.
3. The camera control method according to claim 2, characterized in that: The method for obtaining the video feature data set includes: Detect and repair abnormal frames in surveillance video data, re-encode and optimize the repaired surveillance video data, and then obtain optimized surveillance video data; perform target detection on the optimized surveillance video data and extract feature data of all targets to obtain a video feature data set; the specific steps are: S31, converting each frame of the real-time video stream data into a grayscale image, respectively calculating the brightness and contrast of each frame; calculating the abnormality index E based on the degree of deviation of the brightness and contrast; the abnormality index E is calculated by the abnormality limit formula, and the specific mathematical formula is: Wherein, L is the brightness of each frame of real-time video stream data; is the mean brightness of the entire real-time video stream data; σ L is the standard deviation of the brightness of the entire real-time video stream data; C is the contrast of each frame of the real-time video stream data; is the mean value of the contrast of the entire real-time video stream data; σ C is the standard deviation of the contrast of the entire real-time video stream data; α is the weight coefficient of the brightness deviation; β is the weight coefficient of the contrast deviation; based on the abnormality index E, a threshold φ is set to determine the abnormal frame. If E≤φ is satisfied, the frame is determined to be a normal frame; if E>φ is satisfied, the frame is determined to be an abnormal frame; S32, repairing the abnormal frame using a block matching method, selecting a previous frame adjacent to the abnormal frame as a reference frame; dividing the abnormal frame and the reference frame into r small blocks, each of which has a fixed size of M×N pixels; searching for the best matching block in the reference frame for each small block of the abnormal frame; obtaining the best matching block by using an absolute error and criterion, and selecting a small block with the smallest absolute error as the best matching block; calculating a motion vector from the best matching block to the corresponding small block of the abnormal frame; replacing the best matching block in the reference frame with the corresponding small block in the abnormal frame according to the motion vector, thereby repairing the abnormal frame; S33, decoding the repaired real-time video stream data according to the encoding format in the video stream metadata, and re-encoding it into a unified standard format; optimizing the re-encoded real-time video stream data through gamma correction, and performing time series synchronization according to the timestamp in the video stream metadata, accurately matching each frame in the optimized real-time video stream data with the timestamp, and thereby obtaining optimized surveillance video data; S34. Use the target detection algorithm to perform frame-by-frame target detection on the optimized surveillance video data, extract features of the detected targets, and obtain feature data of the targets; collect feature data of all targets, and then obtain a video feature data set.
4. The camera control method according to claim 3, characterized in that: The method for obtaining the comprehensive feature data set includes: Perform rank correlation analysis on the video feature data set, parameter feature data set and environmental feature data set: sort all feature data in the video feature data set, parameter feature data set and environmental feature data set, and assign ranks; select any data from different feature data sets, and combine them into non-repeating feature data pairs; calculate the difference between the ranks of two data in the feature data pair, calculate the rank correlation coefficient based on the difference, and then obtain the rank correlation coefficient of all feature data pairs; The video feature dataset, the parameter feature dataset and the environment feature dataset are fused by a weighted formula to obtain a comprehensive feature dataset, where the video feature dataset is denoted as P1, the parameter feature dataset is denoted as P2, and the environment feature dataset is denoted as P3; The weighting formula is: SV = P1·λ1+P2·λ2+P3·λ3; where SV is the comprehensive feature data set; λ1 is the weight coefficient of the video feature data set; λ2 is the weight coefficient of the parameter feature data set; λ3 is the weight coefficient of the environment feature data set; the weight coefficients of the video feature data set, the parameter feature data set and the environment feature data set are obtained through the coefficient adjustment formula; the coefficient adjustment formula is: Among them, λ′ is the adjusted weight coefficient; N′ is the number of feature data pairs that appear in the data set P k The number of data in; k is the index of the feature data set; ρ is the rank correlation coefficient; P k is the kth feature data set; s k is the proportion of the kth feature data set data in all feature data pairs; (σ′) 2 is the variance of the rank correlation coefficients of all feature data pairs; G(ρ) is the sum of the absolute values of the rank correlation coefficients of all feature data pairs; K is the total number of feature data pairs; δ is the normalization factor.
5. The camera control method according to claim 4, characterized in that: The method for constructing an intelligent early warning model comprises: The data set is divided into a training set, a validation set, and a test set; the sample set is a subset of the data set, and each sample set includes a historical comprehensive feature data set and a corresponding abnormal probability; a generative adversarial network is used to build an intelligent early warning model, which includes a generator and a discriminator; the generator is used to generate simulated data similar to the input real data, and the discriminator is used to distinguish the simulated data from the input real data; the training set is used to train the model, initialize the parameters of the generator and the discriminator, and train the generator and the discriminator in turn and iterate continuously until the discriminator cannot distinguish the simulated data generated by the generator from the input real data and stops the iteration; Use the validation set to evaluate the performance of the model and tune the model parameters until the model performance reaches the preset stopping condition to obtain a trained intelligent early warning model; use the test set to evaluate the performance of the model in the prediction task, input the current comprehensive feature data set into the trained intelligent early warning model to obtain the anomaly probability.
6. The camera control method according to claim 5, characterized in that: The method of comparing the predicted abnormal probability with a preset abnormal probability threshold to determine whether an abnormal situation occurs includes: If the predicted abnormal probability is less than or equal to the preset abnormal rate threshold, it is determined that no abnormal situation has occurred; If the predicted abnormal probability is greater than the preset abnormal rate threshold, an abnormal situation is determined to have occurred, triggering the intelligent early warning mechanism, automatically sending an alarm and collecting abnormal feature data.
7. The camera control method according to claim 6, characterized in that: The abnormal feature data includes time feature data, space feature data, type feature data, environment change data and motion trajectory data.
8. The camera control method according to claim 7, characterized in that: The method for constructing a control strategy model comprises: The data set is divided into a training set and a test set for model training and performance verification; the sample set is a subset of the data set, and each sample set includes historical abnormal feature data and corresponding control instructions; the control strategy model is a decision tree model, which is constructed using the CART algorithm; the input data of the model is the historical abnormal feature data, and the output label of the model is the control instruction; The decision tree model is trained on the training data set. The decision tree divides nodes and generates regular paths based on historical abnormal feature data. Each path represents a set of feature combinations and ultimately points to a control instruction. The decision tree model is optimized through pruning methods to reduce overfitting. The accuracy of the decision tree model is verified on the test set, and the performance of the model in generating control instructions for different abnormal situations in actual scenarios is evaluated. The hyperparameters of the decision tree model are tuned according to performance feedback until the preset stopping conditions are reached, thereby obtaining a trained control strategy model. The current abnormal feature data is input into the trained control strategy model to obtain control instructions.
9. The camera control method according to claim 8, characterized in that: The method for real-time tracking of abnormal situations by executing predicted control instructions through a camera includes: the camera executes the predicted control instructions on hardware, and adjusts the angle and focal length of the camera according to the relative position and distance changes of the abnormal situation; adjusts the abnormal feature data according to the real-time changes of the abnormal situation and tracking feedback information, and updates the control instructions of the camera in real time.
10. A camera control system, used to implement the camera control method according to any one of claims 1 to 9, characterized in that: include: Data acquisition module, used to collect multi-dimensional perception data of the camera; The camera's multi-dimensional perception data includes surveillance video data, camera parameter data, and environmental perception data; The data processing module obtains abnormal frames in the surveillance video data by calculating the abnormality degree index; Repair abnormal frames and detect targets in surveillance video data to obtain video feature data sets; Processing the camera parameter data and the environment perception data to obtain a parameter feature data set and an environment feature data set; The data fusion module performs rank correlation analysis and fusion on the video feature data set, parameter feature data set and environmental feature data set to obtain a comprehensive feature data set; The intelligent detection module is used to build an intelligent early warning model, which includes a generator and a discriminator. Based on the adversarial training of the generator and the discriminator, a trained intelligent early warning model is obtained. The comprehensive feature data set is input into the trained intelligent early warning model to predict the abnormal probability. The risk warning module compares the predicted abnormal probability with the preset abnormal probability threshold to determine whether an abnormal situation has occurred; if no abnormal situation has occurred, monitoring will continue; If an abnormal situation occurs, the intelligent early warning mechanism will be triggered, automatically sending an alarm and collecting abnormal feature data; The automatic control module is used to build a control strategy model, input abnormal feature data into the control strategy model, and predict control instructions; the predicted control instructions are executed through the camera to track abnormal situations in real time.
Citation Information
Patent Citations
IoT camera control method and control system
CN113114950A