Camera control method and system based on multi-sensor data fusion
Through multi-sensor data fusion and deep Q network evaluation, the camera system can achieve intelligent adaptive adjustment in complex environments, solving the problems of poor adaptability and performance degradation of traditional camera systems in complex environments, and improving detection accuracy and system reliability.
Patent Information
- Application Number
- CN202510755633.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-19
AI Technical Summary
Existing security camera systems have poor adaptability in complex environments and cannot effectively cope with complex scenarios such as sudden changes in lighting, target occlusion, and multi-target intersections. Traditional multi-sensor fusion methods lack adaptability and collaborative control, resulting in degraded detection accuracy and performance.
A camera control method based on multi-sensor data fusion is adopted. The sensor performance is evaluated through a deep Q network, and environment-adaptive weight allocation and anomaly identification are performed to achieve intelligent adjustment of camera parameters. The camera parameters can be dynamically adjusted by combining multi-dimensional anomaly detection and closed-loop feedback mechanism.
It improves the detection accuracy and reliability of the system in complex environments, avoids the impact of single sensor failure, and realizes intelligent adaptive adjustment of camera parameters and continuous optimization of system performance.
Smart Images

Figure CN120676248A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of camera control technology, and in particular to a camera control method and system based on multi-sensor data fusion. Background Art
[0002] Existing security camera systems often rely on fixed parameter settings or simple automatic adjustment mechanisms, which are unable to effectively cope with complex and changing environmental conditions. This is particularly true in complex scenarios such as sudden changes in illumination, object occlusion, and the intersection of multiple objects. The systems' detection accuracy and adaptability are severely insufficient. Furthermore, traditional methods lack the ability to perceive environmental changes and are unable to dynamically optimize sensor configuration and camera parameters based on real-time environmental conditions, resulting in significant performance degradation in dynamic environments.
[0003] While multi-sensor fusion technology offers new solutions to these challenges, existing fusion methods generally employ fixed weight allocation strategies, failing to adapt to changes in sensor performance under varying environmental conditions. Traditional Kalman filtering and weighted averaging fusion methods lack effective anomaly identification and isolation mechanisms when faced with sensor degradation, failure, or environmental interference, which can easily lead to a cascading decline in overall system performance. Furthermore, there is a lack of coordinated control mechanisms between sensor subsystems, and a lack of integration between camera parameter adjustments and multi-sensor fusion results, preventing the formation of a unified intelligent control strategy. Summary of the Invention
[0004] The present invention provides a camera control method and system based on multi-sensor data fusion, which solves the technical defect of poor adaptability of traditional methods in complex environments and realizes intelligent adaptive adjustment of focal length, exposure and pan / tilt angle.
[0005] In a first aspect, the present invention provides a camera control method based on multi-sensor data fusion, the camera control method based on multi-sensor data fusion comprising: Collecting a multi-sensor synchronous data set, and inputting the multi-sensor synchronous data set into a deep Q network to perform sensor performance evaluation to obtain a performance score set; Performing environment adaptive weight allocation based on the performance score set to obtain confidence weight parameters and fusion detection threshold parameters for each sensor; Perform abnormal sensor identification on the multi-sensor synchronous data set according to the confidence weight parameter and the fusion detection threshold parameter to obtain a normal working sensor identifier and an abnormal sensor isolation instruction; The multi-sensor synchronous data set is weightedly fused based on the normal working sensor identifier and the confidence weight parameter to obtain target detection confidence data, and a camera parameter adjustment instruction of the RGB camera is generated according to the target detection confidence data.
[0006] In combination with the first aspect, in a first implementation of the first aspect of the present invention, collecting a multi-sensor synchronous dataset and inputting the multi-sensor synchronous dataset into a deep Q network for sensor performance evaluation to obtain a performance score set includes: Perform hardware clock synchronization and software timestamp calibration on the output signals of the RGB camera, infrared thermal imager, PIR motion sensor, and acoustic sensor to obtain a multi-sensor synchronized dataset. Calculating the ratio of the output signal power spectral density and the noise power spectral density of each sensor in the multi-sensor synchronous data set to obtain a sensor comprehensive signal-to-noise ratio value; Performing illuminance quantification calculation based on the image brightness histogram of the RGB camera and the temperature gradient change rate of the infrared thermal imager in the multi-sensor synchronous data set to obtain the current ambient light illuminance value; Performing inter-frame motion analysis on the continuous frame image data in the multi-sensor synchronized data set and performing speed analysis in combination with the trigger frequency data of the PIR motion sensor to obtain a target motion speed value, and combining the sensor comprehensive signal-to-noise ratio value, the current ambient light illumination value, and the target motion speed value into a three-dimensional environmental state vector; The three-dimensional environment state vector is input into a deep Q network to perform sensor performance evaluation to obtain a performance score set.
[0007] In combination with the first aspect, in a second implementation of the first aspect of the present invention, inputting the three-dimensional environment state vector into a deep Q network for sensor performance evaluation to obtain a performance score set includes: The three-dimensional environmental state vector is input into the deep Q network to perform environmental adaptability analysis to obtain the coefficient matrix of the influence of environmental conditions on the working performance of each sensor; Performing environmental sensitivity calculation based on the influence coefficient matrix to obtain an environmental adaptive feature vector; Based on the environmental adaptation feature vector, the imaging quality of the RGB camera under current lighting conditions, the detection accuracy of the infrared thermal imager under current temperature gradients, the response sensitivity of the PIR motion sensor under current motion speeds, and the recognition accuracy of the acoustic sensor under current signal-to-noise ratio conditions are quantitatively evaluated to obtain the environmental adaptation performance value of each sensor; The environmental adaptability performance values of the sensors are normalized by probability distribution to obtain a performance score set.
[0008] In combination with the first aspect, in a third implementation of the first aspect of the present invention, inputting the three-dimensional environmental state vector into a deep Q network for environmental adaptability analysis to obtain a coefficient matrix of influence of environmental conditions on the working performance of each sensor includes: Performing a multi-layer perceptron calculation on the three-dimensional environment state vector through a state encoder of a deep Q network to obtain an environment state encoding vector; Performing a state-action value evaluation on the environment state encoding vector by using a Q-value function calculator of the deep Q network to obtain a four-dimensional sensor Q-value vector; The environmental impact quantification module of the deep Q network is used to perform matrix transformation and weight distribution calculation on the Q value vector of the four-dimensional sensor to obtain the influence coefficient matrix of the environmental conditions on the working performance of each sensor.
[0009] In combination with the first aspect, in a fourth implementation of the first aspect of the present invention, performing environment-adaptive weight allocation based on the performance score set to obtain the confidence weight parameter and fusion detection threshold parameter of each sensor includes: Inputting the performance score set into the deep deterministic policy gradient algorithm to calculate the environmental adaptability coefficient to obtain the environmental adaptability coefficient data of each sensor; Performing a weighted confidence calculation based on the environmental adaptability coefficient data of each sensor and the performance score set to obtain a confidence weight parameter of each sensor; An environmental complexity quantification calculation is performed based on the three-dimensional environmental state vector to obtain an environmental complexity index value, and a basic detection threshold is corrected based on the environmental complexity index value to obtain a fusion detection threshold parameter.
[0010] In combination with the first aspect, in a fifth implementation of the first aspect of the present invention, the performing abnormal sensor identification on the multi-sensor synchronized data set according to the confidence weight parameter and the fusion detection threshold parameter to obtain a normal working sensor identifier and an abnormal sensor isolation instruction includes: Perform multi-dimensional anomaly detection on the RGB camera, infrared thermal imager, PIR motion sensor, and acoustic sensor based on the multi-sensor synchronous data set to obtain an anomaly detection result matrix; Quantitatively calculate the degree of abnormality based on the abnormality detection result matrix to obtain a numerical value of the degree of abnormality of each sensor; Compare and judge the abnormality level score value of each sensor with the fusion detection threshold parameter to obtain the abnormality level classification result of each sensor; Screening normal sensors based on the abnormality level classification results of each sensor to obtain normal working sensor identifiers; An abnormal sensor isolation instruction is generated according to the abnormality level classification results of each sensor and the confidence weight parameter.
[0011] In combination with the first aspect, in a sixth implementation of the first aspect of the present invention, performing weighted fusion on the multi-sensor synchronized data set based on the normal working sensor identifier and the confidence weight parameter to obtain target detection confidence data, and generating a camera parameter adjustment instruction for the RGB camera based on the target detection confidence data, includes: Filtering the multi-sensor synchronous data set for valid data according to the normal working sensor identifier to obtain a valid sensor data set; Performing a weighted fusion calculation based on the valid sensor data set and the confidence weight parameter to obtain target detection confidence data; Performing camera multi-parameter collaborative calculation based on the target detection confidence data to obtain three-dimensional parameter control data; Instruction packaging and timing coordination are performed based on the three-dimensional parameter control data to obtain camera parameter adjustment instructions.
[0012] In combination with the first aspect, in a seventh implementation of the first aspect of the present invention, performing camera multi-parameter collaborative calculation based on the target detection confidence data to obtain three-dimensional parameter control data includes: Inputting the target detection confidence data into the camera parameter adaptive adjustment module for data analysis and feature extraction to obtain a comprehensive feature vector of the target state; Performing depth of field analysis and temperature gradient calculation based on the target state comprehensive feature vector to obtain a focus adjustment parameter; Performing brightness distribution statistics and exposure feedback calculations based on the target state comprehensive feature vector to obtain exposure compensation parameters; The target state comprehensive feature vector is input into the target trajectory prediction algorithm for Kalman filtering and motion prediction calculation to obtain the gimbal angle control parameters, and the focal length adjustment parameters, the exposure compensation parameters and the gimbal angle control parameters are combined into three-dimensional parameter control data.
[0013] In combination with the first aspect, in an eighth implementation of the first aspect of the present invention, the camera control method based on multi-sensor data fusion further includes: Obtaining image data after the camera parameter adjustment instruction is executed, and performing image quality assessment and feature extraction to obtain an image feature vector; Performing a weighted fusion calculation based on the image feature vector and the target detection confidence data to obtain a fused feature description vector; Performing care target recognition and positioning calculation based on the fused feature description vector to obtain a care target detection result; Performing accuracy and latency analysis based on the care target detection results to obtain system performance evaluation indicators; Based on the system performance evaluation indicators, the learning rate, reward function weight and network structure parameters of the deep Q network are adjusted online, and the parameter configuration of the environment adaptive weight distribution strategy is updated to obtain an updated deep Q network.
[0014] In a second aspect, the present invention provides a camera control system based on multi-sensor data fusion, the camera control system based on multi-sensor data fusion comprising: An acquisition module is used to acquire a multi-sensor synchronous data set and input the multi-sensor synchronous data set into a deep Q network for sensor performance evaluation to obtain a performance score set; A weight allocation module, configured to perform environment-adaptive weight allocation based on the performance score set to obtain confidence weight parameters and fusion detection threshold parameters for each sensor; an abnormal sensor identification module, configured to identify abnormal sensors on the multi-sensor synchronous data set according to the confidence weight parameter and the fusion detection threshold parameter, and obtain a normal working sensor identifier and an abnormal sensor isolation instruction; A generation module is used to perform weighted fusion on the multi-sensor synchronous data set based on the normal working sensor identification and the confidence weight parameter to obtain target detection confidence data, and generate a camera parameter adjustment instruction for the RGB camera according to the target detection confidence data.
[0015] In the technical solution provided by the present invention, through the real-time calculation of the three-dimensional environmental state vector and the environmental adaptability analysis of the deep Q network, the system can dynamically evaluate the working performance of each sensor under the current environmental conditions according to the changes in the sensor signal-to-noise ratio, ambient light illumination and target motion speed, achieving a technical breakthrough from fixed weights to environmental adaptive weight distribution, and solving the technical defect of poor adaptability of traditional methods in complex environments. By adopting a multi-dimensional anomaly detection mechanism, through the comprehensive evaluation of data consistency detection, timing stability detection, response delay detection and output range detection, combined with a hierarchical isolation strategy, predictive identification and progressive isolation of sensor anomalies are achieved, effectively avoiding the impact of a single sensor failure on the performance of the entire system, and significantly improving the reliability and continuity of the system. By performing multimodal feature fusion of RGB image features, infrared temperature features, PIR motion features and acoustic spectrum features, combined with the dynamic adjustment of confidence weight parameters, intelligent weighted fusion of multi-sensor data is achieved, overcoming the technical limitations of the insufficient accuracy of traditional single sensor control methods and the lack of intelligence of traditional multi-sensor simple fusion methods. By combining a target distance estimation algorithm with binocular vision and infrared temperature gradient calculation, an ambient light analysis algorithm with a closed-loop feedback mechanism, and a target trajectory prediction algorithm with Kalman filtering for collaborative calculation of three-dimensional parameters, we achieve intelligent adaptive adjustment of focal length, exposure, and gimbal angle, resolving the existing problem of poor overall performance caused by the independent operation of each subsystem. A complete closed-loop feedback optimization mechanism, from sensor data input to detection result output, has been established. Through system performance evaluation and online adjustment of the deep Q network parameters using a gradient descent algorithm, we achieve autonomous and continuous improvement of system performance, overcoming the technical limitation of existing technologies where system performance decays over time and cannot be self-repaired. This provides technical support for the long-term stable operation of the intelligent care system. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0017] Figure 1 A flowchart of a camera control method based on multi-sensor data fusion provided in an embodiment of the present application; Figure 2 This is a schematic block diagram of the structure of a camera control system based on multi-sensor data fusion provided in an embodiment of the present application. DETAILED DESCRIPTION
[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0019] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may change based on actual circumstances.
[0020] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0021] It should be further understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0022] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments may be combined with each other.
[0023] See also Figure 1 , Figure 1 A flowchart of a camera control method based on multi-sensor data fusion provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the camera control method based on multi-sensor data fusion provided in the embodiment of the present application includes steps S100 to S400.
[0024] Step S100: Collect a multi-sensor synchronous data set, and input the multi-sensor synchronous data set into a deep Q network to perform sensor performance evaluation to obtain a performance score set; It is understandable that the execution subject of the present invention can be a camera control system based on multi-sensor data fusion, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking the server as the execution subject as an example.
[0025] Specifically, unified hardware clock synchronization and software timestamp calibration are performed on the output signals of the connected RGB camera, infrared thermal imager, PIR motion sensor, and acoustic sensor. Hardware clock synchronization aligns the sampling times of various sensors at the hardware level by setting a unified reference clock signal source, ensuring millisecond-level synchronization accuracy. Software timestamp calibration further assigns a unified time stamp to the collected data and uses a time drift correction algorithm to eliminate hardware residual errors, resulting in a synchronized multi-sensor dataset. Based on the synchronized multi-sensor dataset, the signal power characteristics of each sensor's output signal are analyzed. The comprehensive signal-to-noise ratio (SNR) for each sensor is calculated by calculating the ratio between the signal power spectral density and the corresponding noise power spectral density. The power spectral density describes the energy distribution of the signal across different frequency components, while the noise power spectral density represents the energy level of the noise component in the signal. The ratio of the two effectively characterizes the clarity and effectiveness of the signal in the current environment. By performing these ratio calculations on the image signal from the RGB camera, the thermal imaging signal from the infrared thermal imager, the motion detection signal from the PIR motion sensor, and the audio signal from the acoustic sensor, the comprehensive SNR for each sensor is calculated. To obtain ambient illumination information, statistical features are extracted from the brightness histograms of images captured by the RGB cameras in the synchronized dataset. The distribution of pixels within each brightness interval within the image is analyzed, and combined with the temperature gradient change rate of the infrared thermal imager, a comprehensive assessment of the scene's illumination conditions is performed. The brightness histogram reflects the overall brightness level and its uniformity, while the temperature gradient change rate complements the thermal distribution characteristics caused by illumination. The combined quantification of the two effectively improves the accuracy of depicting complex lighting environments and calculates the current ambient illumination value. Target velocity is then analyzed. Inter-frame motion analysis techniques are applied to consecutive frames of image data from the multi-sensor synchronized dataset, and the optical flow method is used to extract the pixel displacement vector field of the moving target within the image, estimating the target's motion trajectory and velocity characteristics. Furthermore, combined with the trigger frequency data from the PIR motion sensor, a supplementary correction is performed based on the relationship between trigger frequency and target appearance probability, improving the accuracy and robustness of velocity estimation. Inter-frame optical flow provides high-spatial-resolution motion information, while the PIR sensor provides high-temporal-resolution motion trigger information. These two complement each other and are combined to calculate the target's velocity. The sensor's comprehensive signal-to-noise ratio value, the current ambient light illumination value, and the target motion speed value are combined to form a three-dimensional environmental state vector that describes the current detection environment state.The three-dimensional environmental state vector is input into the deep Q network for sensor performance evaluation. The deep Q network uses a neural network model with a multi-layer perceptron structure. The input layer receives the environmental state vector. After nonlinear mapping and feature extraction through three hidden layers, the environmental feature space is gradually converted into the output space of sensor performance scores. The output layer gives the performance score of each sensor in the current environment, and finally obtains a set of performance scores for the RGB camera, infrared thermal imager, PIR motion sensor and acoustic sensor.
[0026] The three-dimensional environmental state vector is input into a deep Q-network for environmental adaptability analysis. Through the network's complex feature mapping and policy learning mechanisms, the inherent correlation between the environmental state vector and performance changes of various sensors is explored. The network outputs a matrix of environmental influence coefficients on the performance of each sensor. Each element of this matrix characterizes the intensity of the impact of a specific environmental factor on the performance of a particular sensor, forming a high-dimensional environmental-performance correlation map. Environmental sensitivity calculation is performed based on the information within the influence coefficient matrix. This calculation uses a weighted integration of the influence coefficients to reflect the sensitivity of each sensor to different environmental changes. This calculation extracts an environmental adaptability feature vector that comprehensively describes the impact of environmental changes on sensor performance. Based on this environmental adaptability feature vector, the performance of different sensors in the current environment is quantitatively evaluated. For RGB cameras, image quality is quantitatively analyzed based on current lighting conditions. Image quality assessment is based on a comprehensive calculation of brightness uniformity, image clarity, and noise density, reflecting the degree of image degradation in low-light or overexposure conditions. For infrared thermal imagers, target detection accuracy is calculated by analyzing the response characteristics of temperature gradients under current environmental conditions, assessing their imaging resolution in high-noise environments or against low-temperature target backgrounds. For PIR motion sensors, trigger sensitivity is assessed based on target velocity data, with particular attention paid to the changing trends in detection capabilities for high- and low-speed targets to ensure that changes in velocity do not significantly reduce sensitivity. For acoustic sensors, recognition accuracy is assessed based on the current comprehensive signal-to-noise ratio (SNR), analyzing changes in recognition performance when background noise increases or speech signals weaken. This results in a comprehensive environmental adaptability performance value for each sensor. The resulting environmental adaptability performance values are then normalized using a probability distribution. During normalization, each performance value is mapped to a standard range of 0 to 1 using a softmax function or a minimum-maximum normalization method, transforming all sensor performance indicators into a set of probabilistic scores.
[0027] The state encoder module within the deep Q-network processes the input three-dimensional environment state vector. Based on a multi-layer perceptron architecture, this state encoder comprises multiple fully connected layers. It uses nonlinear activation functions to perform multi-level feature extraction and high-dimensional mapping on the input environment state vector. Layer-by-layer transformations transform the original three-dimensional vector into a high-dimensional, dense environment state encoding vector. The Q-value function calculator within the deep Q-network takes the environment state encoding vector as input and performs state-action value assessment. The Q-value function calculator employs the principles of Q-learning in deep reinforcement learning, combining the current encoded state vector with a predefined action space to infer the expected value of each action in the current environment. Since this system targets four sensor types (RGB camera, infrared thermal imager, PIR motion sensor, and acoustic sensor), the action space is defined as the set of actions associated with each of these four sensors. The Q-value function calculator outputs a four-dimensional sensor Q-value vector. Each element in this four-dimensional vector represents the value assessment of the corresponding sensor in the detection task under the current environment, reflecting the degree to which different environments favor or inhibit the performance of different sensors. The environmental impact quantification module, based on a deep Q-network, performs a matrix transformation on the four-dimensional sensor Q-value vector. By introducing a pre-trained mapping matrix, the raw Q-values are transformed linearly or nonlinearly to capture the interdependencies between sensors and the interactions between environmental factors. This matrix transformation not only adjusts the adaptability of individual sensors in specific environments but also comprehensively reflects how environmental changes induce performance trade-offs in a multi-sensor system. The environmental impact quantification module then performs a weighting calculation based on the transformed data. Taking into account environmental complexity, sensor characteristics, and dynamic performance, it employs an adaptive weighting strategy to assign different importance weights to each sensor, forming a comprehensive matrix of environmental influence coefficients on each sensor's performance. Each row of this matrix corresponds to an environmental state component (signal-to-noise ratio, illumination, motion speed), and each column corresponds to a sensor. Each element in the matrix quantifies the specific impact of a specific environmental factor on the performance of a particular sensor.
[0028] Step S200: Performing environment adaptive weight allocation based on the performance score set to obtain confidence weight parameters and fusion detection threshold parameters for each sensor; Specifically, the performance score set is input into the deep deterministic policy gradient algorithm framework. Unlike traditional Q-learning methods in discrete action spaces, this algorithm employs a continuous action space optimization strategy, enabling dynamic output adjustment within a more nuanced numerical range. This makes it suitable for weight allocation problems requiring high-precision adjustments. The deep deterministic policy gradient algorithm uses the performance score set as state input and generates environmental adaptability coefficient data for each sensor through the policy network. These coefficients directly reflect the performance responsiveness and adaptability of each sensor under the current environmental conditions, enabling sensitive perception and rapid adjustment to environmental changes. A weighted confidence score is calculated based on the obtained environmental adaptability coefficient data for each sensor and the original performance score set. This process uses the environmental adaptability coefficient as a weighting factor, multiplying it by each sensor's performance score. The results are then normalized to ensure that the sum of all sensor confidence weight parameters is 1. This calculation method numerically emphasizes sensors with good environmental adaptability while suppressing the impact of sensors with degraded performance in the current environment, forming a dynamic, flexible, and context-aware confidence weight allocation mechanism. If a sensor's performance score under current lighting, temperature, or noise conditions is high, and its environmental adaptability coefficient is also in the high range, its weighted confidence weight will be increased accordingly. Conversely, sensors with poor performance or that are sensitive to environmental changes will automatically have their weights reduced, thus enabling the multi-sensor fusion system to intelligently adapt to environmental changes. Furthermore, the system performs a quantification of environmental complexity based on the input three-dimensional environmental state vector—a combined vector of signal-to-noise ratio, illumination, and motion speed. This quantification not only considers the extreme values or rates of change of individual factors but also comprehensively considers the interactions among them. For example, the combined effects of low signal-to-noise ratio, high-speed background motion, and complex lighting variations pose challenges to the system's detection capabilities. By projecting the environmental state vector into feature space and combining it with empirical parameters or a complexity function model derived through reinforcement learning, an environmental complexity index is calculated. This value reflects the comprehensive challenge posed by the current environment to the sensor system's detection task. Based on this environmental complexity index, the existing basic detection threshold is adjusted. The basic detection threshold is a static parameter set for an ideal or standard environment and cannot flexibly adapt to dynamic environmental changes. Therefore, it is dynamically adjusted based on environmental complexity. When environmental complexity increases—that is, when the environment becomes more severe and detection becomes more difficult—the system automatically raises the detection threshold to minimize false detections and missed detections. Conversely, in a simple and stable environment, the detection threshold is appropriately lowered to improve detection sensitivity. This correction process introduces an adjustment coefficient, based on the product of environmental complexity and the basic threshold, or a nonlinear mapping relationship, to ultimately output a fused detection threshold parameter, ensuring the system consistently maintains optimal detection performance under various environmental conditions.
[0029] Step S300: identifying abnormal sensors on the multi-sensor synchronous data set according to the confidence weight parameter and the fusion detection threshold parameter, and obtaining a normal working sensor identifier and an abnormal sensor isolation instruction; Specifically, multi-dimensional anomaly detection is performed on the RGB camera, infrared thermal imager, PIR motion sensor, and acoustic sensor based on a synchronized multi-sensor dataset. This multi-dimensional anomaly detection process integrates multiple dimensions, including temporal feature analysis, frequency domain feature extraction, and spatial feature distribution. Specifically, it includes submodules such as data consistency detection, temporal stability detection, response delay detection, and output range detection. Each submodule calculates characteristic indicators reflecting abnormal behavior based on the physical characteristics and output signal form of each sensor. Ultimately, the detection results of each submodule are aggregated to form an anomaly detection result matrix containing the detection data of the four sensor categories. Each row in the matrix corresponds to a sensor type, and each column corresponds to an anomaly detection indicator. Based on the anomaly detection result matrix, the degree of anomaly is quantitatively calculated for each sensor. This process generates an anomaly score by weighting the detection indicators of each dimension or fusing them using a multi-dimensional anomaly scoring function. The score reflects the health of the corresponding sensor and its degree of deviation from normal status. Each sensor's anomaly score is then compared with the fused detection threshold parameter. Based on the comparison results, the system categorizes abnormality levels, setting specific multi-level abnormality classification criteria. For example, sensors below a baseline threshold are considered normal, those slightly above the threshold are considered mildly abnormal, and those significantly above the threshold are classified as moderately or severely abnormal. Based on the abnormality classification results for each sensor, a normal sensor screening operation is performed, filtering out all sensors deemed normal or with abnormality levels below the set safety threshold, thereby obtaining a list of properly functioning sensors in the current environment. After the normal sensor screening is complete, an abnormal sensor isolation instruction is generated based on the abnormality classification results and corresponding confidence weight parameters for each sensor. This isolation instruction goes beyond simply marking isolation or not, and formulates a graded isolation strategy based on the combined determination of abnormality level and confidence weight. For example, for sensors with high confidence but mild abnormality, their weight may be reduced rather than directly isolated. For sensors with severe abnormality and low confidence, full isolation is implemented, triggering the system's adaptive reconstruction mechanism to dynamically adjust the weight distribution of the remaining sensors to ensure stable overall system performance.
[0030] Step S400: weighted fusion is performed on the multi-sensor synchronous data set based on the normal working sensor identification and confidence weight parameters to obtain target detection confidence data, and a camera parameter adjustment instruction of the RGB camera is generated according to the target detection confidence data.
[0031] Specifically, valid data is filtered from the current multi-sensor synchronized dataset based on the identification of properly functioning sensors. By comparing the identity information of each sensor in the synchronized dataset with the identification of properly functioning sensors, all sensor data marked as abnormal or isolated is filtered out, retaining only sensor outputs determined to be normal, forming a valid sensor dataset containing high-confidence and reliable data. Based on this valid sensor dataset and the corresponding confidence weight parameters, a weighted fusion calculation is performed. During the weighted fusion process, each properly functioning sensor data stream is assigned a different fusion weight based on its confidence weight, with data with higher weights receiving a larger proportion in the final fusion result. The specific calculation method uses a weighted summation or weighted averaging strategy, combining the target detection results of each sensor according to the confidence weights to form a comprehensive target detection confidence data. Based on the target detection confidence data, multi-camera parameter collaborative calculation is performed to determine the imaging control parameters that need to be adjusted. By comprehensively considering the three core parameters of focal length, exposure compensation, and gimbal rotation angle, dynamic derivation is performed based on the target distance, ambient lighting changes, and target motion trajectory contained in the target confidence information. Target distance information is used for focus adjustment, adjusting the camera lens's focal length to achieve a clear image of the target. Ambient lighting changes are addressed through exposure compensation, ensuring moderate brightness and rich detail in the image. The target's trajectory is calculated by predicting its future position and, combined with motion prediction algorithms such as the Kalman filter, adjusting the pan / tilt tilt angles to achieve continuous camera tracking. The calculated results of these three parameters are combined into three-dimensional parameter control data. Based on this three-dimensional parameter control data, command encapsulation and timing coordination are performed to generate parameter adjustment commands that comply with the camera control protocol. The command encapsulation process encodes the focus, exposure, and pan / tilt tilt angle adjustment commands in a specific communication protocol format, ensuring accurate recognition and execution by the camera hardware. Furthermore, to avoid conflicts or delays between multiple parameter adjustments, timing coordination is performed. The order and execution time windows of each control command are rationally arranged based on the priority and execution dependencies of the parameter changes, ensuring smooth transitions during the parameter adjustment process without causing image jitter or target loss.
[0032] The target detection confidence data is input into the camera parameter adaptive adjustment module for depth data parsing and feature extraction. This module parses the confidence data and extracts multiple sub-indicators related to target detection, including target location information, confidence score, target category characteristics, and environmental status indicators. These sub-features are then combined using a feature fusion algorithm to form a high-dimensional vector that comprehensively reflects the target state, known as the target state comprehensive feature vector. Based on the target state comprehensive feature vector, targeted depth of field analysis and temperature gradient calculation are performed to derive the focal length adjustment parameters that need to be adjusted. Depth of field analysis calculates the optimal focal plane position based on the target's spatial location and size characteristics. The focal length range that the camera lens needs to adjust is calculated based on the image resolution and background complexity. Furthermore, the temperature gradient information provided by the infrared thermal imager is used to analyze the thermal distribution differences between the target and the background. This allows for further refinement of the depth of field adjustment strategy, ensuring that the target remains in a clear image range under varying thermal environments. By integrating these two components, the focus adjustment parameters are output to guide the camera lens for precise focusing. Brightness distribution statistics and exposure feedback calculations are performed on the same target state comprehensive feature vector to determine appropriate exposure compensation parameters. The image captured by the RGB camera is analyzed for brightness distribution. By statistically analyzing the histogram characteristics of the image pixel brightness, the current image brightness balance and brightness deviation are assessed. The current lighting change rate is calculated, combined with the ambient lighting indicator information in the detection confidence level. This information is then linked to the camera's built-in exposure feedback mechanism to dynamically correct the base exposure setting. By adjusting the exposure time or gain control, the system rapidly adapts to the current lighting environment, outputting exposure compensation parameters to ensure clear and detailed images under varying lighting conditions. Simultaneously, the target state's comprehensive feature vector is input into the target trajectory prediction algorithm module, which performs Kalman filter-based target motion state estimation and future trajectory prediction. The Kalman filter utilizes the current target's position, velocity, and acceleration information, combined with historical motion data, to recursively predict the target's position at several future moments. Based on the predicted trajectory data, the camera's pan / tilt (PTZ) horizontal and vertical angle adjustments are calculated to obtain the PTZ angle control parameters. These parameters enable the camera to continuously track the moving target, keeping the target centered in the frame, and improving capture stability and continuity in dynamic scenes. The focus adjustment parameters, exposure compensation parameters, and PTZ angle control parameters are combined to generate three-dimensional parameter control data.
[0033] When the camera executes parameter adjustment commands, updated image data is obtained in real time. The newly generated image data includes the imaging effects after adjusting the focus, exposure compensation, and gimbal angle. It also reflects the current ambient lighting changes, the target's motion state, and the overall imaging system's adjustment response. This batch of image data undergoes image quality assessment and feature extraction, using metrics such as image clarity, contrast, brightness balance, and texture detail fidelity. A comprehensive evaluation model quantifies the overall image quality. A deep convolutional feature extraction network is used to extract a high-dimensional image feature vector. This feature vector effectively captures core visual features such as the target object's shape, boundary, and texture while preserving fine-grained image information. A weighted fusion calculation is performed based on the image feature vector and target detection confidence data. This weighted fusion strategy dynamically adjusts the weight distribution of each feature dimension in the image feature vector based on the target confidence score, location coordinates, and environmental label information contained in the confidence data. This ensures that the fused feature description vector more accurately reflects the target object's salient characteristics and spatial distribution patterns in the current environment. Through feature enhancement and suppression mechanisms, key information is highlighted amidst information redundancy and noise interference, generating a more discriminative and environmentally adaptable fused feature description vector. Based on the fused feature description vector, care target recognition and localization are performed. The recognition module combines a deep convolutional neural network with an object detection framework, such as a modified YOLO or Faster R-CNN network. It uses the fused feature vector to perform category determination and bounding box regression, accurately identifying the target category, such as personnel, equipment, or specific care objects. Simultaneously, the localization module precisely calculates the target's coordinate position and bounding box extent on the image plane. Based on the recognition and localization results, accuracy and latency analysis are performed. Accuracy analysis primarily calculates key metrics such as detection accuracy, false positive rate, and missed detection rate by counting the number of true positives, false negatives, and detected cases. Latency analysis, on the other hand, calculates the overall response delay from image acquisition to target recognition and localization based on processing time data at each stage of the detection process, reflecting the system's real-time performance under varying workloads and environmental conditions. Based on these system performance evaluation metrics, the deep Q network is online tuned and adaptively optimized. Based on the accuracy trend and latency fluctuations, the learning rate of the deep Q network is dynamically adjusted to enable the network to adapt to changes in the difficulty of detection tasks caused by environmental changes while maintaining stable learning. At the same time, the reward function weight is adjusted according to the changes in detection accuracy rewards and false alarm penalties to optimize the strategy selection tendency in the Q value update process. If the long-term performance indicator fluctuation exceeds the threshold, the structural parameters of the deep Q network are fine-tuned, including the number of hidden layer neurons, activation function configuration and regularization parameters, to enhance the model's generalization ability in complex environments.Through this series of online adjustment mechanisms, the system can dynamically update the environment-adaptive weight allocation strategy, continuously optimize the sensor fusion weights and detection threshold parameters, and ultimately obtain an updated deep Q network.
[0034] In an embodiment of the present invention, through real-time calculation of the three-dimensional environmental state vector and environmental adaptability analysis using a deep Q network, the system can dynamically evaluate the performance of each sensor under current environmental conditions based on changes in the sensor signal-to-noise ratio, ambient light intensity, and target motion speed. This achieves a technological breakthrough from fixed weights to environmentally adaptive weight allocation, resolving the technical drawback of traditional methods, which suffer from poor adaptability in complex environments. A multi-dimensional anomaly detection mechanism is employed, which, through comprehensive evaluation of data consistency detection, timing stability detection, response delay detection, and output range detection, combined with a hierarchical isolation strategy, enables predictive identification and progressive isolation of sensor anomalies, effectively avoiding the impact of a single sensor failure on overall system performance and significantly improving system reliability and continuity. By performing multimodal feature fusion of RGB image features, infrared temperature features, PIR motion features, and acoustic spectrum features, combined with dynamic adjustment of confidence weight parameters, intelligent weighted fusion of multi-sensor data is achieved, overcoming the technical limitations of traditional single-sensor control methods, which suffer from insufficient precision, and traditional multi-sensor simple fusion methods, which lack intelligence. By combining a target distance estimation algorithm with binocular vision and infrared temperature gradient calculation, an ambient light analysis algorithm with a closed-loop feedback mechanism, and a target trajectory prediction algorithm with Kalman filtering for collaborative calculation of three-dimensional parameters, we achieve intelligent adaptive adjustment of focal length, exposure, and gimbal angle, resolving the existing problem of poor overall performance caused by the independent operation of each subsystem. A complete closed-loop feedback optimization mechanism, from sensor data input to detection result output, has been established. Through system performance evaluation and online adjustment of the deep Q network parameters using a gradient descent algorithm, we achieve autonomous and continuous improvement of system performance, overcoming the technical limitation of existing technologies where system performance decays over time and cannot be self-repaired. This provides technical support for the long-term stable operation of the intelligent care system.
[0035] In a specific embodiment, the process of executing step S100 may specifically include the following steps: Perform hardware clock synchronization and software timestamp calibration on the output signals of the RGB camera, infrared thermal imager, PIR motion sensor, and acoustic sensor to obtain a multi-sensor synchronized dataset. The sensor's comprehensive signal-to-noise ratio is obtained by calculating the ratio of the output signal power spectrum density and noise power spectrum density of each sensor in the multi-sensor synchronous data set. The current ambient light value is obtained by performing quantification calculation of the illumination based on the image brightness histogram of the RGB camera and the temperature gradient change rate of the infrared thermal imager in the multi-sensor synchronous dataset. Inter-frame motion analysis is performed on the continuous frame image data in the multi-sensor synchronous data set, and speed analysis is performed in combination with the trigger frequency data of the PIR motion sensor to obtain the target motion speed value. The sensor's comprehensive signal-to-noise ratio value, the current ambient light value, and the target motion speed value are combined into a three-dimensional environmental state vector. The three-dimensional environment state vector is input into the deep Q network for sensor performance evaluation to obtain a performance score set.
[0036] Specifically, during the system construction phase, a unified hardware clock source is connected to each sensor. A high-precision synchronization clock module provides a standardized timing reference for all sensors. A dedicated synchronization clock chip, such as one based on the IEEE 1588 Precision Time Protocol or a higher-level local reference clock system, is used to ensure time consistency between sensors at the millisecond or microsecond level, preventing discrepancies in data fusion calculations caused by inconsistent sampling times. Furthermore, to account for minor time offsets caused by temperature drift, line delays, or other uncontrollable factors during actual operation, a software-level timestamp calibration mechanism further refines synchronization. Each sensor data stream is assigned a high-precision timestamp during acquisition. A sliding window-based time drift detection algorithm monitors minute differences in arrival times between sensor data in real time. Interpolation correction or time alignment algorithms are used to realign the data on the time axis, resulting in a synchronized multi-sensor dataset. Based on the multi-sensor synchronized dataset, the power characteristics of each sensor's output signal are analyzed. The power spectral density and noise power spectral density are calculated, and the sensor's comprehensive signal-to-noise ratio (SNR) is calculated based on the ratio between the two. Power spectral density (PSD) is calculated using fast Fourier transform technology, converting the sensor's time-domain signal into a frequency-domain representation. The energy distribution of each frequency component is then analyzed to measure the signal's spectral characteristics. Noise power spectral density (PSD) is calculated by collecting background noise data under static or no-signal input conditions and performing frequency-domain analysis to determine the noise power distribution at each frequency point. The signal-to-noise ratio (SNR) spectrum curve is calculated by comparing the signal power spectral density to the noise power spectral density at each frequency point. This curve is then integrated or averaged to obtain a single numerical value for the comprehensive SNR. The comprehensive SNR directly reflects the reliability and clarity of the sensor's output signal in the current environment. Ambient illumination quantification is performed on data from the RGB camera and infrared thermal imager in the multi-sensor synchronized dataset. The brightness histogram of the continuous image sequence provided by the RGB camera calculates the distribution of pixels at each grayscale level, thereby assessing the overall image brightness balance and average brightness level. Parameters such as the mean, standard deviation, and peak position of the brightness histogram are extracted as preliminary features. Combined with the temperature field information provided by the infrared thermal imager, specifically the rate of change of the temperature gradient between the target and the background, the thermal distribution characteristics caused by the current illumination are inferred. By comprehensively analyzing the brightness distribution and temperature gradient changes, the ambient light intensity and its changing trend are accurately quantified, thereby calculating the current ambient illumination value. Inter-frame motion analysis is performed on the continuous image frames in the multi-sensor synchronized dataset, and the target motion speed is analyzed in combination with the trigger frequency data of the PIR motion sensor. Inter-frame motion analysis uses classic computer vision techniques such as optical flow to calculate the target's movement speed and direction on the image plane by extracting the pixel-level displacement vector field of adjacent frames.The optical flow density field can reflect the overall displacement trend of the target area. Combined with feature point matching or target area tracking methods, it further improves the accuracy and robustness of motion estimation. Furthermore, the trigger frequency information provided by the PIR motion sensor—the number of times a moving target is detected per unit time—can complement the motion data from image analysis in the temporal domain. By weightedly fusing the inter-frame displacement velocity of consecutive image frames with the PIR sensor's trigger frequency, a reliable target velocity value is calculated, comprehensively reflecting the dynamic characteristics of target motion in the current monitoring scene. The sensor's comprehensive signal-to-noise ratio, the current ambient light intensity, and the target velocity value are combined to form a three-dimensional environmental state vector. This 3D environmental state vector is input into a pre-trained deep Q-network for sensor performance evaluation. As a reinforcement learning model, the deep Q-network has the ability to learn optimal action-value functions from complex input spaces. The network's input layer receives the 3D environmental state vector. Through nonlinear mapping across multiple hidden layers, high-level features are gradually extracted. The Q-value output layer then calculates the expected value of each action (i.e., the performance score of each sensor) under the current environmental state. The Q value reflects the reliability and quality of each sensor in performing the target detection task in the current environment. After normalization, a performance score set is finally obtained. Each score value in the performance score set corresponds to a sensor and represents its performance in the current environment.
[0037] In a specific embodiment, the step of inputting the three-dimensional environment state vector into the deep Q network for sensor performance evaluation to obtain a performance score set may specifically include the following steps: The three-dimensional environmental state vector is input into the deep Q network for environmental adaptability analysis, and the coefficient matrix of the influence of environmental conditions on the working performance of each sensor is obtained; Calculate environmental sensitivity based on the influence coefficient matrix and obtain the environmental adaptive eigenvector; Based on the environmental adaptation feature vector, the imaging quality of the RGB camera under current lighting conditions, the detection accuracy of the infrared thermal imager under current temperature gradients, the response sensitivity of the PIR motion sensor under current motion speeds, and the recognition accuracy of the acoustic sensor under current signal-to-noise ratio conditions are quantitatively evaluated to obtain the environmental adaptation performance value of each sensor. The environmental adaptability performance values of each sensor are normalized by probability distribution to obtain a performance score set.
[0038] Specifically, the three-dimensional environmental state vector is input into a deep Q-network for environmental adaptability analysis. The deep Q-network includes a state encoder module, which uses a multi-layer perceptron architecture to perform nonlinear mapping and high-dimensional feature extraction on the input low-dimensional environmental state vector. The state encoder converts the environmental state vector into a high-dimensional environmental state encoding vector through layer-by-layer weighted transformations and activation functions. After obtaining the environmental state encoding vector, the system uses the Q-value function calculator module within the deep Q-network to evaluate the state-action value (Q-value) of the vector. The action space is defined as a set of actions related to improving the performance of four sensor types: optimizing RGB camera image quality, improving infrared thermal imager detection accuracy, enhancing PIR motion sensor response sensitivity, and improving acoustic sensor recognition accuracy. The Q-value function calculator calculates the expected performance improvement for different action options based on the current environmental state, outputting a four-dimensional sensor Q-value vector. Each element of the Q-value vector represents the expected benefit of optimizing a particular sensor's performance under the current environment. The deep Q-network then invokes the environmental impact quantification module to perform matrix transformation and weight assignment on the four-dimensional sensor Q-value vector, generating a matrix of coefficients that determine the impact of environmental conditions on the performance of each sensor. The matrix transformation operation introduces a weight matrix for environmental and sensor interactions to model the coupling effects between environmental factors and the non-independence between sensor response characteristics, further refining the specific extent to which different environmental variables affect each sensor's performance. The weight distribution is dynamically adjusted based on the changing trends of the characteristic components in the environmental state encoding vector, ensuring that the influence coefficient matrix can reflect the actual impact of environmental changes on sensor performance in real time. This results in an environmental influence coefficient matrix with high accuracy and dynamic adaptability. Based on the environmental influence coefficient matrix, environmental sensitivity calculations are performed to comprehensively assess the performance stability and responsiveness of each sensor under the current environment. A weighted sum or weighted average is taken per row of the influence coefficient matrix, and the weight distribution characteristics of each environmental factor in the matrix are combined to extract an environmental adaptation eigenvector reflecting the sensitivity to environmental changes. Each component of this eigenvector corresponds to a sensor type. A larger value indicates a more sensitive sensor to environmental changes, while a smaller value indicates a higher stability under environmental changes. Based on the environmental adaptation eigenvector, the performance of each sensor under specific environmental conditions is quantitatively evaluated.For RGB cameras, the current illumination value is combined with the corresponding components in the adaptive eigenvector to comprehensively calculate image quality indicators. Image quality assessment primarily includes parameters such as image clarity, brightness uniformity, and noise density. Imaging performance is quantified using techniques such as brightness histogram uniformity and edge sharpness detection. For infrared thermal imagers, detection accuracy is quantified by combining the current temperature gradient change rate with the eigenvector components. Key indicators evaluated include thermal resolution, equivalent thermal noise temperature difference, and target boundary sharpness. For PIR motion sensors, response sensitivity is calculated by combining the target motion speed value with environmental characteristic components. Performance indicators such as trigger sensitivity, false trigger rate, and response delay are evaluated. For acoustic sensors, recognition accuracy is quantified by combining the comprehensive signal-to-noise ratio value with the characteristic components. These indicators include speech recognition rate, background noise suppression ratio, and dynamic response range. Each performance evaluation is calculated using a model to obtain specific environmental adaptability performance values. These values reflect the performance and reliability of each sensor under specific environmental conditions. The environmental adaptability performance values of each sensor are normalized by probability distribution to eliminate the incomparability of performance values due to different sensor types and indicator dimensions. The performance indicators of different sensors are mapped to a standardized interval between 0 and 1 through a unified standard, and finally a standardized performance score set is obtained. Each score corresponds to a sensor. The higher the value, the better the performance of the sensor in the current environment, and it is suitable to be assigned a higher fusion weight and decision priority.
[0039] In a specific embodiment, the step of inputting the three-dimensional environmental state vector into the deep Q network for environmental adaptability analysis to obtain the coefficient matrix of the influence of environmental conditions on the working performance of each sensor can specifically include the following steps: Through the state encoder of the deep Q network, the multi-layer perceptron calculation is performed on the three-dimensional environment state vector to obtain the environment state encoding vector; The state-action value of the environment state encoding vector is evaluated by the Q-value function calculator of the deep Q network to obtain the four-dimensional sensor Q-value vector; Through the environmental impact quantification module of the deep Q network, the matrix transformation and weight distribution calculation of the four-dimensional sensor Q value vector are performed to obtain the coefficient matrix of the impact of environmental conditions on the working performance of each sensor.
[0040] Specifically, a deep Q-network state encoder performs multi-layer perceptron computation on the 3D environment state vector. Within the state encoder, the multi-layer perceptron consists of several fully connected layers, configured as three or four layers, with the number of neurons in each layer gradually decreasing to achieve feature compression and abstraction. The input layer receives the original 3D vector. The first fully connected layer performs computations to map the 3D vector to a higher-dimensional feature space, such as 128 or 256 dimensions, to increase feature representation capability. Each fully connected layer is followed by a nonlinear activation function, using the ReLU (rectified linear unit) activation function to ensure good nonlinear fitting across different intervals and avoid the vanishing gradient problem. After these multiple transformations, the vector is continuously extracted and compressed, gradually removing redundant information and enhancing deep environmental features related to sensor performance. Ultimately, the output layer produces a low-dimensional but semantically rich environmental state encoding vector. This encoding vector not only retains the essential environmental information of the input vector but also, through the neural network's learning mechanism, extracts the underlying complex nonlinear relationship between the environment and sensor performance. The state-action value of the environmental state encoding vector is evaluated using the deep Q-network's Q-value function calculator. The Q-value function calculator is a neural network-based function approximator. Its task is to predict the expected reward of each possible action based on the input environmental state encoding. In this system, the action space is defined as the set of actions associated with optimizing the performance of four different sensors: optimizing the image quality of the RGB camera, improving the detection accuracy of the infrared thermal imager, enhancing the response sensitivity of the PIR motion sensor, and improving the recognition accuracy of the acoustic sensor. The Q-value function calculator takes the environmental state encoding vector as input and, through a series of linear transformations and nonlinear activation operations, outputs a four-dimensional vector, with each component corresponding to the Q-value of a specific sensor. This process involves several hidden layers with a moderate number of hidden nodes to ensure sufficient expressive power while avoiding overfitting. Each layer uses ReLU or LeakyReLU activation functions to increase the network's nonlinearity. The output layer is a linear layer that directly generates a four-dimensional output vector. Each element of this four-dimensional vector represents the expected reward of optimizing the corresponding sensor under the current environmental state, thereby quantifying the impact of the environmental state on the performance potential of each sensor. The four-dimensional sensor Q-value vector is then input into the environmental impact quantification module within the deep Q-network for processing. The environmental impact quantification module performs a matrix transformation on the Q-value vector to reveal the specific coefficient relationship between environmental conditions and the performance of each sensor. This matrix transformation is based on a pre-trained or dynamically learned transformation matrix of size 4×4, which is used to simulate the influence and interaction between environmental factors among sensors.The matrix transformation process involves performing a matrix multiplication between the four-dimensional Q-value vector and the 4×4 transformation matrix. The result is a new four-dimensional vector. Each element not only reflects the environmental adaptability characteristics of the corresponding sensor but also comprehensively considers the interaction effects of other sensors in the current environment. For example, when the lighting environment is unstable, the image quality of the RGB camera degrades, which may cause the system to rely on data from the infrared thermal imager or acoustic sensor. Therefore, the weights of the infrared thermal imager and acoustic sensor need to increase simultaneously. This complementary and coupled relationship is modeled through the matrix transformation. After the matrix transformation, the environmental impact quantification module performs a weight allocation calculation to refine the quantitative assessment of the environmental impact of each sensor. The weight allocation calculation is based on the matrix-transformed vector. By introducing a normalization operation, such as Softmax normalization, the environmental impact weight corresponding to each sensor is normalized to between 0 and 1, ensuring that the sum of the weights of all sensors is 1. After the matrix transformation and weight allocation, a matrix of coefficients of environmental conditions affecting the performance of each sensor is obtained.
[0041] In a specific embodiment, the process of executing step S200 may specifically include the following steps: The performance score set is input into the deep deterministic policy gradient algorithm to calculate the environmental adaptability coefficient and obtain the environmental adaptability coefficient data of each sensor; A weighted confidence calculation is performed based on the environmental adaptability coefficient data and performance score set of each sensor to obtain the confidence weight parameter of each sensor; The environmental complexity is quantitatively calculated based on the three-dimensional environmental state vector to obtain the environmental complexity index value, and the basic detection threshold is corrected based on the environmental complexity index value to obtain the fusion detection threshold parameter.
[0042] Specifically, the performance score set is input into the Deep Deterministic Policy Gradient (DDPG) algorithm to calculate the environmental adaptability coefficient. DDPG is a reinforcement learning algorithm based on deterministic policy gradients. It is suitable for policy optimization problems in continuous action spaces and, therefore, is well-suited for the refined calculation of sensor environmental adaptability coefficients. During this process, the performance score set serves as the state input to the actor network module within DDPG. The actor network receives the sensor performance scores and maps them through a multi-layer neural network, outputting a set of continuous action values. The action space is defined as the environmental adaptability coefficient of each sensor. The actor network consists of three to four fully connected layers, with a gradually decreasing number of hidden layer nodes to compress features. Each layer uses a nonlinear activation function, such as ReLU, to enhance the model's nonlinear fitting capabilities. It outputs the adaptability coefficient of each sensor in the current environmental state. The adaptability coefficient represents the weight of each sensor's ability to adjust based on environmental changes. A higher value indicates greater adaptability to the environment, and therefore warrants a higher weight in data fusion. Simultaneously, the critic network estimates the value of an actor based on its current state and the action it outputs. It uses the Bellman equation to update its Q-value and continuously optimizes the actor network parameters to ensure that updates to the adaptability coefficient improve system performance. A weighted confidence score is calculated based on each sensor's environmental adaptability coefficient data and performance score set. A new confidence base score is generated by multiplying each sensor's performance score by its corresponding adaptability coefficient. Each sensor's weighted base score is normalized, for example using max-min normalization or softmax normalization, to ensure that the sum of the confidence weight parameters for each sensor is 1. This minimizes the impact of numerical discrepancies and ensures a controllable and reasonable distribution of sensor data contributions during the subsequent multi-sensor data fusion process. This weighted confidence calculation allows the system to dynamically adjust the contribution weight of each sensor and optimize the fusion strategy in real time based on environmental changes, thereby maintaining high detection performance and system stability under varying environmental conditions. Environmental complexity is quantified based on the three-dimensional environmental state vector. The three-dimensional environment state vector, composed of the sensor's comprehensive signal-to-noise ratio, the current ambient illumination, and the target's motion speed, is a crucial representation of the current environment's characteristics. The core goal of quantifying environmental complexity is to assess the difficulty the current environment poses to sensor detection tasks. Therefore, the three-dimensional environment state vector undergoes feature analysis and quantitative modeling. A weighted feature fusion strategy is employed, identifying a decrease in signal-to-noise ratio, extreme changes in illumination (too dark or too bright), and an increase in motion speed as signs of increasing environmental complexity. A complexity evaluation function is constructed, for example, by taking the inverse of the signal-to-noise ratio and normalizing it to represent noise intensity, normalizing the absolute value of the illumination deviation from the standard value to measure the severity of illumination changes, and normalizing the motion speed to reflect the target's dynamics.The three indicators are combined through a weighted summation to obtain an environmental complexity index value. A larger value indicates a more complex environment and a higher difficulty level for the sensor detection task. Based on the environmental complexity index value, the basic detection threshold is modified to obtain the fusion detection threshold parameter. The basic detection threshold is the static detection threshold set by the system under standard conditions to distinguish between target and non-target states. The detection threshold is dynamically adjusted based on environmental complexity to improve the system's adaptability. A correction function is designed based on the complexity index value, using either linear or nonlinear mapping. When complexity increases, the detection threshold is appropriately increased to reduce the false detection rate; when complexity decreases, the threshold is appropriately lowered to improve sensitivity.
[0043] In a specific embodiment, the process of executing step S300 may specifically include the following steps: Based on the multi-sensor synchronous data set, multi-dimensional anomaly detection is performed on RGB cameras, infrared thermal imagers, PIR motion sensors, and acoustic sensors to obtain the anomaly detection result matrix; Quantify the degree of abnormality based on the abnormality detection result matrix to obtain the abnormality score value of each sensor; Compare the abnormality score value of each sensor with the fusion detection threshold parameter to obtain the abnormality level classification result of each sensor; Based on the abnormal level classification results of each sensor, normal sensors are screened to obtain the normal working sensor identification; Generate abnormal sensor isolation instructions based on the abnormal level classification results and confidence weight parameters of each sensor.
[0044] Specifically, based on a synchronized multi-sensor dataset, multi-dimensional anomaly detection is performed for four heterogeneous sensor types. For RGB cameras, detection focuses on image clarity anomaly detection, brightness distribution deviation detection, and noise density anomaly detection. Analysis is performed using methods such as image blurriness metrics (such as gradient variance), brightness histogram uniformity testing, and image noise level estimation. For infrared thermal imagers, detection focuses on temperature distribution uniformity detection, abnormal hot spot detection, and temperature gradient stability assessment. Methods include statistical temperature distribution variance, hot spot clustering anomaly analysis, and gradient change rate fluctuation detection. For PIR motion sensors, detection dimensions include trigger frequency anomaly detection, false trigger rate assessment, and response delay monitoring. Anomaly identification is achieved by statistically analyzing the deviation of trigger frequency from historical baseline data within a period, analyzing the effectiveness and false alarm rate of trigger events, and measuring the time delay from trigger to signal output. For acoustic sensors, the focus is on signal energy anomaly detection, background noise level change detection, and spectral purity analysis. Specifically, anomaly monitoring is performed using methods such as detecting abnormal fluctuations in the acoustic signal energy envelope, background noise signal mean square error analysis, and calculating the signal spectrum sharpness index. Using the multi-dimensional anomaly detection method described above, a complete set of anomaly detection results is established for each sensor. These results are then aggregated into a unified anomaly detection result matrix, where each row corresponds to a sensor, each column to an anomaly detection metric, and each matrix element represents a specific detection score or anomaly determination value. After obtaining the anomaly detection result matrix, the matrix is subjected to an anomaly degree quantification calculation. The core of anomaly degree quantification lies in integrating the results of multiple detection metrics into a single anomaly score, reflecting the current overall health of each sensor. A weighted fusion approach is used for comprehensive evaluation, with weights assigned based on the importance and sensitivity of different anomaly detection metrics. For example, for an RGB camera, the weight for image clarity anomalies is set to 0.4, for noise density anomalies to 0.3, and for brightness deviation anomalies to 0.3. Multiple detection metrics are then fused into a single anomaly score using a weighted average. The weighting coefficients are obtained using historical statistical data or empirical tuning based on the sensor's specific application scenarios. A higher comprehensive anomaly score indicates a greater deviation from normal operating conditions and lower system reliability. This quantification process generates an anomaly degree score for each sensor, constructing an anomaly score vector. The anomaly level score of each sensor is compared with the system's set fusion detection threshold parameters to perform anomaly classification. The fusion detection threshold parameters are derived from the dynamic correction value of the environmental complexity quantification results. They can flexibly adjust the judgment threshold according to the current environmental changes, improving the adaptability and accuracy of classification.Anomaly classification is based on a three-level standard: when the anomaly score is less than 0.5 times the fusion detection threshold, it is considered "normal"; when the anomaly score is between 0.5 and 1 times the fusion detection threshold, it is considered "mildly abnormal"; and when the anomaly score exceeds the fusion detection threshold, it is considered "severely abnormal." This grading mechanism refines the status level of abnormal sensors and enables the development of differentiated handling strategies for different levels. Based on the anomaly level classification results for each sensor, normal sensor screening is performed. All sensors classified as "normal" or "mildly abnormal" are considered normal operating units and included in the list of normal operating sensor identifiers. Sensors labeled "severely abnormal" are excluded from the normal sensor set and are subject to further isolation or replacement. Based on the anomaly level classification results and the confidence weight parameters of each sensor, an isolation instruction for the abnormal sensor is generated. The isolation instruction generation module comprehensively considers the sensor's anomaly level and its confidence weight in the fusion calculation to form a fine-grained isolation decision mechanism. For sensors with severe anomalies and high confidence weights, the system directly issues a complete isolation command, completely eliminating their data stream and preventing high-weight anomaly data from seriously interfering with the overall fusion results. For sensors with mild anomalies or low confidence weights, the system adopts a de-emphasis strategy, appropriately reducing their contribution to the fusion calculation rather than completely isolating them, in order to maintain the system's multi-source redundancy and detection flexibility. The isolation command also includes sensor recheck time settings and status monitoring cycle configurations, ensuring that isolated sensors can be reintegrated into the system detection framework after environmental changes or self-test recovery, enhancing the system's self-healing capabilities and long-term stability.
[0045] In a specific embodiment, the process of executing step S400 may specifically include the following steps: Filter the multi-sensor synchronous data set for valid data according to the normal working sensor identification to obtain the valid sensor data set; Perform weighted fusion calculation based on valid sensor data sets and confidence weight parameters to obtain target detection confidence data; Based on the target detection confidence data, the camera multi-parameter collaborative calculation is performed to obtain the three-dimensional parameter control data; Based on the three-dimensional parameter control data, command encapsulation and timing coordination are performed to obtain camera parameter adjustment commands.
[0046] Specifically, valid data is filtered from the multi-sensor synchronized dataset based on the identifiers of properly functioning sensors. The system reads the list of properly functioning sensor identifiers output by the anomaly detection and screening module and extracts the data streams corresponding to these identifiers from the synchronized dataset to form a new valid sensor dataset. During this screening process, time synchronization is adhered to to ensure that all sensor data matches within the same time window, avoiding fusion errors caused by timing misalignment. The data streams of isolated anomalous sensors are completely removed during this step, ensuring that no abnormal noise or distorted data is contaminated during subsequent processing, thus ensuring high data reliability and consistency from the source. After obtaining a valid sensor dataset, a weighted fusion calculation is performed on the filtered data based on confidence weight parameters. Based on the confidence weight of each sensor, different fusion ratios are assigned to its contribution. During this process, the detection outputs of each sensor are normalized to bring the data outputs to the same scale. For example, this is achieved through normalization or Z-score standardization, eliminating fusion obstacles caused by different physical dimensions of the data. The data from each sensor is weighted and summed according to the confidence weight parameter. Sensors with high confidence levels contribute more to the fusion result, while sensors with low confidence levels contribute less. Through weighted fusion, the system leverages the strengths of different sensors in different environments to maximize the accuracy and robustness of detection results, generating target detection confidence data. This target detection confidence data includes not only a confidence score for the target's presence, but also target location information, category prediction results, and stability indicators for the detection environment, forming a multi-dimensional, structured fusion detection output. Based on the target detection confidence data, multi-parameter collaborative calculations are performed on the camera. Based on the detection confidence information, the camera's imaging parameters are dynamically adjusted to adapt to environmental changes and target motion characteristics. This collaborative calculation covers three core parameters: focus adjustment, exposure compensation, and gimbal angle control. The system uses the target distance and clarity information in the target detection confidence data to perform depth of field analysis, and combines the current optical characteristics of the camera to calculate the optimal focal length adjustment parameters to ensure that the target is always within the optimal focal plane range of the imaging system, avoiding image blur due to distance changes; by analyzing the light stability index and the current ambient brightness change rate in the detection data, exposure compensation calculation is performed, and the camera's exposure time and gain settings are dynamically adjusted to ensure moderate image brightness and clear details, and maintain imaging quality even when lighting conditions change drastically; the system combines target position information with motion trajectory prediction results, and uses prediction algorithms such as Kalman filtering to calculate the target's position at several future moments, thereby determining the horizontal and vertical angle adjustment values of the gimbal, and adjusting the camera's direction in real time to ensure that the target is always in the center of the picture.The focus adjustment parameters, exposure compensation parameters, and pan / tilt angle control parameters output by the three computational submodules are combined into three-dimensional parameter control data through a unified data structure, describing the camera's optimal control strategy for the current inspection task. Instruction encapsulation and timing coordination are performed based on this three-dimensional parameter control data to generate parameter adjustment instructions that can be directly recognized and executed by the camera hardware. The system encapsulates this three-dimensional parameter control data into a standardized instruction format in accordance with the camera control protocol specification. Each instruction includes key information such as parameter category, target value, execution priority, and valid time window, ensuring a rigorous and complete instruction structure. The instruction encapsulation process considers instruction conflicts and resource contention, particularly when adjusting multiple parameters simultaneously, where there is overlap between focus adjustment and pan / tilt rotation. Therefore, the system performs timing coordination, establishing a reasonable execution order and scheduling schedule based on the priorities and dependencies of the various parameter adjustments. Pan / tilt angle adjustment, as it involves target tracking, has a higher priority, followed by exposure compensation due to its significant impact on image quality. Focus adjustment is prioritized as long as it does not affect tracking or exposure. The system ensures physical coordination of multi-parameter adjustments by setting time slices and synchronized triggering mechanisms, preventing imaging fluctuations or target loss caused by out-of-order execution. Camera parameter adjustment instructions are sent to the camera control unit in real time, driving the camera hardware to dynamically adjust to environmental changes and target motion characteristics.
[0047] In a specific embodiment, the execution step of performing camera multi-parameter collaborative calculation based on target detection confidence data to obtain three-dimensional parameter control data may specifically include the following steps: Input the target detection confidence data into the camera parameter adaptive adjustment module for data analysis and feature extraction to obtain the target state comprehensive feature vector; The depth of field analysis and temperature gradient calculation are performed in combination with the comprehensive characteristic vector of the target state to obtain the focus adjustment parameters; Combined with the target state comprehensive feature vector, brightness distribution statistics and exposure feedback calculation are performed to obtain exposure compensation parameters; The comprehensive feature vector of the target state is input into the target trajectory prediction algorithm for Kalman filtering and motion prediction calculation to obtain the gimbal angle control parameters, and the focus adjustment parameters, exposure compensation parameters and gimbal angle control parameters are combined into three-dimensional parameter control data.
[0048] Specifically, the target detection confidence data is input into the camera parameter adaptive adjustment module for data parsing and feature extraction. The camera parameter adaptive adjustment module integrates a high-dimensional feature processing unit, which can parse the sub-features in the detection confidence data and extract key feature quantities that are highly relevant to imaging adjustment, such as the pixel size of the target, the temperature difference between the target and the background, the target center point position coordinates, the image brightness histogram statistics, the image clarity gradient value, etc. These features are normalized and encoded to construct a unified format target state comprehensive feature vector. Based on the target state comprehensive feature vector, depth of field analysis and temperature gradient calculation are performed to derive the camera's focal length adjustment parameters. Depth of field analysis relies on the size and clarity characteristics of the target in the image, combined with the current optical characteristic parameters of the camera, such as focal length range, aperture size and pixel density, and reversely infers the estimated distance from the current target to the camera through the depth of field formula. At the same time, the image clarity index is referred to to determine whether there is a focal plane offset. The system combines target temperature information provided by the infrared thermal imager with analysis of the temperature gradient distribution between the target and the background. By calculating the rate of change of the temperature gradient in the target area, it infers whether the current focus needs to be shifted to ensure clear imaging of the target under thermal imaging characteristics. Through depth of field analysis and temperature gradient optimization, it derives a precise focus adjustment parameter. This parameter guides the camera lens to adjust the focal length, ensuring optimal imaging of the target at varying spatial distances and ambient temperature variations, avoiding out-of-focus blur and thermal imaging out-of-focus issues. After calculating the focus adjustment parameter, the system then performs brightness distribution statistics and exposure feedback calculations based on the integrated feature vector of the same target state to obtain exposure compensation parameters. The system analyzes the image brightness histogram, statistically analyzing the pixel brightness distribution characteristics, including brightness mean, standard deviation, and skewness coefficient, to assess the overall image brightness level and distribution balance. If the brightness histogram is concentrated in the low grayscale area, the image is underexposed; conversely, if it is concentrated in the high grayscale area, it indicates overexposure. The system combines information on the rate of change of external light provided by the ambient light sensor to dynamically estimate the impact of the current lighting environment on image brightness. It then integrates the resulting brightness statistics with real-time light change data and adjusts the exposure compensation value using an adaptive exposure feedback algorithm. This compensation value is reflected in extending or shortening the exposure time, or in some cases, adjusting the gain to achieve brightness balance. The resulting exposure compensation parameters effectively improve image clarity and detail reproduction under various lighting conditions, ensuring the imaging system's excellent visual stability in complex environments. The target state's comprehensive feature vector is input into the target trajectory prediction algorithm module, where Kalman filtering and motion prediction calculations are performed to determine the gimbal angle control parameters. The Kalman filter uses the target's current position information, velocity vector, and acceleration estimate, combined with continuous observation data over a time series, to recursively predict the target's position at several future moments.The Kalman filter effectively removes noise interference during target motion and maintains the continuity and stability of trajectory estimation in the face of occlusion, partial loss, and other issues. Based on the predicted future target position, the system combines the camera's mounting position and motion freedom parameters to calculate the required gimbal horizontal rotation angle (yaw angle) and vertical pitch angle (pitch angle) adjustments, forming the gimbal angle control parameters. The focal length adjustment parameters, exposure compensation parameters, and gimbal angle control parameters are combined into three-dimensional parameter control data. This three-dimensional parameter control data uses a structured format and contains three fields: focal length adjustment value, exposure compensation value, and gimbal angle adjustment value. Each field is accompanied by additional information such as the target value, change rate limit, priority identifier, and effective time window to support command parsing and scheduling by the camera control module during actual execution.
[0049] In a specific embodiment, executing the camera control method based on multi-sensor data fusion further includes the following steps: Obtain image data after the camera parameter adjustment instruction is executed, and perform image quality assessment and feature extraction to obtain image feature vectors; Perform weighted fusion calculation based on image feature vector and target detection confidence data to obtain a fused feature description vector; Perform care target recognition and positioning calculation based on the fused feature description vector to obtain the care target detection result; Based on the detection results of the care target, the accuracy and latency are analyzed to obtain the system performance evaluation indicators; Based on the system performance evaluation indicators, the learning rate, reward function weight and network structure parameters of the deep Q network are adjusted online, and the parameter configuration of the environment adaptive weight distribution strategy is updated to obtain the updated deep Q network.
[0050] Specifically, after the camera completes adjustments according to parameter adjustment instructions, the latest image data is captured in real time. Since this image data reflects the camera's perception of the environment and target after adjustments to focus, exposure, and pan / tilt angle, it provides valuable insights into the effectiveness of camera adjustments. After capture, the image data is input into the image quality assessment module for quality analysis. Evaluation metrics cover multiple dimensions, including image clarity, brightness balance, contrast, noise level, and detail richness. Standard image quality assessment methods such as gradient variance, structural similarity index, signal-to-noise ratio, and brightness histogram mean and variance are used to quantify the overall image quality from multiple perspectives. After image quality assessment is complete, the deep feature extraction module extracts high-dimensional features from the image. This feature extraction is based on a pretrained convolutional neural network. Through forward propagation, the output of the intermediate feature layer is extracted to produce a structured image feature vector. This vector contains low-level information such as texture, edges, and color distribution, as well as high-level semantic features such as target shape, spatial layout, and background complexity, comprehensively reflecting the perceptual characteristics of the current image. A weighted fusion calculation is performed based on the image feature vector and target detection confidence data. During the weighted fusion process, the two feature vectors are scaled and normalized to minimize the impact of feature distribution differences on the fusion effect. Based on the detection confidence, which includes the target presence probability, detection accuracy, and ambient lighting information, an adaptive weighting coefficient is set, giving higher confidence detection results a higher weight in the fusion process and lower confidence detection results a lower weight. A weighted calculation generates a fused feature description vector. This fused feature description vector is used for care target recognition and localization. The recognition module employs a deep learning detection network, such as a modified YOLO or CenterNet architecture, and uses the fused feature vector for end-to-end object classification and bounding box regression. The system performs joint inference based on target category and location features to identify key objects in the current image, such as elderly people, children, and medical equipment. The target's specific location coordinates in the image are determined through regression prediction. During the localization process, a non-maximum suppression algorithm is used to remove redundant detection results, ensuring that each target corresponds to only one optimal detection box, improving detection accuracy and stability. This process outputs a care target detection result, including the target category, location bounding box, detection confidence, and key attribute labels. Based on the results of supervised object detection, accuracy and latency analysis are performed to generate system performance evaluation metrics. Accuracy analysis calculates the accuracy, false positive rate, and missed detection rate based on the number of true positives, false negatives, and false positives detected. By comparing the detection results with a set of manually annotated true labels, the system quantitatively evaluates the recognition performance of the current detection model in real-world environments. Latency analysis focuses on the total time from image acquisition to target detection completion, calculating the execution time of each submodule, such as image preprocessing, feature extraction, fusion calculation, and detection inference, to comprehensively calculate the system's average response time.By periodically compiling these metrics and generating performance curves, the system can dynamically monitor its performance trends and identify potential performance degradation risks. Based on the system performance evaluation metrics, the system enters the online tuning and optimization phase of the Deep Q Network (DQN). Based on accuracy and latency data, the system dynamically adjusts key hyperparameters of the DQN. The learning rate is adaptively adjusted based on accuracy trends. During periods of stable improvement in detection performance, the learning rate is appropriately reduced to converge to a more optimal solution, and during periods of performance degradation, the learning rate is appropriately increased to escape local optima and maintain training activity. The reward function weights are adjusted based on the false alarm rate and missed detection rate, increasing the weight of rewards for correct detections, reducing the lag in penalties for false detections, and improving the model's sensitivity to detection accuracy. If detection performance fluctuates significantly over a long period of time and the response time exceeds a set threshold, the DQN network structure is fine-tuned, including increasing the hidden layer depth, adjusting the number of hidden units, or changing the activation function, to improve the model's expressiveness and ability to adapt to complex environmental changes. After the online tuning of the Deep Q Network is complete, the updated network is redeployed to the environment-adaptive weight allocation strategy module. The updated DQN outputs more accurate and stable sensor weight allocation results based on the latest environmental state input, realizing dynamic optimization of sensor information fusion strategy under environmental changes.
[0051] See also Figure 2 , Figure 2 The structure schematic block diagram of the camera control system 200 based on multi-sensor data fusion provided in the embodiment of the present application is as follows: Figure 2 As shown, the camera control system 200 based on multi-sensor data fusion includes: An acquisition module 210 is configured to acquire a multi-sensor synchronous data set and input the multi-sensor synchronous data set into a deep Q network for sensor performance evaluation to obtain a performance score set; The weight allocation module 220 is used to perform environment adaptive weight allocation based on the performance score set to obtain the confidence weight parameter and fusion detection threshold parameter of each sensor; An abnormal sensor identification module 230 is used to identify abnormal sensors on a multi-sensor synchronous data set according to a confidence weight parameter and a fusion detection threshold parameter, and obtain a normal working sensor identifier and an abnormal sensor isolation instruction; The generation module 240 is used to perform weighted fusion on the multi-sensor synchronous data set based on the normal working sensor identification and confidence weight parameters to obtain target detection confidence data, and generate camera parameter adjustment instructions for the RGB camera according to the target detection confidence data.
[0052] Through the collaborative efforts of these components, real-time calculation of the three-dimensional environmental state vector and environmental adaptability analysis using a deep Q-network, the system dynamically evaluates the performance of each sensor under current environmental conditions based on changes in sensor signal-to-noise ratio, ambient light intensity, and target motion speed. This represents a technological breakthrough, moving from fixed weights to adaptive weight allocation, addressing the poor adaptability of traditional approaches in complex environments. A multi-dimensional anomaly detection mechanism, combined with a hierarchical isolation strategy, enables predictive identification and progressive isolation of sensor anomalies through comprehensive evaluation of data consistency, timing stability, response delay, and output range, effectively mitigating the impact of a single sensor failure on overall system performance and significantly improving system reliability and continuity. By integrating multimodal feature fusion of RGB image features, infrared temperature features, PIR motion features, and acoustic spectrum features, combined with dynamic adjustment of confidence weight parameters, the system achieves intelligent weighted fusion of multi-sensor data, overcoming the technical limitations of traditional single-sensor control methods, which suffer from insufficient accuracy, and the lack of intelligence of traditional multi-sensor fusion methods. By combining a target distance estimation algorithm with binocular vision and infrared temperature gradient calculation, an ambient light analysis algorithm with a closed-loop feedback mechanism, and a target trajectory prediction algorithm with Kalman filtering for collaborative calculation of three-dimensional parameters, we achieve intelligent adaptive adjustment of focal length, exposure, and gimbal angle, resolving the existing problem of poor overall performance caused by the independent operation of each subsystem. A complete closed-loop feedback optimization mechanism, from sensor data input to detection result output, has been established. Through system performance evaluation and online adjustment of the deep Q network parameters using a gradient descent algorithm, we achieve autonomous and continuous improvement of system performance, overcoming the technical limitation of existing technologies where system performance decays over time and cannot be self-repaired. This provides technical support for the long-term stable operation of the intelligent care system.
[0053] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0054] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.
[0055] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A camera control method based on multi-sensor data fusion, characterized in that: include: Collecting a multi-sensor synchronous data set, and inputting the multi-sensor synchronous data set into a deep Q network to perform sensor performance evaluation to obtain a performance score set; Performing environment adaptive weight allocation based on the performance score set to obtain confidence weight parameters and fusion detection threshold parameters for each sensor; Perform abnormal sensor identification on the multi-sensor synchronous data set according to the confidence weight parameter and the fusion detection threshold parameter to obtain a normal working sensor identifier and an abnormal sensor isolation instruction; The multi-sensor synchronous data set is weightedly fused based on the normal working sensor identifier and the confidence weight parameter to obtain target detection confidence data, and a camera parameter adjustment instruction of the RGB camera is generated according to the target detection confidence data.
2. The camera control method based on multi-sensor data fusion according to claim 1, characterized in that: The multi-sensor synchronous data set is collected and input into the deep Q network to perform sensor performance evaluation to obtain a performance score set, including: Perform hardware clock synchronization and software timestamp calibration on the output signals of the RGB camera, infrared thermal imager, PIR motion sensor, and acoustic sensor to obtain a multi-sensor synchronized dataset. Calculating the ratio of the output signal power spectral density and the noise power spectral density of each sensor in the multi-sensor synchronous data set to obtain a sensor comprehensive signal-to-noise ratio value; Performing illuminance quantification calculation based on the image brightness histogram of the RGB camera and the temperature gradient change rate of the infrared thermal imager in the multi-sensor synchronous data set to obtain the current ambient light illuminance value; Performing inter-frame motion analysis on the continuous frame image data in the multi-sensor synchronized data set and performing speed analysis in combination with the trigger frequency data of the PIR motion sensor to obtain a target motion speed value, and combining the sensor comprehensive signal-to-noise ratio value, the current ambient light illumination value, and the target motion speed value into a three-dimensional environmental state vector; The three-dimensional environment state vector is input into a deep Q network to perform sensor performance evaluation to obtain a performance score set.
3. The camera control method based on multi-sensor data fusion according to claim 2, characterized in that: The three-dimensional environment state vector is input into the deep Q network to perform sensor performance evaluation to obtain a performance score set, including: The three-dimensional environmental state vector is input into the deep Q network to perform environmental adaptability analysis to obtain the coefficient matrix of the influence of environmental conditions on the working performance of each sensor; Performing environmental sensitivity calculation based on the influence coefficient matrix to obtain an environmental adaptive feature vector; Based on the environmental adaptation feature vector, the imaging quality of the RGB camera under current lighting conditions, the detection accuracy of the infrared thermal imager under current temperature gradients, the response sensitivity of the PIR motion sensor under current motion speeds, and the recognition accuracy of the acoustic sensor under current signal-to-noise ratio conditions are quantitatively evaluated to obtain the environmental adaptation performance value of each sensor; The environmental adaptability performance values of the sensors are normalized by probability distribution to obtain a performance score set.
4. The camera control method based on multi-sensor data fusion according to claim 3, characterized in that: The three-dimensional environmental state vector is input into the deep Q network to perform environmental adaptability analysis to obtain the influence coefficient matrix of the environmental conditions on the working performance of each sensor, including: Performing a multi-layer perceptron calculation on the three-dimensional environment state vector through a state encoder of a deep Q network to obtain an environment state encoding vector; Performing a state-action value evaluation on the environment state encoding vector by using a Q-value function calculator of the deep Q network to obtain a four-dimensional sensor Q-value vector; The environmental impact quantification module of the deep Q network is used to perform matrix transformation and weight distribution calculation on the Q value vector of the four-dimensional sensor to obtain the influence coefficient matrix of the environmental conditions on the working performance of each sensor.
5. The camera control method based on multi-sensor data fusion according to claim 4, characterized in that: The step of performing environment adaptive weight allocation based on the performance score set to obtain confidence weight parameters and fusion detection threshold parameters for each sensor includes: Inputting the performance score set into the deep deterministic policy gradient algorithm to calculate the environmental adaptability coefficient to obtain the environmental adaptability coefficient data of each sensor; Performing a weighted confidence calculation based on the environmental adaptability coefficient data of each sensor and the performance score set to obtain a confidence weight parameter of each sensor; An environmental complexity quantification calculation is performed based on the three-dimensional environmental state vector to obtain an environmental complexity index value, and a basic detection threshold is corrected based on the environmental complexity index value to obtain a fusion detection threshold parameter.
6. The camera control method based on multi-sensor data fusion according to claim 1, characterized in that: The performing abnormal sensor identification on the multi-sensor synchronous data set according to the confidence weight parameter and the fusion detection threshold parameter to obtain a normal working sensor identifier and an abnormal sensor isolation instruction includes: Perform multi-dimensional anomaly detection on the RGB camera, infrared thermal imager, PIR motion sensor, and acoustic sensor based on the multi-sensor synchronous data set to obtain an anomaly detection result matrix; Quantitatively calculate the degree of abnormality based on the abnormality detection result matrix to obtain a numerical value of the degree of abnormality of each sensor; Compare and judge the abnormality level score value of each sensor with the fusion detection threshold parameter to obtain the abnormality level classification result of each sensor; Screening normal sensors based on the abnormality level classification results of each sensor to obtain normal working sensor identifiers; An abnormal sensor isolation instruction is generated according to the abnormality level classification results of each sensor and the confidence weight parameter.
7. The camera control method based on multi-sensor data fusion according to claim 1, characterized in that: The weighted fusion of the multi-sensor synchronous data set based on the normal working sensor identifier and the confidence weight parameter to obtain target detection confidence data, and generating a camera parameter adjustment instruction for the RGB camera according to the target detection confidence data includes: Filtering the multi-sensor synchronous data set for valid data according to the normal working sensor identifier to obtain a valid sensor data set; Performing a weighted fusion calculation based on the valid sensor data set and the confidence weight parameter to obtain target detection confidence data; Performing camera multi-parameter collaborative calculation based on the target detection confidence data to obtain three-dimensional parameter control data; Instruction packaging and timing coordination are performed based on the three-dimensional parameter control data to obtain camera parameter adjustment instructions.
8. The camera control method based on multi-sensor data fusion according to claim 7, characterized in that: The performing of camera multi-parameter collaborative calculation based on the target detection confidence data to obtain three-dimensional parameter control data includes: Inputting the target detection confidence data into the camera parameter adaptive adjustment module for data analysis and feature extraction to obtain a comprehensive feature vector of the target state; Performing depth of field analysis and temperature gradient calculation based on the target state comprehensive feature vector to obtain a focus adjustment parameter; Performing brightness distribution statistics and exposure feedback calculations based on the target state comprehensive feature vector to obtain exposure compensation parameters; The target state comprehensive feature vector is input into the target trajectory prediction algorithm for Kalman filtering and motion prediction calculation to obtain the gimbal angle control parameters, and the focal length adjustment parameters, the exposure compensation parameters and the gimbal angle control parameters are combined into three-dimensional parameter control data.
9. The camera control method based on multi-sensor data fusion according to claim 1, characterized in that: The camera control method based on multi-sensor data fusion also includes: Obtaining image data after the camera parameter adjustment instruction is executed, and performing image quality assessment and feature extraction to obtain an image feature vector; Performing a weighted fusion calculation based on the image feature vector and the target detection confidence data to obtain a fused feature description vector; Performing care target recognition and positioning calculation based on the fused feature description vector to obtain a care target detection result; Performing accuracy and latency analysis based on the care target detection results to obtain system performance evaluation indicators; Based on the system performance evaluation indicators, the learning rate, reward function weight and network structure parameters of the deep Q network are adjusted online, and the parameter configuration of the environment adaptive weight distribution strategy is updated to obtain an updated deep Q network.
10. A camera control system based on multi-sensor data fusion, characterized in that: The method for executing a camera control method based on multi-sensor data fusion according to any one of claims 1 to 9 comprises: An acquisition module is used to acquire a multi-sensor synchronous data set and input the multi-sensor synchronous data set into a deep Q network for sensor performance evaluation to obtain a performance score set; A weight allocation module, configured to perform environment-adaptive weight allocation based on the performance score set to obtain confidence weight parameters and fusion detection threshold parameters for each sensor; an abnormal sensor identification module, configured to identify abnormal sensors on the multi-sensor synchronous data set according to the confidence weight parameter and the fusion detection threshold parameter, and obtain a normal working sensor identifier and an abnormal sensor isolation instruction; A generation module is used to perform weighted fusion on the multi-sensor synchronous data set based on the normal working sensor identification and the confidence weight parameter to obtain target detection confidence data, and generate a camera parameter adjustment instruction for the RGB camera according to the target detection confidence data.
Citation Information
Cited By
Security protection method and device based on AI, equipment and storage medium
CN121246595A
Newborn image record detection method, system and equipment based on artificial intelligence and medium
CN122224397A