Imaging control method, training method, system and computer readable storage medium
By dynamically adjusting the optical component parameters and imaging control model, and combining optical information and sensor information, the ambiguity problem of the imaging system in low-visibility environments is solved, clear imaging and accurate target recognition are achieved, and the system's intelligent perception and real-time response capabilities are enhanced.
Patent Information
- Application Number
- CN202510824886.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-12
AI Technical Summary
The existing imaging system produces blurred images in low-visibility environments, laser sensing equipment lacks texture and color information, and the fusion of optical imaging and laser sensing data is affected by temperature, resulting in ranging errors and fusion delays, which cannot meet the needs of intelligent perception.
By dynamically adjusting the parameters of optical components, combining imaging control models and dual-branch neural networks, fusing optical information and sensor information, dynamically adjusting weights, optimizing focal length and feature fusion, clear imaging and accurate target recognition are achieved.
Improve the clarity and accuracy of the imaging system under adverse weather conditions, reduce false alarms and missed alarms, and enhance the system's intelligent perception and real-time response capabilities.
Smart Images

Figure CN120640118A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to an imaging control method, a training method, a system, and a computer-readable storage medium. Background Art
[0002] With the advancement of technologies such as security monitoring, autonomous driving, and drone navigation, imaging systems are widely used in these fields. Clear imaging systems can retain more detailed information, enhancing target feature recognition in security monitoring and understanding of road signs and the surrounding environment in autonomous driving. However, the lack of texture and color information in some existing devices significantly limits comprehensive scene perception. Furthermore, in low-visibility environments such as rain and fog (for example, visibility <100 meters), water droplets and particles in the fog and rain scatter and absorb the laser beam, weakening the signal strength and accuracy received by the sensor components. This leads to significant deviations in ranging results, affecting their reliable application in adverse weather conditions. The integration of optical and sensor components in specific scenarios also presents numerous challenges, failing to meet the growing demand for intelligent perception. Summary of the Invention
[0003] In view of this, the purpose of the embodiments of the present application is to provide an imaging control method, a training method, a system and a computer-readable storage medium to improve the problem of dynamic target imaging blur in the prior art.
[0004] In a first aspect, an embodiment of the present application provides an imaging control method, which is applied to an imaging control system, and the method includes: adjusting the optical component based on the motion information and optical information of the dynamic target to obtain a monitoring image; wherein the optical component is used to obtain the optical information; using an imaging control model to analyze the monitoring image of the dynamic target, and obtaining a fused image through the analysis results of the analysis model; wherein the motion information includes the distance of the dynamic target; and the optical information includes the motion image of the dynamic target and the physical parameters of the optical component.
[0005] In the above implementation process, the parameters of the optical components are adjusted based on the motion and optical information of the dynamic target to obtain a monitoring image. A clear monitoring image is a prerequisite for imaging control. Only when the acquired monitoring image retains sufficient information can the subsequent imaging control model analyze the dynamic target and better extract and process the image features. After processing, a fused image is obtained, which can reflect the actual conditions in the specific application scenario, more accurately identify dynamic targets, reduce false positives and false negatives, improve the overall performance of the imaging system, and further provide the ability and possibility of generalized analysis based on monitoring images.
[0006] In one embodiment of the present application, the optical component is adjusted based on the motion information and optical information of the dynamic target to obtain a monitoring image, including: acquiring a first image and a second image; calculating the motion speed of the dynamic target based on the position change of the dynamic target in the first image and the second image; wherein the first image is a motion image of the dynamic target at a first moment, and the second image is a motion image of the dynamic target at a second moment.
[0007] In this implementation, by comparing the pixel position changes of the target object in two adjacent frames and combining them with the time interval between image acquisitions, the target's motion displacement in the image is determined, and the target's motion speed is then calculated. This eliminates the need for dedicated monitoring equipment to measure dynamic targets, and even in the case of a large number of dynamic targets, each target can be processed separately, completing the monitoring function of the imaging system more efficiently.
[0008] In one embodiment of the present application, the optical component is adjusted based on the motion information and optical information of the dynamic target to obtain the monitoring image, and also includes: obtaining the distance of the dynamic target recorded by the sensing component; calculating the adjusted focal length of the optical component through the first formula; wherein the first formula is established based on the distance of the dynamic target, the motion speed of the dynamic target and the physical parameters of the optical component; comparing the adjusted focal length with the current focal length of the optical component, and adjusting the physical parameters of the optical component according to the comparison result to obtain the monitoring image; wherein the physical parameters of the optical component include the focal length.
[0009] In the above implementation, the focal length of the optical component is adjusted based on the distance of the dynamic target, the distance of the dynamic target, and the physical parameters of the optical component. It is easy to understand that adjusting the focus to keep the subject within the focal length will achieve a clear image. Tracking different dynamic targets and adjusting the optical component's ability to capture their images based on their varying distances and speeds can achieve optimal imaging results in different application scenarios, thereby obtaining more scene information.
[0010] In one embodiment of the present application, the monitoring image of the dynamic target is analyzed using an analysis model, and a fused image is obtained through the analysis result of the imaging control model, including: obtaining image information of the monitoring image, the image information of the monitoring image having a first weight; wherein the image information includes the grayscale, thermal radiation information and texture features of the monitoring image; obtaining the motion information obtained by the sensing component, the motion information having a second weight; wherein the motion information includes: the contour and structural information of the dynamic target; according to the image information of the monitoring image and the weight of the motion information obtained by the sensing component, feature fusion is performed based on a second formula; wherein the second formula is established based on the image information, the motion information, the first weight and the second weight.
[0011] In the above implementation process, in order to fully integrate the image information data obtained by the optical component and the motion information obtained by the sensor component, different weights are used to dynamically adjust the contribution ratio of the dynamic target image captured by the optical component and the motion information of the dynamic target obtained by the sensor component in the fused image according to different scenarios and task requirements. This can improve the intelligent perception effect and obtain computer vision images with more complete information and easy to analyze.
[0012] In one embodiment of the present application, the motion information acquired by the sensing component also includes: the acceleration, angular velocity and position of the dynamic target; the method also includes: converting the coordinates of the dynamic target in the sensing component to actual coordinates based on the acceleration, angular velocity and position of the dynamic target; calculating the posture of the dynamic target based on the acceleration of the dynamic target and the angular velocity of the dynamic target; when obtaining the posture of the dynamic target, correcting the short-term posture changes of the dynamic target according to the position of the dynamic target, and predicting the motion trend of the dynamic target.
[0013] In the above implementation process, the acceleration, angular velocity, and position of the dynamic target are collected to provide basic information for calibration. This solves the problem of coordinate system offset caused by installation deviation of sensor components and / or optical components, posture changes, or environmental changes, and provides a spatial consistency foundation for the subsequent fusion of image information data and dynamic target motion information data.
[0014] In a second aspect, an embodiment of the present application also provides an imaging control model training method, wherein the imaging control model training method is configured to execute the imaging control method described above; the imaging control model includes: a dual-branch neural network and a lightweight optimization network; the method includes: extracting image information of the monitoring image of the dynamic target and the motion information of the dynamic target by the dual-branch neural network; wherein the image information includes an image feature vector, and the motion information includes a motion feature vector; the image feature vector passes through an image weight matrix, and the motion feature vector passes through a motion information weight matrix to output a fusion feature vector; when the dual-branch neural network training is completed, the lightweight optimization network analyzes the dual-branch neural network; based on the set importance threshold, the importance score of each channel in the dual-branch neural network is calculated, and the channels are screened.
[0015] In the above implementation, a dual-branch neural network is used to process multimodal data and fuse feature vectors. This is primarily used in model building and training, which falls under the data fusion and model inference component. The dual-branch neural network extracts image information from monitoring images of dynamic targets, as well as motion information from dynamic targets, leveraging its powerful ability to extract image features and its advantages in processing motion information data. By passing the image feature vector through an image weight matrix and the motion feature vector through a motion information weight matrix, a fused feature vector is output. This takes into account scene characteristics (such as visibility and target distance) or dynamic adjustments through reinforcement learning, improving robustness in complex environments and providing the core technical support for dynamic matching and multimodal data fusion.
[0016] In one embodiment of the present application, the dual-branch neural network extracts image features of a monitoring image of a dynamic target, as well as motion information features of the dynamic target, including: adjusting the image features and motion information features during the training process of the dual-branch neural network through a loss function; wherein the loss function adopts a weighted summation method, the image features include image clarity, and the motion information features include predicted motion errors.
[0017] In the above implementation process, the image features and motion information features in the dual-branch neural network training process are adjusted by the loss function, which can jointly optimize the focal length prediction error (MAE) of the optical component and the fused image clarity (SSIM) in subsequent processing.
[0018] In one embodiment of the present application, setting the importance threshold and calculating the importance score of each channel in the dual-branch neural network include: determining the importance threshold according to the requirements of the imaging control training model; wherein the requirements of the imaging control training model include: performance requirements, hardware resource limitations and preliminary experimental results; and adjusting the dual-branch neural network based on the screening of the channels.
[0019] In the above implementation process, the features extracted from some channels may be redundant or contribute little to target recognition. Therefore, these channels need to be screened based on performance requirements, hardware resource limitations, and previous experimental results. This is to solve the problems of large number of parameters, high computational complexity, insufficient computing resources, and slow running speed when directly deployed on embedded devices while maintaining model accuracy.
[0020] In a third aspect, an embodiment of the present application further provides an imaging control system, wherein the imaging control system is configured to execute the imaging control method described above or the imaging control model training method described above; the imaging control system includes: a sensing component, an optical component, a memory, and a processor; the sensing component is configured to collect motion information; wherein the motion information includes the distance of the dynamic target; the optical component is configured to collect optical information; wherein the optical information includes the motion image of the dynamic target and the physical parameters of the optical component; program instructions are stored in the memory, and when the processor runs the program instructions, it executes the steps of any implementation method of the imaging control method described in the first aspect or the imaging control model training method described in the second aspect.
[0021] In this implementation, the synergistic effects of sensors, optical components, memory, and processors enable the system to acquire clearer target images during long-range detection, providing more reliable visual information for applications such as security monitoring and autonomous driving. In dynamic scenarios, such as moving vehicles and people running, the system can clearly capture the target's motion, enhancing the system's ability to monitor and track dynamic targets.
[0022] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the steps of any implementation method of the imaging control method described in the first aspect or the imaging control model training method described in the second aspect are executed. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0024] Figure 1 A first flow chart of the imaging control method provided in an embodiment of the present application; Figure 2 A second flow chart of the imaging control method provided in an embodiment of the present application; Figure 3 A third flow chart of the imaging control method provided in an embodiment of the present application; Figure 4 A fourth flow chart of the imaging control method provided in an embodiment of the present application; Figure 5 A first flow chart of the imaging control training method provided in an embodiment of the present application; Figure 6 A second flow chart of the imaging control training method provided in an embodiment of the present application; Figure 7 A third flow chart of the imaging control training method provided in an embodiment of the present application; Figure 8 A schematic diagram of the imaging control training method system provided in an embodiment of the present application.
[0025] Icons: 10-electronic device; 11-memory; 12-storage controller; 13-processor; 14-peripheral interface; 15-input and output unit; 16-display unit. DETAILED DESCRIPTION
[0026] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of them. Based on the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the embodiments of the present application.
[0027] With the development of technologies such as security monitoring, autonomous driving and drone navigation, optical imaging technology and laser sensing technology, as important means of perception, each play a unique role, but also have significant limitations.
[0028] Optical imaging technology suffers from reduced resolution at long distances. When the target distance exceeds 50 meters, the modulation transfer function (MTF) falls below 0.3. MTF is a key metric that measures an imaging system's ability to transmit signals of varying spatial frequencies. Lower MTF values impair image clarity and detail reproduction, meaning that the outlines and features of distant targets are difficult to clearly visualize. Furthermore, the fixed focal length design results in insufficient depth of field, making it impossible to simultaneously and clearly image targets at varying distances, limiting the application of optical imaging in complex scenarios. This problem is particularly acute when dealing with dynamic targets. When the target speed exceeds 5 m / s, the peak signal-to-noise ratio (PSNR) of the image drops by at least 15 dB. PSNR measures image quality; higher values indicate reduced noise and higher clarity. A significant decrease in PSNR indicates that dynamic targets will appear severely blurred in infrared images, making it difficult to accurately identify and track their motion.
[0029] Laser sensing devices primarily acquire target distance information by emitting and receiving laser beams, generating point cloud depth data. However, this data format only reflects the spatial location of the target and cannot provide rich visual information, such as the target's texture and color. In many application scenarios, such as identifying target features in security monitoring and understanding road signs and the surrounding environment in autonomous driving, the lack of texture and color information greatly limits the overall perception of the scene. In low-visibility environments such as rain and fog (for example, visibility <100 meters), the ranging error of lidar exceeds 10%. Water droplets and particles in rain and fog scatter and absorb the laser beam, weakening the signal strength and reducing the accuracy of the signal received by the laser sensing device. This results in significant deviations in the ranging results, affecting its reliable application in adverse weather conditions.
[0030] Furthermore, when optical imaging data is fused with laser sensor data, the traditional calibration of the fusion of optical imaging data and laser sensor data is significantly affected by temperature. When the temperature fluctuates by ±20°C, the focal length offset is ≥5%. Inaccurate focal length will lead to increased registration errors between the image and point cloud data, affecting the fusion effect and, in turn, reducing the system's perception accuracy of the target. At the same time, due to differences in data acquisition frequency and processing time between optical imaging equipment and laser sensors, the fusion delay exceeds 50ms. In application scenarios with extremely high real-time requirements, such as autonomous driving, a delay of 50ms may cause the vehicle to miss the best decision-making opportunity, posing a safety hazard.
[0031] First, the embodiments of the present application provide an imaging control method that is applied to fields such as security monitoring, autonomous driving, and drone navigation. It can provide strong support for the imaging control system to more accurately analyze and process data, thereby improving the performance and accuracy of the imaging control system.
[0032] See also Figure 1 , Figure 1 This is a first flow chart of the imaging control method provided in an embodiment of the present application.
[0033] Methods include: S110: Based on the motion information and optical information of the dynamic target, adjust the optical component to obtain a monitoring image; wherein the optical component is used to obtain optical information; wherein the optical information includes the motion image of the dynamic target and physical parameters of the optical component.
[0034] It's understandable that the resolution of optical components is a key parameter for measuring image clarity. High resolution (such as 4K) provides clearer images and can capture more details, such as facial features and license plate numbers, which is particularly important for security monitoring. The low-light sensitivity parameter determines the device's imaging capabilities in low-light environments. The lower the value, the better the optical components perform in low light, which is very important for nighttime monitoring. The lens focal length parameter affects the monitoring coverage and detail capture capabilities. Short-focal-length lenses are suitable for wide-angle monitoring, but details may not be clear enough; long-focal-length lenses can focus on distant details, but the field of view is limited. Choosing a reasonable focal length can meet different target acquisition requirements.
[0035] In this case, adjusting parameters such as image quality, exposure, and white balance can improve the clarity of the captured image. Furthermore, when operating in different environments, it may be necessary to adjust the parameters of the optical components to adapt to environmental changes. In other words, proper parameter settings can expand the imaging range and quality, reduce blind spots, and improve the reliability of the imaging system.
[0036] Optionally, the optical component may be an infrared thermal imager, a night vision imaging device, or a panoramic camera, an intelligent tracking camera, a waterproof camera, or other device that can be used to monitor the surrounding environment.
[0037] S120: Analyze the monitoring image of the dynamic target using the imaging control model, and obtain a fused image based on the analysis results of the analysis model; wherein the motion information includes the distance of the dynamic target; and the optical information includes the motion image of the dynamic target and the physical parameters of the optical components.
[0038] Understandably, a single sensor often cannot capture all of a scene's information. Fusion images combine the strengths of each source image, compensating for individual deficiencies and providing a more comprehensive description of the scene. Fusion significantly improves detail clarity, contrast, and visual quality, producing images more consistent with human visual perception and facilitating manual analysis or subsequent computer vision processing. In situations where partial sensor failure or dynamic changes occur (such as occlusion or extreme weather), fused images provide backup information, ensuring data integrity and system stability.
[0039] For the specific process of imaging control system to obtain monitoring images by adjusting optical components based on the motion information and optical information of dynamic targets, please refer to Figure 2 , Figure 2 This is a second flow chart of the imaging control method provided in an embodiment of the present application.
[0040] In one embodiment of the present application, adjusting the optical components based on the motion information and optical information of the dynamic target to obtain a monitoring image includes: S111: Acquire a first image and a second image; S112: Calculating a motion speed of the dynamic target based on a position change of the dynamic target in the first image and the second image; wherein the first image is a motion image of the dynamic target at a first moment, and the second image is a motion image of the dynamic target at a second moment.
[0041] The above process uses continuous frame differencing to calculate the motion speed of dynamic targets. Continuous frame differencing is a commonly used target motion detection method that calculates the target's motion speed by comparing the position changes of the target in two adjacent frames.
[0042] Optionally, the displacement vector of a dynamic target or optical component can be calculated by comparing the position changes between two or more consecutive frames of point cloud data, thereby inferring velocity. After aligning the point clouds of adjacent frames and establishing a corresponding relationship, the rigid body transformation (rotation and translation matrices) is calculated based on the registration results. The displacement components are extracted and, combined with the time interval, the instantaneous velocity is obtained by differential displacement.
[0043] Alternatively, the optical flow method can be used to calculate the motion speed of dynamic objects. This assumes that the grayscale value of the same pixel in adjacent frames remains unchanged and that the pixel motion in the local area is similar. The motion is estimated using the least squares method within the local area of the image, and a system of linear equations is solved to calculate the motion vector.
[0044] In one embodiment of the present application, adjusting the optical components based on the motion information and optical information of the dynamic target to obtain the monitoring image further includes: S113: Acquire the distance of the dynamic target recorded by the sensor component; S114: Calculating an adjusted focal length of the optical component using a first formula; wherein the first formula is established based on a distance to the dynamic target, a moving speed of the dynamic target, and physical parameters of the optical component; S115: Compare the adjusted focal length with the current focal length of the optical component, and adjust the physical parameters of the optical component according to the comparison result to obtain a monitoring image; wherein the physical parameters of the optical component include the focal length.
[0045] It is understood that to ensure a clear image of the observed object on the imaging plane, the position of the image will change as the distance between the object and the lens of the optical assembly changes. By adjusting the distance between the lens of the optical assembly and the imaging plane, the image of the object or dynamic target formed through the lens can be ensured to fall precisely on the photosensitive element or receiving plane, thereby obtaining a clear image.
[0046] Optionally, the sensing component may be a lidar, a millimeter-wave radar, a camera combined with additional sensors to form a multi-sensor fusion system, or an ultrasonic monitoring component may be used.
[0047] In the above implementation process, through in-depth research on the relationship between target distance, movement speed and focal length, the first formula, namely the dynamic focal length matching formula, was established:
[0048] Indicates the dynamically calculated optimal focal length (unit: mm). It is dynamically determined based on parameters such as the real-time ranging value of the sensor component and the target movement speed, enabling the optical component to achieve the best imaging effect in different scenarios.
[0049] The sensor component continuously measures the distance to the target and provides real-time distance information for focal length adjustment.
[0050] is the target's velocity (in m / s), calculated by successive frame differencing. Frame differencing is a commonly used method for detecting target motion. It calculates the target's velocity by comparing the target's position between two consecutive frames.
[0051] The motion blur suppression coefficient, determined through extensive experimental calibration, ranges from 0.2 to 0.5. This coefficient is used to balance the effects of target distance and motion speed on focal length, effectively suppressing blur when imaging moving targets.
[0052] The calculated optimal focal length The actual focal length of the optical component If the two are not equal, it means that the focus needs to be adjusted, and the system will calculate the focus adjustment amount. ,Right now .
[0053] After completing the focus adjustment and update, the imaging control system will determine whether to continue focusing based on the set conditions. For example, the imaging control system may be set to recalculate and adjust the focus at regular intervals, or trigger refocusing when a large change in the distance or motion state of the target object in the scene is detected. If continued focusing is required, the imaging control system will return to the step where the sensor component collects distance data and start the next round of dynamic focusing operation; if continued focusing is not required, the entire dynamic focusing process ends. Such a dynamic focusing control process can adjust the focal length in real time according to changes in the actual scene, ensuring that the imaging control system always obtains clear images and effectively improves imaging quality.
[0054] Optionally, a temperature drift compensation circuit can be designed to account for the impact of temperature changes on the optical system and motor performance. With a temperature sensitivity of ≤0.001% / °C, it monitors ambient temperature changes in real time and compensates for focal length adjustments accordingly, effectively reducing focal length shifts caused by temperature fluctuations and improving system reliability and stability. For dynamic scenes, timely data fusion can keep pace with rapid changes in the target.
[0055] In one embodiment of the present application, the monitoring image of the dynamic target is analyzed using the analysis model, and the specific process of obtaining the fused image through the analysis results of the imaging control model is shown in FIG. Figure 3 , Figure 3 This is a third flow chart of the imaging control method provided in an embodiment of the present application.
[0056] S121: Acquire image information of a monitoring image, where the image information of the monitoring image has a first weight; wherein the image information includes grayscale, thermal radiation information, and texture features of the monitoring image; S122: Acquire motion information acquired by the sensing component, where the motion information has a second weight; wherein the motion information includes: contour and structural information of the dynamic target; S123: Perform feature fusion based on a second formula according to the weights of the image information of the monitoring image and the motion information obtained by the sensor component; wherein the second formula is established based on the image information, the motion information, the first weight, and the second weight.
[0057] In the above implementation process, in order to fully integrate the advantages of infrared images and lidar point cloud data, the contribution ratio of the image features of the optical components and the motion information of the sensor components in the fused image can be dynamically adjusted according to different scenarios and task requirements.
[0058] In this embodiment, an infrared camera is used as an optical component and a lidar is used as a sensing component. The contribution ratio of the edge features of the infrared image and the lidar point cloud in the fused image is dynamically adjusted, and a second formula, namely the multimodal feature fusion formula, is proposed:
[0059] in : Dynamic weight coefficient (satisfying ). Represents the grayscale value of the infrared image, which contains the thermal radiation information and certain texture features of the target. It is the edge feature of the lidar point cloud, reflecting the contour and structural information of the target.
[0060] By processing the laser radar point cloud data with edge detection algorithms, these edge features are extracted. and ( ) can be achieved through a dual strategy of scene feature perception and task goal driving. The specific implementation details are as follows: In terms of scene feature perception, when adjusting according to environmental visibility, in high visibility scenes (such as sunny days, visibility > 500 meters), The weight is increased to 0.6-0.8, giving priority to strengthening the edge features of the lidar point cloud, because its ranging error is small and the edge of the point cloud can accurately reflect the target geometry; in low visibility scenes (such as fog and night, for example, visibility <100 meters), The weight is increased to 0.7-0.9, highlighting the thermal characteristics of infrared images and using their ability to penetrate rain and fog to compensate for the sparse point cloud problem caused by scattering of LiDAR. When adjusting according to target distance and motion status, close-range static targets (such as monitoring distance <30 meters, <1m / s)Use ( :β=0.5:0.5), balance texture and structure information; long-range dynamic targets (such as monitoring distance 50-200 meters, >5m / s), Increases linearly with distance (e.g. at a distance of 100 meters =0.7, 200 m =0.9 to cope with the problem of reduced resolution of long-distance infrared images. When adjusting according to scene complexity, simple background scenes (such as empty parking lots) use =0.4,β=0.6, quickly locate the target outline; complex occlusion scenes (such as shopping malls) use =0.8, =0.2, using the thermal imaging characteristics of infrared images to penetrate partial obstructions.
[0061] In terms of task goal driving, weights are dynamically generated through a dual-branch neural network. A weight generation subnetwork is added to the dual-branch network, and two feature vectors are input. and , the output is and (satisfy + =1), optimize weights according to classification or detection tasks during training, and generate weights by forward propagation per frame during real-time inference, which takes ≤10ms; dynamic tuning using reinforcement learning is used to define the state including environment, target and evaluation indicators, so as to discretize For actions (such as 0.1, 0.3, ..., 0.9), detection confidence or tracking error is used as the reward function, and offline training generates a policy table for the embedded system to look up and obtain the optimal weight combination.
[0062] In engineering implementation, the scene mode switch is set to switch through hardware trigger (such as the automatic driving mode default =0.7, security mode default =0.6, integrated environmental sensors monitor parameters in real time, update weights every 100ms and trigger fusion calculations synchronously, and the monitoring terminal visualizes the weight ratio and historical change curve. In typical scenarios, such as daytime driving on highways ( =0.3, =0.7, preset rules + lidar priority), nighttime city street security ( =0.8, =0.2, automatically generated by deep learning models), etc., all use multi-dimensional mechanisms to balance the advantages of infrared and lidar to improve intelligent perception effects.
[0063] Before adjusting the optical components and obtaining monitoring images based on the motion and optical information of dynamic targets, there are some details on data acquisition. Figure 4 , Figure 4 This is the fourth flow chart of the imaging control method provided in an embodiment of the present application.
[0064] In one embodiment of the present application, the motion information acquired by the sensing component further includes: acceleration, angular velocity, and position of the dynamic target; and the method further includes: S101: Converting the coordinates of the dynamic target in the sensor component to actual coordinates based on the acceleration, angular velocity and position of the dynamic target; S102: Calculating the posture of the dynamic target based on the acceleration and angular velocity of the dynamic target; S103: When the posture of the dynamic target is acquired, the short-term posture change of the dynamic target is corrected according to the position of the dynamic target, and the movement trend of the dynamic target is predicted.
[0065] In this implementation, data registration and calibration are essential, as even slight deviations in the sensor assembly's installation position and attitude can affect data accuracy. Based on the dynamic target's acceleration, angular velocity, and location, the extended Kalman filter algorithm calculates the LiDAR's attitude information (pitch, yaw, and roll). Rotation matrix and translation vector operations are used to calculate the transformation relationship (homogeneous transformation matrix) between the sensor assembly coordinate system and the global coordinate system, converting the dynamic target's coordinates in the sensor assembly to their actual coordinates.
[0066] In a second aspect, an embodiment of the present application further provides an imaging control model training method, wherein the imaging control model training method is configured to execute the above-mentioned imaging control method.
[0067] See also Figure 5 , Figure 5 This is a first flow chart of the imaging control training method provided in an embodiment of the present application.
[0068] The imaging control model includes: a dual-branch neural network and a lightweight optimization network; the methods include: S210: extracting image information of the monitoring image of the dynamic target and motion information of the dynamic target by a dual-branch neural network; wherein the image information includes an image feature vector, and the motion information includes a motion feature vector; S220: Pass the image feature vector through the image weight matrix and the motion feature vector through the motion information weight matrix to output a fused feature vector; S230: When the dual-branch neural network training is completed, the lightweight optimization network analyzes the dual-branch neural network; S240: Based on the set importance threshold, calculate the importance score of each channel in the two-branch neural network and screen the channels.
[0069] In the above implementation, taking an infrared camera as the optical component and a lidar as the sensor component, the system can be composed of a CNN (convolutional neural network) for infrared and a PointNet for lidar. These branches are used to process multimodal data and fuse feature vectors. These branches are primarily used in model building and training, which falls under the data fusion and model inference components. Parallel CNN and PointNet branches process infrared images and lidar point cloud data, respectively, ultimately fusing feature vectors to achieve tasks such as object detection and focal length prediction. In its original state, a trained two-branch neural network typically has a large number of parameters and high computational complexity. Direct deployment on embedded devices can lead to problems such as insufficient computing resources and slow execution. Once the two-branch neural network is trained, a lightweight optimization network analyzes it and selects channels based on their contribution to the final output.
[0070] Specifically, the image feature vector passes through the image weight matrix, and the motion feature vector passes through the motion information weight matrix. The output layer formula for the output fusion feature vector is:
[0071] in, and are the weight matrices of the infrared branch and the lidar branch respectively, and is the feature vector extracted by each branch network.
[0072] It's easy to understand that before the dual-branch neural network can extract image information and motion information from monitoring images of dynamic targets, data must be collected and labeled. Data collection and labeling are crucial steps, directly impacting the effectiveness and performance of model training.
[0073] For data collection, a comprehensive data collection scope is determined based on the diversity of application scenarios, such as security monitoring, autonomous driving, and drone navigation. For autonomous driving, for example, this encompasses diverse road conditions, including highways, city streets, and rural roads; varying weather conditions, including sunny, rainy, and foggy days; and varying lighting environments, including daytime and nighttime. Based on the characteristics of the data collection scenario, the appropriate platform is selected and the corresponding data collection equipment deployed.
[0074] In the above implementation process, since intelligent monitoring scenarios are dynamic, the characteristics of noise and outliers may change over time and with environmental changes. Therefore, real-time monitoring of data changes is necessary. This can be achieved by analyzing statistical features: calculating statistics such as the mean, variance, and standard deviation. The mean reflects the average level of the data, while the variance and standard deviation measure the degree of dispersion. Data visualization can also be used: plotting the point cloud data collected by the lidar into scatter plots, line graphs, and other tools. Scatter plots provide a visual overview of the data distribution. If data points suddenly become dispersed, deviate from the normal distribution area, or isolated outliers appear, this indicates a data change. Alternatively, model-based prediction comparison can be used: a data prediction model, such as a time series model, is established to predict current data based on historical data. The predicted values are compared with the actual collected data. If the difference exceeds the set error range, this indicates an abnormal data change. The filtering algorithm parameters can be adjusted accordingly based on the actual situation. When a new interference source appears in the monitoring area, such as a high-power electromagnetic device being turned on nearby, causing the lidar data noise to increase significantly, the standard deviation of the Gaussian filter can be appropriately increased to enhance the noise suppression capability; and when the environmental interference decreases, the standard deviation can be appropriately reduced to retain more target detail information, thereby always maintaining the effectiveness of the filtering processing and ensuring the accuracy of the depth information.
[0075] In one embodiment of the present application, when an infrared camera is used as an optical component and a laser radar is used as a sensing component, the laser radar may be affected by the reflected light of the surrounding environment, electromagnetic interference, etc., resulting in noise and abnormal data. Therefore, it is necessary to pre-process the data. At this time, a statistical filtering algorithm is used, such as setting the standard deviation of the Gaussian filter to 0.1 meters, which can smooth the collected point cloud data. In intelligent monitoring scenarios, the laser radar is easily affected by the reflected light of the surrounding environment, electromagnetic interference, etc., resulting in noise and abnormal data. Gaussian filtering is based on statistical principles and makes the data smoother and effectively removes noise by performing weighted averaging on the values within the data neighborhood. In outdoor monitoring scenarios, there may be interference factors such as swaying leaves and changes in light and shadow. After filtering, the distance jump points caused by these interferences can be removed, making the acquired scene depth information more stable and accurate, ensuring that the monitoring system can accurately identify the real target object and reduce the occurrence of false alarms. The standard deviation determines the shape of the Gaussian function. A smaller standard deviation means less smoothing of the data, preserving more detail but with less noise removal. A larger standard deviation results in smoother data but may lose some detail. In this scenario, a standard deviation of 0.1 meter demonstrates good adaptability across a variety of common scenarios. It effectively removes noise and abnormal data caused by environmental interference, making the acquired scene depth information more stable and accurate. It also preserves the details of the target object to a certain extent, ensuring that the monitoring system can accurately identify the target object and reducing the occurrence of false alarms.
[0076] Regarding data logging, a comprehensive data storage system should be established, storing data in specific formats. Common formats include HDF5 and TFRecord. HDF5 (Hierarchical Data Format 5) is suitable for storing large amounts of complex, heterogeneous data. It organizes data in a hierarchical manner, forming a file system-like structure with the concepts of datasets and groups. TFRecord is the binary file format used for storing data in the TensorFlow framework. It serializes data into Example Protocol Buffers (protobufs) and stores multiple examples in a single file. This format offers the advantage of tight integration with the TensorFlow ecosystem. During model training, data can be efficiently read through TensorFlow's data reading API. It also supports data parallel processing in distributed training scenarios, helping to improve training efficiency. In practical applications, the appropriate format will be selected based on data characteristics, intended use, and processing tools. Custom formats can also be designed to meet unique system requirements.
[0077] In the data labeling section, the specific content of the labeling task should be clarified based on the application scenario and model training requirements. In security monitoring scenarios, labeling tasks may include identifying the category, location, and bounding box of target objects (such as people, vehicles, suspicious items, etc.), as well as labeling target behaviors (such as walking, running, staying, etc.). To ensure the consistency and accuracy of labeling, detailed labeling specifications should be formulated. For target objects of different categories, unified labeling standards should be stipulated, such as the method for determining the object boundary and the rules for drawing the labeling box. When labeling pedestrians, the minimum circumscribed rectangular labeling box is drawn with the outer contour of the pedestrian as the boundary, and the coordinate representation of the labeling box is clarified. For overlapping targets or partially occluded targets in complex scenes, special labeling rules are formulated to guide labelers to perform accurate labeling. Semi-automatic labeling not only improves labeling efficiency, but also provides a reference for manual labeling, reducing the workload of manual labeling.
[0078] In one embodiment of the present application, using an infrared camera as the optical component and a lidar as the sensing component, a convolutional neural network (CNN)-based target detection algorithm is used to detect targets in infrared images and fused images, automatically generating annotation boxes and category predictions for target objects. For lidar point cloud data, a point cloud segmentation algorithm is used to identify point cloud clusters of different objects and perform preliminary annotation. In the point cloud data annotation of autonomous driving scenarios, the algorithm can automatically identify point clouds of objects such as vehicles, pedestrians, and road facilities, and mark the corresponding categories.
[0079] In one embodiment of the present application, although semi-automatic labeling can quickly generate labeling results, manual review and correction are still required to ensure the accuracy of the labeling. Manual labelers carefully check the results of semi-automatic labeling according to the labeling specifications. For targets that are incorrectly labeled by the algorithm, such as misidentifying a vehicle as a pedestrian, the labeler manually corrects the error; for targets that are incompletely labeled, such as partially obscured objects, the labeler completes the labeling box based on experience and image details. In complex scenes, manual labelers use their professional knowledge and judgment to accurately label unclear or difficult-to-identify areas. When labeling targets in night infrared images, the labeler needs to combine point cloud data and surrounding environment information to determine the true category and location of the target. During the manual review and correction process, the labeler also checks the consistency of the labeling to ensure that the same type of targets are labeled consistently in different images. A labeling quality assessment mechanism is established to quantitatively evaluate the quality of the labeled data. The degree of match between the labeling results and the actual situation is evaluated by calculating indicators such as the accuracy, recall rate, and F1 value of the labeling.
[0080] See also Figure 6 , Figure 6 This is a second flow chart of the imaging control training method provided in an embodiment of the present application.
[0081] In one embodiment of the present application, a dual-branch neural network extracts image information of a monitoring image of a dynamic target and motion information of the dynamic target, including: S211: Adjusting the image features and motion information features during the dual-branch neural network training process through a loss function; wherein the loss function adopts a weighted summation method, the image features include image clarity, and the motion information features include predicted motion errors.
[0082] In the above implementation, the loss function jointly optimizes the focal length prediction error (MAE, mean absolute error) and image clarity (SSIM, structural similarity index), and the formula is:
[0083] in, and is a weight coefficient used to balance the importance of focal length prediction error and image clarity in the loss function. and They are the predicted focal length and the real focal length, and are the predicted image and the real image.
[0084] Optionally, this process is trained on an NVIDIA A100 GPU with a batch size of 32 and an initial learning rate of 1e-4.
[0085] Specifically, when an infrared camera is used as an optical component, a lidar is used as a sensing component, or in similar situations, the input of the infrared branch is a long-wave infrared image (wavelength range 8-14μm, resolution dynamically adjusted with the scene), and its core structure includes multiple layers of two-dimensional convolutional layers with padding (such as 3×3 convolution kernels, the first layer in the example configuration is 64 3×3 convolution kernels, step size 1, activation function ReLU, followed by BatchNorm and maximum pooling; the second layer is 128 3×3 convolution kernels, step size 1, activation function ReLU, followed by BatchNorm and maximum pooling), and gradually extracts spatial features such as edges, textures, thermal radiation distribution, and high-level semantic features (such as target contours and structures) through layer deepening. The end is connected to 1-2 layers of fully connected layers to flatten the feature map output by the convolution layer into a one-dimensional feature vector. (e.g. dimension is 1024).
[0086] The input of the lidar branch is lidar point cloud data (three-dimensional coordinates (X, Y, Z) and intensity information, in the format of N×D). Its core structure first uses a multi-layer perceptron (MLP) to extract point-by-point features for each point. The formula is (in are the point cloud coordinates and intensity, is the intermediate feature, is a point-by-point feature vector), and then the global features are aggregated through Max Pooling. The formula is To extract the global geometric features of the point cloud (such as target shape, edge structure), and finally map the global features into a one-dimensional feature vector through the fully connected layer (e.g. dimension is 1024).
[0087] In the feature fusion and output layer, the feature vectors of the two branches are linearly combined through the weight matrix by weighted summation. The formula is: (in and is a learnable weight matrix, Softmax is used for classification tasks, and is replaced by a linear layer for regression tasks). The loss function jointly optimizes the focal length prediction error (MAE) and image clarity (SSIM), and the formula is
[0088] This structure has the advantages of multimodal complementarity. CNN is good at extracting texture and thermal features of infrared images, and PointNet is good at processing the geometric structure of lidar point clouds. The combination of the two can make up for the shortcomings of a single sensor. Through dual-branch parallel computing and hardware acceleration (such as NVIDIA A100 GPU), an inference delay of ≤15ms can be achieved to meet real-time requirements. The weight matrix can be dynamically adjusted through scene features (such as visibility, target distance) or reinforcement learning to improve robustness in complex environments. It is the core technical support for realizing dynamic focal length matching and multimodal data fusion.
[0089] See also Figure 7 , Figure 7 This is the third flow chart of the imaging control training method provided in an embodiment of the present application.
[0090] In one embodiment of the present application, an importance threshold is set and an importance score of each channel in a two-branch neural network is calculated, including: S241: Determine an importance threshold based on requirements of the imaging control training model; wherein the requirements of the imaging control training model include: performance requirements, hardware resource limitations, and previous experimental results; In the implementation described above, an appropriate threshold is determined based on the model's performance requirements, hardware resource limitations, and previous experimental results. Setting the threshold is a key step in channel pruning, as it determines which channels are removed. If the threshold is set too high, too many important channels may be removed, significantly reducing model accuracy. If the threshold is set too low, the model's computational and parameter requirements may not be effectively reduced. Through multiple experiments, the accuracy, computational complexity, and parameter requirements of the model can be evaluated at different thresholds. By observing the performance changes on key metrics, a threshold can be selected that strikes a good balance between reducing parameters and maintaining accuracy. The importance score of each channel is calculated, and an appropriate threshold is set. Channels with scores below the threshold are removed. This significantly reduces the model's computational and parameter requirements without significantly compromising model accuracy.
[0091] S242: Adjust the dual-branch neural network based on the channel screening situation.
[0092] In the above implementation, the calculated importance score for each channel is compared with a set threshold, and channels with scores below the threshold are removed from the model. When removing channels, the model structure and parameters need to be adjusted to ensure correctness in subsequent computations. For convolutional layers, removing channels will also change the number of input channels in subsequent layers, requiring readjustment of the weight parameters of subsequent layers to ensure the integrity of the model's computational logic. After removing channels, the model is fine-tuned. Retraining the model using the previously prepared dataset allows the model to adapt to the structural changes caused by channel pruning and restore and optimize model performance. During fine-tuning, hyperparameters such as the learning rate and number of training epochs are adjusted to achieve optimal training results. After fine-tuning, a comprehensive model evaluation is conducted using a test dataset to examine the model's performance on key metrics (such as precision, recall, and F1 score) to verify that the model meets the goal of significantly reducing the model's computational complexity and parameter count without significantly compromising model accuracy. If model performance does not meet expectations, adjustments and optimizations should be made to the threshold setting and pruning algorithm selection, and the above steps should be repeated until the model performance meets the requirements. This significantly reduces the amount of computation and parameters required for the model without significantly compromising model accuracy. In practice, careful channel pruning can reduce the number of model parameters by 30% to 50%, while keeping the performance loss on key metrics below 5%.
[0093] Specifically, taking the case where the optical component uses an infrared camera and the sensor component uses a lidar, we first collect and organize the data sets used for model training and evaluation, covering infrared image data in various scenarios to ensure the diversity and representativeness of the data. Load the trained two-branch neural network model, in which the infrared branch contains a convolutional layer, and the model has complete parameter and structural information, providing a basis for subsequent channel pruning operations. Use task-driven structured channel pruning (TSP-CP) to analyze each channel in the convolutional layer of the infrared branch. In terms of channel importance scoring, the weighted norm of the fusion task loss is used, and the formula is ,in and They are the infrared image focus sensitive area and the laser radar motion speed sensitive point cloud, is the correlation coefficient between the channel output and the task area; in the dynamic adjustment of the pruning threshold, the pruning threshold table is pre-stored according to the scene label (such as "static long distance", "dynamic close distance", "rain and fog"), and the D and Adjust the threshold using a table lookup and fine-tune with the temperature drift compensation coefficient. For post-pruning fine-tuning, use a joint task loss function to enforce focus prediction accuracy and image clarity. Determine the appropriate threshold based on the model's performance requirements, hardware resource limitations, and preliminary experimental results.
[0094] Quantization techniques, while reducing data storage and computational complexity, can also result in a certain loss of accuracy. To balance accuracy and computational efficiency, careful quantization parameter adjustment and calibration are required. This can be achieved by optimizing quantization strategies, such as dynamic quantization and mixed-precision quantization. For example, dynamic quantization settings can be tailored to the model structure and data characteristics, selecting an appropriate quantization granularity.
[0095] In one embodiment of the present application, quantization can be performed on a per-layer basis for different layers in a neural network, or on a per-channel or per-neuron basis within each layer. For convolutional layers, per-channel quantization can more carefully account for differences in the distribution of data across different channels; while quantizing fully connected layers on a per-neuron basis can potentially reduce quantization overhead while maintaining accuracy.
[0096] In another embodiment of the present application, dynamic quantization adjusts the quantization parameters according to the dynamic range of the data. In actual settings, it is necessary to determine how to dynamically track the range of the data. A common method is to count the maximum and minimum values of the data through a sliding window, and the choice of window size is critical. If the window is too large, it may not be able to adapt to data changes in time; if the window is too small, the statistical results may be unstable. For example, when processing continuous image data, a time window of appropriate size can be set to update the dynamic range of the data in real time to adjust the quantization parameters and ensure that the quantized values can better retain the data characteristics. Dynamic quantization requires timely updating of quantization parameters to adapt to data changes. This requires determining the update frequency. If the update is too frequent, the computational burden will increase; if the update frequency is too low, data changes cannot be tracked in time, affecting the quantization effect.
[0097] In an additional embodiment of the present application, in actual applications, the update frequency can be adjusted based on the rate of change of the data. During periods of relatively stable data changes, the update frequency can be appropriately reduced; during periods of drastic data changes, the update frequency can be increased. For example, in an autonomous driving scenario, when the vehicle is traveling on a stable road, the update frequency can be set lower; when encountering complex road conditions, such as intersections with dense traffic and frequent movement, the update frequency can be increased.
[0098] Optionally, evaluating changes in model inference accuracy is a key step in ensuring model performance. Use a standard test dataset: Select a representative standard test dataset, such as the ImageNet dataset, commonly used in image recognition, or the COCO dataset, used in object detection. Use the same test dataset before and after quantization to perform inference tests on the model. By comparing the model's prediction results on the dataset before and after quantization, such as metrics like accuracy in classification tasks and mean average precision (mAP) in object detection tasks, you can intuitively understand the changes in model inference accuracy.
[0099] Optionally, focus on key metrics for different application scenarios and tasks. In image segmentation tasks, the intersection over union (IoU) is often used to measure the accuracy of the model's segmentation of target objects. Calculate the IoU value between the segmentation results predicted by the model on the test data and the true labels before and after quantization, and compare the changes. If the IoU value decreases significantly after quantization, it means that the model's accuracy in object segmentation has been affected, and the quantization strategy may need to be adjusted. In speech recognition tasks, the word error rate (WER) is an important metric. By comparing the WER of the model on the test speech data before and after quantization, the model's accuracy in speech content recognition can be evaluated.
[0100] Optionally, to ensure the reliability of the evaluation results and avoid errors caused by accidental factors, multiple repeated experiments are performed. In each experiment, the same quantization method and parameter settings are used, but different random seeds are used to initialize the model training process. Statistical analysis is performed on the model inference accuracy indicators obtained from multiple experiments, and the average value and standard deviation are calculated. If the average value of the model inference accuracy after multiple experiments is significantly different from that before quantization, and the standard deviation is small, it means that the quantization method has a relatively stable and significant impact on the model inference accuracy; if the standard deviation is large, it indicates that the experimental results are greatly affected by random factors, and it is necessary to further optimize the experimental conditions or increase the number of experiments.
[0101] Optionally, samples in the test dataset can be classified according to different features. For example, in image recognition tasks, this can be done by image category, size, and lighting conditions; in object detection tasks, this can be done by object size, occlusion, and other factors. The inference accuracy of the model on different sample types can be calculated before and after quantization, analyzing how the model's performance changes under different circumstances. You may find that the model's accuracy decreases significantly on certain sample types, such as reduced detection accuracy for small objects. This requires optimizing the quantization strategy or further improving the model structure to improve its adaptability to different sample types.
[0102] Optionally, for tasks involving data such as images and videos, changes in model inference accuracy can be intuitively evaluated through visualization. For image classification tasks, the model's predictions for certain images before and after quantization can be visualized, and the predicted category labels can be compared with the actual image content to observe whether the model has made any misjudgments. For object detection tasks, the target boxes detected by the model before and after quantization can be plotted on the image to verify their accuracy in position and size, as well as any missed or false detections. Visualization can more intuitively identify problems with the model after quantization, providing a basis for subsequent improvements.
[0103] In a third aspect, the present application also provides an imaging control system, see Figure 8 , Figure 8 A schematic diagram of the imaging control training method system provided in an embodiment of the present application.
[0104] In which, the imaging control system is configured to execute the above-mentioned imaging control method or the above-mentioned imaging control model training method; the imaging control system includes: a sensing component, an optical component, a memory and a processor; the sensing component is configured to collect motion information; wherein the motion information includes the distance of the dynamic target; the optical component is configured to collect optical information; wherein the optical information includes the motion image of the dynamic target and the physical parameters of the optical component; program instructions are stored in the memory, and when the processor runs the program instructions, the steps of any implementation method in the above-mentioned imaging control method or the above-mentioned imaging control model training method are executed.
[0105] The imaging control system 10 may include a memory 11, a storage controller 12, a processor 13, a peripheral interface 14, an input and output unit 15, and a display unit 16. It will be understood by those skilled in the art that Figure 8 The structure shown is only for illustration and does not limit the structure of the imaging control system 10. For example, the imaging control system 10 may also include Figure 8 More or fewer components than shown, or with Figure 8 Different configurations shown.
[0106] The aforementioned memory 11, storage controller 12, processor 13, peripheral interface 14, input / output unit 15, and display unit 16 are electrically connected to each other, directly or indirectly, to enable data transmission or interaction. For example, these components may be electrically connected to each other via one or more communication buses or signal lines. The aforementioned processor 13 is used to execute the executable modules stored in the memory.
[0107] The memory 11 may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. The memory 11 is used to store programs, and the processor 13 executes the programs after receiving execution instructions. The method executed by the imaging control system 10 defined by the process disclosed in any embodiment of the present application can be applied to or implemented by the processor 13. Preferably, ≥4GB of LPDDR4 memory is provided to support multimodal data caching. Sufficient memory ensures that infrared images, lidar point cloud data, and intermediate calculation results can be effectively stored and quickly accessed during data processing, avoiding data loss and processing delays.
[0108] The above-mentioned processor 13 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 13 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The various methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. Preferably, an embedded AI chip with a computing power ≥ 4TOPS is selected, such as the Horizon J5 chip. This type of chip has powerful artificial intelligence computing capabilities and can quickly process infrared images and lidar point cloud data to meet the system's requirements for real-time and computing performance.
[0109] The peripheral interface 14 couples various input / output devices to the processor 13 and the memory 11. In some embodiments, the peripheral interface 14, the processor 13, and the memory controller 12 can be implemented in a single chip. In other embodiments, they can be implemented in separate chips.
[0110] The input and output unit 15 is used to provide data input to the user. The input and output unit 15 can be, but is not limited to, a mouse and a keyboard.
[0111] The display unit 16 provides an interactive interface (e.g., a user operation interface) between the imaging control system 10 and the user or is used to display image data for the user's reference. In this embodiment, the display unit can be a liquid crystal display or a touch display. If it is a touch display, it can be a capacitive touch screen or a resistive touch screen that supports single-point and multi-point touch operations. Supporting single-point and multi-point touch operations means that the touch display can sense touch operations generated simultaneously from one or more locations on the touch display, and pass the sensed touch operations to the processor for calculation and processing. In this embodiment of the present application, the display unit 16 can display the fused image provided by the imaging system.
[0112] Then, the imaging control system of the embodiment of the present application is introduced according to the actual implementation steps of the imaging control method.
[0113] First, the optical and sensor components have been described above. In this embodiment, the preferred optical component is a long-wave infrared camera with a wavelength range of 8-14μm. This wavelength range is highly sensitive to thermal radiation and can effectively capture the target's thermal signature. Its focal length is electrically adjustable from 10-300mm, allowing for flexible adjustment to suit imaging requirements at varying distances. A frame rate of ≥30fps ensures high temporal resolution of acquired infrared images, meeting the requirements for real-time monitoring of dynamic targets. The preferred sensor component is a solid-state lidar with a wavelength of 905nm, which offers compact design and high reliability. Its ranging accuracy can reach ±2cm / 100m, enabling precise measurement of target distance information. The scanning frequency directly impacts the temporal resolution of point cloud data. In intelligent surveillance, the movement of people and objects within a scene is frequent. Increasing the lidar scanning frequency from 20Hz to 50Hz can significantly increase the amount of point cloud data acquired per unit time.
[0114] The infrared lens can optionally be coated with an anti-reflection coating with a transmittance of ≥95% at 8-14μm. The anti-reflection coating reduces infrared light reflection from the lens surface and increases infrared light transmittance, thereby enhancing infrared image quality and improving the system's target detection capabilities.
[0115] At the hardware level, solid-state lidar uses phased array technology (OPA) to replace traditional mechanical rotating parts. It uses integrated circuits to precisely adjust the phase difference of multiple laser emitting units to reduce the electronic scanning angle interval of the laser beam; at the same time, it increases the density of laser emission / receiving channels and improves the unit angle beam sampling rate within the same field of view.
[0116] At the signal processing level, beamforming algorithms are used to enhance signal strength in specific directions and suppress sidelobe interference to compress the effective beamwidth. (For example, by optimizing beamforming weights, the effective beamwidth of a lidar can be compressed from 0.1° to 0.05°, improving the resolution of adjacent targets.) Multi-frame data fusion and super-resolution reconstruction algorithms utilize continuous multi-frame scan data (such as a continuous point cloud at a 50Hz scan frequency) to generate higher-resolution point cloud data. This allows for the reconstruction of finer target outlines from the continuous scan data. Gaussian filtering algorithms and IMU / GPS-assisted calibration are combined to remove the effects of noise and installation deviations. Real-time calibration of the conversion relationship between the lidar coordinate system and the global coordinate system (with a registration accuracy of ±5 cm) ensures the baseline accuracy of angle measurement, indirectly improving the actual angular resolution.
[0117] Before using the lidar, the lidar coordinate system must be calibrated. The inertial measurement unit (IMU) collects the lidar's acceleration and angular velocity data in real time, and the global positioning system (GPS) obtains the device's geographic location (longitude, latitude, and altitude). Using the IMU data, the lidar's attitude information (pitch, yaw, and roll) is calculated using an extended Kalman filter algorithm. Combining the GPS position information with the attitude information calculated by the IMU, the transformation relationship (homogeneous transformation matrix) between the lidar coordinate system and the global coordinate system is calculated using rotation matrix and translation vector operations. The point cloud data is then converted from the lidar coordinate system to the global coordinate system, with registration accuracy controlled within ±5 cm.
[0118] Specifically, a high-frequency real-time calibration mechanism (updated every 100 milliseconds) and a dual-sensor collaborative compensation strategy (IMU + GPS) have been combined to create an optimized solution that differs from traditional static calibration. The system uses hardware trigger signals to ensure strict alignment (time synchronization error ≤ 1ms) of the sampling times of the IMU (a high-precision model such as the Bosch BMI088, with angular velocity noise density ≤ 0.01° / √h, acceleration noise ≤ 100μg, and support for sampling frequencies exceeding 100Hz), GPS (using an RTK-GPS module such as the u-blox ZED-F9P, achieving centimeter-level positioning accuracy), and lidar. The IMU's quaternion attitude is converted into a rotation matrix, and the global coordinates provided by the GPS are used as translation vectors to construct a homogeneous transformation matrix. At the same time, a complementary filtering algorithm is used to fuse IMU and GPS data, the high-frequency output of the IMU is used to correct short-term posture changes, and the cumulative error of the IMU is corrected through the low-frequency but globally consistent position information of the GPS. The iterative closest point (ICP) algorithm is triggered every 100 milliseconds to match the current frame lidar point cloud with the global map to optimize the transformation matrix parameters, and the point cloud registration accuracy is controlled within ±5 cm, significantly improving the real-time performance and accuracy of multi-sensor fusion.
[0119] Through the collaborative work of IMU and GPS, the computer achieves precise calibration of the lidar coordinate system, greatly improving the quality and availability of point cloud data, providing strong support for the intelligent monitoring system to more accurately analyze and process data, thereby improving the performance and accuracy of the intelligent monitoring system.
[0120] For example, in urban street monitoring, as time goes by and the environment changes, the sensor may experience slight displacement and angle changes. Real-time calibration can ensure that the point cloud data acquired at different times can be accurately matched, thereby providing a reliable depth information foundation for subsequent target tracking and behavior analysis, ensuring that the monitoring system can continuously and accurately monitor the target.
[0121] Secondly, the high-precision stepper motor in the dynamic focus module achieves precise focal length adjustment, using a repeatability accuracy of ≤0.01mm. This motor accurately drives the movement of the infrared lens based on the control signal, ensuring accurate and stable focal length adjustment.
[0122] The temperature drift compensation circuit takes into account the impact of temperature changes on the optical system and motor performance, and a temperature drift compensation circuit is designed.
[0123] Its temperature sensitivity is ≤0.001% / ℃, and it can monitor changes in ambient temperature in real time and make corresponding compensation for focal length adjustments, effectively reducing focal length offset caused by temperature fluctuations and improving system reliability and stability.
[0124] In dynamic scenarios, timely data fusion can keep pace with rapid changes in targets. For example, in monitoring traffic intersections, vehicles and pedestrians move at high speeds. Delayed fusion results in a lag in acquired information. Real-time fusion of LiDAR point cloud data and infrared image data allows the monitoring system to grasp the target's position and state changes in real time. If a vehicle violates traffic regulations or a pedestrian runs a red light, the system can react quickly, issuing an alarm or initiating relevant processing procedures. This improves the monitoring system's real-time responsiveness and effectively reduces the risk of accidents.
[0125] In complex environments, such as city streets, factors such as light changes and obstructions can affect data accuracy. Choosing the right fusion time allows the system to fuse data when it is less subject to interference. For example, in the evening, when the light gradually dims, both infrared images and lidar data may be affected to a certain extent. However, if they are fused in time when the data quality is relatively good, the complementarity of the two data can be utilized to reduce the impact of noise and interference. LiDAR can compensate for the problem of decreased resolution of infrared images at long distances, and infrared images can provide texture information for lidar point cloud data. The combination of the two can more accurately identify target objects, reduce the occurrence of false alarms and missed alarms, and improve the overall performance of the monitoring system.
[0126] In embedded deployment, selecting the appropriate embedded hardware platform and optimizing its adaptation are key to ensuring the stable operation of the fusion imaging system in an embedded environment. Given the high computing power requirements of the fusion imaging system, an embedded AI chip with powerful artificial intelligence computing capabilities was selected. For example, the Horizon J5 chip, with a computing power of ≥4TOPS, provides efficient parallel computing capabilities to meet the real-time processing requirements of infrared imagery and LiDAR point cloud data. This chip features a dedicated hardware acceleration unit optimized for deep learning algorithms, enabling rapid execution of key operations such as convolution and matrix multiplication, effectively improving model inference speed. Equipped with ≥4GB of LPDDR4 memory to support multimodal data caching. Infrared imagery and LiDAR point cloud data processing requires a large amount of storage space. Sufficient memory ensures fast data read and write speeds, avoiding processing delays caused by data transmission bottlenecks. Furthermore, high-speed storage media, such as UFS 3.0, is selected to further improve data storage and read speeds, ensuring the system's real-time performance. The performance of optical components directly impacts the quality of data acquisition. Optimize the infrared lens and apply an antireflection coating to ensure a transmittance of ≥95% at 8-14μm. This reduces infrared light reflection from the lens surface, improving infrared image clarity and signal-to-noise ratio. Ensure that the optical components closely match the mechanical structure of the embedded device to ensure stable and precise installation and avoid imaging errors caused by vibration or displacement. Design appropriate hardware interfaces to enable high-speed data transmission between components. Use high-speed serial interfaces, such as USB 3.0 or PCIe, to connect the infrared camera, lidar, and processor, ensuring fast and stable data transmission to the processor for processing. Furthermore, reserve expansion interfaces to allow for the subsequent addition of additional sensors or functional modules based on actual needs, enhancing system scalability.
[0127] An embodiment of the present application also provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the steps of any implementation method of the above-mentioned imaging control method or the above-mentioned imaging control model training method are executed.
[0128] In summary, the present application provides an imaging control method, a training method, a system and a computer-readable storage medium, which relate to the field of computer vision technology. The method includes: adjusting the optical components based on the motion information and optical information of the dynamic target to obtain a monitoring image; wherein the optical component is used to obtain optical information; using the imaging control model to analyze the monitoring image of the dynamic target, and obtaining a fused image through the analysis results of the analysis model; wherein the motion information includes the distance of the dynamic target; the optical information includes the motion image of the dynamic target and the physical parameters of the optical component. The method can adjust the optical components in real time according to the changes in the actual scene, ensuring that the system always obtains clear images, effectively improving the imaging quality, and providing more reliable visual information for applications such as security monitoring and autonomous driving. On this basis, the motion state of the target can be clearly captured, which improves the system's monitoring and tracking capabilities for dynamic targets.
[0129] In a static target test (at a distance of 100 meters), the modulation transfer function (MTF) of a conventional infrared system was only 0.25, while the proposed system improved this to 0.48, resulting in a 92% increase in resolution. This means that at a distance of 100 meters, the proposed system can more clearly visualize the details of static targets, enabling more accurate identification of target features in security surveillance scenarios, for example. Furthermore, the ranging error is reduced from 0.1 meters in conventional solutions to 0.05 meters, improving target position measurement accuracy. The system consumes 8.2W, achieving low power consumption while maintaining high performance.
[0130] For dynamic vehicle testing (15 m / s), conventional infrared systems achieve an MTF of 0.12 at a target speed of 15 m / s, resulting in poor image clarity. This new system improves the MTF to 0.35, enhancing dynamic target clarity by 191%. This allows for clear capture of vehicle motion and characteristics, even during rapid travel. The ranging error is 0.08 meters, enabling accurate tracking of vehicle position changes. The system consumes only 9.5 W, meeting the requirements for real-time monitoring in dynamic scenarios.
[0131] In foggy and rainy environments (visibility 50 meters), the MTF of conventional infrared systems is 0.08, resulting in a short effective detection range. This new system improves the MTF to 0.28, tripling the effective detection range and significantly enhancing target detection capabilities in adverse weather conditions. The ranging error is 0.15 meters, and the system power consumption is 10.1W. This system maintains high accuracy while achieving reliable operation in foggy and rainy environments.
[0132] The imaging control system is capable of operating normally within a wide temperature range of -40°C to 85°C. This feature allows the system to operate stably in extremely cold or hot environments, making it suitable for a variety of complex application scenarios, such as polar expeditions and desert monitoring. With an IP68 protection rating, it is dustproof, waterproof, and salt spray-proof. In harsh outdoor environments, it can effectively protect the optical and electronic components within the system from dust, moisture, and salt spray, thereby improving the reliability and service life of the system. It can withstand random vibrations of 5-2000Hz / 15g. In vibrating environments such as vehicle driving and drone flight, the system can maintain a stable operating state, ensuring the accuracy of data acquisition and processing.
[0133] Since the principle of solving the problem by the computer-readable storage medium in the embodiment of the present application is similar to the embodiment of the aforementioned imaging control method and / or imaging control training method system, the implementation of the device in this embodiment can refer to the description in the embodiment of the above-mentioned method, and the repeated parts will not be repeated.
[0134] In the several embodiments provided in this application, it should be understood that the disclosed devices can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices according to the multiple embodiments of the present application. In this regard, each box in the block diagram can represent a module, a program segment or a part of a code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram, and the combination of the block diagrams, can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0135] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0136] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard drives, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks.
[0137] The foregoing is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures.
[0138] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.
Claims
1. An imaging control method, characterized in that: The method comprises: Based on the motion information and optical information of the dynamic target, the optical component is adjusted to obtain a monitoring image; wherein the optical component is used to obtain the optical information; Analyzing the monitoring image of the dynamic target using an imaging control model, and obtaining a fused image based on the analysis results of the analysis model; The motion information includes the distance of the dynamic target; and the optical information includes the motion image of the dynamic target and the physical parameters of the optical component.
2. The method according to claim 1, characterized in that The step of adjusting the optical component based on the motion information and optical information of the dynamic target to obtain a monitoring image includes: Acquire a first image and a second image; calculating a moving speed of the dynamic object according to a position change of the dynamic object in the first image and the second image; The first image is a motion image of the dynamic target at a first moment, and the second image is a motion image of the dynamic target at a second moment.
3. The method according to claim 2, characterized in that The step of adjusting the optical components based on the motion information and optical information of the dynamic target to obtain a monitoring image further includes: acquiring the distance of the dynamic target recorded by the sensing component; Calculating the adjusted focal length of the optical component by using the first formula; wherein the first formula is established based on the distance of the dynamic target, the movement speed of the dynamic target, and the physical parameters of the optical component; The adjusted focal length is compared with the current focal length of the optical component, and the physical parameters of the optical component are adjusted according to the comparison result to obtain the monitoring image; wherein the physical parameters of the optical component include the focal length.
4. The method according to claim 1, wherein The analyzing the monitoring image of the dynamic target by using the analysis model and obtaining a fused image based on the analysis result of the imaging control model includes: Acquire image information of the monitoring image, wherein the image information of the monitoring image has a first weight; wherein the image information includes grayscale, thermal radiation information, and texture features of the monitoring image; Acquire the motion information acquired by the sensing component, wherein the motion information has a second weight; wherein the motion information includes: contour and structure information of the dynamic target; performing feature fusion based on a second formula according to the weight of the image information of the monitoring image and the motion information acquired by the sensor component; The second formula is established according to the image information, the motion information, the first weight and the second weight.
5. The method according to claim 1, wherein in, The motion information acquired by the sensing component further includes: the acceleration, angular velocity and position of the dynamic target; the method further includes: Converting the coordinates of the dynamic target in the sensor component to actual coordinates according to the acceleration, angular velocity and position of the dynamic target; Calculating the posture of the dynamic target according to the acceleration of the dynamic target and the angular velocity of the dynamic target; When the posture of the dynamic target is acquired, the short-term posture change of the dynamic target is corrected according to the position of the dynamic target, and the movement trend of the dynamic target is predicted.
6. A method for training an imaging control model, characterized in that: in, The imaging control model training method is configured to execute the imaging control method according to any one of claims 1 to 5; The imaging control model includes: a dual-branch neural network and a lightweight optimization network; The method comprises: The dual-branch neural network extracts image information of a monitoring image of a dynamic target and motion information of the dynamic target; wherein the image information includes an image feature vector, and the motion information includes a motion feature vector; Outputting a fused feature vector by passing the image feature vector through an image weight matrix and the motion feature vector through a motion information weight matrix; When the dual-branch neural network training is completed, the lightweight optimization network analyzes the dual-branch neural network; Based on the set importance threshold, the importance score of each channel in the two-branch neural network is calculated, and the channels are screened.
7. The method according to claim 6, characterized in that The dual-branch neural network extracts image information of a monitoring image of a dynamic target and motion information of the dynamic target, including: The image features and motion information features in the dual-branch neural network training process are adjusted through a loss function; wherein the loss function adopts a weighted summation method, the image features include image clarity, and the motion information features include predicted motion errors.
8. The method according to claim 6, characterized in that The step of setting the importance threshold and calculating the importance score of each channel in the dual-branch neural network includes: Determining the importance threshold according to the requirements of the imaging control training model; wherein the requirements of the imaging control training model include: performance requirements, hardware resource limitations, and previous experimental results; Based on the screening of the channels, the dual-branch neural network is adjusted.
9. An imaging control system, characterized in that: The imaging control system is configured to execute the imaging control method according to any one of claims 1 to 5 or the imaging control model training method according to claims 6 to 8; the imaging control system comprises: a sensing component, an optical component, a memory, and a processor. The sensing component is configured to collect motion information; wherein the motion information includes the distance of the dynamic target; The optical component is configured to collect optical information; wherein the optical information includes a motion image of the dynamic target and physical parameters of the optical component; The memory stores program instructions, and when the processor runs the program instructions, the steps of the method according to any one of claims 1 to 8 are executed.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are executed.
Citation Information
Patent Citations
Photographing method, terminal and computer readable storage medium
CN107920211A
Exposure time adjusting method and device and camera
CN110753178A
Video processing method and device, electronic equipment and readable storage medium
CN113286194A
Image acquisition method and equipment
CN113709353A
Shutter value adjusting method and device, storage medium and electronic equipment
CN116264641A
Cited By
Cunninghamia lanceolata forest-oriented periodic tending scheduling method, device and equipment and storage medium
CN121581333A