Hunting camera imaging quality optimization method based on multi-mode data fusion

Through multi-modal data fusion technology, the hunting camera has achieved a deep understanding of the shooting scene and intelligent parameter optimization, which solves the problem of unstable imaging quality in existing technologies and improves image clarity and adaptability.

CN121304463BActive Publication Date: 2026-03-27NINGBO JINSHENGXIN IMAGE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-11
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing hunting camera imaging technology cannot deeply understand the shooting scene and lacks the ability to extract high-level semantic information from multi-source data and make intelligent decisions, resulting in improper configuration of imaging parameters, leading to problems such as blurry images, improper exposure, or loss of target details.

Method used

By collecting multi-source heterogeneous data (visible light images, thermal infrared images, laser ranging point clouds, and inertial measurement unit data), spatiotemporal alignment and registration are performed, imaging quality-related parameters and scene context description information are extracted, and fusion decision calculations are performed using an imaging optimization knowledge base to generate global imaging parameter optimization instructions and adjust the camera imaging component parameters.

Benefits of technology

It achieves a deep understanding of complex shooting scenes and precise parameter optimization, ensuring the stability and clarity of image quality in various environments, avoiding parameter conflicts, and improving image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121304463B_ABST
    Figure CN121304463B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image quality optimization, and discloses a hunting camera imaging quality optimization method based on multi-mode data fusion. The method comprises the following steps: collecting visible light images, thermal infrared images, laser ranging point clouds and inertial measurement unit data, and performing space-time alignment and registration to form synchronous multi-mode data. Through joint feature extraction on the synchronous data, a controllable imaging parameter set and scene context description information are separated out. For each to-be-optimized parameter, the system takes the current value and the scene context as a query condition, retrieves multiple candidate strategies from an imaging optimization knowledge base, calculates an optimal strategy through fusion decision, and finally generates and executes a global imaging parameter optimization instruction set to control the camera to complete image capture. The application realizes intelligent optimization of imaging parameters based on deep scene understanding, and improves the image quality and adaptability of the hunting camera in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image quality optimization, in particular to a hunting camera imaging quality optimization method based on multi-mode data fusion. BACKGROUND

[0002] The existing hunting camera imaging technology mainly relies on single type sensor data or simple double sensor linkage mechanism for parameter adjustment. The common method is to switch day / night mode according to the intensity of ambient light, or to start shooting according to the trigger signal of passive infrared sensor. These methods can only respond to limited and single-dimensional environmental changes. More advanced solutions may combine visible light and thermal infrared data, but their combination usually stays at the data level, lacking deep understanding of the imaging scene and intelligent decision-making ability of imaging parameters.

[0003] The core defect of the prior art is that the parameter adjustment logic is based on preset fixed threshold or simple rules, which cannot adapt to the complex and changeable shooting scenes in the wild. When facing the comprehensive situation of target motion speed mutation, complex environmental lighting conditions, meteorological interference or low target and background contrast, the system based on simple rules is difficult to make the optimal imaging parameter configuration, resulting in problems such as blurred image, improper exposure, focusing failure or loss of target details. This is essentially due to the fact that the system fails to extract context information describing the nature of the scene from multi-source data, and also lacks a strategy library containing rich imaging experience for intelligent decision-making.

[0004] The hunting camera needs a method that can deeply understand the shooting scene and dynamically and finely adjust multiple imaging parameters based on the understanding. The current technology cannot jointly extract high-level semantic information from multi-mode data to describe the scene state, and cannot convert it into specific and coordinated optimization parameter instructions. At the same time, the parameter optimization process lacks flexibility and cannot select the most suitable solution from multiple possible optimization strategies according to the specific scene context, which restricts the further improvement of imaging quality. SUMMARY

[0005] The purpose of the present application is to provide a hunting camera imaging quality optimization method based on multi-mode data fusion to solve the problems raised in the background.

[0006] To achieve the above purpose, the present application provides a hunting camera imaging quality optimization method based on multi-mode data fusion, which comprises:

[0007] Collecting multi-source heterogeneous data generated by the hunting camera in the working state, the multi-source heterogeneous data at least including visible light image, thermal infrared image, laser ranging point cloud and inertial measurement unit data;

[0008] The collected multi-source heterogeneous data are spatio-temporally aligned and registered to generate synchronized multi-modal data with consistent spatio-temporal reference;

[0009] The synchronized multi-modal data are jointly feature-extracted to separate a set of controllable imaging parameters directly related to imaging quality and scene context description information reflecting the state of the shooting scene;

[0010] The set of controllable imaging parameters is traversed, and for each imaging parameter to be optimized, the following operations are performed: based on the imaging parameter to be optimized and the scene context description information, a plurality of candidate optimization strategies are retrieved from a preset imaging optimization knowledge base;

[0011] The plurality of candidate optimization strategies retrieved are fused and decided based on the scene context description information as a basis to select an optimal optimization strategy for the imaging parameter to be optimized;

[0012] According to the optimal optimization strategies selected for all the imaging parameters to be optimized, a global imaging parameter optimization instruction set is generated;

[0013] The global imaging parameter optimization instruction set is executed to adjust the working parameters of the internal imaging components of the hunting camera, and the camera is triggered to capture images, and finally the optimized hunting camera images are obtained.

[0014] Preferably, the collected multi-source heterogeneous data are spatio-temporally aligned and registered to generate synchronized multi-modal data with consistent spatio-temporal reference, including:

[0015] The visible light image and the thermal infrared image are feature point-based image registered to eliminate parallax and generate a pixel-level aligned dual-spectrum image pair;

[0016] The laser ranging point cloud data is projected to the image coordinate system of the aligned dual-spectrum image pair to establish a correspondence between each image pixel and depth information;

[0017] According to the timestamp of the inertial measurement unit data, the posture of the camera at the image capture moment is estimated and compensated to ensure that all data correspond to the same spatio-temporal observation point.

[0018] Preferably, the synchronized multi-modal data are jointly feature-extracted to separate a set of controllable imaging parameters directly related to imaging quality and scene context description information reflecting the state of the shooting scene, including:

[0019] The brightness distribution, contrast, spectral features and potential moving target regions are extracted from the aligned dual-spectrum image pair;

[0020] The average distance, distance change gradient and target relative size of the shooting scene are extracted from the projected depth information;

[0021] fusing the camera pose estimated from the inertial measurement unit data to determine whether the scene is predominantly static or has a high degree of relative motion;

[0022] generating scene context description information based on the extracted spectral features, depth features and motion features; and defining the current adjustable parameters of the camera as a set of controllable imaging parameters.

[0023] Preferably, based on the imaging parameter to be optimized and the scene context description information, a plurality of candidate optimization strategies are retrieved from a pre-set imaging optimization knowledge base, including:

[0024] using the imaging parameter to be optimized as a primary query key and using the scene context description information as an auxiliary filtering condition;

[0025] performing a fuzzy matching query based on semantic similarity in the imaging optimization knowledge base to preliminarily filter out candidate optimization strategy entries related to the current imaging context;

[0026] ranking the matching degrees of the primary query key and the auxiliary filtering condition with the content of the knowledge base entries and returning a plurality of candidate optimization strategies with the highest matching degrees.

[0027] Preferably, the scene context description information is used as a basis for discriminant decision calculation on the plurality of candidate optimization strategies retrieved, including:

[0028] encoding the scene context description information into a fixed-dimensional scene feature vector;

[0029] encoding the content description information of each candidate optimization strategy into a corresponding strategy feature vector;

[0030] calculating the semantic correlation degree between the scene feature vector and each strategy feature vector;

[0031] combining the semantic correlation degree and the matching degree returned by the knowledge base to calculate a comprehensive confidence degree for each candidate optimization strategy through a weighted decision model;

[0032] selecting the candidate optimization strategy with the highest comprehensive confidence degree as the optimal optimization strategy for the imaging parameter to be optimized.

[0033] Preferably, the scene context description information is encoded into a fixed-dimensional scene feature vector, including:

[0034] normalizing the brightness features, distance features and motion features contained in the scene context description information;

[0035] using a multi-layer perception model to perform non-linear transformation and dimension reduction on the normalized multi-dimensional features;

[0036] The reduced dimension feature sequence is input into a sequence encoder to capture the dependency between the features, and the final hidden state of the sequence encoder is the scene feature vector.

[0037] Preferably, the semantic correlation between the scene feature vector and each strategy feature vector is calculated, including:

[0038] The scene feature vector is dot multiplied with a strategy feature vector to obtain an initial correlation score.

[0039] The initial correlation score is input into a single-layer neural network for scaling and normalization, and a semantic correlation score between zero and one is output.

[0040] All strategy feature vectors are traversed, and the dot product and neural network processing are repeated to obtain a set of semantic correlation scores corresponding to each candidate optimization strategy.

[0041] Preferably, a comprehensive confidence is calculated for each candidate optimization strategy by a weighted decision model, including:

[0042] A dynamic weight value is assigned to the semantic correlation score, and a fixed weight value is assigned to the matching degree returned by the knowledge base.

[0043] The weighted semantic correlation score and the weighted knowledge base matching degree are added to obtain the comprehensive confidence of each candidate optimization strategy.

[0044] The dynamic weight value is adjusted according to the complexity of the scene context description information. The more complex the scene, the higher the dynamic weight value.

[0045] Preferably, a global imaging parameter optimization instruction set is generated according to the optimal optimization strategies selected for all imaging parameters to be optimized, including:

[0046] Check whether there is a parameter conflict between the optimal optimization strategies selected for each imaging parameter to be optimized.

[0047] If there is a parameter conflict, adjust the conflicting strategies according to the preset priority rules to ensure the internal consistency of the final instruction set.

[0048] Convert the adjusted and consistent optimal optimization strategies into specific parameter setting commands executable by the hunting camera, and combine them into a global imaging parameter optimization instruction set according to the execution order.

[0049] Preferably, the global imaging parameter optimization instruction set is executed to adjust the working parameters of the internal imaging components of the hunting camera and trigger the camera to capture images, including:

[0050] The global imaging parameter optimization instruction set is sent to the corresponding control unit of the hunting camera in sequence;

[0051] The control unit adjusts the aperture, shutter speed, sensitivity, focusing distance and image signal processing parameters according to the received commands;

[0052] After all the parameters are adjusted, a capture instruction is sent to the camera trigger system to obtain a single frame or multiple frames of optimized hunting camera images.

[0053] Compared with the prior art, the beneficial effects of the present application are:

[0054] By jointly extracting features from multi-modal data to generate scene context description information, the system can overcome the limitations of single sensor perception and achieve deep understanding of the shooting scene. This understanding is not simply a superposition of data, but a high-level semantic description that integrates target distance, motion vector, thermal radiation characteristics and environmental light characteristics. It provides an accurate and rich basis for subsequent optimization decisions, enabling the system to accurately perceive complex scene states. This deep scene perception capability is the fundamental prerequisite for precise imaging parameter optimization, ensuring that the subsequent strategy selection is highly matched with the current real shooting needs.

[0055] The fusion decision mechanism based on the imaging optimization knowledge base improves the parameter optimization from fixed rule automation to intelligence based on scene understanding. For each parameter to be optimized, the system retrieves multiple candidate strategies and uses scene context for fusion calculation to dynamically select the optimal strategy. This process simulates experience judgment and can handle complex and even contradictory scene information to make fine and robust decisions. This method improves the accuracy and adaptability of parameter configuration, enabling the camera to automatically generate a parameter combination close to professional standards in various complex environments, overcoming the unstable imaging quality problem caused by rigid rules in traditional methods.

[0056] The combination of the deep fusion generated scene context and the knowledge base driven decision mechanism realizes the collaborative optimization of imaging parameters. The system synchronously selects the best strategy for focus, exposure, shutter speed and other parameters under unified scene understanding, avoiding conflicts between parameters. This overall optimization ensures that the final imaging achieves the best balance in terms of sharpness, brightness, noise control and other aspects, improving the overall quality and usability of the image, especially in complex outdoor environments where image quality is critical. BRIEF DESCRIPTION OF DRAWINGS

[0057] Figure 1 The working principle diagram of the hunting camera imaging quality optimization method based on multi-modal data fusion described in the present application;

[0058] Figure 2Flowchart for spatio-temporal alignment and registration

[0059] Figure 3 Flowchart for candidate optimization strategy retrieval

[0060] Figure 4 Flowchart for dynamic weight adjustment and decision comprehensive influence correlation analysis

[0061] Figure 5 Flowchart for candidate strategy confidence composition analysis DETAILED DESCRIPTION

[0062] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the present application.

[0063] Please refer to Figure 1 The present application provides a hunting camera imaging quality optimization method based on multi-mode data fusion, which comprises the following steps: a hunting camera generates visible light images, thermal infrared images, laser ranging point clouds and inertial measurement unit data synchronously in a working state. These data differ in physical form and information dimension, and jointly constitute the data basis for subsequent optimization. After collection, the system immediately performs strict spatio-temporal alignment and registration processing on these multi-source heterogeneous data, the purpose of which is to eliminate the data deviation caused by the inconsistency of different sensor collection time and spatial position, so as to generate a set of synchronous multi-mode data with consistent reference in time and space. The processing flow enters the feature extraction stage, and the system performs joint analysis on the synchronous multi-mode data, from which two types of key information are separated out: one type is a set of controllable imaging parameters that can be adjusted by the camera itself and are directly related to the imaging quality, such as aperture, shutter, sensitivity, etc.; the other type is scene context description information used to describe the objective state of the current shooting scene, such as scene brightness, target distance, whether there is motion, etc.

[0064] The system enters the core optimization decision-making link, which will traverse each of the to-be-optimized parameters in the controllable imaging parameter set. For each parameter, the system will retrieve multiple candidate optimization strategies from a pre-constructed imaging optimization knowledge base based on the identity of the parameter and the current scene context description information. The system will again use the scene context description information as the basis for decision-making to calculate and fuse the candidate strategies, and through calculation and comparison, select an optimal optimization strategy for the current to-be-optimized parameter. When the optimal strategies for all to-be-optimized parameters are selected, the system will generate a global imaging parameter optimization instruction set based on these strategies, which ensures the coordination between parameter adjustments. The hunting camera will execute this instruction set to adjust the working parameters of its internal imaging components and trigger the image capture action, thereby obtaining high-quality images after optimization.

[0065] Embodiment 1: refer to Figure 2 In specific implementation, the spatio-temporal alignment and registration process starts with the collaborative processing of visible light images and thermal infrared images. Due to the differences in physical location and imaging principles between visible light sensors and thermal infrared sensors, there are obvious parallax and geometric distortion between the two images directly collected. The feature point-based image registration method is the core means to eliminate these differences. The registration algorithm extracts a set of feature points with saliency and stability from the visible light image. Feature point extraction can be achieved using the scale-invariant feature transform algorithm or the oriented fast and robust feature point detection algorithm. In specific implementation, the feature point extraction process focuses on high-contrast regions such as corner points and edge intersection points in the image to ensure that feature points can be stably detected on multispectral images. Similarly, the registration algorithm also performs the same feature point detection process at the corresponding position on the thermal infrared image, thereby obtaining another set of feature points. After obtaining two sets of feature points, the system needs to establish the correspondence between the visible light image feature points and the thermal infrared image feature points, which is called feature matching. Feature matching is usually completed by calculating the Hamming distance or Euclidean distance between feature point descriptors. Descriptors are numerical vectors used to represent patch information around feature points. The matching algorithm finds the nearest feature point in the thermal infrared image as the candidate matching point for each feature point in the visible light image based on the descriptor distance. To eliminate false matching pairs, the system uses a random consistency verification algorithm to filter the preliminary matching results. The random consistency verification algorithm can effectively eliminate false matches caused by noise or repetitive textures, thereby obtaining a set of accurate and reliable matching feature point pairs. These correct matching point pairs provide a data basis for calculating the spatial transformation relationship between the two images.

[0066] With the correct matching feature points, the system can estimate an optimal spatial transformation model, which is usually represented by a homography matrix. The homography matrix describes the projection transformation relationship from one image plane to another, and the values of the elements of the homography matrix can be calculated by solving an overdetermined equation set. In specific implementation, the least squares method is a commonly used method to solve the homography matrix, which can minimize the projection error of all matching point pairs after transformation. After calculating the homography matrix, the system applies it to the original thermal infrared image to resample and geometrically transform the thermal infrared image, so that every pixel point on the thermal infrared image is accurately aligned with the corresponding pixel point on the visible light image in spatial position, and finally generates a pixel-level aligned dual-spectrum image pair. Pixel-level alignment means that the same pixel coordinates point to the same physical point in the scene in both images.

[0067] After completing the registration of the visible light image and the thermal infrared image, the system begins to process the laser ranging point cloud data. The point cloud data obtained by the laser range finder is a set of scattered points in three-dimensional space, and each point contains three-dimensional coordinate information. In order to fuse the three-dimensional point cloud with the two-dimensional image, the point cloud data needs to be projected into the two-dimensional image coordinate system. The projection process depends on the internal and external parameters of the camera. The internal parameters of the camera include focal length, principal point coordinates, etc., which describe the internal geometric model of the camera; the external parameters of the camera describe the rotation and translation relationship between the laser radar coordinate system and the camera coordinate system. In specific implementation, the system uses the pre-calibrated camera-laser radar joint calibration parameters to convert the coordinates of each three-dimensional laser point to the image pixel coordinate system through the perspective projection transformation formula. For laser points falling within the image boundary, the system records their corresponding pixel coordinates and depth values; for pixel positions without laser point projection, interpolation algorithms can be used to estimate the depth values from adjacent laser points to establish a correspondence between each image pixel point and depth information, and generate a depth map completely aligned with the visible light image.

[0068] The processing of IMU data is closely related to the time of image capture. IMU continuously outputs angular rate and acceleration data at a fixed frequency, and each set of data is accompanied by an accurate timestamp. The system needs to record the global timestamp of each frame of image exposure accurately, and then according to this timestamp, the angular rate and acceleration values at the moment of image exposure are interpolated from the IMU data stream. In specific implementation, linear interpolation method is sufficient to meet the accuracy requirements in most cases. After obtaining the IMU data at the moment of exposure, the system estimates the orientation of the camera in space through a pose solving algorithm. The pose solving algorithm can use complementary filtering or Kalman filtering to fuse the integral of angular rate and the gravity direction information measured by the accelerometer, and calculate the roll angle, pitch angle and yaw angle of the camera relative to the inertial reference frame. These attitude angles describe the accurate orientation of the camera at the moment of image capture. The system uses these attitude angles to perform geometric compensation, such as rotation correction, on the already aligned dual-spectrum images and the depth map obtained by projection, to eliminate image distortion caused by slight camera shaking, and to ensure that all visible light images, thermal infrared images, depth information and attitude information correspond to the same spatiotemporal observation point, complete the synchronization of multi-modal data, and generate synchronized multi-modal data with consistent spatiotemporal reference.

[0069] The joint feature extraction process acts on the synchronized multi-modal data, aiming to separate the imaging quality related parameters and scene description information. From the aligned dual-spectrum image pair, the system extracts multiple features in parallel. The brightness distribution feature is obtained by calculating the gray level histogram of the visible light image, which counts the number of pixels at different brightness levels in the image and can reflect the overall illumination level of the scene. The contrast feature is obtained by calculating the standard deviation within the local window of the image. High contrast areas usually correspond to areas with rich texture. The extraction of spectral features depends on the difference in pixel values between the visible light and thermal infrared channels. For the same pixel position in the dual-spectrum image pair, the system calculates the ratio or difference between the visible light brightness value and the thermal infrared brightness value, which can distinguish objects of different materials. The detection of potential moving target regions can be achieved by comparing the dual-spectrum image pair between consecutive frames. The system calculates the difference between corresponding pixels between the current frame and the previous frame image. The region whose difference exceeds the preset threshold is marked as a potential moving target region.

[0070] From the projected depth information, the system extracts distance-related features. The average distance is obtained by averaging the depth values of all valid pixels in the whole depth map, which reflects the approximate distance between the target object and the camera. The distance variation gradient is obtained by calculating the Sobel operator gradient magnitude of the depth map in the horizontal and vertical directions, and the region with large gradient magnitude corresponds to the edge part with sharp depth variation in the scene. The calculation of the target relative size needs to combine the potential motion target region. The system first locates the motion target region in the depth map, then counts the number of pixels in the region, and combines the average distance information to estimate the approximate physical size of the motion target in the real world.

[0071] The camera pose estimated from the inertial measurement unit data is used to determine the motion state of the scene. The system buffers the camera pose data of consecutive frames of images, including the roll angle, pitch angle and yaw angle. By calculating the change amount of the pose angle between consecutive frames, the angular velocity of the camera itself can be obtained. If the angular velocity is continuously below a lower threshold, it can be inferred that the scene is mainly static and the camera is in a stable state; if the angular velocity exceeds the threshold, it indicates that there is a sharp relative motion in the scene, which may be caused by camera shaking or fast movement of the target.

[0072] Based on the above-mentioned extracted spectral features, depth features and motion features, the system performs feature fusion to generate scene context description information. The scene context description information is a structured data vector, and each dimension of the vector represents a scene attribute, such as overall brightness level, average contrast, thermal infrared saliency, scene average depth, depth complexity, motion intensity flag, etc. These quantitative feature values together constitute a comprehensive description of the state of the shooting scene. At the same time, the system reads the list of all currently available imaging parameters from the camera's configuration file, including aperture size, shutter speed, sensitivity, focus distance, white balance mode, image sharpening intensity, etc. These parameters are explicitly defined as the controllable imaging parameter set. The controllable imaging parameter set and the scene context description information together serve as input for the subsequent optimization decision process.

[0073] Embodiment 2: see Figure 3In a specific implementation, the process of retrieving multiple candidate optimization strategies from the preset imaging optimization knowledge base based on the imaging parameter to be optimized and the scene context description information is a structured information query and matching process. The imaging optimization knowledge base is a pre-constructed structured database that stores a large number of historical imaging optimization cases. Each case record contains three main fields: an imaging parameter field, a scene condition description field, and an optimization strategy field. The imaging parameter field records the specific imaging parameter name and its specific value or state that is adjusted in the case; the scene condition description field describes the scene state when the optimization strategy is triggered in the form of a structured feature vector or a set of keywords, such as lighting conditions, target distance, motion conditions, etc.; the optimization strategy field records the parameter adjustment suggestions or control instructions taken for the scene in detail. When the system starts optimizing a certain imaging parameter in the set of controllable imaging parameters, for example, the "shutter speed" needs to be optimized, then this "shutter speed" parameter is specified as the main query key. The role of the main query key is to lock the record subset in the knowledge base that is directly related to the current parameter to be optimized. At the same time, the scene context description information generated in the previous step is specified as an auxiliary filtering condition, which is a comprehensive scene state summary. The role of the auxiliary filtering condition is to further refine the selection based on the matching degree of the scene on the basis of the record subset preliminarily selected by the main query key.

[0074] The system needs to find the records in this preliminarily selected record set that are closest to the current actual situation in terms of scene conditions. This is the fuzzy matching query stage based on semantic similarity. The core of the fuzzy matching query is to calculate the semantic similarity between the current scene context description information and the scene condition description field of each record in the knowledge base. In a specific implementation, in order to realize semantic similarity calculation, the scene context description information and the scene condition description in the knowledge base need to be converted into a calculable representation. A feasible way is to represent both as vectors in a high-dimensional space. The scene context description information itself may already be a numerical feature vector. For the scene condition description field of the records in the knowledge base, if its storage form is a structured feature vector, it can be directly used for calculation; if the storage form is a text keyword, it needs to be converted into a semantic vector through a word embedding model. Semantic similarity is quantified by calculating the cosine similarity or Euclidean distance between two vectors. The system calculates a similarity score for each record preliminarily selected, and the higher the score, the more similar the scene described by the record to the current scene. Then, the system sets a similarity threshold, and only those records with a similarity score higher than the threshold are retained, thereby completing the auxiliary filtering based on scene conditions and preliminarily selecting candidate optimization strategy entries related to the current imaging context.

[0075] In a specific implementation, the candidate optimization strategy entries obtained after semantic similarity filtering can still differ in quantity and matching quality. The system needs to rank these entries to identify the best suggestions that are most likely to apply to the current situation. The basis for ranking is the overall matching degree of the main query key and the auxiliary filtering conditions with the content of the knowledge base entries. The matching degree is a comprehensive score that combines the matching accuracy of the main query key and the semantic similarity of the auxiliary filtering conditions. The matching accuracy of the main query key is binary, i.e., whether the record contains the key, but since it has been ensured that all candidate records contain the main query key in the preliminary screening, the accuracy weight of the main query key can be considered a constant in actual ranking. The semantic similarity score of the auxiliary filtering conditions constitutes the main variable for ranking. The system ranks the candidate optimization strategy entries in order of descending semantic similarity score. In some embodiments, the calculation of the matching degree can be more complex, and additional weighting factors such as the historical success rate of the strategy, the applicable environmental range of the strategy, etc. metadata, which are also stored in the imaging optimization knowledge base, are associated with each record. The system calculates a final matching degree comprehensive score for each candidate record.

[0076] It can be understood that the imaging optimization knowledge base can be very large, and returning all matching records is not realistic in terms of computational efficiency and decision-making efficiency. Therefore, the system usually sets an upper limit N on the number of returned records. After completing the matching degree ranking, the system selects the top N records with the highest matching degree as the final output. This N value is a configurable parameter, which can be set to 5 or 10, for example, the purpose is to provide a moderate number of high-quality candidate strategy sets for subsequent fusion decision-making calculations. These top N candidate optimization strategy entries with the highest matching degree contain a variety of possible solutions extracted from historical experience for the current imaging parameters to be optimized and the current scene state, which will be passed to the next step for deeper decision analysis. Optionally, the implementation of the imaging optimization knowledge base can be based on a relational database management system, which uses its powerful indexing and query optimization functions to speed up the retrieval process; it can also be based on a special vector database to efficiently handle high-dimensional vector similarity searches. The whole retrieval process emphasizes semantic similarity rather than exact matching, which enables the system to cope with complex and variable scene conditions in the field environment that may not be accurately recorded in the knowledge base, demonstrating the adaptability and robustness of the method.

[0077] In some embodiments, the construction of the imaging optimization knowledge base is a continuous learning process. The system can add each case after a successful optimization, including the final imaging parameter settings, the scene context description information at that time, and the obtained image quality evaluation feedback, as a new record to the imaging optimization knowledge base. This mechanism enables the imaging optimization knowledge base to evolve continuously, gradually covering more diverse scenarios, thereby improving the accuracy of future retrieval and optimization decisions. It can be understood that the initial data of the imaging optimization knowledge base can come from the preset parameter library of the camera manufacturer, the experience rules of professional photographers, or by a large number of test calibration in a controllable environment. The efficiency of the retrieval process is crucial for the real-time response of the hunting camera, so the design of the index structure and the optimization of the similarity calculation algorithm are important implementation details that need to be considered.

[0078] Embodiment 3: In specific implementation, the use of scene context description information as a basis for fusion decision calculation on the retrieved multiple candidate optimization strategies is the core step of determining the optimal strategy. Scene context description information is a comprehensive numerical description of the current shooting scene generated in the previous step, and candidate optimization strategies are multiple possible solutions retrieved from the imaging optimization knowledge base. The goal of fusion decision calculation is to select a strategy that best fits the current scene context and is most likely to improve image quality from these candidate solutions as the final choice. To achieve this goal, it is first necessary to convert the scene context description information and the content description information of each candidate optimization strategy into a form that can be mathematically compared, i.e., a fixed-dimension feature vector. This process is called feature encoding. Scene context description information usually contains multiple dimensional feature values, such as brightness level, average distance, motion intensity, etc. These feature values may have different dimensions and value ranges. The first step of feature encoding is to normalize these original features, which scales each feature value to a uniform interval, such as between zero and one, to eliminate the influence of different dimensions on subsequent calculations. The multi-dimensional features after normalization form an initial feature sequence.

[0079] In a specific implementation, the normalized multi-dimensional feature sequence is encoded into a fixed-dimensional scene feature vector, usually by a neural network model. A typical encoding procedure is to first input the normalized feature sequence into a multi-layer perceptron model. The multi-layer perceptron model is a feed-forward neural network, which contains an input layer, one or more hidden layers, and an output layer. The number of neurons in the input layer is the same as the dimension of the normalized feature, the hidden layers are responsible for nonlinear transformation and feature interaction of the input feature, and the output layer outputs a feature vector with lower dimension than the input, thereby realizing dimension reduction. Through its hierarchical structure and nonlinear activation function, the multi-layer perceptron model can learn the complex combination relationship between the original features and extract higher-level and more abstract scene semantic information. It can be understood that the weight parameters of the multi-layer perceptron model need to be obtained through supervised training of a large amount of historical data, and the training target is to make the encoded scene feature vector well predict the optimization strategy to be adopted in the scene.

[0080] In some embodiments, in order to further capture the dependency relationship that may exist in the normalized multi-dimensional feature sequence, such as the correlation between the brightness feature and the spectral feature, the system will further input the feature sequence after dimension reduction by the multi-layer perceptron model into a sequence encoder. The sequence encoder can adopt a recurrent neural network structure, such as a long short-term memory network or a gated recurrent unit. The recurrent neural network has an internal state and can process sequence data and capture the forward and backward dependency relationship between elements in the sequence. The dimension-reduced feature sequence is input into the recurrent neural network one by one, and the hidden state of the recurrent neural network is updated with the input of the sequence, so as to fuse the information of all previous features. When the entire feature sequence is input, the final hidden state of the recurrent neural network, i.e. the hidden state vector at the last time step, is extracted as the final scene feature vector. This scene feature vector not only contains the static attributes of the scene, but also captures the dynamic association between different attributes, forming a deep representation of the scene context.

[0081] For each candidate optimization strategy retrieved from the imaging optimization knowledge base, its content description information also needs to be encoded to generate a corresponding strategy feature vector. The content description information of a candidate optimization strategy can be in the form of text, such as “suggest using larger aperture and lower shutter speed in low-light static scene”, or in the form of structured parameter suggestion list. In specific implementation, if the content is text description, a word embedding model combined with recurrent neural network or self-attention mechanism can be used to generate the strategy feature vector; if the content is structured data, the encoding method is similar to the scene context description information, which is transformed and reduced by a multi-layer perception model. The dimension of the strategy feature vector needs to be consistent with the dimension of the scene feature vector, so as to perform vector operation between them later. Finally, each candidate optimization strategy corresponds to a strategy feature vector, which encodes the scene conditions applicable to the strategy and the core content of the strategy itself.

[0082] After obtaining the scene feature vector and a set of strategy feature vectors, the system needs to calculate the semantic correlation degree between the scene feature vector and each strategy feature vector. Semantic correlation degree is a quantitative indicator for measuring the matching degree of a particular candidate optimization strategy with the current scene. The first step of calculating the semantic correlation degree is to perform vector dot product operation. For the scene feature vector and the strategy feature vector of a certain candidate optimization strategy , the dot product operation is defined as:

[0083]

[0084] Wherein: is the dimension of the vector, and are the th component of the vector. The result of the dot product operation is a scalar, called initial correlation score. The initial correlation score reflects the similarity of the direction of the two vectors, and the closer the direction of the two vectors, the larger the dot product value.

[0085] The size of the initial correlation score is directly affected by the vector length, and its numerical range is uncertain, which is not convenient for direct comparison and use. Therefore, the initial correlation score needs to be scaled and normalized to map it to a standardized range. In specific implementation, this is achieved by a single-layer neural network. This single-layer neural network usually only contains a fully connected layer followed by a Sigmoid activation function. The initial correlation score is input into this single-layer neural network, the fully connected layer performs a linear transformation on it, and the Sigmoid function compresses the transformed value to between zero and one, outputting a scalar value between zero and one, which is the final semantic correlation score The output of sigmoid function can be intuitively interpreted as the correlation probability, the value closer to 1 means the higher correlation. The system will traverse all the strategy feature vectors, and repeat the above dot product operation and single-layer neural network processing for each strategy feature vector, so as to obtain a set of semantic correlation scores corresponding to each candidate optimization strategy. Alternatively, the semantic correlation can also be calculated by using cosine similarity instead of dot product, and the cosine similarity calculates the cosine value of the angle between vectors, which is in the range of -1 to 1, and can also be mapped to the interval of 0 to 1 by transformation.

[0086] The last step of fusion decision calculation is to calculate a comprehensive confidence for each candidate optimization strategy, and select according to the comprehensive confidence. The comprehensive confidence integrates two aspects of information: on the one hand, the semantic correlation score calculated in real time, which reflects the adaptability of the strategy to the current specific scene; on the other hand, the matching degree returned by the imaging optimization knowledge base, which reflects the reliability of the strategy in historical experience. In specific implementation, a weighted decision model is used to calculate the comprehensive confidence. The weighted decision model assigns a dynamic weight value to the semantic correlation score , and assigns a fixed weight value to the matching degree returned by the knowledge base. The fixed weight value is a constant set based on prior knowledge. The dynamic weight value is not fixed, but will be dynamically adjusted according to the complexity of the scene context description information. The complexity of the scene can be quantified by calculating the information entropy of the scene feature vector or the variance of the feature value. The more complex the scene, the higher the dynamic weight value is set, which means that more reliance is placed on the semantic correlation calculated in real time for the current complex scene in decision making. The formula for calculating the comprehensive confidence is as follows:

[0087]

[0088] , where represents the semantic correlation score , and represents the matching degree returned by the knowledge base. The system will calculate the comprehensive confidence

[0089] of each candidate optimization strategy, and then select the candidate optimization strategy with the highest comprehensive confidence as the optimal optimization strategy of the current imaging parameter to be optimized.In a specific implementation, the semantic correlation between the scene feature vector and the policy feature vector is a key step in the fusion decision calculation, and the semantic correlation quantifies the matching degree of a specific optimization strategy and the current scene context. The calculation process starts with vector dot product operation. For the scene feature vector after encoding and the policy feature vector of a candidate optimization strategy, the system performs dot product operation. Dot product operation multiplies the values in the corresponding dimensions of the two vectors, and then sums all the products to obtain an initial correlation score. The initial correlation score reflects the alignment degree of the two vectors in direction. If the directions of the scene feature vector and the policy feature vector are closer, the initial correlation score obtained by the dot product operation is larger, indicating that the strategy is more strongly associated with the current scene in the semantic level. It can be understood that the result of dot product operation is a scalar value without normalization, and its absolute value is affected by the length of the original vector, so the initial correlation scores calculated between different strategy pairs may be in different numerical intervals and are not directly comparable.

[0090] The initial correlation score needs to be processed subsequently to be converted into a standardized measure. The system inputs the initial correlation score into a specially designed single-layer neural network for scale adjustment and normalization. The structure of this single-layer neural network is very simple, usually containing only one fully connected layer neuron and one activation function. The fully connected layer neuron performs a linear transformation on the initial correlation score, and the linear transformation involves a weight coefficient and a bias term. The values of the weight coefficient and the bias term are determined in the model training stage. The purpose of linear transformation is to adjust the scale and offset of the initial correlation score to make its distribution adapt to the subsequent activation function. The value after linear transformation is sent to the Sigmoid activation function, which has the property of mapping any real number to the interval [0, 1], and its output is a floating-point number between 0 and 1. This output value processed by the Sigmoid function is the final semantic correlation score. The closer the semantic correlation score is to 1, the higher the semantic correlation is, and the closer it is to 0, the lower the semantic correlation is. The system will iterate through all the policy feature vectors, repeat the above dot product operation and single-layer neural network processing process for each policy feature vector, and generate a corresponding semantic correlation score for each candidate optimization strategy, forming a set of correlation indicators for comparison.

[0091] In specific implementations, assigning appropriate weights to the semantic correlation score and the knowledge base matching degree is the core of the weighted decision model. The semantic correlation score is assigned a dynamic weight value, and the key characteristic of the dynamic weight value is that it is not fixed but is adaptively adjusted according to the complexity of the scene context description information. The complexity of the scene needs to be quantified by a computable index. One feasible method is to calculate the information entropy of the scene feature vector. A high scene feature vector information entropy value indicates that the vector carries a large amount of information and is uniformly distributed, corresponding to a complex scene state. Another method is to calculate the variance of each original feature value in the scene context description information. A large variance indicates that the scene feature fluctuates dramatically and has a high complexity.

[0092] The matching degree returned by the knowledge base is assigned a fixed weight value, which is a constant set according to prior knowledge during system design or initialization. The fixed weight value reflects the basic confidence level of the reliability of historical data in the imaging optimization knowledge base. The fixed weight value and the dynamic weight value jointly determine the relative importance of the semantic correlation score and the knowledge base matching degree in the final decision. In some embodiments, the initial values of the fixed weight value and the dynamic weight value can be determined by grid search on a validation set to optimize the decision effect. The dynamic weight value and the fixed weight value need to satisfy certain constraint relationships, such as ensuring that their sum is one or that their ratio is within a reasonable range, to avoid the weight of one side being too high or too low, leading to an unbalanced decision.

[0093] The weighted decision model calculates the comprehensive confidence by weighted summation. The calculation formula of the comprehensive confidence is: the comprehensive confidence is equal to the sum of the weighted semantic correlation score and the weighted knowledge base matching degree. The specific calculation process is to multiply the semantic correlation score by the dynamic weight value to obtain the weighted semantic correlation score, multiply the matching degree returned by the knowledge base by the fixed weight value to obtain the weighted knowledge base matching degree, and then add the two weighted values to obtain the comprehensive confidence of the candidate optimization strategy. The system will independently perform this weighted summation calculation for each candidate optimization strategy, thereby assigning a comprehensive confidence score to each strategy. The candidate optimization strategy with the highest comprehensive confidence score will be selected as the optimal optimization strategy for the current imaging parameter to be optimized. Alternatively, the weighted decision model can also use a more complex mechanism, such as using a small feedforward neural network to replace the weighted summation formula. The input of the neural network is the semantic correlation score and the knowledge base matching degree, and the output is the comprehensive confidence. The parameters of the neural network are obtained by training historical decision data.

[0094] To show the dynamic weight adjustment strategy more clearly, refer to Table 1, which provides a specific configuration illustrating the correspondence between scene complexity, dynamic weight, and fixed weight. It is emphasized that the threshold and weight values in Table 1 are illustrative, and need to be calibrated through experiments in actual systems.

[0095] Table 1: Scene complexity and weight configuration mapping table

[0096]

[0097] It can be understood that the role of the weighted decision model is to balance the adaptability of real-time scene analysis and the reliability of historical experience data. In the case of low scene complexity, the scene pattern is relatively simple and fixed, and the reliability of historical experience (knowledge base matching degree) is high, so it is reasonable to give it a higher fixed weight value. In the case of high scene complexity, the current scene may contain more uncertainty or belong to rare cases, and the direct matching degree of historical experience may decrease, so increasing the dynamic weight value and relying more on the semantic correlation degree calculated based on real-time data helps to make decisions that are more in line with the current special situation. This dynamic weight mechanism enhances the robustness and adaptability of the optimization system when dealing with diverse and complex field environments. Alternatively, the weight allocation strategy can also be learned and optimized online based on feedback results after strategy execution, so that the decision model can continuously improve itself over time.

[0098] Refer to Figure 4 With scene feature variance (complexity quantification indicator) as the horizontal axis, the correlation between dynamic weight, fixed weight, and decision-making comprehensive influence score is shown simultaneously, and the low, medium, and high complexity regions are divided by background color. Specifically, when the scene feature variance is <0.1 (low complexity region), the dynamic weight stabilizes at about 0.3, the fixed weight maintains 0.7, and the decision-making comprehensive influence score is about 0.6; as the variance enters the 0.1-0.3 interval (medium complexity region), the dynamic weight gradually rises to 0.5, the fixed weight simultaneously decreases to 0.5, and the decision-making comprehensive influence score gently rises to 0.7; when the variance is ≥0.3 (high complexity region), the dynamic weight jumps to 0.7 and remains stable, the fixed weight decreases to 0.3, and the decision-making comprehensive influence score rapidly rises to about 0.8. The threshold lines in the figure clearly define the boundaries of different complexity regions, intuitively presenting the correlation rule that "the higher the scene complexity, the higher the dynamic weight proportion and the higher the decision-making comprehensive influence score", and quantifying the gain effect of dynamic weight adjustment on decision-making results.

[0099] In an embodiment, generating the global imaging parameter optimization instruction set is a critical step to integrate all the independent decisions and make sure they are executable. The process starts with a consistency check on the selected optimal optimization strategies for each imaging parameter to be optimized. Assume that after the aforementioned fusion decision computation, the system has selected the following three optimal optimization strategies for three key imaging parameters: strategy A selected for the "aperture" parameter suggests "set the aperture value to f / 2.8 to get more light in", strategy B selected for the "shutter speed" parameter suggests "set the shutter speed to 1 / 1000s to freeze fast moving objects", and strategy C selected for the "sensitivity" parameter suggests "keep the sensitivity ISO value at 800 to control image noise". The system needs to check whether there is a parameter conflict among the aperture strategy A, shutter speed strategy B, and sensitivity strategy C. A parameter conflict means that the setting suggestions of different parameters cannot be satisfied simultaneously in a physical or logical way, or the simultaneous satisfaction will cause a significant decline in imaging quality. In an embodiment, the conflict detection is based on a predefined conflict rule base of imaging parameters, which contains the knowledge of the mutual constraint relationship among imaging parameters. For example, a typical conflict rule can be: "when the ambient light intensity is lower than a threshold L, if the shutter speed is higher than S max and the aperture is smaller than F min, it will cause underexposure". The system will substitute the selected aperture strategy A, shutter speed strategy B, sensitivity strategy C, and the currently measured ambient light intensity value into the rules in the conflict rule base for matching verification.

[0100] If a parameter conflict is detected, for example, in the case of a weak ambient light, the combination of the high shutter speed strategy B (1 / 1000s) and the small aperture strategy A (f / 2.8) may not be able to meet the light requirement for correct exposure even with the ISO sensitivity strategy C (ISO 800), the system will adjust the conflicting strategies according to the pre-defined priority rules. The pre-defined priority rules explicitly state which imaging quality attributes have higher priority when different imaging objectives cannot be satisfied simultaneously. For example, the priority rules can state: “correct exposure has higher priority than motion blur control, motion blur control has higher priority than depth of field control”. According to this rule, in the above conflict, “correct exposure” has the highest priority, therefore the motion blur control objective (corresponding to high shutter speed strategy B) or the depth of field control objective (corresponding to small aperture strategy A) that conflicts with it has lower priority and needs to be adjusted. The adjustment can be in the form of modifying the parameter suggestion value of the shutter speed strategy B or the aperture strategy A while ensuring correct exposure. For example, the system can reduce the shutter speed from 1 / 1000s to 1 / 250s, or increase the aperture from f / 2.8 to f / 2.0, or fine tune both parameters and recalculate whether the adjusted combination can meet the exposure requirement and conflict is eliminated. Alternatively, in complex conflict situations, the system can also abandon the currently selected strategy and replace it with the second highest ranked strategy in the fusion decision that has no conflict. The goal of the adjustment is to ensure that all parameter setting commands in the final instruction set are logically self-consistent and physically realizable, thus guaranteeing the internal consistency of the final instruction set.

[0101] After all detected parameter conflicts are eliminated, the system starts to convert the adjusted and self-consistent optimal optimization strategies into specific parameter setting commands executable by the hunting camera. The conversion process relies on a command mapping table that defines the mapping relationship between the abstract suggestion value of each imaging parameter and its corresponding specific control instruction recognizable by the hardware driver layer of the hunting camera. For example, the abstract value “f / 2.8” suggested by the aperture strategy A will be converted to a sequence of instructions sent to the aperture stepper motor controller by querying the command mapping table, which can include the pulse number corresponding to the target aperture blade position. The abstract value “1 / 1000s” suggested by the shutter speed strategy B will be converted to a specific register write value for the camera shutter controller. The abstract value “ISO 800” suggested by the ISO sensitivity strategy C will be converted to a specific voltage control code word for the image sensor analog gain channel. This conversion translates high-level optimization strategies into low-level hardware operation instructions.

[0102] The system combines these specific parameter setting commands into a global imaging parameter optimization instruction set in execution order. The determination of execution order needs to consider the dependency and timing requirements of camera hardware initialization and parameter setting. For example, a common execution order can be: first send focus motor control command to focus, then set image sensor ISO, then set aperture size, and finally set shutter speed and trigger exposure. Such an order can avoid transient effects caused by improper parameter setting timing, such as triggering the shutter before the aperture has been contracted in place. The global imaging parameter optimization instruction set is logically an ordered command list, each command containing the address identifier of the target hardware unit and the specific control data.

[0103] The execution process of the global imaging parameter optimization instruction set is the final link of the optimization method acting on physical devices. The system sends each parameter setting command in the global imaging parameter optimization instruction set to the corresponding control unit of the hunting camera in turn through the control bus inside the hunting camera. The control bus can be an I2C bus, an SPI bus, or other types of embedded system internal communication bus. The aperture setting command is sent to the aperture control unit, which is usually a stepper motor driver that accurately controls the position of the aperture blade according to the received pulse command, thereby changing the size of the aperture aperture. The shutter speed setting command is sent to the shutter control unit, which controls the movement speed of the mechanical shutter curtain or the exposure time of the electronic shutter according to the received register value. The ISO setting command is sent to the control interface of the image sensor itself to adjust the gain of the internal analog amplifier of the sensor. The focus distance command is sent to the focus motor drive unit to move the focus lens group in the lens to the specified position. Various image signal processing parameters, such as white balance gain, color matrix coefficient, sharpening intensity, etc., are sent to the corresponding configuration registers of the image processing chip.

[0104] The system needs to monitor the completion of important parameter settings, such as waiting for the return of the "positioning complete" signal from the aperture control unit, or waiting for the "parameter configuration success" response returned by the image processing chip. After all parameter adjustment commands have been sent and the confirmation feedback of successful execution has been received, the system determines that all parameter adjustments are complete. At this time, the system sends a capture instruction to the camera trigger system. The capture instruction can instruct the camera to take a single frame of image, or can instruct the camera to take multiple frames of image continuously at a certain interval. The camera trigger system starts the complete image capture process, including sensor exposure, signal readout, analog-to-digital conversion, and subsequent image processing pipeline. Finally, the hunting camera outputs the image data after multi-modal data fusion and parameter optimization. It can be understood that the entire execution process needs to be completed in a very short time to meet the rapid response requirements of the hunting camera for dynamic targets. Alternatively, the system can preliminarily evaluate the imaging effect after executing the complete imaging parameter optimization instruction set and capturing the image, and add the optimized parameter settings, scene context description information, and evaluation results of this time to the imaging optimization knowledge base as a new case record, realizing self-learning and continuous optimization of the system.

[0105] Referring to Figure 5 In the fusion decision calculation of the imaging parameters of the hunting camera, the comprehensive confidence evaluation of the candidate strategies needs to combine the knowledge base matching degree and the semantic association degree. In specific operation, the confidence of different candidate optimization strategies (such as strategies A to E) is supported by the "knowledge base matching degree" of the blue column, the "semantic association degree" of the orange column, and the "comprehensive confidence" of the green column after weighted calculation. Taking strategy C (ISO 800) as an example, its semantic association degree (orange column) is the best among all strategies, combined with a higher knowledge base matching degree (blue column), and finally forms a higher comprehensive confidence; and the knowledge base matching degree and the semantic association degree of strategy D (aperture f / 4.0) are at a lower level, so the comprehensive confidence is significantly lower than that of other strategies. In the parameter configuration process, the weight of the semantic association degree is dynamically adjusted according to the scene complexity, and the knowledge base matching degree adopts a fixed weight. The comprehensive confidence obtained by the weighted sum of the two is the core basis for the selection of the candidate strategy.

[0106] It should be noted that, in this text, relational terms such as first and second are used merely to distinguish one entity or action from another, without necessarily requiring or implying any such actual relationship or order between such entities or actions. Moreover, the terms "include", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device.

[0107] While embodiments of the present application have been shown and described, it is to be understood that various modifications, substitutions, replacements and variations can be made to these embodiments without departing from the principles and spirit of the present application, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A multi-modal data fusion based hunting camera imaging quality optimization method, characterized in that, The method comprises the following steps executed in sequence: Collecting multi-source heterogeneous data generated by the hunting camera in a working state, the multi-source heterogeneous data at least including visible light images, thermal infrared images, laser ranging point clouds, and inertial measurement unit data; Performing spatio-temporal alignment and registration on the collected multi-source heterogeneous data to generate synchronized multi-modal data with consistent spatio-temporal reference; Performing joint feature extraction on the synchronized multi-modal data to separate out a controllable imaging parameter set directly related to imaging quality and scene context description information reflecting the state of the shooting scene, including: Extracting brightness distribution, contrast, spectral features, and potential moving target regions from the aligned dual-spectrum image pair; extracting average distance, distance variation gradient, and target relative size of the shooting scene from the projected depth information; fusing the camera pose estimated from the inertial measurement unit data to determine whether the scene is mainly static or has intense relative motion; based on the extracted spectral features, depth features, and motion features, generating scene context description information; and simultaneously defining the current adjustable parameters of the camera as the controllable imaging parameter set; Iterating through the controllable imaging parameter set, for each imaging parameter to be optimized, performing the following operations: based on the imaging parameter to be optimized and the scene context description information, retrieving multiple candidate optimization strategies from a pre-set imaging optimization knowledge base; Using the scene context description information as the basis for discrimination, performing fusion decision calculation on the retrieved multiple candidate optimization strategies to select an optimal optimization strategy for the imaging parameter to be optimized; According to the optimal optimization strategies selected for all the imaging parameters to be optimized, generating a global imaging parameter optimization instruction set; Executing the global imaging parameter optimization instruction set to adjust the working parameters of the internal imaging components of the hunting camera and triggering the camera to capture images, finally obtaining optimized hunting camera images.

2. The multi-modal data fusion based hunting camera imaging quality optimization method of claim 1, wherein, Performing spatio-temporal alignment and registration on the collected multi-source heterogeneous data to generate synchronized multi-modal data with consistent spatio-temporal reference, including: Performing feature point-based image registration on the visible light images and the thermal infrared images to eliminate parallax and generate a pixel-level aligned dual-spectrum image pair; Projecting the laser ranging point cloud data to the image coordinate system of the aligned dual-spectrum image pair to establish a correspondence between each image pixel and depth information; According to the timestamp of the inertial measurement unit data, estimating and compensating the camera pose at the image capture moment to ensure that all data correspond to the same spatio-temporal observation point.

3. The multi-modal data fusion based hunting camera imaging quality optimization method of claim 1, wherein, Based on the imaging parameter to be optimized and the scene context description information, retrieving multiple candidate optimization strategies from a pre-set imaging optimization knowledge base, including: Taking the imaging parameter to be optimized as the main query key and the scene context description information as the auxiliary filtering condition; Performing fuzzy matching query based on semantic similarity in the imaging optimization knowledge base to preliminarily filter out candidate optimization strategy entries related to the current imaging context; According to the matching degree of the main query key and the auxiliary filtering condition with the content of the knowledge base entries, sorting and returning multiple candidate optimization strategies with the highest matching degree.

4. The multi-modal data fusion based hunting camera imaging quality optimization method of claim 3, wherein, The scene context description information is used as a basis for discrimination to perform fusion decision calculation on the retrieved multiple candidate optimization strategies, including: The scene context description information is encoded into a fixed-dimensional scene feature vector; The content description information of each candidate optimization strategy is encoded into a corresponding strategy feature vector; The semantic correlation degree between the scene feature vector and each strategy feature vector is calculated; A comprehensive confidence degree of each candidate optimization strategy is calculated through a weighted decision model in combination with the semantic correlation degree and the matching degree returned by the knowledge base; The candidate optimization strategy with the highest comprehensive confidence degree is selected as the optimal optimization strategy of the imaging parameter to be optimized.

5. The multi-modal data fusion based hunting camera imaging quality optimization method of claim 4, wherein, The scene context description information is encoded into a fixed-dimensional scene feature vector, including: The brightness feature, distance feature and motion feature contained in the scene context description information are normalized; A multi-layer perception model is used to perform nonlinear transformation and dimension reduction on the normalized multi-dimensional features; The dimension-reduced feature sequence is input into a sequence encoder to capture the dependency relationship between the features, and the final hidden state of the sequence encoder is the scene feature vector.

6. The multi-modal data fusion based hunting camera imaging quality optimization method of claim 5, wherein, The semantic correlation degree between the scene feature vector and each strategy feature vector is calculated, including: The scene feature vector and a strategy feature vector are subjected to dot product operation to obtain an initial correlation score; The initial correlation score is input into a single-layer neural network for scale adjustment and normalization, and a semantic correlation degree score between zero and one is output; All strategy feature vectors are traversed, and the above dot product and neural network processing are repeated to obtain a set of semantic correlation degree scores corresponding to each candidate optimization strategy.

7. The multi-modal data fusion based hunting camera imaging quality optimization method of claim 6, wherein, A comprehensive confidence degree of each candidate optimization strategy is calculated through a weighted decision model, including: A dynamic weight value is assigned to the semantic correlation degree score, and a fixed weight value is assigned to the matching degree returned by the knowledge base; The weighted semantic correlation degree score and the weighted knowledge base matching degree are added to obtain the comprehensive confidence degree of each candidate optimization strategy; The dynamic weight value is adjusted according to the complexity of the scene context description information. The more complex the scene, the higher the dynamic weight value.

8. The multi-modal data fusion based hunting camera imaging quality optimization method of claim 1, wherein, A global imaging parameter optimization instruction set is generated according to the optimal optimization strategies selected for all imaging parameters to be optimized, including: It is checked whether there is a parameter conflict between the optimal optimization strategies selected for each imaging parameter to be optimized; If there is a parameter conflict, the conflicting strategies are adjusted according to a preset priority rule to ensure the internal consistency of the final instruction set; The adjusted and consistent optimal optimization strategies are converted into specific parameter setting commands executable by the hunting camera, and are combined into a global imaging parameter optimization instruction set in execution order.

9. The multi-modal data fusion based hunting camera imaging quality optimization method of claim 8, wherein, The global imaging parameter optimization instruction set is executed to adjust the working parameters of the internal imaging components of the hunting camera and trigger the camera to capture images, including: Each parameter setting command in the global imaging parameter optimization instruction set is sent to the corresponding control unit of the hunting camera in turn; The control unit adjusts the aperture, shutter speed, sensitivity, focus distance and image signal processing parameters according to the received commands; After all the parameters are adjusted, send a capture instruction to the camera triggering system to obtain a single frame or multiple frames of optimized hunting camera images.

Citation Information

Patent Citations

  • Unmanned aerial vehicle shooting system control method based on adaptive optimization

    CN120610469A

  • Industrial vision adaptive illumination compensation system and method based on Bayesian optimization

    CN120725942A