An intelligent video analysis method and system that adapts to environmental changes
Through the intelligent video analysis method that adapts to environmental changes, local linear embedding or isometric mapping method is used to reduce the dimensionality of image features, generate optimal enhancement parameters, and combine visual continuity and semantic features to maintain time-sequence consistency constraints, and introduce reinforcement learning strategies for dynamic updates, solving the problem of unstable image processing effects under environmental changes in the existing technology, and improving the robustness and recognition accuracy of the video analysis system.
Patent Information
- Application Number
- CN202510619896.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The existing video analysis system lacks adaptive enhancement and feedback adjustment mechanisms in complex environments, resulting in unstable image processing effects and affecting the accuracy of identification and analysis.
Through the intelligent video analysis method that adapts to environmental changes, local linear embedding or isometric mapping method is used to reduce the dimensionality of image features, generate optimal enhancement parameters, and combine visual continuity and semantic features to maintain time-sequence consistency constraints, and introduce reinforcement learning strategies for dynamic updates.
It realizes adaptive enhancement of the system in complex environments, improves image quality stability and recognition accuracy, and enhances the robustness and adaptability of the system.
Smart Images

Figure CN120126061B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent image processing, and in particular to an intelligent video analysis method and system that is adaptive to environmental changes. Background Art
[0002] The widespread deployment of video analytics systems in scenarios such as security surveillance, traffic sensing, and smart terminals has placed higher demands on the intelligence and adaptability of image processing. Many current video enhancement and analysis solutions still rely on static parameter settings and cannot effectively adjust to changes in the external environment. In situations with sudden changes in lighting conditions, obstructions from rain and fog, or at night, system processing performance becomes unstable, and image clarity and detail integrity are difficult to guarantee, directly impacting the accuracy of subsequent recognition and analysis.
[0003] While image enhancement technology is widely used to improve image quality, most methods rely solely on fixed mappings or empirical models to generate enhancement strategies, lacking the ability to dynamically respond to the current environmental state. This approach can easily lead to unbalanced enhancement results. While image brightness may be improved, details may be lost or noise may be amplified, making the enhanced image difficult to use for practical analysis tasks.
[0004] Traditional systems often lack effective feedback mechanisms between enhancement and analysis. The processing process is a one-time operation, and once the enhancement strategy is set, it cannot be adjusted based on the analysis results. This limits the system's ability to respond to external changes. Consistency and structural preservation during image processing are also often overlooked. Enhanced images are prone to semantic drift and edge distortion, further reducing the robustness and accuracy of the analysis system. Summary of the Invention
[0005] In response to the deficiencies of the prior art, the present invention provides an intelligent video analysis method and system that is adaptive to environmental changes, which solves the problem that the existing video analysis system lacks adaptive enhancement and feedback adjustment mechanisms in complex environments.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: an intelligent video analysis method that adapts to environmental changes, comprising the following steps:
[0007] S1. extracting image features of each frame from an input video frame sequence, wherein the image features include brightness, contrast, color histogram, and high-dimensional features related to the scene;
[0008] S2. Using local linear embedding or isometric mapping to reduce the dimensionality of image features to obtain a low-dimensional feature representation for representing the image environment state;
[0009] S3. Generate corresponding image enhancement parameters based on the mapping relationship between the low-dimensional environment state representation of the current frame and the preset target enhancement state, wherein the mapping relationship is determined based on the optimal path in the low-dimensional space;
[0010] S4. Enhance the current frame image according to the image enhancement parameters, and in the enhancement process, constrain the visual continuity and semantic features of the image to maintain temporal consistency in a joint optimization manner, wherein the joint optimization step includes constructing a loss function, and the loss function includes:
[0011] The visual continuity term is used to constrain the gradient changes of the enhanced image to maintain edge stability;
[0012] Semantic feature preservation item, used to control the feature distance between the enhanced image and the original image in the semantic space to be less than a set threshold;
[0013] Temporal consistency term, used to minimize the pixel differences between enhanced images of adjacent frames;
[0014] S5. Use optical flow information to align enhanced images of adjacent frames to ensure consistency between frames in the temporal dimension.
[0015] S6. Introduce a reinforcement learning strategy to dynamically update the enhancement parameters based on the semantic recognition accuracy feedback in the image enhancement results through a deep Q-network model.
[0016] Preferably, the image feature extraction in step S1 includes performing color space conversion, gradient calculation and local texture analysis on the video frame, and combining them to form a high-dimensional feature vector for reflecting the current scene state.
[0017] Preferably, the low-dimensional feature representation after dimensionality reduction in step S2 retains the original structure and semantic relationship of the image, and is used to describe the environmental state of the current image frame, and the representation is used to guide the subsequent generation of enhancement parameters.
[0018] Preferably, the generation of the enhancement parameters in step S3 is obtained by calculating the optimal path from the current state point to the target state point in a low-dimensional embedding space, wherein the path is determined according to the optimal transmission theory and reflects the direction and amplitude of feature changes.
[0019] Preferably, the deep Q network model in step S6 takes the environmental state of the current frame and the enhancement feedback index as input, and outputs the enhancement strategy parameters for updating. The feedback index is the change in accuracy of the enhanced image in the semantic recognition task, and the reward function assigns positive and negative feedback based on the improvement or decrease in accuracy.
[0020] The present invention provides an intelligent video analysis method and system that is adaptive to environmental changes. It has the following beneficial effects:
[0021] 1. By constructing an environmental state modeling module, this invention dynamically captures and predicts environmental change trends, achieving the technical effect of enabling the system to adapt to different environmental conditions in real time. Traditional video processing systems, which rely heavily on static parameters, have poor stability in complex environments and cannot automatically identify external changes. This overcomes the technical shortcomings of such systems, which are environmentally sensitive and lack robustness.
[0022] 2. This invention combines an optimal enhancement map generation module with a consistency reasoning module to establish a more targeted image enhancement mechanism. This technical solution generates the most appropriate enhancement strategy for the current environment and video quality, ultimately significantly improving the visibility of video images. Compared with existing processing methods that mainly rely on fixed enhancement rules, this method avoids over-enhancement and distortion, effectively overcoming the algorithm's weak generalization ability.
[0023] 3. This invention introduces an enhancement feedback adjustment module, forming a closed-loop "enhancement-feedback-reenhancement" system. The system automatically adjusts image processing strategies based on inference results, demonstrating self-optimization capabilities. Unlike existing data enhancement methods, which rely solely on one-way processing, this invention makes the enhancement process responsive and adaptable, resolving the rigidity of traditional enhancement processes and their inability to adapt to changing scenarios.
[0024] 4. Through refined control of the image enhancement execution module, this invention achieves the ability to retain key visual information during the image enhancement process. This not only enhances video clarity but also avoids common issues such as edge blur and color shift. Compared to traditional image enhancement systems that fail in extreme environments such as low light and strong light, this effectively improves image quality stability and enhances the system's adaptability in complex real-world scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a flowchart of the steps of the method;
[0026] Figure 2 Modeling flow charts for environmental states;
[0027] Figure 3 Generate a flow chart for the optimal enhancement mapping;
[0028] Figure 4 This is the overall structural diagram of the system. DETAILED DESCRIPTION
[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the specification of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0030] Example 1
[0031] Please see the attached Figure 1 -Attached Figure 3 , an embodiment of the present invention provides an intelligent video analysis method that adapts to environmental changes, comprising the following steps:
[0032] S1. Extract image features of each frame from the input video frame sequence. Image features include brightness, contrast, color histogram, and high-dimensional features related to the scene.
[0033] Image feature extraction and environmental state modeling are crucial steps in implementing this invention. This process allows us to extract meaningful features from the input video frames and map them into a low-dimensional space, enabling more efficient subsequent image enhancement and processing. The following describes the technical solutions and implementations for image feature extraction and environmental state modeling in detail.
[0034] In this embodiment, the core task of image feature extraction and environmental state modeling is to map the high-dimensional features of the input video frame into a low-dimensional manifold space for further environmental state modeling. This approach reduces computational complexity while preserving important image information, providing effective data support for subsequent image enhancement, time series modeling, and feedback adjustment processes.
[0035] The input video frame sequence contains rich spatiotemporal information, including not only the changes in objects and background in the scene, but also various complex features such as lighting, color, and texture. Therefore, extracting high-dimensional features from the image and performing effective dimensionality reduction are key to subsequent image enhancement.
[0036] As an option, this embodiment uses a manifold learning method to reduce the dimensionality of image features. The advantage of manifold learning is that it can preserve the local geometric structure of image features, avoiding the global information loss that can occur with traditional dimensionality reduction methods. Manifold learning technology preserves the intrinsic relationships between individual features in an image by searching for embedded low-dimensional manifolds within a high-dimensional feature space. This method is particularly suitable for processing data with complex structures or nonlinear relationships.
[0037] Specifically, this example uses Local Linear Embedding (LLE) and Isomap as two typical methods for manifold learning. In practical applications, these two methods can be selected based on actual needs or used in combination. The specific implementation processes of these two methods are described below.
[0038] In one possible implementation, LLE is used to reduce the dimensionality of input image features. LLE assumes that the local neighborhood of a data point is linear in a low-dimensional space, thereby reducing the dimensionality through a linear representation in the neighborhood.
[0039] In this embodiment, the LLE processing process is divided into the following steps:
[0040] Constructing a neighborhood graph: For each frame, we calculate the Euclidean distance between feature vectors to find the K nearest neighbors of each feature point (K is a set constant). These nearest neighbors constitute the local neighborhood of the point.
[0041] Local linear reconstruction: For each data point, the point is reconstructed by other points in the local neighborhood. Under the local linear assumption, the feature point is approximated by the weighted sum of its neighborhood points. The weight coefficients can be obtained by solving an optimization problem to minimize the reconstruction error:
[0042] ;
[0043] in, is a data point, is the weight coefficient, It's a neighbor point.
[0044] After optimizing the weight coefficients, a low-dimensional embedding representation is obtained by solving the eigenvalue problem. This process aims to minimize the reconstruction error between data points to ensure that the local structure of the feature points can be preserved in the low-dimensional space.
[0045] As another option, Isomap can perform dimensionality reduction by preserving the global geometric structure between data points. The main idea of Isomap is to establish the global geometric structure of the data by measuring the shortest path distance between sample points.
[0046] In this embodiment, the processing of Isomap is as follows:
[0047] First, for each image frame’s feature vector, we select its K nearest neighbors (K is a constant) and construct a weighted adjacency graph. The edge weights in the graph are the Euclidean distances between image features.
[0048] Secondly, the classic Dijkstra algorithm is used to calculate the shortest path distance between each pair of data points in the adjacency graph to obtain a distance matrix.
[0049] Finally, the distance matrix is reduced to a low-dimensional manifold using the classic multidimensional scaling (MDS) method. MDS minimizes the distance error between data points, ensuring that the points in the low-dimensional space retain as much global geometric information as possible from the original space.
[0050] In this embodiment, the manifold learning method not only helps reduce the dimensionality of image features, but also preserves the local and global structural information of the image in a low-dimensional space, thereby effectively representing the environmental state. These low-dimensional representations provide a more accurate foundation for subsequent image enhancement and time series modeling.
[0051] In this embodiment, the low-dimensional representation obtained after image feature extraction and dimensionality reduction is used as the basis for environmental state modeling. The environmental state of each frame of image can be represented by its low-dimensional feature vector, and the environmental state trajectory is a state sequence composed of these low-dimensional vectors in chronological order.
[0052] The key to modeling environmental states is mapping the high-dimensional features of an image into a semantically meaningful low-dimensional space. Through manifold learning, the spatial structure of the image is better preserved in the low-dimensional manifold space, allowing the environmental state to more accurately reflect the true semantic features of the image.
[0053] In this embodiment, the low-dimensional representation obtained through manifold learning not only reflects the visual features of the image, but also captures the underlying semantic information in the image. The low-dimensional embedding vector of each image frame is considered to be the environmental state of the image in that frame, and the sequence of these environmental states represents the overall state of the image in the video over time.
[0054] The image feature extraction and environmental state modeling method in this embodiment uses manifold learning technology to reduce the dimensionality of input image features, preserving both the local and global structure of the image. It then constructs an environmental state model through a low-dimensional embedding representation. This method not only effectively reduces computational complexity but also provides accurate environmental state support for subsequent image enhancement, time series modeling, and feedback control.
[0055] Through these technical means, the present invention can adaptively adjust the image enhancement effect when processing video analysis tasks in complex environments, thereby improving the accuracy and robustness of the video analysis system.
[0056] S2. Using local linear embedding or isometric mapping to reduce the dimensionality of image features to obtain a low-dimensional feature representation for representing the image environment state;
[0057] After image feature extraction and environmental state modeling, the system needs to further generate an optimal enhancement map based on the extracted features and modeled states to achieve adaptive image enhancement effects under different environmental conditions. Generating the optimal enhancement map is a core component of the technical solution, linking low-dimensional feature representations with specific enhancement parameters to achieve targeted adjustments to the input image. To achieve this goal, the present invention combines multiple optimization methods, deep learning strategies, and explicit mathematical modeling to obtain the optimal enhancement map.
[0058] In this embodiment, the generation process of the optimal enhancement mapping includes multiple key modules, and these modules are closely connected through intermediate variables and calculation results.
[0059] The input low-dimensional environment state representation is mapped to enhancement parameters through a regression network or optimization module. This process must not only ensure that the enhancement effect matches the environment state, but also balance different objectives (such as brightness, contrast, color saturation, etc.).
[0060] Specifically, this embodiment adopts an optimization framework based on a multi-objective loss function to generate the optimal enhancement mapping.
[0061] As an option, this embodiment introduces the following optimization model:
[0062] First, define the low-dimensional state vector of the input image frame as , the enhancement parameter vector is ,in .here, represents the brightness gain coefficient, represents the contrast adjustment coefficient, Indicates the color saturation adjustment coefficient.
[0063] The optimization objective is defined as minimizing the following multi-objective loss function:
[0064] ;
[0065] in, 、 、L Represent the loss functions of brightness, contrast and color direction respectively, 、 、 is the weight coefficient, which is used to control the trade-off between various objectives.
[0066] In one possible implementation, the brightness loss function Defined as:
[0067] ;
[0068] in, For the enhanced image The brightness value of each pixel, is the target brightness value, is the total number of pixels.
[0069] Contrast loss function Defined as:
[0070] ;
[0071] in, For the enhanced image The local contrast of each pixel, is the target contrast value.
[0072] Color loss function Defined as:
[0073] ;
[0074] in, For the enhanced image The color saturation vector of pixels, is the target saturation vector.
[0075] In some embodiments, the optimization process is implemented in a deep neural network by back propagation, and the network input is , the output is The network structure can adopt a multi-layer perceptron (MLP), in which the feature layer and the output layer are connected through a nonlinear activation function.
[0076] As an alternative, a numerical optimization method based on gradient descent can be used to iteratively calculate:
[0077] ;
[0078] in, is the learning rate, is the gradient under the current enhancement parameters.
[0079] In practical applications, in order to ensure the smoothness of the enhancement effect, a regularization term is introduced into the optimization objective in this embodiment:
[0080] ;
[0081] in, Before Frame enhancement parameters, is the window size, which is used to penalize drastic changes in enhancement parameters between consecutive frames.
[0082] Therefore, the overall optimization goal can be expressed as:
[0083] ;
[0084] in, is the weight coefficient of the regularization term.
[0085] By combining the above modules, this embodiment can generate optimal enhancement mappings for different scenes, different lighting conditions, and different content conditions, providing reliable support for subsequent video enhancement and processing modules.
[0086] This modular design ensures the flexibility and scalability of the system, enabling it to adapt to more complex scenarios.
[0087] In some embodiments, the system can also be expanded to support other enhancement targets according to actual application requirements, such as sharpness, dynamic range adjustment, etc., to further improve the adaptability of the overall system.
[0088] S3. Generate corresponding image enhancement parameters based on the mapping relationship between the low-dimensional environment state representation of the current frame and the preset target enhancement state, and the mapping relationship is determined based on the optimal path in the low-dimensional space;
[0089] After generating the optimal enhancement maps, the system needs to apply these maps to the actual image and further improve the overall enhancement effect through a joint optimization strategy. Image enhancement and joint optimization are key components of the entire technical solution. Their mission is not only to perform map-based enhancement operations but also to globally coordinate and adjust multiple enhancement targets to achieve a consistent and smooth image output. This process is closely integrated with the aforementioned image feature extraction, environmental state modeling, and enhancement map generation, ensuring the integrity and scalability of the technical solution.
[0090] In this embodiment, the image enhancement and joint optimization process includes an image enhancement operation, a joint optimization module, and a global consistency constraint module.
[0091] The image enhancement module adjusts the brightness, contrast, saturation and other parameters of the input image pixel by pixel based on the optimal enhancement map generated above. Specifically, for the input image frame , the enhanced output image frame It can be expressed as:
[0092] ;
[0093] in, represents pixel coordinates, represents the color channel, To enhance the gain factor, is the offset coefficient.
[0094] As an option, a nonlinear adjustment function is introduced in this embodiment to enhance the adjustment effect:
[0095] ;
[0096] in, is a nonlinear activation function, such as the hyperbolic tangent function Or sigmoid function, used to avoid saturation or distortion caused by over-enhancement.
[0097] Specifically, the joint optimization module adopts a multi-objective optimization strategy to adjust the consistency of the enhancement effect in the spatial domain, temporal domain, and perceptual domain. The total loss function is defined as follows:
[0098] ;
[0099] in, represents the spatial consistency loss, represents the time smoothing loss, represents the perceived quality loss, 、 、 is the corresponding weight coefficient.
[0100] In one possible implementation, the spatial consistency loss is defined as:
[0101] ;
[0102] in, represents the image gradient, is the total number of pixels, is the pixel position.
[0103] The temporal smoothing loss is defined as:
[0104] ;
[0105] in, is the enhanced output of the previous frame.
[0106] The perceptual quality loss is defined as:
[0107] ;
[0108] in, represents a pre-trained perceptual feature extractor (e.g., intermediate features in a VGG network), Target high-quality images.
[0109] In some embodiments, a global consistency constraint module is also introduced to ensure that the statistical characteristics (such as histogram distribution, brightness mean, etc.) between multiple frames remain consistent. Define the global constraint loss:
[0110] ;
[0111] in, Represents the color histogram features of an image.
[0112] Specifically, in order to further expand the application scenarios, an interface is reserved in the joint optimization module in this embodiment to access additional enhancement targets (such as noise suppression, sharpness improvement, dynamic range expansion, etc.), thereby enhancing the adaptability and scalability of the system.
[0113] Through the above design, the present invention can not only achieve the enhancement effect of a single frame, but also maintain the stability and consistency of the enhancement effect in a time series, thereby providing a solid technical foundation for multi-frame video processing in complex environments.
[0114] S4. Enhance the current frame image according to the image enhancement parameters, and during the enhancement process, constrain the visual continuity and semantic features of the image to maintain temporal consistency in a joint optimization manner;
[0115] After image enhancement and joint optimization, to further improve the enhancement effect in multi-frame sequences, it is necessary to model the temporal dependencies between adjacent frames and achieve inter-frame alignment by combining optical flow information. Inter-frame temporal modeling and optical flow alignment are key steps in ensuring spatial and temporal consistency during multi-frame processing. They are closely integrated with the aforementioned image feature extraction, environmental state modeling, optimal enhancement map generation, and image enhancement and joint optimization, ensuring the stability and practicality of the overall solution.
[0116] In this embodiment, inter-frame temporal sequence modeling is mainly achieved by learning temporal dependencies using a recurrent neural network (RNN) or a long short-term memory network (LSTM), while an optical flow estimation module is used to achieve pixel-level alignment between adjacent frames.
[0117] The input image sequence can be represented as ,in Represents the time step image frames, is the sequence length.
[0118] Specifically, this embodiment calculates the optical flow field between adjacent frames to estimate pixel motion, where , representing pixels exist arrive The displacement between .
[0119] As an option, this embodiment adopts a pyramid optical flow estimation method. First, the image pyramid is layered from low resolution to high resolution to calculate the coarse to fine optical flow field to improve the accuracy and robustness of the estimation.
[0120] In one possible implementation, the optical flow alignment module uses bilinear interpolation to transform the current frame The eigenvector of Perform the transformation:
[0121] ;
[0122] in, Optical flow field The displacement of the corresponding pixel in .
[0123] To ensure that the timing information is effectively captured, this embodiment introduces a timing modeling module based on the gated recurrent unit (GRU). Define the hidden state , the input is the aligned features , the recursive formula is as follows:
[0124] ;
[0125] in, Represents a gated recurrent unit operation.
[0126] The loss function of time series modeling includes prediction error and time smoothing loss. The prediction error is defined as:
[0127] ;
[0128] in, is the prediction output based on time series modeling, is the target frame, is the number of frames.
[0129] The temporal smoothing loss is defined as:
[0130] ;
[0131] This loss is used to reduce output fluctuations between adjacent frames.
[0132] As an option, in order to improve the adaptability of the model to complex dynamic scenes, this embodiment also introduces an attention mechanism. Specifically, the attention weight is defined as :
[0133] ;
[0134] in, Feature-based The calculated attention score, Used to weight the feature contributions at different time steps.
[0135] In some embodiments, this embodiment also integrates edge preservation constraints based on optical flow to avoid edge blurring caused by motion. Define edge constraint loss:
[0136] ;
[0137] in, represents the gradient operator.
[0138] The comprehensive loss function is defined as:
[0139] ;
[0140] in, 、 are the weight coefficients of time smoothing and edge constraint respectively.
[0141] Through the above modular design, the present invention can effectively utilize temporal dependencies in multi-frame video processing and achieve cross-frame consistency through optical flow alignment, providing technical support for image enhancement in complex dynamic scenes.
[0142] S5. Use optical flow information to align enhanced images of adjacent frames to ensure consistency between frames in the temporal dimension.
[0143] After completing inter-frame temporal modeling and optical flow alignment, a reinforcement learning and policy update module is introduced to further enhance the system's adaptability and the stability of the enhancement effect. This module dynamically optimizes the enhancement policy using an interactive learning framework, enabling the system to continuously adjust its decisions based on varying input environments and output effects. The reinforcement learning and policy update component is closely integrated with the aforementioned image feature extraction, environmental state modeling, optimal enhancement map generation, image enhancement and joint optimization, and inter-frame temporal modeling, providing a feedback-based closed-loop optimization mechanism for the system.
[0144] In this embodiment, the reinforcement learning module is modeled based on a Markov decision process (MDP).
[0145] Define the system state as , which means that at time step The state of the environment. Action Indicates the reinforcement strategy selected in this state. Indicates taking action The immediate feedback obtained after the enhancement is usually related to the perceived quality or downstream task performance. Indicates that the status Next select action The probability distribution of .
[0146] Specifically, in this embodiment, the state transition probability is defined as:
[0147] Indicates that the status Take action Then transfer to the next state probability.
[0148] As an option, this embodiment adopts the policy gradient method in deep reinforcement learning to update the policy.
[0149] Define the policy objective function:
[0150] ;
[0151] in, are policy network parameters, is the discount factor, is the time step length, Indicates about strategy expectations.
[0152] In one possible implementation, policy updates are performed via gradient ascent:
[0153] ;
[0154] in, is the learning rate, is the policy gradient.
[0155] Specifically, the policy gradient is calculated as follows:
[0156] ;
[0157] in, From the time step Starting cumulative discounted returns.
[0158] As an alternative, this embodiment introduces a reinforcement learning method based on the actor-critic architecture. This method uses the actor network Generating Action-Critic Networks Estimated state value.
[0159] Define the loss function of the critic network:
[0160] ;
[0161] in, are the critic network parameters.
[0162] The critic network is used to reduce the variance of the policy gradient and improve learning stability.
[0163] In some embodiments, this embodiment also uses entropy regularization to encourage the exploration of the strategy:
[0164] ;
[0165] The overall optimization goal is:
[0166] ;
[0167] in, is the weight of the entropy regularization term.
[0168] Through the above design, the present invention can realize dynamic strategy updating based on reinforcement learning, effectively improve the adaptability and generalization ability of the image enhancement system in different scenarios and different tasks, and provide comprehensive technical support for complex application scenarios.
[0169] S6. Introduce a reinforcement learning strategy to dynamically update the enhancement parameters based on the semantic recognition accuracy feedback in the image enhancement results through a deep Q-network model.
[0170] In this embodiment, the image enhancement module incorporates a dynamic optimization mechanism based on a deep Q-network. This mechanism continuously explores and optimizes enhancement parameters through a reinforcement learning strategy, aiming to automatically improve the enhancement strategy based on feedback from semantic recognition accuracy. The system architecture primarily comprises a state acquisition unit, an action decision unit, a reward calculation unit, and a parameter update unit, all of which collaborate to dynamically adjust image enhancement parameters.
[0171] The state acquisition unit extracts the current enhanced image features and semantic recognition accuracy data as environmental state inputs to the deep Q-network model. Alternatively, the state acquisition unit can extract enhanced features across multiple dimensions, including but not limited to image brightness, contrast, sharpness, and noise level, and use the accuracy output by the recognition module as a performance feedback metric.
[0172] Specifically, the action decision unit selects an appropriate enhancement operation from a predefined set of actions based on the current state information. The action set can include fine-tuning of parameters, such as increasing or decreasing the amplitude of brightness, contrast, sharpening intensity, etc. The action selection follows a deep Q-learning strategy, and its decision is based on the following formula:
[0173] ;
[0174] in, Indicates that the status Next action The expected cumulative reward when is the current parameter of the Q network, is the immediate reward value, is the discount factor, is the new state after executing the action, For the next action, is the parameter of the target Q network. The model is continuously updated , to achieve strategy improvement.
[0175] The reward calculation unit is used to calculate the instant reward value The reward value is measured by the change in recognition accuracy. Specifically, the reward calculation can be done using the following formula:
[0176] ;
[0177] in, Indicates the recognition accuracy after processing with the current enhancement parameters. Indicates the last recognition accuracy, is the reward scaling factor, which is used to adjust the sensitivity of the reward.
[0178] The parameter update unit dynamically updates the image enhancement parameters based on the output of the deep Q network. The updated parameters serve as the input for the next round of enhancement operations, forming a closed-loop feedback mechanism. The update process meets the following conditions:
[0179] ;
[0180] in, is the learning rate, Is the loss function, the commonly used loss function is the mean square error (MSE) form:
[0181] ;
[0182] In one possible implementation, the system uses an experience replay mechanism to improve learning efficiency. This involves storing historical state, action, reward, and next-state tuples in an experience pool and updating the Q-network parameters through random sampling. This strategy helps break down correlations between samples and improves training stability.
[0183] In some embodiments, a Double Deep Q Network (Double DQN) architecture is introduced to alleviate the Q-value overestimation problem in traditional DQN. Double DQN improves the robustness of parameter updates by separating action selection and Q-value evaluation.
[0184] Example 2
[0185] Please see the attached Figure 4As part of this application, an embodiment of the present invention further provides an intelligent video analysis system that is adaptive to environmental changes, including:
[0186] The environmental state modeling module is used to generate an environmental state model based on environmental change data; the environmental state modeling module includes: an environmental data receiving unit, used to receive environmental data from an environmental change source; a state modeling unit, used to generate an environmental change model based on the environmental data and predict future changes in the environment.
[0187] Within the overall system architecture, the environment state modeling module is a key component connecting image feature extraction, optimal enhancement map generation, and subsequent image enhancement. By accurately modeling the environment state of the input frame sequence, this module provides a stable, low-dimensional, and semantically expressive state representation for enhancement decisions. This module not only ensures the targeted enhancement strategy but also provides reliable state input for modules such as multi-frame time series modeling, optical flow alignment, and reinforcement learning policy updates. Therefore, the environment state modeling module plays a crucial role in connecting the preceding and subsequent stages of the solution.
[0188] In this embodiment, the core goal of the environment state modeling module is to map the input high-dimensional image features into low-dimensional state representations. , so as to serve as input for subsequent modules.
[0189] Input image frame After being processed by the feature extraction module, a high-dimensional feature vector is obtained The feature vector contains a variety of information such as scene brightness, contrast, texture, edge distribution, etc., but it is not suitable to use it directly. As an environmental state, it will lead to excessive dimensionality and redundancy, so dimensionality reduction is required.
[0190] Specifically, this embodiment uses a manifold learning method to Perform dimensionality reduction to obtain state representation .
[0191] As an option, this embodiment uses the local linear embedding (LLE) method for dimensionality reduction.
[0192] In this method, we first need to determine each eigenvector of Nearest neighbors, defined using Euclidean distance:
[0193] ;
[0194] Next, the reconstruction weight is determined by optimizing the following reconstruction error :
[0195] ;
[0196] in, express Neighborhood, for By its neighbors Reconstruction weights.
[0197] Finally, the low-dimensional embedding is obtained by solving the following eigenvalue problem :
[0198] ;
[0199] in, Represents the state vector after dimensionality reduction.
[0200] To ensure that the state vector after dimensionality reduction is semantically interpretable, this embodiment introduces multi-task supervision signals during the training phase, combining image quality assessment indicators and environmental labels (such as lighting categories and weather conditions) to jointly optimize the state modeling module.
[0201] Through the above design, the environmental state modeling module of this embodiment can provide low-dimensional, stable, and semantically rich state representation for downstream modules, playing a key supporting role in the entire enhancement system.
[0202] An optimal enhancement map generation module, used to generate an optimal image enhancement map according to an environment state model;
[0203] The optimal enhancement map generation module includes: a feature extraction unit for extracting image features from video data; an enhancement map generation unit for generating an optimal image enhancement map based on the extracted image features and an environment change model.
[0204] After the environmental state modeling is complete, the system must generate a matching image enhancement strategy based on the current state to achieve targeted and adaptive enhancement operations. The optimal enhancement mapping generation module plays a key role in this process. This module efficiently transforms environmental states into enhancement operations by constructing a nonlinear mapping relationship from the state to the enhancement parameter space. This module not only directly determines the specific execution of the enhancement operation but also provides the necessary parameter support for subsequent image enhancement execution and joint optimization modules. Closely integrated with the environmental state modeling module, it is the core backbone of the enhancement system's end-to-end performance optimization.
[0205] In this embodiment, the optimal enhancement map generation module is designed to represent the state of the current frame or frame sequence according to the , generate a set of optimal enhancement parameters , to control the specific operations of image enhancement.
[0206] State vector It is a low-dimensional feature vector output from the previous module, which has the ability to express factors such as scene brightness, contrast, and detailed structure.
[0207] Specifically, this embodiment uses a multi-layer perceptron (MLP) to construct an enhanced mapping function:
[0208] ;
[0209] in, Indicated by the parameter The neural network mapping function of the control, output A set of enhanced parameters.
[0210] As an option, Includes the following sub-parameters:
[0211] Brightness adjustment factor ; Contrast adjustment factor ; Color temperature adjustment factor ;Detail enhancement kernel scale ; Therefore, the enhancement parameter vector is defined as:
[0212] ;
[0213] In one possible implementation, the MLP contains three hidden layers, each layer uses the ReLU activation function, and the last layer uses Sigmoid or Tanh normalization to enhance the parameter output range.
[0214] In order to improve the adaptability of enhancement parameters to image content, this embodiment introduces a conditional normalization mechanism, using the following structure for each parameter:
[0215] ;
[0216] in, Indicates the enhancer parameters, For the submodules, Represents a normalization function, for example:
[0217] ;
[0218] in, 、 It is a fixed value determined by pre-training and is used to limit the output range.
[0219] In some embodiments, in order to enhance the correlation modeling between parameters, a collaborative attention mechanism is used for feature weighting.
[0220] The collaborative attention weights are obtained as follows:
[0221] ;
[0222] ;
[0223] in, and is a learnable parameter, Represents element-wise multiplication.
[0224] In one possible implementation, this module supports enhanced mapping modeling based on temporal states, using a sliding state sequence:
[0225] ;
[0226] And generate timing enhancement parameters through a timing network (such as GRU):
[0227] ;
[0228] In order to further improve the robustness of the generated parameters, this embodiment uses a regularization term based on KL divergence to limit the range of variation of the enhanced parameter distribution and avoid sudden changes in the generated parameters. The formula is as follows:
[0229] ;
[0230] in, represents the current parameter distribution, is the reference distribution, which is usually set to Gaussian distribution.
[0231] The training of the enhanced mapping function adopts a combination of supervised learning and reinforced feedback:
[0232] The supervision part uses the L2 distance with the expert enhancement parameters as the loss function;
[0233] The enhanced feedback part combines perceptual indicators after image enhancement, such as SSIM, LPIPS or recognition accuracy as reward signals;
[0234] The joint loss function is:
[0235] ;
[0236] in, 、 、 are weight coefficients respectively.
[0237] Through the above structural design, the optimal enhancement map generation module of this embodiment can dynamically derive appropriate enhancement parameters according to the environmental state, providing comprehensive and accurate support for subsequent image enhancement and optimization modules.
[0238] The image enhancement execution module is used to enhance the video data according to the optimal enhancement mapping; the enhancement feedback adjustment module includes: a feedback signal receiving unit, which is used to receive feedback signals from the environment and video data processing; an adjustment decision unit, which is used to generate adjustment decisions based on the feedback signals; an adjustment execution unit, which is used to execute the adjustment decisions and adjust the system behavior; and a feedback adjustment optimization unit, which is used to optimize the adjustment strategy according to the adjustment results.
[0239] After the optimal enhancement map generation module derives the enhancement parameters for the current scene, the image enhancement execution module receives these parameters and, combined with the input image frame or its feature representation, completes the actual image enhancement operation. This module is the specific execution carrier of the enhancement strategy, acting directly on the original or intermediate image data and outputting the enhanced image result. It has an explicit parameter transmission relationship with the previous module and is the key channel for achieving the transition from "strategy derivation" to "visual effect implementation." Therefore, the image enhancement execution module should not only ensure the controllability and adaptability of the enhancement operation, but also maintain consistency with the enhancement parameter space in terms of algorithm design.
[0240] In this embodiment, the core task of the image enhancement execution module is to For the original image or its intermediate representation Perform differentiable enhancement operations to generate enhanced images .
[0241] The image enhancement execution module consists of multiple enhancement operators or enhancement sub-networks, and the parameters of the enhancement operators are provided by the enhancement map generation module.
[0242] Specifically, this embodiment defines the enhancement process in the following form:
[0243] ;
[0244] in, represents the set of enhancement operators, is the image to be enhanced, is the enhancement parameter vector, To enhance the results.
[0245] In one possible implementation, the enhancement process is modeled as a combination of multi-level linear and nonlinear operations. Taking brightness and contrast enhancement as an example, the enhancement operation can be expressed as:
[0246] ;
[0247] in, are image coordinates, represents the contrast factor, Indicates the brightness offset.
[0248] As an option, in order to improve the ability to express enhanced details, this embodiment introduces a multi-scale enhancement structure based on the Laplacian pyramid.
[0249] Enhancement operation at each scale The above expression is:
[0250] ;
[0251] in, Indicates the Layer Laplacian image, is the enhancement operator of the corresponding scale.
[0252] Image reconstruction after multi-scale enhancement:
[0253] ;
[0254] in, represents the upsampling operation, The layer is low resolution.
[0255] In some embodiments, to support complex texture and color adjustments, a convolutional neural network is introduced as an enhancement execution unit. The input of the enhancement sub-network is the concatenation of the image feature map and the enhancement parameters:
[0256] ;
[0257] in, Indicates that the parameter is affected by Enhanced network control, is the image feature.
[0258] In the specific design, It includes multiple deformable convolutional layers, with parameters given by Control, in the form of:
[0259] ;
[0260] in, is the set of convolution kernel positions, is the convolution weight controlled by the enhancement parameter, is the offset.
[0261] To improve the stability and controllability of enhancement, this embodiment also designs a weight normalization strategy to keep the enhancement operation numerically stable.
[0262] The policy is defined as follows:
[0263] ;
[0264] ;
[0265] As an option, to achieve style transfer enhancement, the image enhancement execution module supports the AdaIN structure, namely adaptive instance normalization:
[0266] ;
[0267] in, 、 is the mean and standard deviation of the feature map, 、 are the target mean and standard deviation generated by the enhancement parameters.
[0268] To improve the end-to-end trainability of the system, all enhancement operators are constructed as differentiable modules to ensure that gradients can be passed back to the enhancement mapping module and state modeling module.
[0269] Through the above-mentioned method, the image enhancement execution module in this embodiment can effectively realize the functional transformation of enhancement parameters, complete the adjustment and optimization of the input image, and provide efficient, stable, and well-structured image processing support in the enhancement system.
[0270] The consistency reasoning module is used to perform consistency reasoning analysis on the enhanced image data;
[0271] In modern artificial intelligence systems, a consistency reasoning module is often required to improve the reliability and accuracy of the reasoning process. This module helps the system maintain logical coherence and effectiveness during complex decision-making processes by determining the consistency between input information. In this embodiment, a detailed technical solution design has been developed for this consistency reasoning module, aiming to provide the necessary constraints and judgment mechanisms for the reasoning process, ensuring that the system can avoid contradictory and inconsistent reasoning results during multi-step reasoning.
[0272] In this embodiment, the consistency reasoning module mainly includes an information input receiving unit, an inference constraint setting unit, a contradiction detection unit and an output result generation unit. The units are connected through a data transmission channel and work together to complete the consistency reasoning task. Specifically, the information input receiving unit is responsible for receiving input information from different modules or external systems, and performing preliminary processing on the information to adapt to subsequent reasoning steps. The inference constraint setting unit generates corresponding inference constraints based on predetermined rules or models. These constraints will play a key role in the reasoning process. The contradiction detection unit is used to identify contradictions between information and provide correction suggestions or judgments to terminate the reasoning process. Finally, the output result generation unit outputs the consistency test results or the revised reasoning conclusions based on the final judgment of the reasoning process.
[0273] The information input receiving unit receives external input information through various interfaces. This information can be text, images, or other forms of data. Optionally, the receiving unit can utilize a natural language processing module to perform operations such as word segmentation and semantic analysis on the input text information, ensuring standardized information format and structure.
[0274] Specifically, the inference constraint setting unit sets corresponding inference rules based on the type and attributes of the input information. These rules are not only based on a pre-defined knowledge base but also take into account the influence of contextual information. For example, an inference constraint can be expressed in the form of a mathematical formula, as shown below:
[0275] ;
[0276] in, Represents a set of conditions that meet consistency requirements, Indicates input information The constraints are the specific rules derived from the algorithm. These constraints play a key role in the reasoning process, ensuring that the inherent logical relationship between information is maintained.
[0277] The contradiction detection unit is a core module in the reasoning process. This unit automatically identifies contradictions by gradually analyzing various input information during the reasoning process and proposes corrections using a specific algorithm. For example, when the input information contains logical contradictions, the contradiction detection unit uses a mathematical model to determine the reasoning contradiction and provides a formula for detecting the reasoning contradiction, as shown below:
[0278] ;
[0279] in, Display information and The difference measure between and Represent information separately and No. If the difference measure exceeds the set threshold, it can be considered a contradiction, which triggers the correction mechanism.
[0280] The output result generation unit is responsible for converting the various judgment results from the reasoning process into actionable output information. In some embodiments, the output results include not only the reasoning conclusions but also recommended correction strategies, detailed logs of the reasoning process, and other information. To ensure system stability and operability, the output result generation unit typically needs to consider the differences in requirements in different scenarios.
[0281] In summary, the consistency reasoning module in this embodiment effectively improves the accuracy and reliability of the reasoning process by determining and correcting the consistency of input information. By combining multiple steps such as constraint conditions, contradiction detection, and output generation, this module can provide stable support in complex systems, ensuring that the reasoning process is not disrupted by contradictory information. In future technological iterations, the module's scalability and adaptability will further enhance its application value in various fields.
[0282] Enhanced feedback regulation module, used to adjust system behavior based on consistency reasoning results.
[0283] In complex systems, feedback and regulation mechanisms are key components for ensuring system performance and stability. Through real-time feedback and regulation, the system can continuously optimize its behavior and respond to changes in the external or internal environment. In this embodiment, the design of the enhanced feedback regulation module aims to further enhance the system's adaptability and regulation accuracy. This module enhances the existing feedback mechanism, enabling more precise regulation of system behavior and ensuring efficient operation in complex environments.
[0284] The enhanced feedback regulation module in this embodiment includes a feedback signal receiving unit, a regulation decision unit, a regulation execution unit, and a feedback adjustment optimization unit. These units exchange data through information flow and work together to contribute to the feedback regulation process. Specifically, the feedback signal receiving unit receives feedback signals from within or outside the system and performs necessary processing on these signals. The regulation decision unit makes decisions based on the received feedback signals and preset regulation rules, generating regulation instructions. The regulation execution unit is responsible for executing the regulation decisions and adjusting the relevant system parameters. The feedback adjustment optimization unit optimizes the regulation mechanism based on the regulation results to improve the overall performance of the system.
[0285] The feedback signal receiving unit collects feedback data from multiple sources. These sources can include sensors within the system, user input, and external environmental factors. Optionally, the receiving unit can combine these data sources and employ preprocessing algorithms to perform operations such as denoising and normalizing the received signals, ensuring that the feedback signal meets the system's processing requirements before being input into the adjustment decision unit.
[0286] Specifically, the adjustment decision unit performs decision analysis based on the received feedback signal and preset rules. The decision process can be implemented through a certain mathematical model or algorithm. For example, by establishing a control model to achieve the mapping between signal processing and adjustment decision. The following formula shows an example of the adjustment decision process:
[0287] ;
[0288] in, Indicates the adjustment decision result, Represents the feedback signal, Indicates the current state of the system or previous adjustment information. Function The mapping relationship is defined by the system's regulation rules or models. By continuously adjusting this mapping relationship, the regulation decision unit can achieve precise regulation of the system's behavior.
[0289] The primary function of the control execution unit is to execute actual system adjustments based on the output of the control decision unit. Typically, the control execution unit adjusts system parameters that affect performance. For example, in a control system, the execution unit can adjust control signals or parameter values to achieve precise control of the system state. The operation of the control execution unit can be expressed as follows:
[0290] ;
[0291] in, Indicates time The adjustment parameter value of The adjustment quantity generated by the adjustment decision unit is represented by the adjustment execution unit. The adjustment execution unit adjusts the system parameters according to this formula and applies it to the system operation.
[0292] In summary, the enhanced feedback regulation module in this embodiment provides effective regulation capabilities through precise feedback signal processing, regulation decision-making, and execution mechanisms. During system operation, the module can adjust the system state based on real-time feedback signals and optimize the regulation mechanism based on the feedback results, thereby ensuring that the system maintains optimal performance under different environments. By introducing a self-learning algorithm and optimization mechanism, the enhanced feedback regulation module can continuously improve the regulation effect under changing conditions, showing broad application prospects and good scalability.
[0293] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent video analysis method that adapts to environmental changes, characterized in that: The following steps are involved: S1. extracting image features of each frame from an input video frame sequence, wherein the image features include brightness, contrast, color histogram, and high-dimensional features related to the scene; S2. Using local linear embedding or isometric mapping to reduce the dimensionality of image features to obtain a low-dimensional feature representation for representing the image environment state; S3. Generate corresponding image enhancement parameters based on the mapping relationship between the low-dimensional environment state representation of the current frame and the preset target enhancement state, wherein the mapping relationship is determined based on the optimal path in the low-dimensional space; S4. Enhance the current frame image according to the image enhancement parameters, and in the enhancement process, constrain the visual continuity and semantic features of the image to maintain temporal consistency in a joint optimization manner, wherein the joint optimization step includes constructing a loss function, and the loss function includes: The visual continuity term is used to constrain the gradient changes of the enhanced image to maintain edge stability; Semantic feature preservation item, used to control the feature distance between the enhanced image and the original image in the semantic space to be less than a set threshold; Temporal consistency term, used to minimize the pixel differences between enhanced images of adjacent frames; S5. Use optical flow information to align enhanced images of adjacent frames to ensure consistency between frames in the temporal dimension. S6. Introduce a reinforcement learning strategy to dynamically update the enhancement parameters based on the semantic recognition accuracy feedback in the image enhancement results through a deep Q-network model.
2. The intelligent video analysis method according to claim 1, wherein: The image feature extraction in step S1 includes color space conversion, gradient calculation and local texture analysis of the video frame, and combining them to form a high-dimensional feature vector to reflect the current scene state.
3. The intelligent video analysis method according to claim 1, wherein: The low-dimensional feature representation after dimensionality reduction in step S2 retains the original structure and semantic relationship of the image, and is used to describe the environmental state of the current image frame, and the representation is used to guide the subsequent generation of enhancement parameters.
4. The intelligent video analysis method according to claim 1, wherein: The generation of the enhancement parameters in step S3 is obtained by calculating the optimal path from the current state point to the target state point in the low-dimensional embedding space. The path is determined according to the optimal transmission theory and reflects the direction and amplitude of feature changes.
5. The intelligent video analysis method according to claim 1, wherein: The deep Q network model in step S6 takes the environmental state of the current frame and the enhancement feedback index as input, and outputs the enhancement strategy parameters for updating. The feedback index is the change in accuracy of the enhanced image in the semantic recognition task, and the reward function assigns positive and negative feedback based on the improvement or decrease in accuracy.
6. An intelligent video analysis system that is adaptive to environmental changes, applied to the intelligent video analysis method that is adaptive to environmental changes according to any one of claims 1 to 5, characterized in that: include: Environmental state modeling module, used to generate environmental state model based on environmental change data; An optimal enhancement map generation module, used to generate an optimal image enhancement map according to an environment state model; an image enhancement execution module, configured to perform enhancement processing on the video data according to the optimal enhancement mapping; The consistency reasoning module is used to perform consistency reasoning analysis on the enhanced image data; Enhanced feedback regulation module, used to adjust system behavior based on consistency reasoning results.
7. The intelligent video analysis system capable of adapting to environmental changes according to claim 6, characterized in that: The environmental state modeling module includes: An environmental data receiving unit, configured to receive environmental data from an environmental change source; The state modeling unit is used to generate an environmental change model based on environmental data and predict future changes in the environment.
8. The intelligent video analysis system capable of adapting to environmental changes according to claim 6, characterized in that: The optimal enhancement map generation module includes: A feature extraction unit, configured to extract image features from video data; The enhancement map generation unit is used to generate an optimal image enhancement map according to the extracted image features and the environment change model.
9. The intelligent video analysis system capable of adapting to environmental changes according to claim 6, characterized in that: The enhanced feedback regulation module includes: A feedback signal receiving unit, configured to receive feedback signals from the environment and video data processing; an adjustment decision unit, configured to generate an adjustment decision according to the feedback signal; An adjustment execution unit, configured to execute the adjustment decision and adjust the system behavior; The feedback adjustment optimization unit is used to optimize the adjustment strategy according to the adjustment results.
Citation Information
Patent Citations
Coal production operation scene video AI algorithm training reasoning platform and method
CN119152407A
Intelligent image recognition system and method based on deep learning
CN119963950A