Lossy compression control method based on AI and video monitoring holder

By constructing a normalized motion intensity sequence and a specific event confidence sequence, and combining a weighted bipartite graph with a relational tree, a lossy compression control strategy is generated. This solves the problems of real-time transmission and integrity of multi-source heterogeneous data in remote mountainous monitoring scenarios in existing technologies, and improves data compression efficiency and event transmission.

CN121125979APending Publication Date: 2025-12-12ZHEJIANG RISESUN SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511266452.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing intelligent video surveillance systems struggle to balance data integrity and real-time transmission for critical events in complex scenarios such as remote mountainous areas, low bandwidth, and high packet loss rates during emergency monitoring. Existing methods have also failed to effectively integrate multi-source heterogeneous data for joint decision-making.

Method used

By acquiring sensor data and video data, a normalized motion intensity sequence and a specific event confidence sequence are constructed to quantify the dynamic correlation between sensor and video data. A weighted bipartite graph and a multi-mode association tree are used to optimize data importance assessment and redundancy discrimination. A lossy compression control strategy is generated based on the proportional coefficient and comprehensive value score.

Benefits of technology

It significantly improves data compression efficiency and event transmission integrity in complex monitoring scenarios in remote mountainous areas, providing a universal, value-based compression control method for intelligent IoT monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121125979A_ABST
    Figure CN121125979A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent data compression and transmission, in particular to an AI-based lossy compression control method and a video monitoring holder. An AI-based lossy compression control method comprises the following steps: S1, acquiring sensor data and video data, acquiring N original reading sequences according to the sensor data, and acquiring global motion intensity, local motion intensity and specific event confidence according to the video data; and according to the global motion intensity, the local motion intensity and the specific event confidence, respectively constructing a normalized global motion intensity sequence, a normalized local motion intensity sequence and a normalized specific event confidence sequence. According to the method, multi-modal monitoring data is fused, the dynamic association model is constructed, self-adaptive adjustment of the compression strategy is realized through weighted optimization and value evaluation, the data transmission efficiency and event integrity in a resource-limited scene are improved, and the method is suitable for intelligent monitoring of a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent data compression and transmission, and in particular to an AI-based lossy compression control method and a video monitoring PTZ. Background Technology

[0002] AI-based adaptive lossy compression technology is a key research direction in the field of intelligent video surveillance. It aims to dynamically adjust compression strategies by analyzing content characteristics to improve transmission efficiency in bandwidth-constrained environments. In recent years, various intelligent algorithms have been applied to video compression control, significantly optimizing bitrate allocation and storage resource utilization.

[0003] However, most existing methods rely solely on information within the video for analysis, failing to effectively integrate heterogeneous data from multiple sources, such as IoT sensors, for joint decision-making. In complex scenarios such as remote mountainous areas, low bandwidth, and high packet loss rates in emergency monitoring, existing systems struggle to balance data integrity and real-time transmission for critical events, limiting the reliability of current technologies in important monitoring tasks.

[0004] Therefore, there is an urgent need to build an intelligent system that can integrate multimodal data, evaluate information value in real time, and dynamically generate compression strategies to support stable and efficient data transmission and monitoring needs in resource-constrained scenarios. Summary of the Invention

[0005] To overcome the shortcomings of lack of cross-modal collaboration, this invention provides an AI-based lossy compression control method and a video monitoring PTZ.

[0006] The technical implementation of this invention is: an AI-based lossy compression control method, comprising the following steps: S1: Acquire sensor data and video data; obtain N raw reading sequences based on the sensor data; obtain global motion intensity, local motion intensity, and specific event confidence based on the video data; construct normalized global motion intensity sequence, normalized local motion intensity sequence, and normalized specific event confidence sequence based on the global motion intensity, the local motion intensity, and the specific event confidence, respectively. S2: Calculate the proportionality coefficients for the N original reading sequences and the normalized global motion intensity sequence; calculate the similarity scores for the N original reading sequences and the normalized local motion intensity sequence; S3: Construct a weighted undirected bipartite graph based on the similarity scores, and construct a weighted bipartite graph based on the weighted undirected bipartite graph; construct an association tree based on the weighted bipartite graph; S4: Obtain the comprehensive value score of each node based on the relationship tree; generate a lossy compression control strategy based on the proportional coefficient and the comprehensive value score.

[0007] Preferably, the step of acquiring sensor data and video data, obtaining N raw reading sequences based on the sensor data, and obtaining global motion intensity, local motion intensity, and confidence level of a specific event based on the video data includes: The sensor data is sorted in chronological order to obtain N raw reading sequences; Global motion intensity is obtained by calculating the average difference between all pixels between two consecutive frames; The target region is determined by a target detection algorithm, and the local motion intensity is obtained by calculating the average optical flow within the target region. A lightweight image classification network is used to perform forward propagation on video frames and output the confidence scores for specific events.

[0008] Preferably, the step of constructing a normalized global motion intensity sequence, a normalized local motion intensity sequence, and a normalized specific event confidence sequence based on the global motion intensity, the local motion intensity, and the specific event confidence score, respectively, includes: The global motion intensity, the local motion intensity, and the confidence level of the specific event are normalized to obtain normalized global motion intensity, normalized local motion intensity, and normalized confidence level of the specific event. Based on the normalized global motion intensity, the normalized global motion intensity values ​​corresponding to each frame are sorted by timestamp to construct a normalized global motion intensity sequence. Based on the normalized local motion intensity, the normalized local motion intensity values ​​corresponding to each frame are sorted by timestamp to construct a normalized local motion intensity sequence. Based on the normalized specific event confidence score, the normalized specific event confidence score values ​​corresponding to each frame are sorted by timestamp to construct a normalized specific event confidence score sequence.

[0009] Preferably, the step of calculating the scaling factor from the N original reading sequences and the normalized global motion intensity sequence includes: Based on the N original reading sequences, the reading fluctuation degree of each original reading sequence is calculated. The reading fluctuation degree is obtained by calculating the standard deviation of each original reading sequence within the same time window. Based on the normalized global motion intensity sequence, the standard deviation of the normalized global motion intensity sequence within the corresponding time window is calculated, and the standard deviation of the normalized global motion intensity sequence is used as the global motion intensity fluctuation. Based on the degree of reading fluctuation, the maximum value among all the degrees of reading fluctuation is extracted as the maximum degree of reading fluctuation; Based on the global motion intensity fluctuation, the maximum value among all the global motion intensity fluctuations is extracted as the maximum global motion intensity fluctuation; The ratio of the maximum reading fluctuation to the maximum global motion intensity fluctuation is used as a proportionality coefficient; The proportionality coefficient is used to characterize the extreme proportional relationship between sensor reading fluctuations and global motion intensity fluctuations.

[0010] Preferably, the step of calculating a similarity score from the N original reading sequences and the normalized local motion intensity sequence includes: Calculate the difference between adjacent data points in the N original reading sequences and the normalized local motion intensity sequence respectively to obtain the N original reading change rate sequences and the normalized local motion intensity change rate sequences; The DTW distance between each original reading rate of change sequence and the normalized local motion intensity rate of change sequence is calculated using the DTW algorithm. Based on the DTW distance, a similarity score is obtained using a distance function formula. The similarity score is used to measure the degree of morphological matching between the sensor reading change rate and the video image motion change rate, and to characterize the strength of the dynamic correlation between the two. The distance function formula is as follows: ,in, The similarity score has a range of [0,1], and K is a scaling factor. For the sensor change rate sequence, For video motion rate sequence, This represents the DTW distance.

[0011] Preferably, the step of constructing a weighted undirected bipartite graph based on the similarity score, and constructing a weighted bipartite graph based on the weighted undirected bipartite graph, includes: The left vertex set consists of the original reading sequence of the N sensors, and the right vertex set consists of the normalized local motion intensity sequence. The edge set contains all edges connecting the left and right vertices. Each edge is assigned a weight, the weight of which is equal to the similarity score between the original reading rate of change sequence and the normalized local motion intensity rate of change sequence. Traverse all edges in the weighted undirected bipartite graph. If the edge weight is less than a preset correlation threshold, remove the edge from the edge set to generate a weighted bipartite graph, where the retained edges represent reliable sensor-video motion association pairs.

[0012] Preferably, the step of constructing the association tree based on the weighted bipartite graph includes: If the weighted bipartite graph is characterized by multiple left vertices connected to the same right vertex, it is defined as a many-to-one pattern. If the weighted bipartite graph is characterized by a left vertex connected to multiple right vertices, it is defined as a one-to-many pattern; Construct an association tree based on the many-to-one pattern and the one-to-many pattern: If any original reading sequence is taken as the root node, then all normalized local motion intensity sequences connected to the original reading sequence in the weighted bipartite graph are taken as child nodes, and an association tree with the original reading sequence as the root is constructed and defined as the first association tree. If any normalized local motion intensity sequence is taken as the root node, then all the original reading sequences connected to the normalized local motion intensity sequence in the weighted bipartite graph are taken as child nodes, and an association tree with the normalized local motion intensity sequence as the root is constructed and defined as the second association tree. The original reading sequences are clustered and grouped according to the sensor type semantics to obtain the clustering results; taking any sensor cluster as the root node, all the normalized local motion intensity sequences associated with the sensor cluster are taken as child nodes to construct an association tree rooted at the sensor cluster and defined as the third association tree.

[0013] Preferably, obtaining the comprehensive value score of each node based on the association tree includes: Traverse all video leaf nodes in the first association tree, the second association tree, and the third association tree; Based on the normalized confidence sequence for a specific event, extract all confidence values ​​for the corresponding time interval within the current decision time window, and use them as the normalized confidence sequence value for the specific event. Extract the normalized specific event confidence sequence value corresponding to each video leaf node within the current decision time window, and take the maximum value within the current decision time window as the current event confidence score of the video leaf node; The current event confidence score is assigned to the video leaf node, and the current event confidence score is propagated upwards to the parent node according to the association relationship of the association tree; The final event confidence score of the parent node is the maximum value of the current event confidence scores of all child nodes; The comprehensive value score for each node is calculated based on the basic value and the confidence score of the current event.

[0014] Preferably, the lossy compression control strategy generated based on the proportional coefficient and the comprehensive value score includes: If the proportional coefficient is greater than the preset abnormal threshold, then for all sensor nodes in the association tree, the upper limit of the compression ratio is calculated according to the proportional coefficient, and the final compression ratio of the sensor node is set to not exceed the upper limit of the compression ratio. For all video leaf nodes in the association tree, the final compression rate of the video leaf node is directly set to the compression rate mapped by the comprehensive value score of the video leaf node; If the proportional coefficient is less than or equal to the preset abnormal threshold, the final compression ratio of all nodes in the association tree is directly set to the compression ratio mapped by the comprehensive value score of all nodes in the association tree. Based on the final compression ratio of all nodes, a compression control instruction list is generated, which specifies the corresponding compression ratio parameter for each data stream.

[0015] Preferably, a video monitoring pan-tilt unit includes: Edge computing components, multimodal data interface components, and compression execution components; The multimodal data interface component includes: A sensor data interface unit, which is used to connect to and acquire N raw reading sequences of IoT sensors; The video data interface unit is used to connect to and acquire video stream data from surveillance cameras. The edge computing component includes: The feature extraction unit, which is connected to the video data interface unit, is used to obtain global motion intensity, local motion intensity and confidence of specific events based on video stream data, and construct the corresponding normalized sequence. The correlation analysis unit, which connects the feature extraction unit and the sensor data interface unit, is used to calculate the proportional coefficient and similarity score based on the N original reading sequences and the normalized sequence. The strategy generation unit, which is connected to the correlation analysis unit, is used to generate a lossy compression control strategy based on the proportional coefficient and the comprehensive value score. The compression execution component is connected to the strategy generation unit and is used to receive the compression control strategy and perform compression operations on the corresponding data stream; The edge computing component also includes: A weighted bipartite graph construction unit, which is connected to the association analysis unit, is used to construct a weighted undirected bipartite graph and a weighted bipartite graph based on similarity scores; The association tree construction unit, which is connected to the weighted bipartite graph construction unit, is used to construct a first association tree, a second association tree, and a third association tree based on the weighted bipartite graph; The value assessment unit, which connects the association tree construction unit and the feature extraction unit, is used to traverse the video leaf nodes in the association tree, extract the normalized confidence sequence value of a specific event, calculate the confidence score of the current event and propagate it upwards, and finally obtain the comprehensive value score of each node. The strategy generation unit is connected to the value assessment unit and generates a compression control instruction list based on the comprehensive value score and the proportional coefficient.

[0016] Beneficial Effects: Based on the needs of multimodal monitoring data fusion compression and intelligent decision-making, this invention proposes a lossy compression control method. It quantifies the dynamic correlation between sensors and video data by constructing normalized motion intensity sequences and specific event confidence sequences; optimizes the data importance assessment and redundancy discrimination process using weighted bipartite graphs and multimodal association trees; and dynamically generates and adaptively adjusts the compression strategy based on proportional coefficients and comprehensive value scores—reconstructing the data transmission workflow in resource-constrained scenarios through a cross-modal collaborative mechanism. This method significantly improves data compression efficiency and event transmission integrity in complex monitoring scenarios in remote mountainous areas, providing a general, value-based compression control method for intelligent IoT monitoring. Attached Figure Description

[0017] Figure 1 This is a flowchart of an AI-based lossy compression control method according to the present invention. Figure 2 This is a flowchart of the proportional coefficient acquisition method described in this invention; Figure 3 This is a flowchart of the similarity score acquisition method described in this invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Example 1: An AI-based lossy compression control method, such as Figures 1-3 As shown, it includes the following steps: S1-1: Acquire sensor data and video data, obtain N raw reading sequences based on the sensor data, and obtain global motion intensity, local motion intensity, and confidence level of specific events based on the video data; The sensor data is sorted in chronological order to obtain N raw reading sequences; Global motion intensity is obtained by calculating the average difference between all pixels between two consecutive frames; The target region is determined by a target detection algorithm, and the local motion intensity is obtained by calculating the average optical flow within the target region. A lightweight image classification network is used to perform forward propagation on video frames and output the confidence scores for specific events.

[0020] It should be noted that sensor data and video data are acquired from IoT sensors and surveillance cameras, respectively. Sensor data includes time-series readings of temperature, humidity, and wind speed, while video data consists of consecutive frames. The sensor data is sorted chronologically to form an ordered time series for easier subsequent analysis. The resulting N raw reading sequences represent the reading changes from different sensors; for example, a temperature sensor generates a temperature reading sequence. Global motion intensity is defined as the average pixel difference between all video frames, calculated using the following formula: Where G is the global motion intensity, and W and H are the frame width and height. and These are consecutive frames, where i and j are the row and column coordinate indices of the pixels. Local motion intensity is defined as the average optical flow amplitude within the target region. Video frames are processed using target detection algorithms (such as YOLO, SSD, or Faster R-CNN) to identify and define the region of interest. Subsequently, optical flow algorithms (such as Lucas-Kanade or Farneback) are used within this region to calculate the optical flow vector for each pixel, and the average amplitude of all optical flow vectors is taken as the local motion intensity. The calculation formula is as follows: Where M is the total number of pixels in the target area. , These are the optical flow displacement components of the i-th pixel in the horizontal and vertical directions, respectively. This represents the optical flow amplitude of the pixel. Specific event confidence is defined as the probability output of a neural network on the presence of a specific event in a video frame. It is obtained through inference using a lightweight image classification network. For example, using a lightweight convolutional neural network (such as MobileNet or SqueezeNet) for forward inference on a video frame, the confidence score for the presence of a wildfire event in that frame is 0.92. By continuously acquiring the specific event confidence of each frame, a temporal feature of event importance is formed, which is used to assess the dynamic changes and criticality of events in the video content. These features are used to analyze the dynamics of video content and the importance of events, thereby supporting the generation of lossy compression control strategies and ensuring optimized data transmission efficiency in resource-constrained environments.

[0021] Example: At a monitoring point in a mountainous area, a temperature sensor collected readings between 10:00 and 10:05, with a sequence of readings [25.1, 25.3, 25.8, 26.3, 26.7] degrees Celsius, while a wind speed sensor simultaneously output [3.2, 3.5, 4.1, 4.8, 5.3] meters per second. Simultaneously, a camera captured consecutive video frames. The average pixel difference between frame t and frame t-1 was calculated to be 15.2, which was used as the global motion intensity. After defining the canopy area through target detection, the average optical flow amplitude within that area was calculated to be 8.7, which was used as the local motion intensity. A lightweight neural network analyzed the current frame and output a wildfire event confidence level of 0.92.

[0022] S1-2: Based on the global motion intensity, the local motion intensity, and the confidence level of the specific event, construct the normalized global motion intensity sequence, the normalized local motion intensity sequence, and the normalized confidence level sequence of the specific event, respectively; The global motion intensity, the local motion intensity, and the confidence level of the specific event are normalized to obtain normalized global motion intensity, normalized local motion intensity, and normalized confidence level of the specific event. Based on the normalized global motion intensity, the normalized global motion intensity values ​​corresponding to each frame are sorted by timestamp to construct a normalized global motion intensity sequence. Based on the normalized local motion intensity, the normalized local motion intensity values ​​corresponding to each frame are sorted by timestamp to construct a normalized local motion intensity sequence. Based on the normalized specific event confidence score, the normalized specific event confidence score values ​​corresponding to each frame are sorted by timestamp to construct a normalized specific event confidence score sequence.

[0023] It's important to note that normalization scales global motion intensity, local motion intensity, and event-specific confidence to the same numerical range, typically using a min-max normalization formula. This allows for comparison of features at different scales and facilitates subsequent joint analysis. The normalized values ​​for each frame are sorted by timestamp to construct a time series, forming a continuous and standardized feature representation to characterize dynamic changes. The normalized global motion intensity sequence reflects the relative intensity of the overall video motion over time; the normalized local motion intensity sequence represents the relative activity of local motion within the region of interest; and the normalized event-specific confidence sequence describes the continuous change in the probability of a specific event occurring. These sequences provide standardized input for subsequent multimodal correlation analysis and compression decisions.

[0024] Example: Taking a monitoring point in a mountainous area from 10:00 to 10:05 as an example, the maximum global motion intensity collected during this period is 20, and the minimum is 5. The current frame intensity is measured to be 15.2. After minimum-maximum normalization, the normalized global motion intensity value is 0.68. The maximum local motion intensity is 12, and the minimum is 2. The current frame value of 8.7 is normalized to 0.67. The confidence level for specific events is directly taken as the original value of 0.92. The normalized global motion intensity values ​​[0.62, 0.65, 0.68, 0.71, 0.75] corresponding to the five time points from 10:00:01 to 10:00:05 are arranged in chronological order to form a normalized global motion intensity sequence. Simultaneously construct the normalized local motion intensity sequence [0.61, 0.63, 0.67, 0.70, 0.74] and the normalized specific event confidence sequence [0.85, 0.88, 0.92, 0.89, 0.86].

[0025] S2-1: Calculate the scaling factor by combining the N original reading sequences and the normalized global motion intensity sequence; Based on the N original reading sequences, the reading fluctuation degree of each original reading sequence is calculated. The reading fluctuation degree is obtained by calculating the standard deviation of each original reading sequence within the same time window. Based on the normalized global motion intensity sequence, the standard deviation of the normalized global motion intensity sequence within the corresponding time window is calculated, and the standard deviation of the normalized global motion intensity sequence is used as the global motion intensity fluctuation. Based on the degree of reading fluctuation, the maximum value among all the degrees of reading fluctuation is extracted as the maximum degree of reading fluctuation; Based on the global motion intensity fluctuation, the maximum value among all the global motion intensity fluctuations is extracted as the maximum global motion intensity fluctuation; The ratio of the maximum reading fluctuation to the maximum global motion intensity fluctuation is used as a proportionality coefficient; The proportionality coefficient is used to characterize the extreme proportional relationship between sensor reading fluctuations and global motion intensity fluctuations.

[0026] It should be noted that the reference Figure 2 The reading volatility of each raw reading sequence is calculated to quantify the discreteness of the sensor data. Reading volatility is defined as the magnitude of variation in sensor readings within a time window. Before calculating the proportional gain, the reading volatility (standard deviation) of each sensor is normalized to make it a dimensionless value. This can be done, for example, by normalizing using the maximum-minimum value of historical sensor data, or by using Z-score standardization. The standard deviation is calculated using the following formula: ,in, Standard deviation It is a single reading. In this context, 'i' is the index of the i-th data item in the data sequence. Here, N is the mean of readings within the time window, and N is the number of data points. Calculating the standard deviation of the normalized global motion intensity sequence aims to measure the overall motion fluctuation of the video; global motion intensity fluctuation represents the dispersion of the normalized global motion intensity sequence. The maximum value among all the reading fluctuations is extracted as the maximum reading fluctuation and the maximum global motion intensity fluctuation to capture extreme fluctuations in sensor and video data. The ratio of these two values ​​is used as a scaling factor to characterize the most significant proportional relationship between sensor and video fluctuations. The scaling factor is defined as the ratio of the maximum reading fluctuation to the maximum global motion intensity fluctuation, calculated using the following formula: ,in, The scaling factor is (maximum sensor fluctuation / maximum global video fluctuation). It is the maximum fluctuation of the reading. This represents the maximum global motion intensity fluctuation. The scaling factor is used to identify abnormal fluctuation scenarios and guide the adjustment of compression strategies.

[0027] Example: At a mountain monitoring point, between 10:00 and 10:05, the temperature sensor reading sequence is [25.1, 25.3, 25.8, 26.3, 26.7] degrees Celsius, and the wind speed sensor reading sequence is [3.2, 3.5, 4.1, 4.8, 5.3] meters per second. The standard deviations of the two sensor sequences within a 5-minute time window are calculated: temperature reading fluctuation is 0.64, wind speed is 0.78, and the maximum reading fluctuation is 0.78. The normalized global motion intensity sequence is [0.62, 0.65, 0.68, 0.71, 0.75], and the standard deviation of the normalized global motion intensity sequence, 0.048, is taken as the global motion intensity fluctuation. The ratio of the maximum reading fluctuation to the maximum global motion intensity fluctuation is 16.25. This ratio indicates that the sensor reading fluctuation is much higher than the video global motion fluctuation, suggesting an abnormal event. The sensor data compression rate needs to be limited to preserve details.

[0028] S2-2: Calculate the similarity score for the N original reading sequences and the normalized local motion intensity sequence; Calculate the difference between adjacent data points in the N original reading sequences and the normalized local motion intensity sequence respectively to obtain the N original reading change rate sequences and the normalized local motion intensity change rate sequences; It should be noted that, in order to further eliminate the influence of different sensor physical dimensions and numerical ranges on morphological comparison, the N original reading change rate sequences are standardized before calculating the DTW distance. The N original reading change rate sequences are converted into a distribution with a mean of 0 and a standard deviation of 1, so that all sensor change rate sequences are on the same numerical scale, thereby ensuring that the DTW distance accurately reflects the morphological similarity of the sequences rather than absolute numerical differences.

[0029] The DTW distance between each original reading rate of change sequence and the normalized local motion intensity rate of change sequence is calculated using the DTW algorithm. Based on the DTW distance, a similarity score is obtained using a distance function formula. The similarity score is used to measure the degree of morphological matching between the sensor reading change rate and the video image motion change rate, and to characterize the strength of the dynamic correlation between the two. The distance function formula is as follows: ,in, The similarity score has a range of [0,1], and K is a scaling factor. For the sensor change rate sequence, For video motion rate sequence, This represents the DTW distance.

[0030] It should be noted that the reference Figure 3 The difference between adjacent data points is calculated to capture the instantaneous trend of sensor readings and video motion intensity changes; this difference represents the amount of change in data at adjacent time points. The obtained N raw reading change rate sequences represent the rate of change of each sensor reading over time, and the normalized local motion intensity change rate sequence represents the rate of change of local motion intensity in the video. The distance is calculated using the DTW algorithm to align and compare the morphological differences between the sensor and video change rate sequences on the time axis. The DTW distance represents the minimum regularized path cumulative distance between the two sequences, and the calculation formula is: ,in, For sequence and The dynamic time-normalized distance between them It is a sequence of sensor change rates. It is a video motion rate sequence. To align the path, For sequence The i-th point in For sequence The j-th point in Minimize the value for all regular paths. Distance function formula. The DTW distance is mapped to a similarity score between 0 and 1, where K is a scaling factor used to adjust the degree of influence of distance on the score. The higher the score, the stronger the dynamic correlation between the sensor and the video motion.

[0031] Example: At a monitoring point in a mountainous area, between 10:00 and 10:05, the temperature sensor reading sequence is [25.1, 25.3, 25.8, 26.3, 26.7] degrees Celsius, the wind speed sensor reading sequence is [3.2, 3.5, 4.1, 4.8, 5.3] meters per second, and the normalized local motion intensity sequence is [0.61, 0.63, 0.67, 0.70, 0.74]. First, the difference between adjacent data points is calculated, resulting in the temperature change rate sequence [0.2, 0.5, 0.5, 0.4], the wind speed change rate sequence [0.3, 0.6, 0.7, 0.5], and the video motion change rate sequence [0.02, 0.04, 0.03, 0.04]. The sensor change rate sequences are then standardized to a form with a mean of 0 and a standard deviation of 1. The DTW distance between the temperature change rate sequence and the video motion change rate sequence calculated by the DTW algorithm is 1.24. Taking the scaling factor k=1, the similarity score obtained by substituting into the distance function formula is 0.82, indicating that the temperature sensor reading change and the local motion intensity change pattern of the video are highly matched and have strong dynamic correlation.

[0032] S3-1: Construct a weighted undirected bipartite graph based on the similarity scores, and construct a weighted bipartite graph based on the weighted undirected bipartite graph; The left vertex set consists of the original reading sequence of the N sensors, and the right vertex set consists of the normalized local motion intensity sequence. The edge set contains all edges connecting the left and right vertices. Each edge is assigned a weight, the weight of which is equal to the similarity score between the original reading rate of change sequence and the normalized local motion intensity rate of change sequence. Traverse all edges in the weighted undirected bipartite graph. If the edge weight is less than a preset correlation threshold, remove the edge from the edge set to generate a weighted bipartite graph, where the retained edges represent reliable sensor-video motion association pairs.

[0033] It's important to note that the difference between a weighted undirected bipartite graph and a weighted bipartite graph lies in whether the edges are filtered; the former includes all connections, while the latter only retains edges with weights exceeding a preset threshold. The constructed weighted undirected bipartite graph represents the potential correlations between all sensor sequences and video motion sequences, where edge weights indicate the correlation strength. The weighted bipartite graph further refines reliable strong correlation pairs. The construction principle of the weighted bipartite graph is to filter noise and weak correlations by removing low-weight edges. The preset correlation threshold is usually set based on historical data or experimental experience to distinguish between significant and accidental correlations. Sensor-video motion correlation pairs represent data channels where changes in sensor readings are highly synchronized with local video motion. These correlation pairs indicate a mutually interpretable dependency between sensor events and changes in video content, providing a basis for subsequent compression decisions.

[0034] Example: At a monitoring point in a mountainous area, between 10:00 and 10:05, the similarity score between the temperature sensor change rate sequence and the video motion change rate sequence is 0.82, while the wind speed sensor's similarity score is 0.45. When constructing a weighted undirected bipartite graph, the left vertex set contains the temperature sensor sequence and the wind speed sensor sequence, and the right vertex set is the normalized local motion intensity sequence. The initial edge set contains two edges: the edge from the temperature sensor to the video motion has a weight of 0.82, and the edge from the wind speed sensor to the video motion has a weight of 0.45. Assuming a preset correlation threshold of 0.5, after traversing all edges, the edge with a weight of 0.45 is removed, generating a weighted bipartite graph that retains only the associated edge between the temperature sensor and the video motion. This retained edge represents a reliable sensor-video motion association pair, indicating a significant correlation between temperature changes and local motion intensity changes in the video.

[0035] S3-2: Construct an association tree based on the weighted bipartite graph; If the weighted bipartite graph is characterized by multiple left vertices connected to the same right vertex, it is defined as a many-to-one pattern. If the weighted bipartite graph is characterized by a left vertex connected to multiple right vertices, it is defined as a one-to-many pattern; Construct an association tree based on the many-to-one pattern and the one-to-many pattern: If any original reading sequence is taken as the root node, then all normalized local motion intensity sequences connected to the original reading sequence in the weighted bipartite graph are taken as child nodes, and an association tree with the original reading sequence as the root is constructed and defined as the first association tree. If any normalized local motion intensity sequence is taken as the root node, then all the original reading sequences connected to the normalized local motion intensity sequence in the weighted bipartite graph are taken as child nodes, and an association tree with the normalized local motion intensity sequence as the root is constructed and defined as the second association tree. The original reading sequences are clustered and grouped according to the sensor type semantics to obtain the clustering results; taking any sensor cluster as the root node, all the normalized local motion intensity sequences associated with the sensor cluster are taken as child nodes to construct an association tree rooted at the sensor cluster and defined as the third association tree.

[0036] It's important to note that the many-to-one pattern represents multiple sensor reading sequences simultaneously associated with the same video motion sequence, indicating that multiple sensor events jointly trigger the same video activity. The one-to-many pattern, on the other hand, represents one sensor reading sequence associated with multiple video motion sequences, indicating that a single sensor event triggers motion in multiple video regions. These two patterns are used to comprehensively capture the complex relationships between sensors and video data. The first association tree is a tree structure built with a single sensor sequence as the root node and all associated video motion sequences as child nodes, representing all video motion responses corresponding to that sensor event. The second association tree is a tree structure built with a single video motion sequence as the root node and all associated sensor sequences as child nodes, representing all potential sensor event sources that trigger that video motion. The third association tree is a tree structure built with a cluster of similar sensors as the root node and all associated video motion sequences as child nodes, representing the overall association pattern between a group of sensors of the same type and video motion. These three tree structures characterize the dependencies between multimodal data from different dimensions.

[0037] Example: In the weighted bipartite graph of a mountain monitoring point from 10:00 to 10:05, there are associated edges between the temperature sensor sequence and the video motion sequence. When constructing the first association tree, the temperature sensor sequence is used as the root node, and the video motion sequence connected to the temperature sensor sequence is used as a child node, forming a tree structure. When constructing the second association tree, the video motion sequence is used as the root node, and the temperature sensor sequence is used as a child node of the video motion sequence. Based on the sensor type semantics, temperature sensors and wind speed sensors are clustered into an environmental monitoring cluster. A third association tree is constructed using this cluster as the root node, and the video motion sequences associated with the sensor cluster are used as child nodes. These three association trees characterize the dependencies between sensor data and video motion data from different dimensions, providing a structured foundation for subsequent value assessment.

[0038] S4-1: Obtain the comprehensive value score of each node based on the aforementioned relationship tree; Traverse all video leaf nodes in the first association tree, the second association tree, and the third association tree; Based on the normalized confidence sequence for a specific event, extract all confidence values ​​for the corresponding time interval within the current decision time window, and use them as the normalized confidence sequence value for the specific event. Extract the normalized specific event confidence sequence value corresponding to each video leaf node within the current decision time window, and take the maximum value within the current decision time window as the current event confidence score of the video leaf node; The current event confidence score is assigned to the video leaf node, and the current event confidence score is propagated upwards to the parent node according to the association relationship of the association tree; The final event confidence score of the parent node is the maximum value of the current event confidence scores of all child nodes; The comprehensive value score for each node is calculated based on the basic value and the confidence score of the current event.

[0039] It should be noted that the video leaf node represents the video motion sequence node (i.e., the normalized local motion intensity sequence) located at the end of the association tree. The maximum value of the normalized event-specific confidence sequence within the current decision time window is extracted as the current event confidence score, used to capture the importance level of the video node at the moment of the most significant event. The normalized event-specific confidence sequence value refers to all confidence values ​​within the current decision time window extracted from the pre-constructed normalized event-specific confidence sequence, based on the timestamp corresponding to the video leaf node, forming a subsequence. Each value in this subsequence is the normalized event-specific confidence sequence value, used to reflect the change in event confidence of the video node within the current decision period, providing a data foundation for subsequent event scoring. In the three association trees, the current event confidence score propagates from bottom to top through a tree structure. Each parent node receives the scores of all direct child nodes and takes the maximum value of all direct child node scores as its own final event confidence score. The principle of taking the maximum value is to adopt the most conservative strategy, ensuring that any high-importance event of any child node can be inherited by the parent node. The comprehensive value score is defined as the final importance measure that integrates the node's basic value and the event confidence score, representing the overall value weight of the node in the compressed decision-making process. The calculation formula is as follows: = in, For the comprehensive value score, B is the basic value, which is a preset static weighting coefficient used to characterize the inherent importance level of different sensor types or video data streams. The basic value is determined based on prior knowledge or historical data statistics. C is the confidence score of the current event. These are the balancing weighting coefficients.

[0040] Example: In the association tree of a monitoring point in a mountainous area during the period from 10:00 to 10:05, the normalized confidence score sequence value of the specific event corresponding to the video leaf node is [0.85, 0.88, 0.92, 0.89, 0.86]. The maximum value of 0.92 is taken as the current event confidence score. This score is propagated upwards to the parent node (temperature sensor sequence), and the final event confidence score of the parent node's temperature sensor sequence is also 0.92. Assuming the base value is 0.7 and the balancing weight coefficient α is 0.3, the comprehensive value score of the temperature sensor node is calculated according to the formula as 0.7 × 0.3 + 0.92 × 0.7 = 0.854. This score reflects the overall importance of the node in the compression decision; the higher the value, the higher the priority of data retention.

[0041] S4-2: Generate a lossy compression control strategy based on the proportional coefficient and the comprehensive value score.

[0042] If the proportional coefficient is greater than the preset abnormal threshold, then for all sensor nodes in the association tree, the upper limit of the compression ratio is calculated according to the proportional coefficient, and the final compression ratio of the sensor node is set to not exceed the upper limit of the compression ratio. For all video leaf nodes in the association tree, the final compression rate of the video leaf node is directly set to the compression rate mapped by the comprehensive value score of the video leaf node; If the proportional coefficient is less than or equal to the preset abnormal threshold, the final compression ratio of all nodes in the association tree is directly set to the compression ratio mapped by the comprehensive value score of all nodes in the association tree. Based on the final compression ratio of all nodes, a compression control instruction list is generated, which specifies the corresponding compression ratio parameter for each data stream.

[0043] It should be noted that the preset anomaly threshold is usually set through historical data statistical analysis or experimental experience to distinguish between normal and abnormal fluctuations. When the proportional coefficient is greater than the preset anomaly threshold, it indicates an extreme proportional relationship between sensor reading fluctuations and global video motion fluctuations, pointing to an abnormal event. In this case, limiting the upper limit of the sensor node compression rate is used to prevent the loss of critical abnormal data due to over-compression; video leaf nodes directly map the compression rate based on the comprehensive value score to maintain content adaptability. When the proportional coefficient is less than or equal to the preset anomaly threshold, it indicates that the system is in a normal fluctuation state. In this case, all nodes directly map the compression rate based on the comprehensive value score to achieve balanced resource allocation. The core principle of this technology lies in dynamically identifying the scene state through multimodal data correlation and achieving differentiated compression control based on value assessment, thereby optimizing overall transmission efficiency while ensuring the integrity of critical information. The upper limit of the compression rate is calculated as: upper limit of compression rate = 1 / proportional coefficient, used to limit the maximum compression degree when sensor data fluctuates abnormally, avoiding the loss of critical information. The compression rate mapping rule is: compression rate = 1 - comprehensive value score, that is, the higher the comprehensive value score (the more important the data), the lower the compression rate to retain more information; conversely, a higher compression rate is used.

[0044] Example: At a mountain monitoring point between 10:00 and 10:05, the scaling factor is 16.25, and the preset anomaly threshold is 10. Since the scaling factor is greater than the threshold, the upper limit of the compression rate calculated for the temperature sensor node is 1 / 16.25 ≈ 6%, and the final compression rate of the temperature sensor node is set to not exceed 6%. The comprehensive value score of the video leaf node, 0.854, is mapped to a compression rate of 15%, and this value is directly set. A compression control instruction list is generated: temperature sensor data stream compression rate 6%, video data stream compression rate 15%. This strategy ensures the integrity of sensor data under abnormal conditions, while adjusting the video compression level according to the content value.

[0045] Example 2, based on Example 1, provides a video monitoring pan-tilt unit, comprising: Edge computing components, multimodal data interface components, and compression execution components; The multimodal data interface component includes: A sensor data interface unit, which is used to connect to and acquire N raw reading sequences of IoT sensors; The video data interface unit is used to connect to and acquire video stream data from surveillance cameras. The edge computing component includes: The feature extraction unit, which is connected to the video data interface unit, is used to obtain global motion intensity, local motion intensity and confidence of specific events based on video stream data, and construct the corresponding normalized sequence. The correlation analysis unit, which connects the feature extraction unit and the sensor data interface unit, is used to calculate the proportional coefficient and similarity score based on the N original reading sequences and the normalized sequence. The strategy generation unit, which is connected to the correlation analysis unit, is used to generate a lossy compression control strategy based on the proportional coefficient and the comprehensive value score. The compression execution component is connected to the strategy generation unit and is used to receive the compression control strategy and perform compression operations on the corresponding data stream; The edge computing component also includes: A weighted bipartite graph construction unit, which is connected to the association analysis unit, is used to construct a weighted undirected bipartite graph and a weighted bipartite graph based on similarity scores; The association tree construction unit, which is connected to the weighted bipartite graph construction unit, is used to construct a first association tree, a second association tree, and a third association tree based on the weighted bipartite graph; The value assessment unit, which connects the association tree construction unit and the feature extraction unit, is used to traverse the video leaf nodes in the association tree, extract the normalized confidence sequence value of a specific event, calculate the confidence score of the current event and propagate it upwards, and finally obtain the comprehensive value score of each node. The strategy generation unit is connected to the value assessment unit and generates a compression control instruction list based on the comprehensive value score and the proportional coefficient.

[0046] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A lossy compression control method based on AI, characterized in that, Includes the following steps: S1: Acquire sensor data and video data; obtain N raw reading sequences based on the sensor data; obtain global motion intensity, local motion intensity, and specific event confidence based on the video data; construct normalized global motion intensity sequence, normalized local motion intensity sequence, and normalized specific event confidence sequence based on the global motion intensity, the local motion intensity, and the specific event confidence, respectively. S2: Calculate the proportionality coefficients for the N original reading sequences and the normalized global motion intensity sequence; calculate the similarity scores for the N original reading sequences and the normalized local motion intensity sequence; S3: Construct a weighted undirected bipartite graph based on the similarity scores, and construct a weighted bipartite graph based on the weighted undirected bipartite graph; construct an association tree based on the weighted bipartite graph; S4: Obtain the comprehensive value score for each node based on the aforementioned relationship tree; A lossy compression control strategy is generated based on the aforementioned proportional coefficient and the aforementioned comprehensive value score.

2. The AI-based lossy compression control method according to claim 1, characterized in that, The acquisition of sensor data and video data, obtaining N raw reading sequences based on the sensor data, and obtaining global motion intensity, local motion intensity, and confidence levels for specific events based on the video data, includes: The sensor data is sorted in chronological order to obtain N raw reading sequences; Global motion intensity is obtained by calculating the average difference between all pixels between two consecutive frames; The target region is determined by a target detection algorithm, and the local motion intensity is obtained by calculating the average optical flow within the target region. A lightweight image classification network is used to perform forward propagation on video frames and output the confidence scores for specific events.

3. The AI-based lossy compression control method according to claim 1, characterized in that, The step of constructing a normalized global motion intensity sequence, a normalized local motion intensity sequence, and a normalized specific event confidence sequence based on the global motion intensity, the local motion intensity, and the specific event confidence score, respectively, includes: The global motion intensity, the local motion intensity, and the confidence level of the specific event are normalized to obtain normalized global motion intensity, normalized local motion intensity, and normalized confidence level of the specific event. Based on the normalized global motion intensity, the normalized global motion intensity values ​​corresponding to each frame are sorted by timestamp to construct a normalized global motion intensity sequence. Based on the normalized local motion intensity, the normalized local motion intensity values ​​corresponding to each frame are sorted by timestamp to construct a normalized local motion intensity sequence. Based on the normalized specific event confidence score, the normalized specific event confidence score values ​​corresponding to each frame are sorted by timestamp to construct a normalized specific event confidence score sequence.

4. The AI-based lossy compression control method according to claim 1, characterized in that, The calculation of the scaling factor from the N original reading sequences and the normalized global motion intensity sequence includes: Based on the N original reading sequences, the reading fluctuation degree of each original reading sequence is calculated. The reading fluctuation degree is obtained by calculating the standard deviation of each original reading sequence within the same time window. Based on the normalized global motion intensity sequence, the standard deviation of the normalized global motion intensity sequence within the corresponding time window is calculated, and the standard deviation of the normalized global motion intensity sequence is used as the global motion intensity fluctuation. Based on the degree of reading fluctuation, the maximum value among all the degrees of reading fluctuation is extracted as the maximum degree of reading fluctuation; Based on the global motion intensity fluctuations, the maximum value among all global motion intensity fluctuations is extracted as the maximum global motion intensity fluctuation. The ratio of the maximum reading fluctuation to the maximum global motion intensity fluctuation is used as a proportionality coefficient; The proportionality coefficient is used to characterize the extreme proportional relationship between sensor reading fluctuations and global motion intensity fluctuations.

5. The AI-based lossy compression control method according to claim 1, characterized in that, The step of calculating a similarity score from the N original reading sequences and the normalized local motion intensity sequence includes: Calculate the difference between adjacent data points in the N original reading sequences and the normalized local motion intensity sequence respectively to obtain the N original reading change rate sequences and the normalized local motion intensity change rate sequences; The DTW distance between each original reading rate of change sequence and the normalized local motion intensity rate of change sequence is calculated using the DTW algorithm. Based on the DTW distance, a similarity score is obtained using a distance function formula. The similarity score is used to measure the degree of morphological matching between the sensor reading change rate and the video image motion change rate, and to characterize the strength of the dynamic correlation between the two. The distance function formula is as follows: ,in , where is the similarity score, with a value range of [0,1], and K is the scaling factor. For the sensor change rate sequence, For video motion rate sequence, This represents the DTW distance.

6. The AI-based lossy compression control method according to claim 1, characterized in that, The step of constructing a weighted undirected bipartite graph based on the similarity score, and constructing a weighted bipartite graph based on the weighted undirected bipartite graph, includes: The left vertex set consists of the original reading sequence of the N sensors, and the right vertex set consists of the normalized local motion intensity sequence. The edge set contains all edges connecting the left and right vertices. Each edge is assigned a weight, the weight of which is equal to the similarity score between the original reading rate of change sequence and the normalized local motion intensity rate of change sequence. Traverse all edges in the weighted undirected bipartite graph. If the edge weight is less than a preset correlation threshold, remove the edge from the edge set to generate a weighted bipartite graph, where the retained edges represent reliable sensor-video motion association pairs.

7. The AI-based lossy compression control method according to claim 1, characterized in that, The construction of the association tree based on the weighted bipartite graph includes: If the weighted bipartite graph is characterized by multiple left vertices connected to the same right vertex, it is defined as a many-to-one pattern. If the weighted bipartite graph is characterized by a left vertex connected to multiple right vertices, it is defined as a one-to-many pattern; Construct an association tree based on the many-to-one pattern and the one-to-many pattern: If any original reading sequence is taken as the root node, then all normalized local motion intensity sequences connected to the original reading sequence in the weighted bipartite graph are taken as child nodes, and an association tree with the original reading sequence as the root is constructed and defined as the first association tree. If any normalized local motion intensity sequence is taken as the root node, then all the original reading sequences connected to the normalized local motion intensity sequence in the weighted bipartite graph are taken as child nodes, and an association tree with the normalized local motion intensity sequence as the root is constructed and defined as the second association tree. The original reading sequences are clustered and grouped according to the sensor type semantics to obtain the clustering results; taking any sensor cluster as the root node, all the normalized local motion intensity sequences associated with the sensor cluster are taken as child nodes to construct an association tree rooted at the sensor cluster and defined as the third association tree.

8. The AI-based lossy compression control method according to claim 7, characterized in that, The step of obtaining a comprehensive value score for each node based on the association tree includes: Traverse all video leaf nodes in the first association tree, the second association tree, and the third association tree; Based on the normalized confidence sequence for a specific event, extract all confidence values ​​for the corresponding time interval within the current decision time window, and use them as the normalized confidence sequence value for the specific event. Extract the normalized specific event confidence sequence value corresponding to each video leaf node within the current decision time window, and take the maximum value within the current decision time window as the current event confidence score of the video leaf node; The current event confidence score is assigned to the video leaf node, and the current event confidence score is propagated upwards to the parent node according to the association relationship of the association tree; The final event confidence score of the parent node is the maximum value of the current event confidence scores of all child nodes; The comprehensive value score for each node is calculated based on the basic value and the confidence score of the current event.

9. The AI-based lossy compression control method according to claim 1, characterized in that, The lossy compression control strategy generated based on the proportional coefficient and the comprehensive value score includes: If the proportional coefficient is greater than the preset abnormal threshold, then for all sensor nodes in the association tree, the upper limit of the compression ratio is calculated according to the proportional coefficient, and the final compression ratio of the sensor node is set to not exceed the upper limit of the compression ratio. For all video leaf nodes in the association tree, the final compression rate of the video leaf node is directly set to the compression rate mapped by the comprehensive value score of the video leaf node; If the proportional coefficient is less than or equal to the preset abnormal threshold, the final compression ratio of all nodes in the association tree is directly set to the compression ratio mapped by the comprehensive value score of all nodes in the association tree. Based on the final compression ratio of all nodes, a compression control instruction list is generated, which specifies the corresponding compression ratio parameter for each data stream.

10. A video monitoring pan-tilt unit, characterized in that, include: Edge computing components, multimodal data interface components, and compression execution components; The multimodal data interface component includes: A sensor data interface unit, which is used to connect to and acquire N raw reading sequences of IoT sensors; The video data interface unit is used to connect to and acquire video stream data from surveillance cameras. The edge computing component includes: The feature extraction unit, which is connected to the video data interface unit, is used to obtain global motion intensity, local motion intensity and confidence of specific events based on video stream data, and construct the corresponding normalized sequence. The correlation analysis unit, which connects the feature extraction unit and the sensor data interface unit, is used to analyze data based on... The proportionality coefficient and similarity score are calculated from the N original reading sequences and the normalized sequence; The strategy generation unit, which is connected to the correlation analysis unit, is used to generate a lossy compression control strategy based on the proportional coefficient and the comprehensive value score. The compression execution component is connected to the strategy generation unit and is used to receive the compression control strategy and perform compression operations on the corresponding data stream; The edge computing component also includes: A weighted bipartite graph construction unit, which is connected to the association analysis unit, is used to construct a weighted undirected bipartite graph and a weighted bipartite graph based on similarity scores; The association tree construction unit, which is connected to the weighted bipartite graph construction unit, is used to construct a first association tree, a second association tree, and a third association tree based on the weighted bipartite graph; The value assessment unit, which connects the association tree construction unit and the feature extraction unit, is used to traverse the video leaf nodes in the association tree, extract the normalized confidence sequence value of a specific event, calculate the confidence score of the current event and propagate it upwards, and finally obtain the comprehensive value score of each node. The strategy generation unit is connected to the value assessment unit and generates a compression control instruction list based on the comprehensive value score and the proportional coefficient.