Image lightweight coding method and system based on deep learning

The deep learning-based image encoding method optimizes urban surveillance image encoding by adapting to diverse environments, reducing data volume and storage needs while maintaining key information.

CN120318345AActive Publication Date: 2025-07-15JIANGSU YUNBO INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510469850.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-15
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

In urban street monitoring systems, traditional image encoding algorithms have high computational complexity, and the encoding speed cannot keep up with the data acquisition speed. The generated image data files are huge in size and lack adaptability in complex environments.

Method used

Using a lightweight image coding method based on deep learning, deep learning models are trained through convolutional neural networks, and generators and discriminators are built for adversarial training. Combined with multi-source data feature templates, image encoding is optimized, simplified images are generated and adjusted to ensure that key information is retained.

Benefits of technology

Quickly analyze image content in complex environments, generate lightweight encoded images, significantly reduce data volume, reduce storage space requirements, and save costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318345A_ABST
    Figure CN120318345A_ABST
Patent Text Reader

Abstract

The invention discloses an image lightweight coding method and system based on deep learning, and belongs to the technical field of deep learning. A deep learning model is trained by using a convolutional neural network, and a current comprehensive environment is identified; constructing a generator and a discriminator, and performing coding optimization on the current security monitoring image by adopting adversarial training and alternate training of the generator and the discriminator to generate a simplified image; inputting the simplified image into an image coding algorithm for coding to obtain a first lightweight coded image; extracting multi-source data features, establishing multi-source data feature templates in different comprehensive environments, and performing comparative analysis to obtain the matching degree of the environments and the templates; determining an information retention condition of a target area in the first lightweight coded image after lightweight coding; and adjusting the first lightweight coded image based on the information retention condition, and coding according to a selected image coding algorithm after adjustment to generate a final lightweight coded image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and specifically to an image lightweight encoding method and system based on deep learning. Background Art

[0002] In the process of smart city construction, urban street security monitoring, as an important line of defense for ensuring residents' safety and maintaining social order, has become increasingly prominent. Urban streets are crisscrossed, with frequent movement of people and vehicles, and complex and changeable public security and traffic conditions. To comprehensively and real-time grasp the street dynamics, a large number of high-definition cameras are deployed at various intersections, road sections, and key public places. These cameras continuously collect a vast amount of image data day and night, which can be used for security monitoring.

[0003] During the peak hours of urban streets, the number of monitoring images generated per unit time is huge. Due to the high computational complexity and cumbersome processing flow of traditional image encoding algorithms, the encoding speed is far from keeping up with the data acquisition speed. Urban street monitoring data needs to be retained for a long time for subsequent analysis and verification. The image data files generated by traditional unoptimized encoding methods are huge in volume and occupy a large amount of storage space. The urban street environment is complex and diverse, facing various adverse factors such as low light, bad weather, and complex light. Traditional encoding technologies lack effective adaptation mechanisms when dealing with these complex environments. Summary of the Invention

[0004] The purpose of the present invention is to provide an image lightweight encoding method and system based on deep learning to solve the problems raised in the prior art.

[0005] To achieve the above purpose, the present invention provides the following technical solutions:

[0006] In the first aspect, the present invention provides an image lightweight encoding method based on deep learning, including:

[0007] Collect historical multi-source data of urban security monitoring, use a convolutional neural network to train a deep learning model, conduct model training, evaluation, and optimization to identify the current comprehensive environment;

[0008] Based on the current comprehensive environment, construct a generator and a discriminator, adopt adversarial training, alternately train the generator and the discriminator, conduct encoding optimization of the current security monitoring image to generate a simplified image; input the simplified image into an image encoding algorithm for encoding to obtain a first lightweight encoded image;

[0009] Collect current multi-source data of urban security monitoring, extract multi-source data features; based on the current comprehensive environment, establish multi-source data feature templates under different comprehensive environments, compare and analyze the multi-source data features with the corresponding multi-source data feature templates to obtain the matching degree between the environment and the template;

[0010] Determine the information retention situation of the target area in the first lightweight encoded image based on the matching degree;

[0011] Adjust the first lightweight encoded image based on the information retention situation, and after adjustment, encode it according to the selected image encoding algorithm to generate the final lightweight encoded image.

[0012] Combined with the first aspect, in the first implementation manner of the first aspect of this application, the collection of historical urban security monitoring multi-source data, using a convolutional neural network to train a deep learning model, and performing model training, evaluation, and optimization to identify the current comprehensive environment, including:

[0013] Collect images under different weather conditions, different time periods, and different scenarios from urban street cameras, and synchronously collect lidar data and millimeter-wave radar data; the lidar data provides three-dimensional point cloud information of the surrounding environment of the street, and the millimeter-wave radar data includes the distance, speed, and angle information of the target object; integrate to form historical urban security monitoring multi-source data;

[0014] Preprocess the historical urban security monitoring multi-source data and divide it into a training set, a validation set, and a test set; select ResNet for the convolutional neural network structure, set the loss function, optimizer, and number of training epochs; input the image data and corresponding sensor data in the training set into the deep learning model, the model calculates the prediction result through forward propagation, calculates the loss value using the loss function according to the difference between the prediction result and the true label; calculate the gradient of the loss value with respect to the model parameters through the backpropagation algorithm, and use the optimizer to update the model parameters to make the prediction result of the model gradually approach the true label; after each round of training, evaluate the performance of the model on the validation set and record the loss value and accuracy on the validation set; perform model evaluation and optimization;

[0015] Real-time collect image data and sensor data around the urban street, preprocess them and input them into the deep learning model, and the model outputs the current comprehensive environment.

[0016] Combined with the first aspect, in the second implementation manner of the first aspect of this application, based on the current comprehensive environment, construct a generator and a discriminator, adopt adversarial training, alternately train the generator and the discriminator, and perform encoding optimization of the current security monitoring image to generate a simplified image, including:

[0017] The encoder-decoder structure is used as the infrastructure of the generator. The image features of the current security monitoring image are extracted by the encoder; the environmental category of the current comprehensive environment is encoded into a one-hot vector, which is concatenated with the image features extracted by the encoder and then used as the conditional input to the generator; skip connections are added between the encoder and decoder of the generator to connect the features of different layers in the encoder to the corresponding layers in the decoder; the generator generates the simplified image to be detected.

[0018] A discriminator based on a convolutional neural network is constructed. The input of the discriminator is the simplified image to be detected and the current security monitoring image. The discriminator consists of multiple convolutional layers, and the features of the image are extracted through convolutional operations; multi-scale feature fusion is used in the discriminator to discriminate the authenticity and quality of the image; the output layer of the discriminator is a binary classification layer, and the sigmoid activation function is used to output a value between 0 and 1, indicating the probability that the simplified image to be detected is the current security monitoring image.

[0019] Prepare the training set, initialize the generator and discriminator, and perform adversarial training to generate simplified images.

[0020] Combined with the first aspect, in the third implementation manner of the first aspect of the present application, the step of preparing the training set, initializing the generator and discriminator, and performing adversarial training to generate simplified images includes:

[0021] The training set uses historical multi-source data of urban security monitoring. Initialize the weight parameters of the generator and discriminator, and compile the generator and discriminator in the TensorFlow framework; determine the number of adversarial training rounds of the generator and discriminator.

[0022] During the training process, monitor the change of the loss values of the generator and discriminator in real time to judge whether they converge; input the current security monitoring image and the corresponding current comprehensive environment into the generator, and the generator outputs a simplified image.

[0023] Combined with the first aspect, in the fourth implementation manner of the first aspect of the present application, the step of inputting the simplified image into an image coding algorithm for encoding to obtain the first lightweight encoded image includes:

[0024] The user specifies the image coding algorithm, sets the coding parameters, and performs data format conversion, including color space conversion and image size adaptation; encode the simplified image according to the coding process of different coding algorithms, and perform MD5 verification after encoding to obtain the first lightweight encoded image.

[0025] Combined with the first aspect, in the fifth implementation manner of the first aspect of the present application, the step of collecting the current multi-source data of urban security monitoring and extracting multi-source data features includes:

[0026] Collect multi-source data of the current urban security monitoring, including the current security monitoring image, the current lidar data, and the current millimeter-wave radar data; extract the features of the current security monitoring image, including color features, texture features, and shape features, extract the features of the current lidar data, including geometric features and motion features, where the geometric features include point cloud density features, object geometric shape features, and surface normal vector features, extract the features of the current millimeter-wave radar data, including target object attribute features and target object motion pattern features, and the target object attribute features include distance features, speed features, and angle features.

[0027] Combined with the first aspect, in the sixth implementation manner of the first aspect of this application, based on the current comprehensive environment, establish multi-source data feature templates under different comprehensive environments, and compare and analyze the multi-source data features with the corresponding multi-source data feature templates to obtain the matching degree between the environment and the template, including:

[0028] Based on the environmental category of the comprehensive environment, conduct statistical analysis to obtain multi-source data feature templates, including security monitoring image feature templates, lidar data feature templates, and millimeter-wave radar data feature templates; compare the current security monitoring image features, current lidar data features, and current millimeter-wave radar data features in the multi-source data features with the security monitoring image feature templates, lidar data feature templates, and millimeter-wave radar data feature templates respectively, and calculate the comprehensive similarity score by setting the weights of different comparison results; determine the matching degree between the current environment and the template according to the comprehensive similarity score.

[0029] Combined with the first aspect, in the seventh implementation manner of the first aspect of this application, based on the matching degree, determine the information retention situation of the target area in the first lightweight encoded image after lightweight encoding, including:

[0030] When the matching degree between the current environment and the template is higher than the set threshold, according to the target area features in the established multi-source data feature template, the DBSCAN algorithm is used in the point cloud data of the current lidar data to locate the target area by setting the neighborhood radius and the minimum number of points parameters; combining the target position provided by the millimeter-wave radar, using the coordinate transformation and matching algorithm, the target position detected by the millimeter-wave radar is associated with the target area in the point cloud data to confirm the accuracy of the target area; for the located target area, its accuracy is cross-validated by multi-source data; according to the target area contour information in the multi-source data feature template, the contour retention situation of the target area in the first lightweight encoded image is evaluated. The specific method is as follows: Key points are extracted on the contour of the target area in the template and the contour of the first lightweight encoded image, and the contour similarity is obtained by calculating the Euclidean distance between the key points. When the contour similarity is higher than the set threshold, it is judged that the contour retention situation is normal; for the color feature and shape feature of the target object, a quantitative evaluation is carried out. The specific method is as follows: For the color feature, the color histogram is used to compare the color of the target object in the first lightweight encoded image with the color in the template; for the shape feature, a shape matching algorithm based on Hu moments is used to calculate the similarity between the shape of the target object in the first lightweight encoded image and the template shape;

[0031] When the matching degree between the current environment and the template is lower than the set threshold, analyze the differences between the multi-source data features and the multi-source data feature template to obtain abnormal data features; according to the abnormal data features, use the Canny algorithm in the first lightweight encoded image, combined with morphological processing, to find the object contour; combine the lidar data and the millimeter-wave radar data, and use the data fusion algorithm based on Kalman filtering to determine the position and range of the target area; use the topological structure matching algorithm based on the contour shape to calculate the contour retention situation; through the data fusion algorithm based on the Bayesian network, fuse the current multi-source data of urban security monitoring, and judge the information retention situation of the target area according to the probability relationship between the current multi-source data of urban security monitoring.

[0032] Combined with the first aspect, in the eighth implementation manner of the first aspect of the present application, based on the information retention situation, adjust the first lightweight encoded image, and after adjustment, encode it according to the selected image encoding algorithm to generate the final lightweight encoded image, including:

[0033] Based on the information retention situation, contour repair, color adjustment, shape adjustment, and texture adjustment are performed. Among them, contour repair uses an image inpainting algorithm based on PatchMatch, color adjustment uses a method based on histogram matching, shape adjustment uses an affine transformation method, and texture adjustment uses a sample-based texture synthesis algorithm. The parameters are optimized according to the characteristics of the adjusted image, and the adjusted image is encoded according to the selected image coding algorithm to generate the final lightweight encoded image.

[0034] In a second aspect, the present invention provides an image lightweight encoding system based on deep learning, including:

[0035] Comprehensive environment recognition module: including: a deep learning model training unit and a comprehensive environment recognition unit; among them, the deep learning model training unit collects historical multi-source data of urban security monitoring, trains a deep learning model using a convolutional neural network, performs model training, evaluation, and optimization, and the comprehensive environment recognition unit recognizes the current comprehensive environment;

[0036] Lightweight encoded image transfer module: including: a simplified image generation unit and a first lightweight encoded image generation unit; among them, the simplified image generation unit constructs a generator and a discriminator based on the current comprehensive environment, adopts adversarial training, alternately trains the generator and the discriminator, performs encoding optimization of the current security monitoring image, and generates a simplified image; the first lightweight encoded image generation unit inputs the simplified image into an image coding algorithm for encoding to obtain a first lightweight encoded image;

[0037] Matching degree calculation module: including: a multi-source data feature extraction unit, a feature template establishment unit, and a matching degree calculation unit; among them, the multi-source data feature extraction unit collects current multi-source data of urban security monitoring and extracts multi-source data features; the feature template establishment unit establishes multi-source data feature templates under different comprehensive environments based on the current comprehensive environment, compares and analyzes the multi-source data features with the corresponding multi-source data feature templates to obtain the matching degree between the environment and the template;

[0038] Information retention situation calculation module: including: an information retention situation calculation unit; among them, the information retention situation calculation unit determines the information retention situation of the target area in the first lightweight encoded image after lightweight encoding based on the matching degree;

[0039] Lightweight encoded image generation module: including: a lightweight encoded image adjustment unit and a lightweight encoded image generation unit; among them, the lightweight encoded image adjustment unit adjusts the first lightweight encoded image based on the information retention situation, and after adjustment, the lightweight encoded image generation unit performs encoding according to the selected image coding algorithm to generate the final lightweight encoded image.

[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0041] 1. The present invention utilizes a deep learning model to accurately identify the current comprehensive environment of urban streets, constructs a generator and a discriminator based on different environmental characteristics, and optimizes the encoding of surveillance images through adversarial training; during peak traffic hours on urban streets, it can quickly analyze the image content in complex scenarios, intelligently adjust the encoding parameters and processes, and rapidly generate simplified images.

[0042] 2. The present invention collects multi-source data under different weather, time periods, and scenarios on urban streets, and establishes a multi-source data feature template; during the actual surveillance process, it flexibly adjusts the encoding strategy according to the matching degree between the current environment and the template.

[0043] 3. The lightweight encoded images generated by the present invention have a significantly reduced data volume while ensuring the integrity of the key information in the images; in terms of storing urban street surveillance data, it can significantly reduce the storage space requirements and save a large amount of storage costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a schematic diagram of the steps of an image lightweight encoding method based on deep learning according to the present invention;

[0045] Figure 2 is a system structure diagram of an image lightweight encoding system based on deep learning according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0047] Embodiment: As Figure 1 - Figure 2 shown, the present invention provides a technical solution,

[0048] As Figure 1 shown in the schematic diagram of the steps of an image lightweight encoding method based on deep learning, the present invention provides an image lightweight encoding method based on deep learning, including:

[0049] Step S100: Collect multi-source data of historical urban security surveillance, train a deep learning model using a convolutional neural network, perform model training, evaluation, and optimization, and identify the current comprehensive environment;

[0050] Specifically, images are collected from urban street cameras under different weather conditions, at different times, and in different scenarios, and lidar data and millimeter-wave radar data are collected synchronously. The lidar data provides three-dimensional point cloud information of the street surrounding environment, and the millimeter-wave radar data includes distance, speed, and angle information of target objects. The historical urban security monitoring multi-source data is integrated and formed.

[0051] Preprocess the historical urban security monitoring multi-source data and divide it into a training set, a validation set, and a test set. Select ResNet for the convolutional neural network structure, and set the loss function, optimizer, and number of training epochs. Input the image data and corresponding sensor data in the training set into the deep learning model. The model calculates the prediction result through forward propagation. According to the difference between the prediction result and the true label, the loss value is calculated using the loss function. Calculate the gradient of the loss value with respect to the model parameters through the backpropagation algorithm, and use the optimizer to update the model parameters to make the prediction result of the model gradually approach the true label. After each round of training, evaluate the performance of the model on the validation set and record the loss value and accuracy on the validation set. Conduct model evaluation and optimization.

[0052] Real-time collect the image data and sensor data around the urban street, and after preprocessing, input them into the deep learning model, and the model outputs the current comprehensive environment.

[0053] In a specific embodiment, 50 representative urban street monitoring cameras are selected in the core area of a medium-sized city, and data is continuously collected for 3 months. 10 lidar devices are deployed at key intersections and sections to collect data synchronously with the image data. 20 millimeter-wave radar sensors are installed in some monitoring areas, and the operating frequency band is 77 GHz.

[0054] Conduct data preprocessing, and divide the preprocessed multi-source data into a training set, a validation set, and a test set according to the ratio of 70%, 15%, and 15%. The training set contains approximately 630,000 samples, and each of the validation set and the test set contains approximately 135,000 samples. Select ResNet-101 as the convolutional neural network structure. The cross-entropy loss function is used as the loss function, the Adam optimizer is selected, the learning rate is set to 0.0001, beta1 = 0.9, and beta2 = 0.999. The number of training epochs is set to 150 rounds.

[0055] During the training process, in each round of training, the image data in the training set and the corresponding lidar and millimeter-wave radar data are input into the deep learning model. The model calculates the prediction results through forward propagation. For example, in the 30th round of training, for a batch of training data containing 128 samples, the cross-entropy loss value is calculated by comparing the predicted comprehensive environment category of the model with the true label. The gradient of the loss value with respect to the model parameters is calculated through the backpropagation algorithm, and the model parameters are updated using the Adam optimizer. During the training process, the loss value and accuracy on the validation set are recorded for each round of training. As the number of training rounds increases, the loss value on the validation set gradually decreases, and the accuracy gradually increases. In the 80th round of training, the loss value on the validation set drops to about 0.4, and the accuracy reaches 85%. By the end of the 150th round of training, the loss value on the validation set stabilizes at about 0.25, and the accuracy reaches 92%. The trained model is evaluated using the test set. In the test set, there are samples with various combinations of weather, time periods, and scenarios. The model predicts 135,000 samples in the test set, and finally obtains an accuracy of 90%, a recall rate of 88%, and an F1 value of 89%. When evaluating different environment categories separately, in the main road scenario on a sunny day, the accuracy of the model is as high as 96% because the data features of this scenario are relatively obvious and the model is easy to learn and identify. While in the secondary road scenario at night in rainy weather, the accuracy is relatively low, at 82%, mainly because the light is dim and the rain interference is large in this scenario, and the data features are complex, increasing the difficulty of model recognition.

[0056] During the actual monitoring process of urban streets, the image data and sensor data around the vehicle are collected in real time. For example, in the evening of a working day, a set of image data and sensor data are collected in real time. After preprocessing, they are input into the deep learning model. The model can quickly output the current comprehensive environment category. During the continuous 1-hour real-time monitoring process, a total of 3,600 environment identifications are carried out, among which 3,200 are accurately identified, and the accuracy reaches 88.9%, which can basically meet the real-time and accuracy requirements of environment identification for urban street security monitoring.

[0057] Step S200: Based on the current comprehensive environment, construct a generator and a discriminator, adopt adversarial training, alternately train the generator and the discriminator, perform encoding optimization on the current security monitoring image to generate a simplified image; input the simplified image into an image encoding algorithm for encoding to obtain a first lightweight encoded image;

[0058] Specifically, an encoder-decoder structure is adopted as the basic architecture of the generator. The image features of the current security monitoring image are extracted by the encoder; the environmental category of the current comprehensive environment is encoded into a one-hot vector, which is concatenated with the image features extracted by the encoder and then used as the conditional input to the generator; skip connections are added between the encoder and the decoder of the generator to connect the features of different layers in the encoder to the corresponding layers in the decoder; the generator generates the simplified image to be detected.

[0059] A discriminator based on a convolutional neural network is constructed. The input of the discriminator is the simplified image to be detected and the current security monitoring image. The discriminator consists of multiple convolutional layers, and the features of the image are extracted through convolutional operations; multi-scale feature fusion is adopted in the discriminator to discriminate the authenticity and quality of the image; the output layer of the discriminator is a binary classification layer, using the sigmoid activation function to output a value between 0 and 1, indicating the probability that the simplified image to be detected is the current security monitoring image.

[0060] Prepare the training set, initialize the generator and the discriminator, and perform adversarial training to generate the simplified image.

[0061] Furthermore, the historical multi-source data of urban security monitoring is used in the training set, the weight parameters of the generator and the discriminator are initialized, and the generator and the discriminator are compiled in the TensorFlow framework; the number of adversarial training rounds of the generator and the discriminator is determined.

[0062] During the training process, the change of the loss values of the generator and the discriminator is monitored in real time to judge whether it converges; the current security monitoring image and the corresponding current comprehensive environment are input into the generator, and the generator outputs the simplified image.

[0063] Furthermore, the user specifies the image encoding algorithm, sets the encoding parameters, and performs data format conversion, including color space conversion and image size adaptation; the simplified image is encoded according to the encoding process of different encoding algorithms, and after encoding, MD5 verification is performed to obtain the first lightweight encoded image.

[0064] In a specific embodiment, based on the encoder-decoder structure, the encoder part uses a series of convolutional layers with 3×3 convolutional kernels and a stride of 2. The current security monitoring image with a size of 224×224 pixels is input. After passing through the first convolutional layer, the image size becomes 112×112 pixels, and the number of channels increases from 3 to 64. As the number of layers increases, the image is finally compressed into a low-dimensional feature representation. The decoder uses deconvolutional layers with 3×3 deconvolutional kernels and a stride of 2 to gradually upsample the low-dimensional features back to the original image size. Skip connections are added between the encoder and the decoder. For example, the features of the second layer of the encoder are concatenated with the input of the corresponding deconvolutional layer in the decoder. If there are 8 categories in the current comprehensive environment, they are encoded into an 8-dimensional one-hot vector, which is concatenated with the features extracted by the encoder in the channel dimension and then input into the generator. The discriminator consists of multiple convolutional layers, with the convolutional kernel size being 3×3 or 5×5 and the stride being set to 1 or 2. The input is the simplified image to be detected and the current security monitoring image. After multiple convolutional operations, the image features are compressed into a fixed-size vector. Multi-scale feature fusion is adopted, and the input image is downsampled at scales of 1 / 2, 1 / 4, and 1 / 8. Features are extracted separately and then concatenated. The output layer is a binary classification layer, and the sigmoid activation function is used to output a value between 0 and 1.

[0065] The image data in the previously collected historical multi-source data of urban security monitoring is used as the training set, and 10,000 images are selected from it, with at least 50 images for each combination of weather, time period, and scene. The weight parameters of the generator and the discriminator are initialized with a Gaussian distribution (mean of 0 and standard deviation of 0.02), and the bias parameters are initialized to 0. In the TensorFlow framework, both the generator and the discriminator use the binary cross-entropy loss function. The optimizer is selected as Adam. The learning rate of the generator is set to 0.0002, the learning rate of the discriminator is set to 0.0001, and beta1 is set to 0.5 for both. The number of adversarial training rounds is set to 300 rounds. The user specifies the JPEG encoding algorithm. The quality factor is set to 70, and the quantization table uses the default standard table. For color space conversion, the RGB image is converted to the YUV color space. In terms of image size adaptation, it is ensured that the size of the simplified image is 224×224 pixels. If the size does not match, it is cropped or padded.

[0066] Model training is carried out. As the number of training rounds increases, the loss values of the generator and the discriminator gradually decrease and tend to be stable. At the 200th round of training, the loss value of the generator drops to about 0.4, and the loss value of the discriminator drops to about 0.35, basically reaching the convergence state. After the training is completed, the new current security monitoring image and the corresponding comprehensive environment are input into the generator. In actual tests, 100 current security monitoring images in different environments are processed, and the generator successfully outputs simplified images. Visually, in the sunny crossroads scene, the simplified images clearly retain key information such as traffic lights, vehicles, and pedestrians, only removing some background detail textures; in the rainy night secondary road scene, although the overall clarity of the images decreases, the road contours, roadside buildings, and moving pedestrians and vehicles are still distinguishable.

[0067] The generated simplified images are encoded in JPEG. After encoding, MD5 verification is performed to ensure data integrity.

[0068] Step S300: Collect multi-source data of current urban security monitoring, extract multi-source data features; based on the current comprehensive environment, establish multi-source data feature templates under different comprehensive environments, and compare and analyze the multi-source data features with the corresponding multi-source data feature templates to obtain the matching degree between the environment and the template.

[0069] Specifically, collect multi-source data of current urban security monitoring, including current security monitoring images, current lidar data, and current millimeter-wave radar data; extract current security monitoring image features, including color features, texture features, and shape features, extract current lidar data features, including geometric features and motion features, where the geometric features include point cloud density features, object geometric shape features, and surface normal vector features, extract current millimeter-wave radar data features, including target object attribute features and target object motion mode features, and the target object attribute features include distance features, speed features, and angle features.

[0070] Furthermore, based on the environmental categories of the comprehensive environment, statistical analysis is carried out to obtain multi-source data feature templates, including security monitoring image feature templates, lidar data feature templates, and millimeter-wave radar data feature templates; compare the current security monitoring image features, current lidar data features, and current millimeter-wave radar data features in the multi-source data features with the security monitoring image feature templates, lidar data feature templates, and millimeter-wave radar data feature templates respectively, calculate the comprehensive similarity score by setting the weights of different comparison results; determine the matching degree between the current environment and the template according to the comprehensive similarity score.

[0071] In a specific embodiment, statistical analysis is performed on the color, texture, and shape features of 250 images in each environment, and statistics such as the mean, median, and standard deviation of various features are calculated to construct a feature template. For example, in the sunny highway scenario, in the image color feature template, the mean of the blue hue in the sky area in the RGB space is [100, 149, 237], and the variance is [10, 12, 15]; in the texture feature template, the mean of the GLCM energy in the road area is 0.12, the mean of the entropy is 3.2, and the mean of the contrast is 0.35; in the shape feature template, the mean of the Hu moments of the vehicle contour is [0.001, 0.002, 0.0015, -0.0005, 0.0003, -0.0002, 0.0001].

[0072] Statistical analysis is performed on the lidar data features in each environment. For example, in the rainy city street scenario, in the point cloud density feature template, the mean of the point cloud density in the building area is 200 points per cubic meter, and the standard deviation is 50 points per cubic meter; in the object geometry feature template, the mean length of the vehicle is 4.5 meters, and the mean width is 1.8 meters; in the surface normal vector feature template, the mean of the main direction of the building wall normal vector is [0, 0, 1] (vertical direction). In the motion feature template, the average speed of the vehicle in this environment is 25 km / h, and the mean acceleration is 0.5 m / s².

[0073] Statistical analysis is performed on the millimeter-wave radar data features in each environment. In the night rural road scenario, in the target object attribute feature template, the mean of the distance feature is 80 meters, the mean of the speed feature is 30 km / h, and the mean of the angle feature is 0° (directly ahead); in the target object motion pattern feature template, the probability of uniform linear motion is 0.7, and the probability of turning motion is 0.2.

[0074] The weights of the comparison results of the current security monitoring image features, current lidar data features, and current millimeter-wave radar data features are set to 0.4, 0.3, and 0.3 respectively. Methods such as Euclidean distance and cosine similarity are used to calculate the similarity between the current features and the template features. For color features, the Euclidean distance of the color histogram is calculated; for texture features, the cosine similarity of the GLCM parameters and LBP pattern distribution is calculated; for shape features, the Euclidean distance of the Hu moments is calculated. For lidar and millimeter-wave radar data features, the corresponding distance or similarity measurement methods are also used.

[0075] Randomly select 50 current security monitoring images and their corresponding lidar and millimeter-wave radar data in each environment for feature comparison. For example, in the scenario of a cloudy urban main road, the Euclidean distance between the color features of a current security monitoring image and the template features is calculated to obtain a similarity of 0.8, the cosine similarity of the texture features is 0.75, and the Euclidean distance similarity of the shape features is 0.85; when the lidar data features are compared with the template, the geometric feature similarity is 0.7, and the motion feature similarity is 0.72; when the millimeter-wave radar data features are compared with the template, the similarity of the target object attribute features is 0.8, and the similarity of the target object motion mode features is 0.78. Calculate the comprehensive similarity score according to the weights: (0.8×0.4 + 0.75×0.4 + 0.85×0.4)×0.4 + (0.7×0.3 + 0.72×0.3)×0.3 + (0.8×0.3 + 0.78×0.3)×0.3 = 0.774.

[0076] Step S400: Based on the matching degree, determine the information retention situation of the target area in the first lightweight encoded image after lightweight encoding;

[0077] Specifically, when the matching degree between the current environment and the template is higher than the set threshold, according to the target area features in the established multi-source data feature template, use the DBSCAN algorithm in the point cloud data of the current lidar data, and by setting the neighborhood radius and minimum point number parameters, locate the target area; combine the target position provided by the millimeter-wave radar, and use the coordinate transformation and matching algorithm to associate the target position detected by the millimeter-wave radar with the target area in the point cloud data to confirm the accuracy of the target area; for the located target area, cross-verify its accuracy through multi-source data; according to the target area contour information in the multi-source data feature template, evaluate the contour retention situation of the target area in the first lightweight encoded image. The specific method is as follows: Extract key points on the contour of the target area in the template and the contour of the first lightweight encoded image, and calculate the Euclidean distance between the key points to obtain the contour similarity. When the contour similarity is higher than the set threshold, it is judged that the contour retention situation is normal; for the color features and shape features of the target object, conduct quantitative evaluation. The specific method is as follows: For the color features, use the color histogram to compare the color of the target object in the first lightweight encoded image with the color in the template; for the shape features, adopt the shape matching algorithm based on Hu moments to calculate the similarity between the shape of the target object in the first lightweight encoded image and the template shape.

[0078] When the matching degree between the current environment and the template is lower than the set threshold, analyze the differences between the multi-source data features and the multi-source data feature template to obtain abnormal data features; according to the abnormal data features, use the Canny algorithm in the first lightweight encoded image, combined with morphological processing, to find the object contour; combine the lidar data and millimeter-wave radar data, and use the data fusion algorithm based on Kalman filtering to determine the position and range of the target area; use the topological structure matching algorithm based on the contour shape to calculate the contour retention; through the data fusion algorithm based on the Bayesian network, fuse the current multi-source data of urban security monitoring, and judge the information retention of the target area according to the probability relationship between the current multi-source data of urban security monitoring.

[0079] In a specific embodiment, the threshold for the matching degree between the current environment and the template is set to 0.8. Higher than this threshold means a high matching degree, and lower means a low matching degree. In the scenario of the main road in the morning on a sunny day, the average matching degree between the current environment and the template is 0.85, which is higher than the threshold. According to the multi-source data feature template, use the DBSCAN algorithm in the current lidar point cloud data to successfully locate multiple target areas. For example, in a certain area, a point cloud cluster is detected, and after calculation, this point cloud cluster conforms to the vehicle target area characteristics. Combining the target position information provided by the millimeter-wave radar, using the coordinate transformation and matching algorithm, associate the target position 70 meters ahead detected by the millimeter-wave radar with the corresponding area in the lidar point cloud data, and confirm that the target area is a moving vehicle. Through multi-source data cross-verification, that is, comparing the visual features of this area in the image with the lidar and millimeter-wave radar data, to confirm the accuracy of the target area. Among 100 localizations, the target area was accurately confirmed 95 times, and the accuracy rate was 95%. Extract key points on the vehicle target area contour in the template and the vehicle contour in the first lightweight encoded image, such as the inflection points of the contour and the points with large curvature changes. By calculating the Euclidean distance between these key points, the contour similarity is obtained. Calculate the vehicle target areas in 100 first lightweight encoded images in this scenario, and the average contour similarity is 0.88, which is higher than the set threshold of 0.75, and it is judged that the contour retention is normal. For example, the average Euclidean distance between the key points of the vehicle contour in a certain first lightweight encoded image and the key points of the template contour is 4 pixels (image resolution is 1920×1080), and the calculated contour similarity is 0.9.

[0080] The color histogram is used to compare the color of the vehicle target object in the first lightweight encoded image with the color in the template. The Bhattacharyya distance of the color histogram is calculated, and the average Bhattacharyya distance of 100 images is 0.13, indicating that the color features are well preserved. For example, the Bhattacharyya distance between the color histogram of a certain vehicle in the template and the color histogram in the first lightweight encoded image is 0.11, indicating that the color features of the two are highly similar. The shape matching algorithm based on Hu moments is used to calculate the similarity between the shape of the vehicle target object in the first lightweight encoded image and the shape of the template. After calculation, the average shape similarity of 100 images is 0.86, indicating that the shape features are well preserved. For example, the Euclidean distance between the Hu moment feature vector of the vehicle shape in a certain image and the Hu moment feature vector of the vehicle shape in the template is small, and the shape similarity is calculated to be 0.88.

[0081] In the scenario of a rainy intersection at noon, the average matching degree between the current environment and the template is 0.72, which is lower than the threshold. By analyzing the differences between the multi-source data features and the template, it is found that the point cloud density of the lidar point cloud data is abnormally reduced in some areas due to rain interference, and the speed and distance information of the target object detected by the millimeter wave radar fluctuate greatly. The Canny algorithm is used in the first lightweight coded image, combined with morphological processing, to find the object contour. The contour retention is calculated using the topological structure matching algorithm based on the contour shape. The algorithm focuses on topological features such as the connectivity, number of holes, and approximate shape of the contour. The target area in the 100 first lightweight coded images in this scenario is calculated, and the average topological structure similarity is 0.73, indicating that the contour retention is acceptable. The current urban security monitoring multi-source data is fused through the data fusion algorithm based on the Bayesian network. According to the probability relationship between the current urban security monitoring multi-source data, the information retention of the target area is judged. In the evaluation of 100 target areas, 70 areas were judged to have good information retention, accounting for about 70%.

[0082] Step S500: adjusting the first light-quantized coded image based on the information retention status, encoding the adjusted image according to the selected image encoding algorithm to generate a final light-quantized coded image.

[0083] Specifically, based on the information retention, contour repair, color adjustment, shape adjustment and texture adjustment are performed; among them, contour repair uses an image repair algorithm based on PatchMatch, color adjustment uses a histogram matching-based method, shape adjustment uses an affine transformation method, and texture adjustment uses a sample-based texture synthesis algorithm; parameters are optimized according to the characteristics of the adjusted image, and the adjusted image is encoded according to the selected image coding algorithm to generate a final lightweight encoded image.

[0084] like Figure 2As shown in the system structure diagram of an image lightweight encoding system based on deep learning, the present invention provides an image lightweight encoding system based on deep learning, including:

[0085] Comprehensive environment recognition module: including: a deep learning model training unit and a comprehensive environment recognition unit; among them, the deep learning model training unit collects historical multi-source data of urban security monitoring, uses a convolutional neural network to train the deep learning model, conducts model training, evaluation and optimization, and the comprehensive environment recognition unit recognizes the current comprehensive environment;

[0086] Lightweight encoding image transfer module: including: a simplified image generation unit and a first lightweight encoding image generation unit; among them, the simplified image generation unit constructs a generator and a discriminator based on the current comprehensive environment, adopts adversarial training, alternately trains the generator and the discriminator, conducts encoding optimization of the current security monitoring image, and generates a simplified image; the first lightweight encoding image generation unit inputs the simplified image into an image encoding algorithm for encoding to obtain a first lightweight encoding image;

[0087] Matching degree calculation module: including: a multi-source data feature extraction unit, a feature template establishment unit and a matching degree calculation unit; among them, the multi-source data feature extraction unit collects current multi-source data of urban security monitoring and extracts multi-source data features; the feature template establishment unit establishes multi-source data feature templates under different comprehensive environments based on the current comprehensive environment, and compares and analyzes the multi-source data features with the corresponding multi-source data feature templates to obtain the matching degree between the environment and the template;

[0088] Information retention situation calculation module: including: an information retention situation calculation unit; among them, the information retention situation calculation unit determines the information retention situation of the target area in the first lightweight encoding image after lightweight encoding based on the matching degree;

[0089] Lightweight encoding image generation module: including: a lightweight encoding image adjustment unit and a lightweight encoding image generation unit; among them, the lightweight encoding image adjustment unit adjusts the first lightweight encoding image based on the information retention situation, and after adjustment, the lightweight encoding image generation unit encodes according to the selected image encoding algorithm to generate a final lightweight encoding image.

[0090] It is obvious to those skilled in the art that the present invention is not limited to the details of the above-described exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, in any respect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Accordingly, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

Claims

1. An image lightweight encoding method based on deep learning, characterized in that Including: Collect multi-source data of historical urban security monitoring, use a convolutional neural network to train a deep learning model, conduct model training, evaluation and optimization, and identify the current comprehensive environment; Based on the current comprehensive environment, construct a generator and a discriminator, adopt adversarial training, alternately train the generator and the discriminator, conduct encoding optimization of the current security monitoring image, and generate a simplified image; Input the simplified image into an image encoding algorithm for encoding to obtain a first lightweight encoded image; Collect multi-source data of current urban security monitoring and extract multi-source data features; Based on the current comprehensive environment, establish multi-source data feature templates under different comprehensive environments, compare and analyze the multi-source data features with the corresponding multi-source data feature templates, and obtain the matching degree between the environment and the template; Based on the matching degree, determine the information retention situation of the target area in the first lightweight encoded image after lightweight encoding; Based on the information retention situation, adjust the first lightweight encoded image, and after adjustment, encode it according to the selected image encoding algorithm to generate a final lightweight encoded image.

2. The image lightweight encoding method based on deep learning according to claim 1, wherein The collection of multi-source data of historical urban security monitoring, using a convolutional neural network to train a deep learning model, conducting model training, evaluation and optimization, and identifying the current comprehensive environment includes: Collect images under different weather conditions, different time periods and different scenarios from street cameras in the city, and synchronously collect lidar data and millimeter-wave radar data; the lidar data provides three-dimensional point cloud information of the surrounding environment of the street, and the millimeter-wave radar data includes distance, speed and angle information of target objects; integrate to form multi-source data of historical urban security monitoring; Preprocess the multi-source data of historical urban security monitoring and divide it into a training set, a validation set and a test set; select the ResNet for the convolutional neural network structure, set the loss function, optimizer and number of training epochs; input the image data and corresponding sensor data in the training set into the deep learning model, the model calculates the prediction result through forward propagation, according to the difference between the prediction result and the true label, use the loss function to calculate the loss value; calculate the gradient of the loss value with respect to the model parameters through the backpropagation algorithm, and use the optimizer to update the model parameters to make the prediction result of the model gradually approach the true label; after each round of training, evaluate the performance of the model on the validation set and record the loss value and accuracy on the validation set; conduct model evaluation and optimization; Real-time collect image data and sensor data around the urban streets, and after preprocessing, input them into the deep learning model, and the model outputs the current comprehensive environment.

3. A method for lightweight encoding of images based on deep learning according to claim 1, characterized in that Based on the current comprehensive environment, construct a generator and a discriminator, adopt adversarial training, alternately train the generator and the discriminator, conduct encoding optimization of the current security monitoring image, and generate a simplified image, including: An encoder-decoder structure is adopted as the infrastructure of the generator. The image features of the current security monitoring image are extracted by the encoder; the environmental category of the current comprehensive environment is encoded into a one-hot vector, which is concatenated with the image features extracted by the encoder and then used as the conditional input to the generator; skip connections are added between the encoder and the decoder of the generator to connect the features of different levels in the encoder to the corresponding levels in the decoder; the generator generates the simplified image to be detected. A discriminator based on a convolutional neural network is constructed. The input of the discriminator is the simplified image to be detected and the current security monitoring image. The discriminator consists of multiple convolutional layers, and the features of the image are extracted through convolutional operations; multi-scale feature fusion is adopted in the discriminator to discriminate the authenticity and quality of the image; the output layer of the discriminator is a binary classification layer, and the sigmoid activation function is used to output a value between 0 and 1, representing the probability that the simplified image to be detected is the current security monitoring image. Prepare the training set, initialize the generator and the discriminator, and perform adversarial training to generate simplified images.

4. The image lightweight encoding method based on deep learning according to claim 3, characterized in that The steps of preparing the training set, initializing the generator and the discriminator, and performing adversarial training to generate simplified images include: The training set uses historical multi-source data of urban security monitoring. Initialize the weight parameters of the generator and the discriminator, and compile the generator and the discriminator in the TensorFlow framework; determine the number of adversarial training rounds of the generator and the discriminator. During the training process, monitor the change of the loss values of the generator and the discriminator in real time to judge whether they converge; input the current security monitoring image and the corresponding current comprehensive environment into the generator, and the generator outputs the simplified image.

5. A method for lightweight encoding of images based on deep learning according to claim 1, characterized in that, The steps of inputting the simplified image into an image coding algorithm for encoding to obtain the first lightweight encoded image include: The user specifies the image coding algorithm, sets the coding parameters, and performs data format conversion, including color space conversion and image size adaptation; encode the simplified image according to the coding process of different coding algorithms, and perform MD5 verification after encoding to obtain the first lightweight encoded image.

6. A lightweight image encoding method based on deep learning according to claim 1, characterized in that The steps of collecting the current multi-source data of urban security monitoring and extracting the multi-source data features include: Collect the current multi-source data of urban security monitoring, including the current security monitoring image, the current lidar data, and the current millimeter wave radar data; extract the features of the current security monitoring image, including color features, texture features, and shape features, extract the features of the current lidar data, including geometric features and motion features, where the geometric features include point cloud density features, object geometric shape features, and surface normal vector features, extract the features of the current millimeter wave radar data, including target object attribute features and target object motion pattern features, and the target object attribute features include distance features, speed features, and angle features.

7. A method for lightweight encoding of images based on deep learning according to claim 1, characterized in that Based on the current comprehensive environment, establish multi-source data feature templates under different comprehensive environments, and compare and analyze the multi-source data features with the corresponding multi-source data feature templates to obtain the matching degree between the environment and the template, including: Based on the environmental categories of the comprehensive environment, statistical analysis is carried out to obtain multi-source data feature templates, including security monitoring image feature templates, lidar data feature templates, and millimeter-wave radar data feature templates; the current security monitoring image features, current lidar data features, and current millimeter-wave radar data features in the multi-source data features are respectively compared with the security monitoring image feature templates, lidar data feature templates, and millimeter-wave radar data feature templates, and the comprehensive similarity score is calculated by setting the weights of different comparison results; according to the comprehensive similarity score, the matching degree between the current environment and the template is determined.

8. A method for lightweight encoding of images based on deep learning according to claim 1, characterized in that Based on the matching degree, determine the information retention situation of the target area in the first lightweight encoded image, including: When the matching degree between the current environment and the template is higher than the set threshold, according to the target area features in the established multi-source data feature template, use the DBSCAN algorithm in the point cloud data of the current lidar data, and by setting the neighborhood radius and minimum number of points parameters, locate the target area; combine the target position provided by the millimeter-wave radar, and use coordinate transformation and matching algorithms to associate the target position detected by the millimeter-wave radar with the target area in the point cloud data to confirm the accuracy of the target area; for the located target area, cross-verify its accuracy through multi-source data; according to the target area contour information in the multi-source data feature template, evaluate the contour retention situation of the target area in the first lightweight encoded image. The specific method is: extract key points on the contour of the target area in the template and the contour of the first lightweight encoded image, and obtain the contour similarity by calculating the Euclidean distance between the key points. When the contour similarity is higher than the set threshold, it is judged that the contour retention situation is normal; for the color features and shape features of the target object, quantitative evaluation is carried out. The specific method is: for color features, use color histograms to compare the colors of the target object in the first lightweight encoded image with the colors in the template; for shape features, use a shape matching algorithm based on Hu moments to calculate the similarity between the shape of the target object in the first lightweight encoded image and the template shape. When the matching degree between the current environment and the template is lower than the set threshold, analyze the differences between the multi-source data features and the multi-source data feature templates to obtain abnormal data features; according to the abnormal data features, use the Canny algorithm in the first lightweight encoded image, combined with morphological processing, to find the object contour; combine the lidar data and millimeter-wave radar data, and use a data fusion algorithm based on Kalman filtering to determine the position and range of the target area; use a topological structure matching algorithm based on contour shape to calculate the contour retention situation; through a data fusion algorithm based on Bayesian network, fuse the current multi-source data of urban security monitoring, and judge the information retention situation of the target area according to the probability relationship between the current multi-source data of urban security monitoring.

9. A method for lightweight encoding of images based on deep learning according to claim 1, characterized in that Based on the information retention situation, adjust the first lightweight encoded image, and after adjustment, encode it according to the selected image encoding algorithm to generate the final lightweight encoded image, including: Based on the information retention situation, contour repair, color adjustment, shape adjustment, and texture adjustment are performed; among them, contour repair uses an image inpainting algorithm based on PatchMatch, color adjustment uses a method based on histogram matching, shape adjustment uses an affine transformation method, and texture adjustment uses a sample-based texture synthesis algorithm; the parameters are optimized according to the characteristics of the adjusted image, and the adjusted image is encoded according to the selected image coding algorithm to generate the final lightweight coded image.

10. An image lightweight encoding system based on deep learning, which uses a method for image lightweight encoding based on deep learning according to any one of claims 1-9, characterized in that, Including: Comprehensive environment recognition module: including: a deep learning model training unit and a comprehensive environment recognition unit; among them, the deep learning model training unit collects historical multi-source data of urban security monitoring, uses a convolutional neural network to train the deep learning model, and conducts model training, evaluation, and optimization. The comprehensive environment recognition unit recognizes the current comprehensive environment; Lightweight coded image transfer module: including: a simplified image generation unit and a first lightweight coded image generation unit; among them, the simplified image generation unit constructs a generator and a discriminator based on the current comprehensive environment, adopts adversarial training, alternately trains the generator and the discriminator, optimizes the encoding of the current security monitoring image, and generates a simplified image; the first lightweight coded image generation unit inputs the simplified image into the image coding algorithm for encoding to obtain the first lightweight coded image; Matching degree calculation module: including: a multi-source data feature extraction unit, a feature template establishment unit, and a matching degree calculation unit; among them, the multi-source data feature extraction unit collects the current multi-source data of urban security monitoring and extracts the multi-source data features; the feature template establishment unit establishes multi-source data feature templates under different comprehensive environments based on the current comprehensive environment, and compares and analyzes the multi-source data features with the corresponding multi-source data feature templates to obtain the matching degree between the environment and the template; Information retention situation calculation module: including: an information retention situation calculation unit; among them, the information retention situation calculation unit determines the information retention situation of the target area in the first lightweight coded image after lightweight coding based on the matching degree; Lightweight coded image generation module: including: a lightweight coded image adjustment unit and a lightweight coded image generation unit; among them, the lightweight coded image adjustment unit adjusts the first lightweight coded image based on the information retention situation. After adjustment, the lightweight coded image generation unit encodes according to the selected image coding algorithm to generate the final lightweight coded image.

Citation Information

Patent Citations

  • Image inpainting method and system based on antagonistic generation neural network

    CN109191402A

  • Medical image enhancement processing method and system based on deep learning

    CN119671884A

  • End-to-end deep generative network for low bitrate image coding

    US20240185473A1

  • Focusing-learning-based CT angiography smart imaging method

    WO2024066711A1

  • Three-dimensional lidar point cloud semantic segmentation method and apparatus based on deep learning

    WO2024130776A1