Data processing method and AI system based on big data
Through big data-based data processing methods combined with deep learning and Transformer models, the problem of poor thermal infrared image stitching ability is solved, high-quality panoramic thermal infrared images are generated, and the accuracy and efficiency of image conversion are improved. It is suitable for a variety of image types and has wide application value.
Patent Information
- Application Number
- CN202510446483.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-04-10
Smart Images

Figure CN120298214B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data processing, and in particular to a data processing method and an AI system based on big data. Background Art
[0002] In modern society, thermal infrared imaging has been widely used in various fields, including military reconnaissance, environmental monitoring, geographic information collection, and agricultural disaster assessment. Thermal infrared imaging provides a method of observation regardless of lighting conditions, revealing the heat distribution and heat generation of objects—information that is crucial in many fields. Drone swarms capture numerous thermal infrared images of the same area while flying. These images are sometimes isolated and fragmentary, so obtaining a complete panoramic view requires stitching them together to form a panoramic thermal infrared image.
[0003] Existing thermal infrared image stitching technology is mainly based on traditional image matching, including image preprocessing, feature extraction, feature matching and image fusion. For example, the patent document with publication number CN113689479A discloses a method for registration of thermal infrared and visible light images of unmanned aerial vehicles. The method includes step 1: determining the reference image and input image of the image pair to be registered, judging the resolution of the reference image and the input image, and directly executing step 2 if the two images have the same resolution and size, otherwise executing step 2 after image preprocessing, the image preprocessing includes: if the resolution is different, sampling the two images to the same resolution, if the size is different, center cropping to the same size; step 2: extracting the directional gradient channel CFOG features of the reference image and the input image respectively, obtaining the CFOG features of each pixel point of the reference image and the CFOG features of each pixel point of the input image, dividing the reference image into multiple non-overlapping blocks, and determining each block obtained The search area of each atomic block is determined according to the maximum deviation between the reference image and the input image. Template matching is performed in the search area of the input image based on the CFOG feature of each atomic block to obtain the similarity map of the atomic block; Step 3: The similarity map of the atomic block is aggregated through multiple layers of local maximum to obtain a pyramid similarity map of blocks of different sizes; Step 4: In the pyramid similarity map, the best matching position corresponding to the block of the current layer is deduced layer by layer starting from the highest layer through the maximum index backtracking to obtain the best matching position corresponding to the atomic block; Step 5: According to the best matching position of the atomic block, its homonymous points are obtained, thereby obtaining a set of homonymous points; Step 6: The homonymous points in the set of homonymous points whose errors exceed the specified error range are deleted, and then the homography matrix is calculated based on the current set of homonymous points to align the reference image and the input image.
[0004] However, this traditional image stitching technology only performs limited stitching on some photos when processing thermal infrared images, resulting in poor stitching capabilities. Summary of the Invention
[0005] To this end, the present invention provides a data processing method and AI system based on big data, which solves the problem of poor splicing ability of thermal infrared images in the prior art.
[0006] To achieve the above objectives, the present invention provides a data processing method based on big data, which includes:
[0007] Acquire a plurality of images to be processed, taken at a collection point in a direction toward a spatial boundary of a space to be imaged, wherein the collection point is set within the space to be imaged, the images to be processed include an actual thermal infrared image and a visible light image corresponding to the actual thermal infrared image, and the images to be processed generate time information, altitude information, and latitude and longitude information based on the time of capture during the capture process;
[0008] For any of the actual thermal infrared images, extracting a first feature of the image to be processed based on the actual thermal infrared image, extracting a second feature of the image to be processed based on the visible light image corresponding to the actual thermal infrared image, splicing the first feature and the second feature to obtain a first splicing feature, and after processing all of the images to be processed, obtaining a first splicing feature set of the images to be processed;
[0009] Extracting a third feature set of the image to be processed based on the time information, extracting a fourth feature set of the image to be processed based on the altitude information, extracting a fifth feature set of the image to be processed based on the latitude and longitude information, splicing the fourth feature set and the fifth feature set to obtain a second spliced feature set, obtaining an enhanced mapping feature set of the third feature set with respect to the second spliced feature set, and splicing the first spliced feature set, the second spliced feature set, and the enhanced mapping feature set to obtain a comprehensive feature set;
[0010] Configure the initial generation model;
[0011] The comprehensive feature set is input into the initial generation model and trained to form a target generation model, and a target panoramic thermal infrared image corresponding to the image to be processed is obtained through the target generation model.
[0012] Furthermore, extracting the first feature of the image to be processed based on the actual thermal infrared image includes:
[0013] Initialize a set of image texture filters with different orientations and sizes;
[0014] Sliding the image texture filter on the thermal infrared image according to the direction and the size;
[0015] Calculating the dot product of the image texture filter in each direction and scale and the image area covered by the image texture filter to obtain a single channel texture feature map;
[0016] Combining the texture feature maps to obtain a multi-channel feature map;
[0017] Inputting the multi-channel feature map into a convolutional neural network, sliding a convolution kernel on the input data after entering the convolution layer so that the convolution kernel covers different image areas of the actual thermal infrared image, calculating the dot product of the image area covered by the convolution kernel and the convolution kernel, taking the result as the output of the convolution layer, and downsampling the output of the convolution layer using an average pooling operation to obtain the first feature;
[0018] Furthermore, extracting the second feature of the image to be processed based on the visible light image includes:
[0019] scaling the visible light image to adjust the visible light image to a set size;
[0020] Normalizing the pixel values of the visible light image to a set range, and converting the visible light image from the RGB color space to the HSV color space;
[0021] Input the visible light image into a convolutional neural network, slide the convolution kernel on the input data after entering the convolution layer, and calculate the dot product between the image area covered by the convolution kernel and the convolution kernel;
[0022] The average pooling operation is used to downsample the output of the convolutional layer to obtain the second feature.
[0023] Furthermore, extracting a third feature set of the image to be processed based on the time information includes:
[0024] Numerically processing the time information to convert the time information into a continuous numerical range containing a plurality of numerical values, wherein the continuous numerical range includes a maximum numerical value and a minimum numerical value;
[0025] The minimum value of all the continuous values is subtracted from each value, and the result is divided by the range of the values to obtain a third feature set.
[0026] Furthermore, extracting a fourth feature set of the image to be processed based on the height information includes:
[0027] Finding the maximum and minimum altitude values in the altitude information;
[0028] A minimum value is subtracted from each height value in the height information, and then divided by a difference between the maximum height value and the minimum height value to obtain a fourth feature set.
[0029] Furthermore, extracting a fifth feature set of the image to be processed based on the latitude and longitude information includes:
[0030] The latitude and longitude information is converted to obtain spherical coordinates, and the spherical coordinates are processed so that each spherical coordinate value is mapped to a range between 0 and 1, thereby obtaining a fifth feature set.
[0031] Furthermore, the reinforcement mapping feature set is obtained by calculating the feature vector scores of the query vector, key vector and value vector through the self-attention mechanism of the Transformer model.
[0032] The query vector, the key vector and the value vector are obtained by multiplying the weight matrix by the second fusion feature set to obtain the query vector, and multiplying the weight matrix by the third feature set to obtain the key vector and the value vector, wherein the weight matrix is randomly initialized based on the calculation process of the Transformer model.
[0033] Furthermore, configuring the initial velocity generation model includes: configuring a generator to receive the comprehensive feature set and output an initial panoramic thermal infrared image corresponding to the comprehensive feature set using a deep convolutional neural network structure;
[0034] configuring a discriminator to determine the probability that the initial panoramic thermal infrared image is the real thermal infrared image using a convolutional neural network structure;
[0035] The generator and the discriminator are optimized using the Wasserstein loss function and the gradient penalty term.
[0036] Furthermore, inputting the comprehensive feature set into the initial generative model for training includes: dividing the data in the comprehensive feature set into multiple batches, using only one batch of data for training at a time, and using the other batches of data as validation sets to verify the performance of the generative model, and obtaining the validation results:
[0037] If the verification result is the first result, increasing the learning rate;
[0038] If the verification result is the second result, reducing the learning rate;
[0039] The first result is an improvement in the performance of the generative model, and the second result is a decrease in the performance of the generative model;
[0040] The generated panoramic thermal infrared image is deformed onto a real panoramic thermal infrared image through geometric transformation, a window of fixed size is selected in the real panoramic thermal infrared image to calculate the local structure similarity index, the window slides within the real panoramic thermal infrared image, gradually covering the real panoramic thermal infrared image, the local structure similarity index value within each window is calculated, all the local structure similarity index values are averaged to obtain a structure similarity index value, and the structure similarity index is used to evaluate the similarity between the generated panoramic thermal infrared image and the real panoramic thermal infrared image. If the similarity is greater than or equal to the preset standard similarity, the generated panoramic thermal infrared image is used as the target panoramic thermal infrared image.
[0041] Another aspect provides a data processing AI system based on big data, the system comprising:
[0042] an acquisition module, configured to acquire a plurality of images to be processed, taken at a collection point in a direction toward a spatial boundary of a space to be imaged, wherein the collection point is set within the space to be imaged, the images to be processed include an actual thermal infrared image and a visible light image corresponding to the thermal infrared image, and the images to be processed generate time information, altitude information, and latitude and longitude information based on the time of capture during the capture process;
[0043] For any of the actual thermal infrared images, a feature extraction module is configured to extract a first feature of the image to be processed based on the actual thermal infrared image, extract a second feature of the image to be processed based on the visible light image corresponding to the actual thermal infrared image, splice the first feature and the second feature to obtain a first spliced feature, and obtain a first spliced feature set of the image to be processed after processing all of the images to be processed;
[0044] a feature stitching module, extracting a third feature set of the image to be processed based on the time information, extracting a fourth feature set of the image to be processed based on the altitude information, extracting a fifth feature set of the image to be processed based on the latitude and longitude information, stitching the fourth feature set and the fifth feature set to obtain a second stitching feature set, obtaining an enhanced mapping feature set of the third feature set with respect to the second stitching feature set, and stitching the first stitching feature set, the second stitching feature set, and the enhanced mapping feature set to obtain a comprehensive feature set;
[0045] Model configuration module, used to configure the initial generation model;
[0046] The target generation module is used to input the comprehensive feature set into the initial generation model and train it to form a target generation model, and generate a target panoramic thermal infrared image corresponding to the image to be processed through the target generation model.
[0047] Compared with existing technologies, the present invention achieves the beneficial effect of transforming the image to be processed into the target panoramic thermal infrared image by using a multimodal deep learning method and a comprehensive feature set to train the initial generation model. This not only improves the accuracy of the image conversion, but also increases the efficiency of generating panoramic thermal infrared images.
[0048] In particular, the present invention adopts an adaptive learning rate adjustment strategy to dynamically adjust the learning rate according to changes in model performance during the training process, which greatly accelerates the convergence speed of the model and shortens the time of model training.
[0049] In particular, by using the structural similarity index to evaluate and verify image quality, it can be ensured that the generated panoramic thermal infrared images have high authenticity and accuracy, which is extremely important for the subsequent use and research of thermal infrared images.
[0050] In particular, the present invention can generate panoramic thermal infrared images from various types and sources of processed images. This improves the applicability of the system and enables more thermal infrared images to be utilized, which has a wide range of application values for thermal infrared images.
[0051] In summary, the present invention can effectively generate high-quality panoramic thermal infrared images, improve the efficiency and quality of image conversion, and has good applicability and wide application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 A flowchart of a data processing method based on big data provided by an embodiment of the present invention;
[0053] Figure 2 This is a structural diagram of the big data-based data processing AI system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0054] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0055] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0056] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside", and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.
[0057] Furthermore, it should be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0058] See also Figure 1 As shown, the data processing method based on big data provided by the embodiment of the present invention includes:
[0059] Step S100: Acquire a plurality of images to be processed, which are taken at a collection point in a direction toward the spatial boundary of the space to be imaged. The collection point is set within the space to be imaged. The images to be processed include actual thermal infrared images and visible light images corresponding to the actual thermal infrared images. The images to be processed form time information, altitude information, and latitude and longitude information based on the shooting time during the shooting process.
[0060] Step S200: For any of the actual thermal infrared images, extract the first feature of the image to be processed based on the actual thermal infrared image, extract the second feature of the image to be processed based on the visible light image corresponding to the actual thermal infrared image, splice the first feature and the second feature to obtain a first splicing feature, and after processing all the images to be processed, obtain a first splicing feature set of the images to be processed.
[0061] Step S300: Extract the third feature set of the image to be processed based on the time information, extract the fourth feature set of the image to be processed based on the altitude information, extract the fifth feature set of the image to be processed based on the latitude and longitude information, splice the fourth feature set and the fifth feature set to obtain a second spliced feature set, obtain an enhanced mapping feature set of the third feature set for the second spliced feature set, and splice the first spliced feature set, the second spliced feature set and the enhanced mapping feature set to obtain a comprehensive feature set.
[0062] Step S400: configuring an initial generation model.
[0063] Step S500: inputting the comprehensive feature set into the initial generation model and training it to form a target generation model, and obtaining a target panoramic thermal infrared image corresponding to the image to be processed through the target generation model.
[0064] Specifically, in step S100, a portion of images captured by the drone swarm's onboard image capture devices during formation flight are collected as images to be processed. A location within the drone swarm's flight trajectory is selected as a collection point, which defines the space to be imaged. All images captured by the drone swarm as it passes through the collection point are collected as images to be processed. These images to be processed include thermal infrared images and their corresponding visible light images. The device also records additional information during the capture process, such as the time, altitude, and latitude and longitude of the capture moment. For example, at a designated collection point (e.g., a point with spatial coordinates (x, y, z)), the drone swarm captures 48 thermal infrared images in all directions. These images are then captured along with 48 corresponding visible light images. Each image is recorded with the corresponding time (13:00:00), altitude (500 meters), and latitude and longitude (45.05°N, 125.98°E).
[0065] Specifically, after completing step S100, a set of images with rich metadata is obtained, including thermal infrared and visible light images, as well as their time, altitude, and latitude and longitude information. This information is used for subsequent feature extraction and data processing to generate a panoramic thermal infrared image. This step converts physical information in real space into a form suitable for digital processing, laying the foundation for subsequent feature extraction, feature splicing, and model training.
[0066] Specifically, in step S200, feature extraction and stitching are performed on the thermal infrared image and the corresponding visible light image. First, a first feature is extracted from each thermal infrared image using an appropriate feature extraction method. Then, a second feature is extracted from the visible light image corresponding to the thermal infrared image. The first and second features are then stitched together to obtain a first stitched feature set. This process is repeated for all images to be processed to obtain a complete set of first stitched features.
[0067] Specifically, after completing step S200, a feature set consisting of the first stitching features of each image to be processed will be obtained. These feature sets contain mixed information of thermal infrared images and visible light images, and provide a rich source of information for the subsequent panoramic thermal infrared image generation process.
[0068] Specifically, extracting the first feature of the image to be processed based on the actual thermal infrared image includes:
[0069] Initialize a set of image texture filters with different orientations and sizes;
[0070] Sliding the image texture filter on the thermal infrared image according to the direction and the size;
[0071] Calculating the dot product of the image texture filter in each direction and scale and the image area covered by the image texture filter to obtain a single channel texture feature map;
[0072] Combining the texture feature maps to obtain a multi-channel feature map;
[0073] Inputting the multi-channel feature map into a convolutional neural network, sliding a convolution kernel on the input data after entering the convolution layer so that the convolution kernel covers different image areas of the actual thermal infrared image, calculating the dot product of the image area covered by the convolution kernel and the convolution kernel, taking the result as the output of the convolution layer, and downsampling the output of the convolution layer using an average pooling operation to obtain the first feature;
[0074] Specifically, the first step is to initialize and apply image texture filters. A set of image texture filters with different orientations and sizes is initialized. These filters are used to extract texture information from thermal infrared imagery. For example, eight filters with different orientations and three different sizes are initialized, for a total of 24 filters. These filters are then slid across the thermal infrared image, and the dot product between each filter and the image area it covers is calculated to generate texture feature maps. These texture feature maps describe the texture characteristics of the image at various orientations and sizes.
[0075] Specifically, a convolutional neural network is then used. The multi-channel texture feature map generated above is input into the convolutional neural network. The network slides a small convolution kernel and calculates the dot product between each region and the kernel. Next, average pooling is used to downsample the output of the convolution layer to obtain the first feature. This feature contains important texture information from the actual thermal infrared image, which serves as the basis for subsequent feature fusion and classification.
[0076] Specifically, extracting the second feature of the image to be processed based on the visible light image includes:
[0077] scaling the visible light image to adjust the visible light image to a set size;
[0078] Normalizing the pixel values of the visible light image to a set range, and converting the visible light image from the RGB color space to the HSV color space;
[0079] Input the visible light image into a convolutional neural network, slide the convolution kernel on the input data after entering the convolution layer, and calculate the dot product between the image area covered by the convolution kernel and the convolution kernel;
[0080] The average pooling operation is used to downsample the output of the convolutional layer to obtain the second feature.
[0081] Specifically, the first step is scaling and normalization. Before the visible light image is input into the convolutional neural network, the size and pixel value need to be adjusted. For example, the image size is scaled to 256x256, so that no matter what the size of the original image is, the image input to the network is of uniform size, which is conducive to subsequent calculations and processing. At the same time, the pixel value is normalized to between 0 and 1, generally by dividing the original pixel value by 255 (the maximum pixel value of an 8-bit image). This can avoid problems such as overflow or gradient disappearance during numerical calculations. The image in the RGB color space is then converted to the HSV color space to capture color information rather than brightness information.
[0082] Specifically, the convolutional neural network is then used. A convolutional neural network is a special type of neural network that can effectively extract local features of an image. In the convolution layer of the network, a small-sized (for example, 3x3 or 5x5) convolution kernel slides over the input image to calculate the dot product between the covered area and the convolution kernel. In this way, the convolution kernel can learn some important features in the image, such as edges and textures. Then, the output of the convolution layer is downsampled using an average pooling operation, which further reduces the size of the features and improves the computational efficiency of the model. The second feature obtained in this step contains the key information of the visible light image and provides important input for subsequent feature fusion and classification prediction.
[0083] Specifically, in step S300, a third feature set is extracted based on time information; a fourth feature set and a fifth feature set are extracted based on altitude information and latitude and longitude information, respectively. The fourth and fifth feature sets are then concatenated to obtain a second concatenated feature set, which is then subjected to enhancement mapping to obtain an enhanced mapping feature set. Finally, the first concatenated feature set, the second concatenated feature set, and the enhanced mapping feature set are concatenated to obtain a comprehensive feature set.
[0084] Specifically, step S300 not only extracts and combines multiple feature sets, but also implements a self-attention mechanism for these features by strengthening the mapped feature set. This self-attention mechanism helps strengthen salient features and weaken irrelevant or minor features, allowing the generation of a comprehensive feature set to focus more on features that are important to the target task. This significantly enhances the expressiveness and resolution of panoramic thermal infrared images, thereby improving overall image quality and accuracy.
[0085] Specifically, extracting the third feature set of the image to be processed based on the time information includes:
[0086] Numerically processing the time information to convert the time information into a continuous numerical range containing a plurality of numerical values, wherein the continuous numerical range includes a maximum numerical value and a minimum numerical value;
[0087] The minimum value of all the continuous values is subtracted from each value, and the result is divided by the range of the values to obtain a third feature set.
[0088] Specifically, time information is converted to numerical values using a mapping rule. For example, in timestamp conversion, time information is converted into a continuous range of values. For example, a time series such as "8:00," "8:01," and "8:02" is converted to 480, 481, and 482, representing the minutes of the day. Then, within this range, the maximum and minimum values are found to calculate the range.
[0089] Specifically, the next step is normalization to generate the third feature set. Each value is subtracted from the minimum value in the range, and then divided by the range: (each value - minimum value) / range. For example, for the minutes above, the range is (482 - 480) = 2, so the normalized values are: 0, 0.5, 1. This is the third feature set, which reflects the relative size and order of the original time information while limiting all values to the range of 0 to 1, ensuring that data of different scales is treated equally.
[0090] The above steps are part of data preprocessing, which aims to convert the raw time information into a format that is easier for machine learning algorithms to process—that is, to digitize and normalize it. This allows us to extract useful features from the raw time information and ensure that these features can be effectively combined with other features in subsequent calculations.
[0091] Specifically, extracting the fourth feature set of the image to be processed based on the height information includes:
[0092] Finding the maximum and minimum altitude values in the altitude information;
[0093] A minimum value is subtracted from each height value in the height information, and then divided by a difference between the maximum height value and the minimum height value to obtain a fourth feature set.
[0094] Specifically, the first step is to find the maximum and minimum values in the altitude information. For example, a dataset representing terrain altitudes might contain values ranging from 200 meters above sea level to 206 meters above sea level. These altitude values are then normalized. Specifically, the minimum value is subtracted from each altitude value, and then divided by the difference between the maximum and minimum values. For example, if an altitude value is 203 meters, the normalized value is (203 - 200) / (206 - 200) = 0.5. This converts all altitude values into values between 0 and 1, forming the fourth feature set.
[0095] Specifically, the above processing results in a feature set corresponding to the original height information, but with values between 0 and 1. This feature set contains all the details of the original height information, but its numerical range has been standardized to accommodate subsequent processing steps. This step preprocesses the height information within the overall processing process, ensuring that this data, along with the other feature sets, can be effectively processed by the subsequent model. Therefore, the above processing not only extracts important information about height but also performs the necessary preprocessing, laying the foundation for subsequent steps.
[0096] Specifically, extracting the fifth feature set of the image to be processed based on the latitude and longitude information includes:
[0097] The latitude and longitude information is converted to obtain spherical coordinates, and the spherical coordinates are processed so that each spherical coordinate value is mapped to a range between 0 and 1, thereby obtaining a fifth feature set.
[0098] Specifically, when extracting the fifth feature set, for example, an image to be processed is taken at 30 degrees north latitude and 120 degrees east longitude. These longitude and latitude information are converted into spherical coordinates and then normalized. The spherical coordinate conversion can understand the geographic location information from a new perspective, and the normalization process maps all location information to a fixed range (such as between 0 and 1), intuitively understanding the geographic location of this location relative to other locations, and thus obtaining the fifth feature set.
[0099] Specifically, the above processing yields a new feature set derived from the latitude and longitude information, with all values between 0 and 1. This feature set helps capture spatial relationships and geographic location, which the original latitude and longitude information cannot directly provide. This step preprocesses the latitude and longitude information, providing essential preparation for subsequent steps.
[0100] Specifically, the comprehensive feature set is composed of the first and second stitching feature sets, and a mapping feature set enhanced by the third feature set. This comprehensive feature set fully integrates the image information and the spatial-temporal information at the time of capture, in order to obtain richer and more accurate information during the generation of panoramic thermal infrared images. In the entire operation process, this step is in the stage of information extraction and fusion. Immediately after the feature extraction and stitching are completed, the next step is to use this comprehensive feature set to generate panoramic thermal infrared images. Therefore, this step determines the quality and accuracy of the generation of panoramic thermal infrared images, and directly affects whether high-quality panoramic thermal infrared images can be successfully and accurately generated.
[0101] Specifically, the reinforcement mapping feature set is obtained by calculating the feature vector scores of the query vector, key vector, and value vector through the self-attention mechanism of the Transformer model.
[0102] The query vector, the key vector and the value vector are obtained by multiplying the weight matrix by the second fusion feature set to obtain the query vector, and multiplying the weight matrix by the third feature set to obtain the key vector and the value vector, wherein the weight matrix is randomly initialized based on the calculation process of the Transformer model.
[0103] Specifically, the Transformer model plays a crucial role in this step, thanks to its unique self-attention mechanism. The self-attention mechanism can capture the dependencies between features, regardless of their positional distance in the sequence, which is of great value for understanding heterogeneous information such as time, altitude, longitude and latitude. Specifically, in the present invention, the characteristics of the self-attention mechanism are utilized to generate an enhanced mapping feature set based on time information, longitude and latitude information, and altitude information. This feature set generation method utilizes the advantages of the Transformer model and can enhance the model's expressive power when processing multi-source data.
[0104] Specifically, by leveraging the advantages of the Transformer model, the present invention can comprehensively understand and analyze the image to be processed from multiple dimensions, enhance the generalization performance of the model, and thus improve the final processing effect.
[0105] Specifically, the query vector, key vector, and value vector are important components of the Transformer model, and the interaction and computation between them is the core of the self-attention mechanism. In this paper, the query vector is primarily used to match appropriate key-value pairs, while the key vector and value vector are used to generate the final output. This weight matrix-based multiplication process allows deep learning of the original features to generate new feature representations.
[0106] Specifically, after this step is completed, a series of deeply encoded features are obtained, capturing the deep information in the image data. This information is embedded in the enhanced mapping feature set. This enhanced mapping feature set not only enhances the expressive power of the original features but also makes them more suitable for subsequent machine learning tasks. This step mainly completes the deep encoding and enhancement of the original features. Leveraging the Transformer model and self-attention mechanism, it provides powerful feature extraction and encoding capabilities, thereby greatly improving the accuracy and reliability of the final processing results.
[0107] Specifically, in step S400, a pre-trained deep learning model, such as a Generative Adversarial Network (GAN), is used to configure the initial generative model. This model can learn the complex distribution of input data and generate new samples that are similar to the training data.
[0108] Specifically, after completing the configuration, an initial generative model is obtained, which serves as the basis for subsequent steps. This initial generative model can generate some basic images similar to the training data. In subsequent steps, through training and optimization of this model, the quality of the generated images is gradually improved and the degree of realism is gradually enhanced.
[0109] configuring a discriminator to determine the probability that the initial panoramic thermal infrared image is the real thermal infrared image using a convolutional neural network structure;
[0110] The generator and the discriminator are optimized using the Wasserstein loss function and the gradient penalty term.
[0111] Specifically, a key step in configuring the initial generative model is setting up the generator and discriminator. The generator is a deep convolutional neural network that receives a comprehensive feature set as input and generates the corresponding initial panoramic thermal infrared image. The comprehensive feature set formed by fusing the first through fifth feature sets mentioned above is used as input and passed through a deep convolutional neural network consisting of multiple convolutional and deconvolutional layers. The output is a thermal infrared image. The discriminator is a convolutional neural network whose task is to evaluate the similarity between the thermal infrared image generated by the generator and the real image, that is, to determine the probability that the generated image is the real image. For example, the discriminator may receive a real thermal infrared image and a generated thermal infrared image and then determine the similarity between the two.
[0112] Specifically, after completing the configuration, an initial generative model is obtained, which serves as the basis for subsequent steps. This initial generative model can generate some basic images similar to the training data. In subsequent steps, through training and optimization of this model, the quality of the generated images is gradually improved and the degree of realism is gradually enhanced.
[0113] Specifically, in step S500, the operation of inputting the comprehensive feature set into the initial generation model and training it is a deep learning task. During this process, the initial generation model gradually learns how to map the comprehensive feature set into a panoramic thermal infrared image.
[0114] Specifically, completing this step yields a target generation model that can generate corresponding panoramic thermal infrared images based on the input comprehensive feature set. This allows for the generation of corresponding panoramic thermal infrared images from pre-processed images at different angles and heights, further enriching observation methods and approaches. This step is crucial for panoramic thermal infrared image generation, as it determines the quality and fidelity of the resulting thermal infrared image.
[0115] Specifically, inputting the comprehensive feature set into the initial generative model for training includes: dividing the data in the comprehensive feature set into multiple batches, using only one batch of data for training at a time, and using the other batches of data as validation sets to verify the performance of the generative model, and obtaining the validation results:
[0116] If the verification result is the first result, increasing the learning rate;
[0117] If the verification result is the second result, reducing the learning rate;
[0118] The first result is an improvement in the performance of the generative model, and the second result is a decrease in the performance of the generative model;
[0119] The generated panoramic thermal infrared image is deformed onto a real panoramic thermal infrared image through geometric transformation, a window of fixed size is selected in the real panoramic thermal infrared image to calculate the local structure similarity index, the window slides within the real panoramic thermal infrared image, gradually covering the real panoramic thermal infrared image, the local structure similarity index value within each window is calculated, all the local structure similarity index values are averaged to obtain a structure similarity index value, and the structure similarity index is used to evaluate the similarity between the generated panoramic thermal infrared image and the real panoramic thermal infrared image. If the similarity is greater than or equal to the preset standard similarity, the generated panoramic thermal infrared image is used as the target panoramic thermal infrared image.
[0120] Specifically, when training the model, computing resources are efficiently utilized by dividing the data into multiple batches, and overfitting is avoided by using a portion of the data as a validation set. For example, if there are 5,000 samples for training, they are divided into 100 batches, each containing 50 samples. During each training session, only one batch of data is used, while the remaining batches are used to validate the model's performance. If the validation results show improved model performance, such as a decrease in error from 0.05 to 0.03, the learning rate is increased, for example, from 0.001 to 0.01, to accelerate model learning. If the validation results show decreased model performance, such as an increase in error from 0.03 to 0.05, the learning rate is decreased, for example, from 0.01 to 0.001, to prevent overfitting. Furthermore, a fixed-size window is slid across the real panoramic thermal infrared imagery, and a local structure similarity index is calculated within each window to assess the similarity between the generated panoramic thermal infrared imagery and the real imagery. For example, use a 10x10 window to divide the real image into several 10x10 regions, then calculate the similarity between these regions and the corresponding regions of the generated image one by one, and finally average all the similarity values to get the overall similarity value.
[0121] Specifically, after completing the above training, a target generation model is obtained, and the similarity between the generated panoramic thermal infrared image and the real image is determined. The significance of this result is that the target generation model can be used to convert the processed image into a panoramic thermal infrared image, and the similarity between the panoramic thermal infrared image and the real image has reached a preset standard, which ensures the authenticity and accuracy of the generated image. The above training method directly determines whether high-quality panoramic thermal infrared images can be obtained, and this method can ensure the authenticity and accuracy of the generated images, which is of great significance for the use and research of thermal infrared imagery.
[0122] See also Figure 2 As shown, another embodiment of the present invention provides a big data-based data processing AI system including:
[0123] An acquisition module 1000 is configured to acquire a plurality of images to be processed, taken at a collection point in a direction toward a spatial boundary of a space to be imaged. The collection point is set within the space to be imaged. The images to be processed include actual thermal infrared images and visible light images corresponding to the thermal infrared images. The images to be processed generate time information, altitude information, and latitude and longitude information based on the time of capture during the capture process.
[0124] The feature extraction module 2000 is configured to extract, for any actual thermal infrared image, a first feature of the image to be processed based on the actual thermal infrared image, extract a second feature of the image to be processed based on the visible light image corresponding to the actual thermal infrared image, concatenate the first and second features to obtain a first concatenated feature, and obtain a first concatenated feature set of the image to be processed after processing all the images to be processed;
[0125] The feature stitching module 3000 extracts a third feature set of the image to be processed based on the time information, extracts a fourth feature set of the image to be processed based on the altitude information, extracts a fifth feature set of the image to be processed based on the latitude and longitude information, stitches the fourth feature set and the fifth feature set to obtain a second stitched feature set, obtains an enhanced mapping feature set of the third feature set with respect to the second stitched feature set, and stitches the first stitched feature set, the second stitched feature set, and the enhanced mapping feature set to obtain a comprehensive feature set.
[0126] Model configuration module 4000, used to configure the initial generation model;
[0127] The target generation module 5000 is used to input the comprehensive feature set into the initial generation model and train it to form a target generation model, and generate a target panoramic thermal infrared image corresponding to the image to be processed through the target generation model.
[0128] Specifically, extracting the first feature of the image to be processed based on the actual thermal infrared image includes:
[0129] Initialize a set of image texture filters with different orientations and sizes;
[0130] Sliding the image texture filter on the thermal infrared image according to the direction and the size;
[0131] Calculating the dot product of the image texture filter in each direction and scale and the image area covered by the image texture filter to obtain a single channel texture feature map;
[0132] Combining the texture feature maps to obtain a multi-channel feature map;
[0133] The multi-channel feature map is input into a convolutional neural network. After entering the convolution layer, the convolution kernel is slid on the input data so that the convolution kernel covers different image areas of the actual thermal infrared image. The dot product of the image area covered by the convolution kernel and the convolution kernel is calculated, and the result is used as the output of the convolution layer. The output of the convolution layer is downsampled using an average pooling operation to obtain the first feature.
[0134] Specifically, extracting the second feature of the image to be processed based on the visible light image includes:
[0135] scaling the visible light image to adjust the visible light image to a set size;
[0136] Normalizing the pixel values of the visible light image to a set range, and converting the visible light image from the RGB color space to the HSV color space;
[0137] Input the visible light image into a convolutional neural network, slide the convolution kernel on the input data after entering the convolution layer, and calculate the dot product between the image area covered by the convolution kernel and the convolution kernel;
[0138] The average pooling operation is used to downsample the output of the convolutional layer to obtain the second feature.
[0139] Specifically, extracting the third feature set of the image to be processed based on the time information includes:
[0140] Numerically processing the time information to convert the time information into a continuous numerical range containing a plurality of numerical values, wherein the continuous numerical range includes a maximum numerical value and a minimum numerical value;
[0141] The minimum value of all the continuous values is subtracted from each value, and the result is divided by the range of the values to obtain a third feature set.
[0142] Specifically, extracting the fourth feature set of the image to be processed based on the height information includes:
[0143] Finding the maximum and minimum altitude values in the altitude information;
[0144] A minimum value is subtracted from each height value in the height information, and then divided by a difference between the maximum height value and the minimum height value to obtain a fourth feature set.
[0145] Specifically, extracting the fifth feature set of the image to be processed based on the latitude and longitude information includes:
[0146] The latitude and longitude information is converted to obtain spherical coordinates, and the spherical coordinates are processed so that each spherical coordinate value is mapped to a range between 0 and 1, thereby obtaining a fifth feature set.
[0147] Specifically, the reinforcement mapping feature set is obtained by calculating the feature vector scores of the query vector, key vector, and value vector through the self-attention mechanism of the Transformer model.
[0148] The query vector, the key vector and the value vector are obtained by multiplying the weight matrix by the second fusion feature set to obtain the query vector, and multiplying the weight matrix by the third feature set to obtain the key vector and the value vector, wherein the weight matrix is randomly initialized based on the calculation process of the Transformer model.
[0149] Specifically, configuring the initial velocity generation model includes: configuring a generator to receive the comprehensive feature set and output an initial panoramic thermal infrared image corresponding to the comprehensive feature set using a deep convolutional neural network structure;
[0150] configuring a discriminator to determine the probability that the initial panoramic thermal infrared image is the real thermal infrared image using a convolutional neural network structure;
[0151] The generator and the discriminator are optimized using the Wasserstein loss function and the gradient penalty term.
[0152] Specifically, inputting the comprehensive feature set into the initial generative model for training includes: dividing the data in the comprehensive feature set into multiple batches, using only one batch of data for training at a time, and using the other batches of data as validation sets to verify the performance of the generative model, and obtaining the validation results:
[0153] If the verification result is the first result, increasing the learning rate;
[0154] If the verification result is the second result, reducing the learning rate;
[0155] The first result is an improvement in the performance of the generative model, and the second result is a decrease in the performance of the generative model;
[0156] The generated panoramic thermal infrared image is deformed onto a real panoramic thermal infrared image through geometric transformation, a window of fixed size is selected in the real panoramic thermal infrared image to calculate the local structure similarity index, the window slides within the real panoramic thermal infrared image, gradually covering the real panoramic thermal infrared image, the local structure similarity index value within each window is calculated, all the local structure similarity index values are averaged to obtain a structure similarity index value, and the structure similarity index is used to evaluate the similarity between the generated panoramic thermal infrared image and the real panoramic thermal infrared image. If the similarity is greater than or equal to the preset standard similarity, the generated panoramic thermal infrared image is used as the target panoramic thermal infrared image.
[0157] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.
[0158] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A data processing method based on big data, characterized in that: include: Acquire a plurality of images to be processed, taken at a collection point in a direction toward a spatial boundary of a space to be imaged, wherein the collection point is set within the space to be imaged, the images to be processed include an actual thermal infrared image and a visible light image corresponding to the actual thermal infrared image, and the images to be processed generate time information, altitude information, and latitude and longitude information based on the time of capture during the capture process; For any of the actual thermal infrared images, extracting a first feature of the image to be processed based on the actual thermal infrared image, extracting a second feature of the image to be processed based on the visible light image corresponding to the actual thermal infrared image, splicing the first feature and the second feature to obtain a first splicing feature, and after processing all of the images to be processed, obtaining a first splicing feature set of the images to be processed; Extracting a third feature set of the image to be processed based on the time information, extracting a fourth feature set of the image to be processed based on the altitude information, extracting a fifth feature set of the image to be processed based on the latitude and longitude information, splicing the fourth feature set and the fifth feature set to obtain a second spliced feature set, obtaining an enhanced mapping feature set of the third feature set with respect to the second spliced feature set, and splicing the first spliced feature set, the second spliced feature set, and the enhanced mapping feature set to obtain a comprehensive feature set; Configure the initial generation model; The comprehensive feature set is input into the initial generation model and trained to form a target generation model, and a target panoramic thermal infrared image corresponding to the image to be processed is obtained through the target generation model.
2. The data processing method based on big data according to claim 1, characterized in that: Extracting the first feature of the image to be processed based on the actual thermal infrared image includes: Initialize a set of image texture filters with different orientations and sizes; Sliding the image texture filter on the thermal infrared image according to the direction and the size; Calculating the dot product of the image texture filter in each direction and scale and the image area covered by the image texture filter to obtain a single channel texture feature map; Combining the texture feature maps to obtain a multi-channel feature map; The multi-channel feature map is input into a convolutional neural network. After entering the convolution layer, the convolution kernel is slid on the input data so that the convolution kernel covers different image areas of the actual thermal infrared image. The dot product of the image area covered by the convolution kernel and the convolution kernel is calculated, and the result is used as the output of the convolution layer. The output of the convolution layer is downsampled using an average pooling operation to obtain the first feature.
3. The data processing method based on big data according to claim 2, characterized in that: Extracting the second feature of the image to be processed based on the visible light image includes: scaling the visible light image to adjust the visible light image to a set size; Normalizing the pixel values of the visible light image to a set range, and converting the visible light image from the RGB color space to the HSV color space; Input the visible light image into a convolutional neural network, slide the convolution kernel on the input data after entering the convolution layer, and calculate the dot product between the image area covered by the convolution kernel and the convolution kernel; The average pooling operation is used to downsample the output of the convolutional layer to obtain the second feature.
4. The data processing method based on big data according to claim 3, characterized in that: Extracting a third feature set of the image to be processed based on the time information includes: Numerically processing the time information to convert the time information into a continuous numerical range containing a plurality of numerical values, wherein the continuous numerical range includes a maximum numerical value and a minimum numerical value; The minimum value of all the continuous values is subtracted from each value, and the result is divided by the range of the values to obtain a third feature set.
5. The data processing method based on big data according to claim 4, characterized in that: Extracting a fourth feature set of the image to be processed based on the height information includes: Finding the maximum and minimum altitude values in the altitude information; A minimum value is subtracted from each height value in the height information, and then divided by a difference between the maximum height value and the minimum height value to obtain a fourth feature set.
6. The data processing method based on big data according to claim 5, characterized in that: Extracting a fifth feature set of the image to be processed based on the latitude and longitude information includes: The latitude and longitude information is converted to obtain spherical coordinates, and the spherical coordinates are processed so that each spherical coordinate value is mapped to a range between 0 and 1, thereby obtaining a fifth feature set.
7. The data processing method based on big data according to claim 6, characterized in that: The enhanced mapping feature set is obtained by calculating the feature vector scores of the query vector, key vector, and value vector through the self-attention mechanism of the Transformer model; The query vector, the key vector and the value vector are obtained by multiplying the weight matrix by the second fusion feature set to obtain the query vector, and multiplying the weight matrix by the third feature set to obtain the key vector and the value vector, wherein the weight matrix is randomly initialized based on the calculation process of the Transformer model.
8. The data processing method based on big data according to claim 7, characterized in that: Configuring the initial velocity generation model includes: configuring a generator to receive the comprehensive feature set and output an initial panoramic thermal infrared image corresponding to the comprehensive feature set using a deep convolutional neural network structure; configuring a discriminator to determine the probability that the initial panoramic thermal infrared image is a real thermal infrared image using a convolutional neural network structure; The generator and the discriminator are optimized using the Wasserstein loss function and the gradient penalty term.
9. The data processing method based on big data according to claim 8, characterized in that: Inputting the comprehensive feature set into the initial generative model for training includes: dividing the data in the comprehensive feature set into multiple batches, using only one batch of data for training at a time, and using the other batches of data as validation sets to verify the performance of the generative model, and obtaining the validation results: If the verification result is the first result, increasing the learning rate; If the verification result is the second result, reducing the learning rate; The first result is an improvement in the performance of the generative model, and the second result is a decrease in the performance of the generative model; The initial panoramic thermal infrared image is deformed onto a real panoramic thermal infrared image through geometric transformation, a window of fixed size is selected in the real panoramic thermal infrared image to calculate the local structural similarity index, the window slides within the real panoramic thermal infrared image, gradually covering the real panoramic thermal infrared image, the local structural similarity index value within each window is calculated, all the local structural similarity index values are averaged to obtain a structural similarity index value, and the structural similarity index is used to evaluate the similarity between the initial panoramic thermal infrared image and the real panoramic thermal infrared image. If the similarity is greater than or equal to a preset standard similarity, the generated panoramic thermal infrared image is used as the target panoramic thermal infrared image.
10. An AI system using the data processing method based on big data according to any one of claims 1 to 9, characterized in that: include: an acquisition module, configured to acquire a plurality of images to be processed, taken at a collection point in a direction toward a spatial boundary of a space to be imaged, wherein the collection point is set within the space to be imaged, the images to be processed include an actual thermal infrared image and a visible light image corresponding to the thermal infrared image, and the images to be processed generate time information, altitude information, and latitude and longitude information based on the time of capture during the capture process; For any of the actual thermal infrared images, a feature extraction module is configured to extract a first feature of the image to be processed based on the actual thermal infrared image, extract a second feature of the image to be processed based on the visible light image corresponding to the actual thermal infrared image, splice the first feature and the second feature to obtain a first spliced feature, and obtain a first spliced feature set of the image to be processed after processing all of the images to be processed; a feature stitching module, extracting a third feature set of the image to be processed based on the time information, extracting a fourth feature set of the image to be processed based on the altitude information, extracting a fifth feature set of the image to be processed based on the latitude and longitude information, stitching the fourth feature set and the fifth feature set to obtain a second stitching feature set, obtaining an enhanced mapping feature set of the third feature set with respect to the second stitching feature set, and stitching the first stitching feature set, the second stitching feature set, and the enhanced mapping feature set to obtain a comprehensive feature set; Model configuration module, used to configure the initial generation model; The target generation module is used to input the comprehensive feature set into the initial generation model and train it to form a target generation model, and generate a target panoramic thermal infrared image corresponding to the image to be processed through the target generation model.
Citation Information
Patent Citations
Unmanned aerial vehicle multispectral image registration method
CN113689479A
Intelligent motion recognition method based on Internet of Things technology
CN116907510A
Cross-modal target tracking method and system based on fully convolutional twin network
CN119540283A