Gravity wave spectral image inversion temperature prediction method based on large language model fine tuning

By fine-tuning the large language model and combining gravity wave spectral images and meteorological environment text data, the problem of unstable observation results in traditional methods was solved, and more accurate and stable gravity wave temperature inversion was achieved.

CN120598876APending Publication Date: 2025-09-05NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510674009.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Traditional gravity wave spectral image inversion methods are affected by instrument stability, environmental complexity and atmospheric variability, resulting in large fluctuations in observation results. Existing methods fail to effectively consider the systematic deviations of observation results caused by environmental changes and equipment status.

Method used

A method based on fine-tuning of a large language model is adopted, combined with gravity wave spectral images and meteorological environment text data. Through feature-level fusion strategy and bias adjustment strategy, a multimodal dataset is constructed to perform gravity wave temperature inversion.

Benefits of technology

The accuracy and stability of temperature inversion are improved, noise interference is reduced, the model's command following ability and adaptive performance are enhanced, and more accurate temperature detection is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120598876A_ABST
    Figure CN120598876A_ABST
Patent Text Reader

Abstract

The invention particularly relates to a gravity wave spectral image inversion temperature prediction method based on large language model fine tuning. The method comprises the following steps: collecting gravity wave spectral image data; preprocessing the collected gravity wave spectral image; dimension compression is carried out on the preprocessed gravity wave spectral image; constructing a gravity wave spectral image inversion temperature detection model based on an LLaMA model; inputting the gravity wave spectral image after dimension compression and meteorological environment text data into a gravity wave spectral image inversion temperature detection model based on an LLaMA model for training; and inputting test data into the trained gravity wave spectral image inversion temperature detection model based on the LLaMA model to generate gravity wave spectral image inversion temperature. According to the method, the problems of unconspicuousness of image features and severity of noise interference are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of spectral image anomaly detection, and in particular to a gravity wave spectral image inversion temperature prediction method based on large language model fine-tuning. Background Art

[0002] The temperature of the near-space top region is a crucial parameter for studying dynamic processes at various scales, energy coupling, assessing long-term Earth climate change, and providing environmental support for space activities. Therefore, capturing and observing gravity wave spectral images is a crucial means of obtaining these temperatures. Meteorology typically relies on passive optical remote sensing techniques, using nighttime airglow radiation (such as O₂ and OH) as a light source. Optical imaging or photometry techniques are used to detect the intensity and ratio of its vibrational and rotational energy level spectral lines to infer rotational temperature. However, due to factors such as instrument stability, the complexity of the observing station's environmental conditions, and the variability of the atmospheric environment during acquisition, gravity wave spectral images obtained using traditional methods can exhibit significant fluctuations in observed values. Furthermore, previous studies have mostly employed comparative observations and mathematical fitting corrections to invert temperature, failing to comprehensively account for systematic biases in the observed results due to factors such as ambient meteorological variations and changes in instrument status.

[0003] In response to these issues, Oberheide et al. published a study in 2006 that primarily examined the differences between 15μm CO2 edge radiation and OH*(3,1) rotation temperatures measured by the atmospheric infrared radiometer SABER and the global infrared spectrometer GRIPS. However, differences in the time and spatial location of the SABER and GRIPS measurements resulted in significant errors in the final observational results. Liu et al. primarily compared rotation temperatures obtained from ground-based OH airglow observations and TIMED / SABER satellite data. They found that in some cases, differences between ground-based observations and SABER data were likely related to changes in atmospheric conditions, instrument errors, and model assumptions. They also determined a set of Einstein coefficients to correct the observational results. However, evaluating the Einstein coefficients may require reliance on model assumptions, such as atmospheric homogeneity and stability. These assumptions may not fully hold in the actual atmosphere, thus affecting the reliability of the results.

[0004] In the field of artificial intelligence, large language models (LLMs) have attracted considerable attention for their exceptional capabilities in understanding, reasoning, and generation. During pre-training, large language models acquire a wealth of general linguistic knowledge and patterns. This knowledge serves as a valuable starting point, helping the model better understand and handle domain-specific tasks. In certain domains, high-quality annotated data can be scarce and expensive. Fine-tuning can leverage the knowledge learned by pre-trained models on large-scale data, reducing the need for extensive annotated data and improving task performance, especially in data-limited environments. Lee et al. used BERT and its variants to classify medical documents and medical records and build a medical knowledge question-answering system. Al-Qurishi et al. fine-tuned the BERT model specifically for Arabic legal texts, aiming to improve the accuracy and efficiency of tasks such as legal document analysis and contract review. Araci introduced a BERT model fine-tuned specifically for the financial sector, improving the accuracy of sentiment analysis in financial news and reports to better support investment decisions and market trend forecasting. These examples demonstrate the feasibility of fine-tuning pre-trained large language models and applying them to specific domains.

[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present invention, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention

[0006] To address the problem that gravity wave temperature inversion is greatly affected by image quality and has low stability, in order to verify the correctness of the inversion results and improve the efficiency of temperature inversion, this paper proposes a gravity wave spectral image inversion temperature prediction method based on large language model fine-tuning. Combining external meteorological environment text data with gravity wave meteorological image data, the large language model is used to perform gravity wave temperature inversion prediction.

[0007] Other features and advantages of the present invention will become apparent from the following detailed description, or may be learned in part by practice of the present invention.

[0008] According to a first aspect of the present invention, a method for temperature prediction based on gravity wave spectral image inversion based on fine-tuning of a large language model is provided, the method comprising:

[0009] Collect gravitational wave spectral image data;

[0010] Preprocessing the collected gravity wave spectrum images;

[0011] Perform dimension compression on the pre-processed gravity wave spectrum image;

[0012] Construct a gravity wave spectral image inversion temperature detection model based on the LLaMA model;

[0013] The dimensionally compressed gravity wave spectral image and meteorological environmental text data were input into the LLaMA-based gravity wave spectral image inversion temperature detection model for training. A feature-level fusion strategy was used to align image features with meteorological environmental text data. A zero-initialized attention mechanism was used to freeze the main parameters of the LLaMA model. A lightweight adapter module and a learnable gating factor were introduced to dynamically adjust the weights of image and text information. A bias adjustment strategy was used to unfreeze the normalization layer, and bias and scaling factors were added to the linear layer to enhance the model's command-following capability.

[0014] The test data is input into the trained gravity wave spectrum image inversion temperature detection model based on the LLaMA model to generate the gravity wave spectrum image inversion temperature.

[0015] In some exemplary embodiments, the collecting of gravity wave spectral image data comprises:

[0016] The forward modeling image is acquired using a mesospheric airglow spectrophotometer, which includes the following modules:

[0017] Airglow spectrum radiation module, based on the HITRAN08 molecular spectrum database, provides O2(0-1) band vibration-rotation spectrum line emission intensity data;

[0018] Atmospheric radiation transfer module, using ARTS software to simulate airglow radiation transfer and intensity attenuation;

[0019] Optical system module, integrating transmittance, filter characteristics, optical distortion and modulation transfer function parameters;

[0020] CCD detector module simulates the conversion process of photon signals to electronic signals.

[0021] In some exemplary embodiments, the preprocessing of the collected gravity wave spectral image specifically includes:

[0022] Remove sensor noise through dark noise image;

[0023] Threshold recognition and median filtering are used to remove interference from cosmic rays and high-energy particles;

[0024] Evaluate and remove moonlight contamination from images based on moon phase and zenith angle.

[0025] In some exemplary embodiments, the dimensionality compression of the pre-processed gravity wave spectral image specifically includes:

[0026] The max-min normalization technique is used to map the 16-bit image to 8-bit.

[0027] In some exemplary embodiments, the feature-level fusion strategy includes:

[0028] Extract image features using the pre-trained CLIP visual encoder;

[0029] Align image features with meteorological context text embedding space through a learnable projection layer;

[0030] Image labels and adaptation cues are injected into different Transformer layers to avoid modality interference.

[0031] In some exemplary embodiments, the deviation adjustment strategy is implemented by:

[0032] Unfreeze all normalization layers in the LLaMA model;

[0033] Added bias and scale factors as learnable parameters to each linear layer in the Transformer.

[0034] In some exemplary embodiments, inputting test data into a trained LLaMA-based gravity wave spectral image inversion temperature detection model to generate gravity wave spectral image inversion temperature includes:

[0035] Based on the images of the test data and their corresponding text prompts, we calculate the image feature representation of each image and encode the text prompts into a token sequence that the model can process.

[0036] Based on the input information, the model gradually generates subsequent text within a predefined maximum generation length. During this period, the temperature parameter is used to control the sampling randomness, and the kernel sampling probability is combined to ensure the quality and diversity of the generated text.

[0037] At each decoding step, the model dynamically updates its predictions based on the previously generated context until a maximum length is reached or an end token is encountered;

[0038] The final output is a gravitational wave image inversion temperature text decoded from the generated tokens, representing a creative interpretation of the original image and prompts.

[0039] According to a second aspect of the present invention, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for temperature prediction by inversion of gravity wave spectral images based on fine-tuning of a large language model described in the first aspect is implemented.

[0040] According to a third aspect of the present invention, a computer program product is provided, on which a computer program is stored. When the computer program is executed by a processor, the temperature prediction method for gravity wave spectral image inversion based on large language model fine-tuning described in the first aspect is implemented.

[0041] According to a fourth aspect of the present invention, there is provided an electronic device, comprising:

[0042] processor; and

[0043] a memory for storing executable instructions of the processor;

[0044] Wherein, the processor is configured to implement the gravity wave spectral image inversion temperature prediction method based on large language model fine-tuning described in the first aspect above by executing the executable instructions.

[0045] The present invention provides a method for temperature prediction based on gravity wave spectral image inversion using large language model fine-tuning. This method uses a large language model as the main framework and, through a joint training paradigm of image text and command follow-up data, fine-tunes the pre-trained model to integrate gravity wave spectral images and meteorological environment data for temperature inversion detection. Specifically, this method offers the following advantages:

[0046] 1. To address the lack of distinct image features and the severity of noise interference, we developed an image preprocessing method. Through a series of filtering, enhancement, and feature extraction steps, this method significantly improves the visualization and recognition of key data in the image. This preprocessing not only provides more accurate and reliable input for subsequent processing steps, but also significantly reduces the interference of noise on anomaly detection results.

[0047] 2. To address the problem that single modal data cannot fully express the complex meteorological environment, a multimodal dataset containing external meteorological environment text data and gravity wave meteorological image data was constructed to provide richer training materials.

[0048] 3. To address the problem that existing models have difficulty in effectively fusing image and text information for accurate temperature inversion, a large language model is adopted as the main body, and the pre-trained model is fine-tuned through a joint training paradigm of image text and command-following data, thereby fusing gravity wave spectral images and meteorological environmental data for more accurate inversion temperature detection.

[0049] 4. To address the difficulties in aligning images with meteorological and environmental data and the potential interference with the model's ability to follow instructions, an innovative feature-level fusion strategy for image information is proposed to integrate image prompts into text prompts to ensure correct alignment between images and meteorological and environmental data while maintaining the model's ability to follow instructions.

[0050] 5. To address the problem that model performance is limited by fixed parameter settings, more learnable parameters are unlocked, including adjusting the bias and scale of the normalization layer to enhance model performance and adaptability.

[0051] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The accompanying drawings are incorporated into and constitute a part of this specification, illustrate embodiments consistent with the present invention, and together with the description, serve to explain the principles of the present invention. Obviously, the drawings described below are only some embodiments of the present invention, and it is clear that those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0053] Figure 1 A diagram of the temperature detection structure of gravity wave spectral image inversion based on large language model fine-tuning in an exemplary embodiment of the present invention;

[0054] Figure 2 Forward model structure diagram and simulation example of the intermediate layer airglow spectrophotometer in an exemplary embodiment of the present invention

[0055] Figure 3 A partial image display of a gravitational wave spectral image dataset in an exemplary embodiment of the present invention;

[0056] Figure 4 is a flow chart of a gravity wave spectral image preprocessing algorithm in an exemplary embodiment of the present invention;

[0057] Figure 5 FIG1 is a structural diagram of feature-level fusion of gravitational wave spectral images in an exemplary embodiment of the present invention;

[0058] Figure 6 This is a flow chart of temperature prediction using gravity wave spectral image inversion in an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0059] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0060] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the blocks shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0061] The execution process of the multimodal gravity wave spectral image inversion temperature algorithm proposed in this invention begins with receiving the input image and associated text data. This initial step is the basis for the algorithm to understand the input content. Then, the algorithm uses a pre-trained visual encoder CLIP model to extract key image features from the image. These features capture the important elements in the image and provide the basis for subsequent semantic understanding and instruction execution. The gravity wave spectral image inversion temperature prediction structure is as follows Figure 1 shown.

[0062] from Figure 1 As can be seen in the figure, the image features extracted by the proposed model are processed through a learnable projection layer. This step is crucial because it allows the image features to be aligned with the embedding space of the meteorological and environmental data, allowing the algorithm to build a bridge between the meteorological and environmental data and the image. The aligned image features are carefully inserted into the early Transformer layers of the large language model, rather than being fused only deep in the model. This feature-level fusion strategy helps the model better integrate image information, thereby reducing interference between image information and text features.

[0063] Specifically, the following steps may be included:

[0064] Step 1: Gravitational wave spectral image data acquisition

[0065] The top region of the mesosphere, as an important energy coupling zone in the Earth's atmosphere, hosts a variety of active dynamic processes, including the phenomenon of gravity waves. Temperature, as an effective tracer of these dynamic activities, plays a key role in revealing the complex physical processes occurring in this region. In view of this, in order to further explore this important scientific field, a new type of efficient gravity wave spectrometer has been designed and implemented for tasks such as gravity wave image capture and temperature monitoring. The gravity wave spectral image dataset used in the present invention is derived from the forward image generated by the mesosphere airglow spectrophotometer, an important optical device in the instrument. The forward model structure of the mesosphere airglow spectrophotometer and the gravity wave spectral image example are as follows: Figure 2 shown.

[0066] like Figure 2 As shown in Figure (a), the forward model of the airglow spectrophotometer consists of four key submodules: the airglow spectral radiation module, the atmospheric radiation transmission module, the optical system module, and the CCD (charge-coupled device) detector module. This forward model simulates how the CCD detector converts atmospheric optical radiation signals into electronic signals, providing theoretical support and practical guidance for instrument design, field observations, and the development of inversion algorithms.

[0067] Airglow Spectral Radiation Module: This module, based on the HITRAN08 molecular spectral database, provides detailed emission intensity data for the vibrational-rotational spectrum of the O2(0-1) band, emitted from multiple energy levels during electronic transitions of oxygen molecules (O2). The data covers the temperature range of 110 to 280 K, with temperature intervals as fine as 0.1 K, ensuring simulation accuracy.

[0068] Atmospheric Radiative Transfer Module: During the radiative transfer simulation process, this paper focuses on the atmospheric transmission of airglow radiation and its intensity attenuation. To this end, the Atmospheric Radiative Transfer Simulator (ARTS) software, jointly developed by the University of Hamburg and Chalmers University, was introduced. The ARTS software is designed to accurately simulate the airglow extinction process along the MASP observation path under clear sky conditions.

[0069] Optical System Module: This module comprehensively considers key factors such as the optical system's overall transmittance, filter characteristics, optical distortion, the flat field coefficient caused by varying light flux, and the modulation transfer function (MTF) within the field of view. These parameters are precisely measured through an in-lab calibration process, ensuring the reliability of the Optical System Module simulation.

[0070] CCD Detector Module: This module focuses on simulating the conversion process from photon signals to electron signals. In the forward model, key performance parameters of the CCD detector, such as responsivity, dark noise, readout noise, and shot noise, are carefully integrated to ensure that the model accurately reflects the actual detection process.

[0071] Based on the above-mentioned structure of the intermediate layer airglow spectrophotometer, the forward simulation diagram is as follows: Figure 2 (b) shows that, except Figure 2 (b), the images taken with different filters are quite different, so two of them are used as the baseline dataset for the three algorithms mentioned in this invention, and the training, testing, and reasoning of the model are all based on this dataset. Figure 3 As shown, the upper part is the large aperture image of gravitational waves, and the lower part is the small aperture image of gravitational waves.

[0072] Step 2: Gravitational wave spectral image preprocessing

[0073] The dataset used in this paper is data collected by the gravity wave instrument during its operation. The dataset was captured continuously for one month starting in September 2018 at the Nanjing University of Information Science and Technology Meteorological Comprehensive Base Platform. In October 2018, the mesosphere airglow spectrophotometer successfully obtained 13 sets of high-quality observation data, with observation time ranging from 7 to 10 hours. These data included 16 dark noise images and 3209 high-quality hyperspectral observation images, each of which was 256×256 in size.

[0074] like Figure 4 As shown in FIG, referring to the processing methods of gravity wave spectrum images in existing literature, the gravity wave spectrum image preprocessing method proposed in the present invention is mainly divided into the following three process modules:

[0075] (1) Dark noise elimination: Dark noise is a form of noise inherent in image sensors under no-light conditions and is an important part of image preprocessing that cannot be ignored. During the experiment, the shutter was closed periodically (e.g., every 30 minutes) and dark shots were taken under the same exposure conditions to obtain dark noise images. These dark noise images actually contain the dark noise and bias of the sensor itself. The superposition effect of the two makes the pixel values ​​fall within the range of 1000 to 3000. These dark noise images are recorded and used for accurate subtraction from the observed image in subsequent image processing, thereby effectively eliminating dark noise and improving the purity and signal-to-noise ratio of the image.

[0076] (2) Elimination of cosmic rays: The impact of cosmic rays and high-energy particles on the observed image is mainly manifested as local overexposed bright spots, which will seriously interfere with the accuracy of the data and subsequent analysis. In order to remove these abnormal bright spots, it is first necessary to subtract the previously obtained dark noise image from the observed image. After that, the threshold setting is used to identify and eliminate these high-brightness pixels. In order to ensure a smooth transition of the image after elimination, for these eliminated points, the pixel values ​​of the surrounding 3x3 pixel matrix are used for median filtering, and the entire image is subjected to an adaptive median filtering method to remove the salt and pepper noise in the image. This process is called adaptive smoothing, which effectively reduces the visual interference and data errors caused by cosmic rays.

[0077] (3) Elimination of moonlight images: In astronomical observations, moonlight not only illuminates the observation target, but also may bring significant background pollution, affecting the purity of the observation data. The degree of moonlight pollution depends on the phase of the moon (full moon, new moon, etc.) and the zenith angle of the moon relative to the observation point. Therefore, for the acquired observation images, it is necessary to evaluate the degree of moonlight influence based on the intensity change of the observed spectral line. If the influence of moonlight is significant, resulting in a significant decrease in the quality of the observation data, these affected images should be screened and eliminated to ensure that the data used in subsequent analysis has a high signal-to-noise ratio and accuracy.

[0078] Step 3: Compression of image dimensions

[0079] The bit depth of meteorological images is generally set to 16 bits. This high bit depth can capture and record richer grayscale information, thereby providing higher accuracy and detail in meteorological observations and analysis. However, when these high-resolution 16-bit images are used to train machine learning models, they inevitably bring a large computational burden and memory requirements. To reduce this burden without losing too much image information, a common processing method is to use max-min normalization technology to effectively map 16-bit images to 8 bits:

[0080]

[0081] In formula (1-1), X is the pixel value of each pixel in the image, X max and X min are the maximum and minimum values ​​of the image pixels, respectively, norm The pixel values ​​are normalized. This step not only reduces the image's bit depth, thereby reducing computational resources, but also maintains the relative integrity and recognizability of key information within the image. Max-min normalization is a linear transformation method that first converts pixel values ​​in an image to between 0 and 1, then further maps this range to the desired bit representation as needed.

[0082] Step 4: Fine-tuning the gravity wave spectral image inversion temperature detection model

[0083] This paper proposes a fine-tuning method based on the LLaMA model, aiming to detect inverted temperature by combining gravity wave spectral images and ambient meteorological temperature data. This method employs a parameter-efficient fine-tuning strategy, introducing a lightweight adapter module and a bias adjustment strategy, effectively improving the model's multimodal reasoning capabilities and command-following performance.

[0084] In order to avoid potential interference between image and meteorological environment data fine-tuning, this paper proposes a feature-level fusion strategy for gravity wave multimodal data. This strategy aims to prevent direct interaction between input image cues and adaptation cues, thereby ensuring the stability and effectiveness of the model in multimodal tasks. Specifically, the input image cues are first encoded through a frozen image encoder with a learnable image projection layer to ensure accurate extraction and representation of image information. The feature-level fusion structure of gravity wave spectral images is shown in Figure 2. Figure 5 shown.

[0085] from Figure 5As can be seen, in the feature-level fusion architecture for gravity wave spectral images, the encoded image tags and adaptation cues are injected separately into different Transformer layers rather than simply fused together. This separate injection method helps avoid conflicts between different modal information and allows the model to independently process image and meteorological and environmental data information at different levels. The shared adaptation cues are inserted into the final L layers to enhance the model's understanding at a high-level semantic level. Furthermore, to further optimize model performance, the input image cues are directly concatenated with the word tags of the first Transformer layer and processed using a zero-initialized attention mechanism. This zero-initialized attention mechanism effectively mitigates instabilities that may occur in the early stages of training, ensuring that the model can better integrate image and meteorological and environmental data information during fine-tuning.

[0086] Step 5: Zero-initialize the attention mechanism

[0087] This model uses a zero-initialized attention mechanism to freeze the entire LLaMA model and introduce only a lightweight adapter module with 1.2 million parameters. The adapter layer is applied to the upper Transformer layer of LLaMA and concatenates a set of learnable soft cues as word prefixes. The adaptation cues are adaptively controlled by learning a gating factor g initialized to zero. The contribution to word chunks enables the model to gradually enhance its command-following ability during training while maintaining its original language generation capabilities. Furthermore, the model has a simple multimodal variant that can combine image and video inputs for multimodal reasoning. For example, when processing images, a pre-trained image encoder is used to extract multi-scale image features, and a learnable projection layer is used to align the image semantics with the meteorological and environmental data embedding space, thereby generating responses based on text and image inputs.

[0088]

[0089] In formula (1-2), and Denote the attention scores of K adaptation hints and M+1 word tags respectively. The activation function tanh(·) is used to convert g l The scale of is adjusted to (-1,1). A separate softmax function ensures that the second term is independent of the adaptation cue and does not use any coefficient to multiply To prevent the pre-trained knowledge from being disturbed, that is, to maintain its original probability distribution. l When it is close to zero, the original pre-trained knowledge of the model can be conveyed to the token for credible generation. Finally, a linear projection layer is used to calculate the output of the lth attention layer.

[0090] This mechanism introduces a set of learnable adaptation cues and prepends them to word tokens in higher Transformer layers, enabling the model to gradually incorporate new meteorological and environmental information without disrupting pre-trained knowledge. This gradual knowledge integration approach helps the model maintain stable temperature inversion predictions even when faced with low-quality gravity wave meteorological images, by better leveraging supplementary information from textual data to correct for deficiencies in the image data. The adaptive nature of the zero-initialized attention mechanism then enables the model to dynamically adjust the influence of adaptation cues based on varying data characteristics and task requirements. When processing complex meteorological imagery and textual data, the model can flexibly balance the information from both, thereby improving the accuracy and reliability of the inversion results. For example, in some cases, textual data may provide more reliable information about meteorological conditions, in which case the model will increase the weight of textual data to ensure the correctness of the inversion results. Furthermore, this mechanism effectively stabilizes model performance in the early stages of training through a zero-gating mechanism, avoiding training fluctuations caused by unstable image quality. This stability is crucial to improving the robustness of the model in practical applications, especially when facing complex and changing meteorological conditions, ensuring that the model continues to provide high-quality temperature inversion results.

[0091] Step 6: Bias Adjustment Strategy

[0092] To address the issue of parameter updates being limited to adaptation cues and gating factors, this model proposes a bias adjustment strategy. Beyond adaptation cues and gating factors, this model further incorporates instruction cues into the large language model. Specifically, all normalization layers in the large language model are unfrozen, and bias and scaling factors are added as learnable parameters to each linear layer in the Transformer.

[0093] y=W·x→y=s·(W·x+b),

[0094] where b=Init(0),s=Init(1).(1-3)

[0095] By initializing the bias factor and scale factor to 0 and 1, respectively, the training process in the early stages is stabilized, enabling the model to achieve excellent instruction following capabilities. The number of newly added parameters only accounts for 0.04% of the entire large language model, maintaining parameter efficiency.

[0096] Step 7: Gravity wave spectrum image inversion temperature detection

[0097] The present invention combines images and text prompts, uses top-p sampling and temperature parameter methods to generate gravity wave spectrum image inversion temperature. Inversion temperature detection is to calculate the similarity between the inversion temperature generated by the model and the inversion temperature generated by MASP. The larger the value, the more similar the values ​​are, that is, the lower the probability of anomaly. After obtaining the similarity score, the gravity wave spectrum image is judged as abnormal by setting the anomaly threshold. The multimodal gravity wave image inversion temperature prediction process is described as follows: Figure 6 shown.

[0098] like Figure 6 The method shown aims to fuse image and text modalities, and realizes text generation based on images and prompts through a multimodal generative model. Specifically, after given a batch of images and their corresponding text prompts, the image feature representation of each image is first calculated, and the text prompt is encoded into a token sequence that can be processed by the model. Subsequently, using this input information, the model gradually generates subsequent text within a predefined maximum generation length. During this period, the temperature parameter is used to regulate the sampling randomness, and the kernel sampling probability is combined to ensure the quality and diversity of the generated text. In each decoding step, the model dynamically updates its prediction based on the previously generated context until the maximum length is reached or the end marker is encountered. The final output is a gravitational wave image inversion temperature text decoded from the generated tokens, representing a creative interpretation of the original image and prompt.

[0099] The main challenges faced by the present invention are the lack of obviousness of image features and the severity of noise interference. In order to overcome these difficulties, the present invention has developed an image preprocessing method, which significantly improves the visualization and recognition capabilities of key data in the image through a series of filtering, enhancement and feature extraction steps. This preprocessing not only provides more accurate and reliable input for subsequent processing steps, but also greatly reduces the interference of noise on abnormal detection results. On this basis, the present invention proposes a multimodal gravity wave spectral image inversion temperature detection method. This method takes a large language model as the main body, and fine-tunes the pre-trained model through a joint training paradigm of image text and instruction following data to fuse gravity wave spectral images and meteorological environment data for inversion temperature detection. In response to the problem that single modality data cannot fully express the complex meteorological environment, a multimodal dataset containing external meteorological environment text data and gravity wave meteorological image data is constructed to provide richer training materials. To address the difficulty of existing models in effectively fusing image and text information for accurate temperature inversion, a large language model was adopted as the main body. The pre-trained model was fine-tuned through a joint training paradigm of image, text, and command-following data, thereby fusing gravity wave spectral images and meteorological and environmental data for more accurate temperature inversion detection. To address the difficulties in aligning images with meteorological and environmental data and the potential interference with the model's command-following ability, an innovative feature-level fusion strategy for image information was proposed, integrating image cues into text cues to ensure correct alignment of images with meteorological and environmental data while maintaining the model's command-following ability. To address the problem that model performance is limited by fixed parameter settings, more learnable parameters were unlocked, including adjusting the bias and scale of the normalization layer to enhance model performance and adaptability.

[0100] Other embodiments of the present invention will readily occur to those skilled in the art after considering the specification and practicing the invention herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the invention being indicated by the claims.

[0101] It should be understood that the present invention is not limited to the exact construction described above and shown in the drawings and that various modifications and variations can be made without departing from the scope thereof, which is limited only by the appended claims.

Claims

1. A temperature prediction method based on gravity wave spectral image inversion and fine-tuning of a large language model, characterized by: The method comprises: Collect gravitational wave spectral image data; Preprocessing the collected gravity wave spectrum images; Perform dimension compression on the pre-processed gravity wave spectrum image; Construct a gravity wave spectral image inversion temperature detection model based on the LLaMA model; The dimensionally compressed gravity wave spectral image and meteorological environmental text data were input into the gravity wave spectral image inversion temperature detection model based on the LLaMA model for training. The image features were aligned with the meteorological environmental text data through a feature-level fusion strategy. A zero-initialized attention mechanism was used to freeze the main parameters of the LLaMA model. A lightweight adapter module and a learnable gating factor were introduced to dynamically adjust the weights of image and text information. A bias adjustment strategy was used to unfreeze the normalization layer, and bias and scaling factors were added to the linear layer to enhance the model's instruction-following capability. The test data is input into the trained gravity wave spectrum image inversion temperature detection model based on the LLaMA model to generate the gravity wave spectrum image inversion temperature.

2. The method according to claim 1, characterized in that The collecting of gravity wave spectral image data is specifically as follows: The forward modeling image is acquired using a mesospheric airglow spectrophotometer, which includes the following modules: Airglow spectrum radiation module, based on the HITRAN08 molecular spectrum database, provides O2(0-1) band vibration-rotation spectrum line emission intensity data; Atmospheric radiation transfer module, using ARTS software to simulate airglow radiation transfer and intensity attenuation; Optical system module, integrating transmittance, filter characteristics, optical distortion and modulation transfer function parameters; CCD detector module simulates the conversion process of photon signals to electronic signals.

3. The method according to claim 1, characterized in that The preprocessing of the collected gravity wave spectrum image specifically includes: Remove sensor noise through dark noise image; Threshold recognition and median filtering are used to remove interference from cosmic rays and high-energy particles; Evaluate and remove moonlight contamination from images based on moon phase and zenith angle.

4. The method according to claim 1, wherein The dimensionality compression of the pre-processed gravity wave spectrum image specifically includes: The max-min normalization technique is used to map the 16-bit image to 8-bit.

5. The method according to claim 1, wherein The feature-level fusion strategy includes: Extract image features using the pre-trained CLIP visual encoder; Align image features with meteorological context text embedding space through a learnable projection layer; Image labels and adaptation cues are injected into different Transformer layers to avoid modality interference.

6. The method according to claim 1, characterized in that The deviation adjustment strategy is implemented in the following way: Unfreeze all normalization layers in the LLaMA model; Added bias and scale factors as learnable parameters to each linear layer in the Transformer.

7. The method according to claim 1, characterized in that Input the test data into the trained LLaMA-based gravity wave spectrum image inversion temperature detection model to generate the gravity wave spectrum image inversion temperature, including: Based on the images of the test data and their corresponding text prompts, we calculate the image feature representation of each image and encode the text prompts into a token sequence that the model can process. Based on the input information, the model gradually generates subsequent text within a predefined maximum generation length. During this period, the temperature parameter is used to control the sampling randomness, and the kernel sampling probability is combined to ensure the quality and diversity of the generated text. At each decoding step, the model dynamically updates its predictions based on the previously generated context until a maximum length is reached or an end token is encountered; The final output is a gravitational wave image inversion temperature text decoded from the generated tokens, representing a creative interpretation of the original image and prompts.

8. The method according to claim 1, characterized in that The method further comprises: The similarity between the generated inversion temperature and the inversion temperature generated by MASP is calculated. The larger the value, the more similar the values ​​are, that is, the lower the probability of anomaly. After obtaining the similarity score, the anomaly threshold is set to judge the anomaly of the gravity wave spectrum image.

9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the temperature prediction method for gravity wave spectral image inversion based on large language model fine-tuning according to any one of claims 1 to 8 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for temperature prediction by inversion of gravity wave spectral images based on large language model fine-tuning according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Baijiu brewing spectral data analysis method and system based on big language model

    CN121543010A