Astronomical rare transient event early warning and monitoring method based on attention mechanism

Through an early warning and monitoring method for astronomical rare transient events based on the attention mechanism, and using data simulation and feature extraction technology, the real-time and accuracy problems of rare event detection in existing technologies are solved, and high-precision, real-time detection and full-process monitoring of rare transient events are achieved, which reduces false alarms and missed alarms and improves the robustness and adaptability of the system.

CN120744332APending Publication Date: 2025-10-03XIAN TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510718260.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing technologies make it difficult to detect rare astronomical transient events in real time and with high precision. Moreover, due to the scarcity of rare event data, the model has poor generalization ability in actual observations, and false alarms and missed alarms are serious, making it difficult to meet the needs of high-precision and real-time detection.

Method used

An early warning and monitoring method for astronomical rare transient events based on the attention mechanism is adopted. Training data is generated through data simulation, and features are extracted using sliding windows and one-dimensional convolution. A deep feature extraction module is constructed. Self-attention and cross-attention mechanisms are combined for feature extraction and classification. The model is trained using the cross-entropy loss function, and the event status is monitored in real time through an early warning inference algorithm.

Benefits of technology

It effectively improves the model's adaptability to different background noises and types of events, realizes high-precision, real-time detection of rare transient events, shortens the warning response time, reduces false alarms and missed alarms, has flexibility and wide applicability, and can achieve precise time series positioning and confidence quantification in astronomical transient events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744332A_ABST
    Figure CN120744332A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of astronomy, and relates to an astronomy rare transient event early warning and monitoring method based on an attention mechanism. The method comprises the following steps: inputting a small amount of real observation data into ChatGPT-4o, analyzing a data mode and a characteristic rule, and performing simulation; segmenting the simulation data into observation fragments by adopting a sliding window, selecting part of the observation fragments for manufacturing training samples, and performing one-dimensional convolution on the observation fragments to form observation tokens; constructing a depth feature extraction module based on an attention mechanism, capturing and observing a long-range dependency relationship in tokens to obtain a final feature vector, obtaining an event state classification probability vector after passing through a classification layer, calculating loss according to a label, and carrying out back propagation on a training model until convergence; a state classification result is obtained from a test sample through a convergence model, and the result reasoning module integrates in real time and conducts reasoning. According to the method, the model adaptability, the generalization ability and the initial recognition accuracy are improved, the early warning response time is shortened, and early warning with the accuracy rate of 98% can be given out within the first 25 observation step lengths of the early stage of event occurrence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of astronomy, and specifically relates to an early warning and monitoring method for astronomical rare transient events based on an attention mechanism. Background Art

[0002] In astrophysics, traditional methods for predicting rare transient events include those based on position offsets and magnetic activity records, but these methods have significant limitations. On the one hand, because transient events can span from hours to years, prediction responses are often delayed, making it difficult to capture critical moments in a timely manner. On the other hand, relying on historical behavioral patterns for predictions in complex background noise can easily lead to high false positive and false negative rates. In addition, because data on rare events themselves is extremely scarce, it is difficult for the model to obtain sufficient samples during training, resulting in insufficient generalization capabilities and limiting its application in actual observations.

[0003] In "2024104466894", a method for classifying motor imagery EEG signals based on multi-scale convolution and self-attention was disclosed. First, a multi-scale residual convolution module was used to extract the multi-scale time domain features of the EEG signal. Then, a spatial convolution module was used to obtain the spatial dimension information of the multi-channel EEG signal. For the extracted EEG spatiotemporal features, the feature dimension was reduced through the average pooling layer to reduce the redundant information in the features and improve the computational efficiency of the model. Then, a self-attention module was used to calculate the global attention score of the EEG spatiotemporal features to obtain the long-distance time dependency of the EEG signal, so that the model can focus on the information in the input features that is most relevant to the motor imagery category. There are problems with applying it to astronomical rare transient events: this method is mainly designed for the classification task of EEG signals, and the extracted spatiotemporal features are different from the features of sudden, non-stationary, and strong noise backgrounds in astronomical transient events, making it difficult to accurately capture the abnormal patterns of rare events; although the self-attention mechanism can model temporal dependencies, it has limited modeling capabilities for the ultra-long time scales required for astronomical observations, which can easily lead to the omission of key change information; in addition, samples of astronomical rare transient events are extremely scarce, and directly applying this method may lead to poor generalization ability of the model in actual observations, serious false alarms and omissions, and it is difficult to meet the needs of high-precision and real-time detection.

[0004] Therefore, a new method is urgently needed to detect rare astronomical transient events in real time and with high precision. Summary of the Invention

[0005] The present invention provides an early warning and monitoring method for astronomical rare transient events based on an attention mechanism to overcome the problem in the existing technology that it is difficult to accurately capture the abnormal patterns of rare events. At the same time, due to the scarcity of rare event data, the model has poor generalization ability and serious false alarms and omissions in actual observations, making it difficult to meet the needs of high-precision and real-time detection.

[0006] To achieve the above objectives, the technical solution of the present invention is: an early warning and monitoring method for astronomical rare transient events based on an attention mechanism, comprising the following steps:

[0007] Step 1: Data simulation: A small amount of real rare astronomical transient event observation data is input into ChatGPT-4o. Through prompts, it conducts an in-depth analysis of the data patterns and characteristic regularities, and provides data simulation methods.

[0008] Step 2: Preprocessing of training data: Use the sliding window method to split the training data obtained in step 1 in chronological order to obtain a series of observation segments. Then use a one-dimensional convolution operation to extract preliminary feature information of each observation segment to form observation tokens.

[0009] Step 3: Event state classification and model training: A deep feature extraction module is constructed based on the attention mechanism. The deep feature extraction module captures the long-range dependencies in the observation tokens obtained in step 2 and obtains the final self-learning feature vector. This vector passes through the output classification layer to obtain the current event state classification probability vector. The loss is then calculated based on the labels, and the model is trained through backpropagation until a converged model is obtained.

[0010] Step 4. Warning reasoning: After preprocessing the test data, the convergence model obtained in step 3 is used to obtain the event status classification results. The result reasoning module integrates the event status classification results in real time and performs reasoning. The reasoning process includes giving warnings in the early stages of the event, giving prompts when the event occurs completely, giving the confidence level of the event at the end of event monitoring, and monitoring the event status in real time throughout the process.

[0011] Furthermore, the deep feature extraction module in the above step 3 uses an encoder structure or an encoder-decoder structure, wherein the encoder structure adopts a self-attention mechanism and the decoder structure adopts a cross attention mechanism.

[0012] Furthermore, in step 3.3 above, the cross entropy loss function used in Alert model training is defined as follows:

[0013]

[0014] The symbols are defined as follows:

[0015] N represents the total number of training samples; M represents the total number of categories; y ic ∈{0,1} is an indicator variable indicating whether sample i belongs to the cth category, if so, it is 1, otherwise it is 0; p ic is the model's predicted probability that sample i belongs to category c.

[0016] Furthermore, the reasoning process in step 4 above specifically includes:

[0017] a. Early stage of the event: When a continuous subsequence of length L and all state values ​​1 (denoted as s1) appears in the state sequence for the first time, it is determined that the event has started at the current time t0 and the flag FS is set. t = True; the confidence score is accumulated by multiplying the mean probability of state 1 in the subsequence by the weighting coefficient a;

[0018] b. Complete occurrence detection phase: If the event start (FS t = True), then continue to determine whether a subsequence (s2) of length M and state values ​​all equal to 2 appears for the first time; if so, record the time point when the event occurs completely as t1, and set FC t = True, and update the confidence score: the mean probability of state 2 is multiplied by the coefficient b and accumulated;

[0019] c. Event monitoring end stage: When the event has started and completely occurred, if a subsequence (s3) with a state value of 3 and a length of N is detected, it means that the event is about to disappear from the observation window, and the flag FP is set. t = True, and the final confidence is obtained by multiplying the probability mean of state 3 by the coefficient c and accumulating them;

[0020] During the entire process of the above ac, the real-time monitoring mechanism continuously outputs the judgment status of the current event at each moment.

[0021] Furthermore, the threshold values ​​L, M, and N are 7, 14, and 4 respectively; and the threshold values ​​a, b, and c are 0.5, 0.25, and 0.25 respectively.

[0022] Furthermore, the simulation process for the microlensing event in step 1 above is as follows:

[0023] The calculation formula of the amplification factor A is as follows:

[0024]

[0025] The magnitude magnorm corresponding to the light curve is calculated as follows:

[0026] magnorm=magnorm base -h·log 10 (A)

[0027] The formula for simulated observation error is as follows:

[0028] σ noise =h·(max(log 10 (A))-min(log 10 (A)))·uniform(0.01,0.10)

[0029] The final noise magnitude expression is:

[0030]

[0031] Furthermore, the simulation process for the stellar flare event in step 1 above is as follows:

[0032] The stellar flare curve consists of two stages: a rapid rise stage and an exponential decay stage. First, the length L is generated. flare A uniform time series t:

[0033] t=linspace(0,1,L flare )

[0034] According to the exponential decay rate r decay Determine the peak time t peak :

[0035]

[0036] At t≤t peak During this period, the magnitude decreases linearly from the baseline value b to the peak value p:

[0037] y rise (t)=interp(t,[0,t peak ],[b,p])

[0038] At t>t peak During this period, the magnitude returned to the baseline value exponentially:

[0039]

[0040] The synthetic light curve is:

[0041]

[0042] Then add the simulation error to get the noisy flare curve:

[0043]

[0044] The background light curve is generated using a Gaussian distribution:

[0045]

[0046] Finally, the flare curve is embedded in the middle region of the background sequence, and the complete light curve is:

[0047]

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] 1) This paper proposes a new method for expanding training samples for rare transient events, which effectively improves the model's adaptability and generalization ability to different background noises and different types of events, ensuring high-precision predictions in different environments.

[0050] 2) The present invention uses a sliding window mechanism to process long-term astronomical data, taking into account both real-time performance and context capture capabilities. Therefore, while maintaining real-time performance, the present invention effectively extracts local and global features and can stably and accurately detect and classify event states.

[0051] 3) The present invention utilizes the attention mechanism to construct a deep feature extraction module and applies it to the detection of rare transient events. Therefore, the present invention can accurately capture key features at long time scales in complex noise backgrounds and sparsely sampled observation data, greatly improving the recognition accuracy of rare transient events in the early stages of their occurrence, greatly shortening the early warning response time, and enabling early detection and full-process monitoring of rare transient events.

[0052] 4) This paper proposes and validates a warning inference algorithm, using it as a warning inference module. This algorithm dynamically integrates the classification results of continuous time series, provides early warnings when an event occurs, generates signals when the event has fully occurred, and provides confidence in the occurrence at the end of event monitoring. It also monitors the event status in real time throughout the entire process. This reduces false positives and missed negatives. This method is rationally designed and clearly structured, enabling precise temporal location and confidence quantification of event occurrences. It is widely applicable to intelligent recognition systems for astronomical transient events.

[0053] 5) The present invention proposes and verifies a complete system for detecting rare astronomical transient events, including a simulation module for rare event training data samples; a preprocessing module for long-term astronomical data; a module for extracting features from long-term dependencies and classifying and setting event states; and an algorithm module for early warning, state monitoring, and event confidence calculation. This system achieves comprehensive improvements in prediction accuracy, response speed, noise immunity, and system robustness in the early warning and full-process monitoring tasks of rare transient events. Compared with traditional methods and single improved methods, the overall performance is significantly enhanced. The system can issue a high-accuracy (up to 98%) warning signal within the first 25 observation time points of an astronomical event and conduct full-process event state monitoring. The modular design of this system gives it strong flexibility. By changing a few parameters in training, it can be applied to event detection under different conditions. It has broad application prospects and provides a new approach for detecting rare astronomical transient events in the time domain. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is the overall framework diagram of the present invention;

[0055] Figure 2 Schematic diagrams of the Alert model constructed in the present invention, wherein Figure (a) shows a schematic diagram of the ALert_E model, and Figure (b) shows a schematic diagram of the Alert_E&D model;

[0056] Figure 3 Schematic diagram of the components of the Alert model constructed in the present invention, wherein Figure (a) is a schematic diagram of the working mechanism of the fully connected layer, Figure (b) is a schematic diagram of the self-attention mechanism, and Figure (c) is a schematic diagram of the cross-attention mechanism;

[0057] Figure 4 This is a schematic diagram of the event status classification set in the present invention, wherein Figure (a) describes event status classification 0, Figure (b) describes event status classification 1, Figure (c) describes event status classification 2, and Figure (d) describes event status classification 3. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0059] Target audience:

[0060] OGLE dataset: Contains photometric monitoring data of about 78.7 million stars over the past 20 years, mainly used for the study of microlensing events.

[0061] GWAC dataset: mainly monitors stellar flare events, with high temporal resolution and large field of view coverage.

[0062] Example 1: This example is targeted at microlensing events. Based on this, the present invention provides an early warning and monitoring method for astronomical rare transient events based on an attention mechanism, which specifically includes the following steps:

[0063] Step 1: Data simulation:

[0064] Step 1.1: Input a small amount of real rare astronomical transient event (microlensing) observation data into ChatGPT-4o. Through prompts, ChatGPT-4o conducts an in-depth analysis of the data patterns and characteristic laws, and provides a data simulation method, as follows:

[0065] Factor A is calculated using the following formula:

[0066] in:

[0067]

[0068] The amplification factor A is used to calculate the magnitude magnorm corresponding to the light curve. The calculation formula is as follows:

[0069] magnorm=magnorm base -h·log 10 (A)

[0070] in:

[0071] magnorm base It represents the base magnitude without lensing effect, and h is the set adjustment factor used to control the magnification.

[0072] To simulate the observation error, Gaussian noise is introduced based on this magnitude, and its standard deviation is:

[0073] σ noise =h·(max(log 10 (A))-min(log 10 (A)))·uniform(0.01,0.10)

[0074] The final noise magnitude expression is:

[0075]

[0076] in Indicates that the mean is 0 and the standard deviation is σ noise Gaussian noise.

[0077] The range of values ​​for the microlensing data simulation parameters is shown in Table 1.

[0078] Table 1

[0079] parameter meaning Value range <![CDATA[μ0]]> Impulse parameters (closest distance) 0.1–0.2 <![CDATA[t E ]]> Einstein time scale 5–100 <![CDATA[magnorm base ]]> Baseline magnitude 4.0–11.0 h Magnification height adjustment factor 0.2–2.5 <![CDATA[σ noise ]]> Gaussian noise amplitude 0.01–0.10 L Time series length 1800–4500

[0080] Step 1.2 generates a large amount of simulation data based on the obtained simulation method and parameter value range, and obtains 100,000 pieces of expanded microlensing training data.

[0081] Step 2: Preprocessing of training data. This step can effectively extract representative features and improve the recognition performance of the model while maintaining the integrity of the event. It specifically includes the following steps:

[0082] Step 2.1: Use the sliding window method to divide the training data obtained in step 1 into a series of observation segments in chronological order: set the sliding window length Δt = 512 to divide the continuous observation time series. Each observation segment is denoted as O (t-Δt,t) , which represents an observation interval from time t-Δt to t. The long time series data obtained in step 1 is segmented moment by moment to obtain a series of observation segments. This time length is optimized to ensure that most transient event signals are fully contained within a single observation segment, thus avoiding information loss caused by event truncation.

[0083] In step 2.2, a one-dimensional convolution operation is used to extract preliminary feature information for each observation segment, forming observation tokens. A one-dimensional convolution operation is applied to each observation segment to extract local temporal features. Specifically, a one-dimensional convolution operation with a kernel size of 4 and a stride of 4 is applied to the observation segment of original length 512, downsampling the time dimension to a feature representation of length 128. Each output value in this process integrates observation information from four consecutive time points, effectively compressing the data while preserving the changing trend of the event.

[0084] To enhance the expressiveness of features, this method uses 64 independent one-dimensional convolution kernels for parallel computation, resulting in a 64-dimensional observation feature vector, or observation tokens. Each observation token extracts a distinct feature pattern from the original segment from a different convolution kernel perspective, resulting in strong semantic separation capabilities.

[0085] The observation tokens, as the main input of the model, have the advantages of high efficiency, compactness, and rich information, providing a well-structured and semantically clear input basis for subsequent model training and event recognition.

[0086] Step 3: Event status classification and model training. The specific steps are as follows:

[0087] Step 3.1: Build a deep feature extraction module based on the attention mechanism:

[0088] See also Figure 2(a) The deep feature extraction module constructed in this embodiment is an encoder structure (Alert_E), which consists of a deep feature extraction module and fully connected classification layers. It is implemented based on the self-attention mechanism and aims to use the multi-layer self-attention mechanism to efficiently extract and fuse deep semantic features in astronomical observation fragments to accurately identify the state category of transient events. Figure 3 (a, b), the implementation methods of full connection and self-attention mechanisms are shown in detail.

[0089] Step 3.2: The deep feature extraction module captures the long-range dependencies in the observation tokens obtained in step 2 and obtains the final feature vector corresponding to the self-learned token. In the Alert_E model, a self-learned token of length 128 is first introduced and concatenated with the preprocessed observation tokens to form a sequence. To enable the model to perceive the relative position information between tokens, all tokens are processed using a positional encoding based on sine and cosine functions. Specifically, for the i-th position in the sequence, the following encoding strategy is used:

[0090] Odd index positions use:

[0091]

[0092] Even index positions use:

[0093]

[0094] The encoded token sequence is input into an encoder consisting of 8 self-attention layers to extract deep features and obtain the final feature vector as follows: Figure 2 (a) Feature extraction part.

[0095] The final feature vector obtained in step 3.3 and step 3.2 is passed through three layers of fully connected layers ( Figure 3 (a)) after the output classification layer is formed, the event state classification probability vector of length 4 at the current moment is obtained ( Figure 2 In the feature analysis and state prediction part (a, b), the event state is determined according to the maximum probability principle. State 0 represents pure background noise. Figure 4 (a); State 1 means the event has just started but is not yet complete. Figure 4 (b); State 2 indicates that the event has completely occurred. Figure 4 (c); State 3 indicates that the event gradually ends. Figure 4 (d).

[0096] In the present invention, the cross entropy loss function is selected as the core optimization criterion for Alert model training.

[0097] The loss function is defined as follows:

[0098]

[0099] The symbols are defined as follows:

[0100] N represents the total number of training samples;

[0101] M represents the total number of categories;

[0102] y ic ∈{0,1} is an indicator variable indicating whether sample i belongs to the cth category, if yes, it is 1, otherwise it is 0;

[0103] p ic is the model's predicted probability that sample i belongs to category c.

[0104] The specific training details of the model in this invention are as follows: First, relying on the NVIDIA DGX workstation equipped with four Tesla V100 graphics processors, a GPU-accelerated training environment based on the PyTorch deep learning framework is built; this environment runs on the official optimized container to ensure high compatibility and high throughput of drivers, library files and hardware instruction sets. During the training phase, all observation sequences are normalized and preprocessed, and then fed into the network in parallel with a batch size of 512 (mini batch). Subsequently, the cross entropy loss function is called to calculate the prediction error, and the Adam optimizer is used to perform gradient backpropagation and update on the parameters; the initial learning rate is set to 1×10 -5 During training, adaptive or segmented attenuation strategies can be employed as needed to balance convergence speed and stability. Full training lasts for 100 epochs, and the validation set is monitored at the end of each epoch to dynamically evaluate model performance, identify overfitting risks in advance, and fine-tune hyperparameters accordingly. Through the aforementioned hardware acceleration, refined hyperparameter configuration, and rigorous periodic verification, this invention significantly shortens model convergence time, improves prediction accuracy and system reliability, and meets the engineering requirements of large-scale real-time observational data processing.

[0105] Train the model according to the above details until a converged model is obtained.

[0106] Step 4: Warning reasoning: After preprocessing the test data, the convergence model obtained in step 3 is used to obtain the event status classification results. The result reasoning module integrates the event status classification results in real time and performs reasoning:

[0107] Step 4.1: After preprocessing the test data, the event status classification results are obtained through the convergence model obtained in step 3:

[0108] The test data mentioned is the OGLE dataset, which contains photometric data from approximately 78.7 million stars over nearly 20 years and is primarily used for studying gravitational microlensing events. The GWAC dataset monitors ultra-white light flares, with high temporal resolution and large field of view.

[0109] The preprocessing is consistent with the training data preprocessing process.

[0110] This step uses the predicted state value at each observation moment as the input of the convergence model to obtain the event state classification result. The event state is divided into four categories: state 0 represents pure background noise, such as Figure 4 (a); State 1 means the event has just started but is not yet complete. Figure 4 (b); State 2 indicates that the event has completely occurred. Figure 4 (c); State 3 indicates that the event gradually ends. Figure 4 (d).

[0111] Step 4.2: Based on the timeline, scan the state sequence to see if there is a subsequence that meets a specific pattern. Inference is performed by combining flag variables and a confidence weighting mechanism to determine the start and end of the event, monitor, and mark it. The reasoning process specifically includes issuing an early warning when the event occurs, issuing a signal when the event has completely occurred, and providing a confidence level at the end of the event monitoring. The event status is monitored in real time throughout the entire process. The specific process includes:

[0112] a. Early stage of the event: When a continuous subsequence of length L and all state values ​​1 (denoted as s1) appears in the state sequence for the first time, it is determined that the event has started at the current time t0 and the flag FS is set. t = True. The confidence score is the sum of the mean probability of state 1 in the subsequence multiplied by the weighting coefficient a.

[0113] b. Complete occurrence detection phase: If the event start (FS t = True), then continue to determine whether a subsequence (s2) of length M and state values ​​all equal to 2 appears for the first time. If so, record the time point when the event occurs completely as t1, and set FC t = True, and update the confidence score: the mean probability of state 2 is multiplied by the coefficient b and accumulated.

[0114] c. Event monitoring end stage: When the event has started and completely occurred, if a subsequence (s3) with a state value of 3 and a length of N is detected, it means that the event is about to disappear from the observation window, and the flag FP is set. t = True, and the final confidence is obtained by multiplying the probability mean of state 3 by the coefficient c and accumulating them.

[0115] During this process, the real-time monitoring mechanism ensures that the current event status is continuously output at every moment. All thresholds (L, M, N) and weighting coefficients (a, b, c) can be adjusted to suit different task requirements. In this embodiment, L, M, and N are 7, 14, and 4, respectively; a, b, and c are 0.5, 0.25, and 0.25, respectively.

[0116] This step combines rule-driven and probabilistic quantification, improving the timeliness and accuracy of event detection while having good interpretability and scalability. It is an efficient early warning strategy for real-time streaming data.

[0117] Step 4.3: Give the confidence level of the event:

[0118] The mean probability of state 1 in the sequence is multiplied by the weighted coefficient a + the mean probability of state 2 is multiplied by the coefficient b + the mean probability of state 3 is multiplied by the coefficient c.

[0119] End the whole process.

[0120] Example 2: This example is aimed at stellar flare events. Based on this, the present invention provides an early warning and monitoring method for astronomical rare transient events based on an attention mechanism. The difference from Example 1 lies in the difference between steps 1 and 3, which are specifically described as follows:

[0121] Step 1: Data simulation:

[0122] Step 1.1: Input a small amount of observational data of rare astronomical transient events (stellar flare fluctuations) into ChatGPT-4o. Through prompts, ChatGPT-4o conducts an in-depth analysis of the data patterns and characteristic laws, and provides a data simulation method, as follows:

[0123] The stellar flare curve consists of two stages: a rapid rise stage and an exponential decay stage. First, the length L is generated. flare A uniform time series t:

[0124] t=linspace(0,1,L flare )

[0125] According to the exponential decay rate r decay Determine the peak time t peak :

[0126]

[0127] At t≤t peak During this period, the magnitude decreases linearly from the baseline value b to the peak value p:

[0128] y rise (t)=interp(t,[0,t peak ],[b,p])

[0129] At t>t peak During this period, the magnitude returned to the baseline value exponentially:

[0130]

[0131] The synthetic light curve is:

[0132]

[0133] Then add the simulation error to get the noisy flare curve:

[0134]

[0135] The background light curve is generated using a Gaussian distribution:

[0136]

[0137] Finally, the flare curve is embedded in the middle region of the background sequence, and the complete light curve is:

[0138] Table 2 Parameter ranges for stellar flare data simulation

[0139] parameter meaning Value range <![CDATA[L flare ]]> Flare duration 50–200 b Baseline magnitude 4.0–11.0 p Peak magnitude b-uniform(0.2,2.5) <![CDATA[r decay ]]> Attenuation ratio 3.2–4.0 <![CDATA[σ flare ]]> Flare noise amplitude 0.01–0.15 <![CDATA[σ background ]]> Background noise amplitude 0.03–0.06 <![CDATA[t insert ]]> Flare insertion location Middle region of the time series L Total length of time series 1800–4500

[0140] Step 1.2 generates a large amount of simulation data based on the obtained simulation method and parameter value range, and obtains 100,000 expanded stellar flare training data.

[0141] In step 3.1 of step 3, a deep feature extraction module is constructed based on the attention mechanism:

[0142] The structure of the deep feature extraction module constructed in this embodiment is an encoder-decoder structure (Alert_E&D), in which the encoder structure is the same as Alert_E, and the decoder structure is constructed based on the cross attention mechanism. The cross attention mechanism is as follows: Figure 3 (c),The difference between the cross-attention mechanism and the self-attention mechanism is that the set learnable vector only provides the Q vector, and the remaining tokens only provide the K and V vectors.

[0143] In the Alert_E&D model, a decoder structure is introduced to enhance the feature fusion effect. In this structure, the learnable token first generates an intermediate feature vector through a fully connected layer. This vector and the token feature set output by the encoder are sent to the cross attention layer to obtain the final feature vector as shown in the figure. Figure 2 (b) Feature extraction part.

[0144] Compared to the Alert_E model, which only uses an encoder structure, the Alert_E&D model further refines and semantically integrates the encoding results by introducing a decoder module, thus having significant advantages in state classification accuracy and model interpretability. The test results are as follows:

[0145] In the early warning task, the Alert_E model achieved 98% accuracy while allowing for an error of 25 observation time steps, with the average warning error concentrated between 9 and 12 time steps. The Alert_E&D model also achieved 98% accuracy under the same conditions, with a slightly better average warning error, concentrated between 8 and 12 time steps.

[0146] Both models demonstrated exceptional performance in the state classification task. Both Alert_E and Alert_E&D achieved 99.8% classification accuracy on the OGLE dataset. On the GWAC dataset, the Alert_E model achieved 98.2% accuracy in event state classification, while the Alert_E&D model achieved 93.1%. Furthermore, in tests using pure noise data, neither model experienced false positives, demonstrating strong noise immunity.

[0147] In summary, the Alert_E and Alert_E&D models proposed in this paper can both achieve high-precision, low-latency early warning and full-process monitoring of astronomical rare transient events, and have excellent generalization performance and practicality.

[0148] The above description is an explanation of the specific implementation of the present invention, rather than a limitation of the present invention. Those skilled in the relevant technical field can also make various equivalent technical solutions without departing from the scope of the present invention, so all equivalent technical solutions should be included in the scope of protection of the present invention.

Claims

1. An attention-based early warning and monitoring method for astronomical rare transient events, characterized by: The following steps are included Step 1: Data simulation: A small amount of real rare astronomical transient event observation data is input into ChatGPT-4o. Through prompts, it conducts an in-depth analysis of the data patterns and characteristic regularities, and provides data simulation methods. Step 2: Preprocessing of training data: Use the sliding window method to split the training data obtained in step 1 in chronological order to obtain a series of observation segments. Then use a one-dimensional convolution operation to extract preliminary feature information of each observation segment to form observation tokens. Step 3: Event state classification and model training: A deep feature extraction module is constructed based on the attention mechanism. The deep feature extraction module captures the long-range dependencies in the observation tokens obtained in step 2 and obtains the final self-learning feature vector. This vector passes through the output classification layer to obtain the current event state classification probability vector. The loss is then calculated based on the labels, and the model is trained through backpropagation until a converged model is obtained. Step 4. Warning reasoning: After preprocessing the test data, the convergence model obtained in step 3 is used to obtain the event status classification results. The result reasoning module integrates the event status classification results in real time and performs reasoning. The reasoning process includes giving warnings in the early stages of the event, giving prompts when the event occurs completely, giving the confidence level of the event at the end of event monitoring, and monitoring the event status in real time throughout the process.

2. The method for early warning and monitoring of astronomical rare transient events based on an attention mechanism according to claim 1, characterized in that: The deep feature extraction module in step three uses an encoder structure or an encoder-decoder structure, wherein the encoder structure adopts a self-attention mechanism and the decoder structure adopts a cross-attention mechanism.

3. The method for early warning and monitoring of astronomical rare transient events based on an attention mechanism according to claim 1, characterized in that: In step 3.3, the cross entropy loss function used in Alert model training is defined as follows: The symbols are defined as follows: N represents the total number of training samples; M represents the total number of categories; y ic ∈{0,1} is an indicator variable indicating whether sample i belongs to the cth category, if so, it is 1, otherwise it is 0; p ic is the model's predicted probability that sample i belongs to category c.

4. The method for early warning and monitoring of astronomical rare transient events based on an attention mechanism according to claim 1, characterized in that: The reasoning process in step 4 specifically includes: a. Early stage of the event: When a continuous subsequence of length L and all state values ​​1 (denoted as s1) appears in the state sequence for the first time, it is determined that the event has started at the current time t0 and the flag FS is set. t = True; the confidence score is accumulated by multiplying the mean probability of state 1 in the subsequence by the weighting coefficient a; b. Complete occurrence detection phase: If the event start (FS t = True), then continue to determine whether a subsequence (s2) of length M and state values ​​all equal to 2 appears for the first time; if so, record the time point when the event occurs completely as t1, and set FC t = True, and update the confidence score: the mean probability of state 2 is multiplied by the coefficient b and accumulated; c. Event monitoring end stage: When the event has started and completely occurred, if a subsequence (s3) with a state value of 3 and a length of N is detected, it means that the event is about to disappear from the observation window, and the flag FP is set. t = True, and the final confidence is obtained by multiplying the probability mean of state 3 by the coefficient c and accumulating them; During the entire process of the above ac, the real-time monitoring mechanism continuously outputs the judgment status of the current event at each moment.

5. The method for early warning and monitoring of astronomical rare transient events based on an attention mechanism according to claim 1, characterized in that: The threshold values ​​L, M, and N are 7, 14, and 4 respectively; the values ​​of a, b, and c are 0.5, 0.25, and 0.25 respectively.

6. The method for early warning and monitoring of astronomical rare transient events based on an attention mechanism according to claim 1, characterized in that: The simulation process for the microlensing event in step 1 is as follows: The calculation formula of the amplification factor A is as follows: The magnitude magnorm corresponding to the light curve is calculated as follows: magnorm=magnorm base -h·log 10 (A) The formula for simulated observation error is as follows: s noise =h·(max(log 10 (A))-min(log 10 (A)))·uniform(0.01,0.10) The final noise magnitude expression is:

7. The method for early warning and monitoring of astronomical rare transient events based on an attention mechanism according to claim 1, characterized in that: The simulation process for the stellar flare event in step 1 is as follows: The stellar flare curve consists of two stages: a rapid rise stage and an exponential decay stage. First, the length L is generated. flare A uniform time series t: t=linspace(0,1,L flare ) According to the exponential decay rate r decay Determine the peak time t peak : At t≤t peak During this period, the magnitude decreases linearly from the baseline value b to the peak value p: y rise (t)=interp(t,[0,t peak ],[b,p]) At t>t peak During this period, the magnitude returned to the baseline value exponentially: The synthetic light curve is: Then add the simulation error to get the noisy flare curve: The background light curve is generated using a Gaussian distribution: Finally, the flare curve is embedded in the middle region of the background sequence, and the complete light curve is: