Meteorological Quality Data Analysis Method and System Based on Multi-Source Data Fusion and AI

By using multi-source data fusion and deep learning techniques in meteorological data analysis, the characteristics of meteorological data and site observation data are extracted and represented, and the problem that existing methods are difficult to effectively utilize these characteristics is solved, achieving more accurate and stable visibility prediction effects.

CN118114201BActive Publication Date: 2025-05-30CHINESE ACAD OF METEOROLOGICAL SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410341383.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-25
Publication Date
2025-05-30
Estimated Expiration
2044-03-25

AI Technical Summary

Technical Problem

Existing methods that combine multi-source data fusion and deep learning are difficult to effectively extract and represent the features of meteorological data and site observation data in meteorological data analysis, and it is difficult to design a reasonable deep learning model to make full use of these features.

Method used

By obtaining the meteorological data of the bet-analyzed system and related site observation data, it is loaded into the trained target multi-source data visibility recognition model, and feature extraction and pattern recognition are used using deep learning technology. The model gradually optimizes its performance in visibility prediction through basic training and model refinement adjustment.

Benefits of technology

The accuracy and stability of visibility prediction are improved, and the effect of meteorological data analysis is improved by making full use of the complementarity and correlation of multi-source data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118114201B_ABST
    Figure CN118114201B_ABST
Patent Text Reader

Abstract

The present application provides a method and system for meteorological quality data analysis based on multi-source data fusion and AI. In the basic training link of the model, first, endogenous learning is respectively carried out on the meteorological data vector representation component and the station data vector representation component. Then, example-driven learning is carried out on the initial multi-source data visibility recognition model that at least covers the initial meteorological data vector representation component and the initial station data vector representation component obtained from endogenous learning, so that the basic training is divided into two links. The endogenous learning link can separately complete the training of a single type of system meteorological data and station observation data to obtain the model ability of feature mining. The combined example-driven learning continuously learns the features of the system meteorological data and the station observation data to complete the visibility level classification. It can fully acquire the knowledge of the basic training knowledge template set, simulate the prior labels, and at the same time simulate the involvement effect between different types of data, improving the accuracy of visibility level classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and specifically, to a meteorological quality data analysis method and system based on multi-source data fusion and AI. Background Art

[0002] With the continuous development of society and the advancement of science and technology, meteorological observation and forecasting technology plays a vital role in many fields such as daily life, agricultural production, transportation, etc. Especially in terms of visibility forecasting, accurate forecasting results are of irreplaceable value in ensuring traffic safety, guiding aviation and navigation, and evaluating the quality of the atmospheric environment. Traditional visibility forecasting methods often rely on a single data source and a simple mathematical model, which makes it difficult to fully utilize the complementarity and correlation of multi-source data, resulting in limited accuracy and stability of the forecast results. In recent years, with the rapid development of artificial intelligence technology, especially the continuous innovation in the field of machine learning, new solutions have been provided for multi-source data fusion and complex pattern recognition.

[0003] In the field of machine learning, deep learning technology has achieved breakthrough results in many fields such as image recognition, speech recognition, and natural language processing with its powerful feature extraction and pattern recognition capabilities. However, in the field of meteorological data analysis, especially in visibility prediction, how to effectively use deep learning technology to improve prediction performance is still a challenging problem due to the complexity, variability, and uncertainty of meteorological data. In order to solve this problem, researchers began to explore methods that combine multi-source data fusion technology with deep learning. By organically fusing meteorological data from different sensors and site observation data, and using deep learning models to mine deep associations and patterns between data, it is expected to improve the accuracy and stability of visibility prediction. However, the existing methods combining multi-source data fusion and deep learning still face some challenges in practical applications. For example, how to effectively extract and represent the features of meteorological data and site observation data? How to design a reasonable deep learning model to make full use of these features? These are all urgent problems to be solved in current research. Summary of the invention

[0004] The object of the present invention is to provide a meteorological quality data analysis method and system based on multi-source data fusion and AI.

[0005] The embodiment of the present application is implemented as follows:

[0006] In a first aspect, an embodiment of the present application provides a method for analyzing meteorological quality data based on multi-source data fusion and AI, including: obtaining meteorological data of a system to be analyzed, and determining observation data of a station to be analyzed related to the meteorological data of the system to be analyzed; loading the meteorological data of the system to be analyzed and the observation data of the station to be analyzed into a trained target multi-source data visibility recognition model to obtain visibility analysis and recognition information output by the target multi-source data visibility recognition model, where the target multi-source data visibility recognition model is a machine learning model obtained by refining and calibrating a multi-source data visibility recognition model after basic training; where the basic training process of the multi-source data visibility recognition model includes the following steps: obtaining a set of basic training knowledge templates, each basic training knowledge template in the set of basic training knowledge templates includes a system meteorological data knowledge template, a station observation data knowledge template related to the system meteorological data knowledge template, and a prior visibility label of the template; according to the system meteorological data knowledge templates included in each knowledge template in the set of basic training knowledge templates, performing multiple endogenous learning on the meteorological data vector representation component to obtain a learned initial meteorological data vector representation component; according to the station observation data knowledge templates related to the system meteorological data knowledge templates included in each knowledge template in the set of basic training knowledge templates, performing multiple endogenous learning on the station data vector representation component to obtain a learned initial station data vector representation component; according to the set of basic training knowledge templates, performing multiple example-driven learning on an initial multi-source data visibility recognition model that at least covers the initial meteorological data vector representation component and the initial station data vector representation component to obtain a multi-source data visibility recognition model after basic training.

[0007] As an implementation manner, the obtaining of the basic training knowledge template set includes: generating a historical meteorological observation information set corresponding to each visibility monitoring area according to the historical meteorological observation information respectively corresponding to each visibility monitoring area; wherein, each historical meteorological observation information includes a historical system meteorological data, historical site observation data related to the historical system meteorological data, and an original visibility prior label; for each visibility monitoring area, after performing example-driven learning according to the historical meteorological observation information corresponding to the visibility monitoring area, obtaining a target auxiliary classifier corresponding to the visibility monitoring area; for each historical system meteorological data, determining respective auxiliary visibility classification information related to the historical system meteorological data according to each target auxiliary classifier, and using the original visibility prior label related to the historical system meteorological data and the augmented label set of the respective auxiliary visibility classification information as the template visibility prior label related to the historical system meteorological data; forming a basic training knowledge template according to the historical system meteorological data, the template visibility prior label related to the historical system meteorological data, and the historical site observation data; and forming the basic training knowledge template set according to each basic training knowledge template.

[0008] As an implementation manner, the generating of the historical meteorological observation information set corresponding to each visibility monitoring area according to the historical meteorological observation information respectively corresponding to each visibility monitoring area includes: forming an initial meteorological observation information set according to each initial meteorological observation information with a monitoring time within a preset time interval in each visibility monitoring area, and each initial meteorological observation information includes an initial system meteorological data, historical site observation data related to the initial system meteorological data, and an original visibility prior label; extracting one initial meteorological observation information one by one from the initial meteorological observation information set, and completing the following single screening according to the extracted initial meteorological observation information until there is no unextracted initial meteorological observation information in the initial meteorological observation information set: respectively determining a meteorological data characterization vector matching factor between the initial system meteorological data in the extracted initial meteorological observation information and the initial system meteorological data in other initial meteorological observation information in the initial meteorological observation information set; removing other initial meteorological observation information in the initial meteorological observation information set whose meteorological data characterization vector matching factor with the extracted initial meteorological observation information reaches a matching factor threshold; using the initial meteorological observation information in the single-screened initial meteorological observation information set as the historical meteorological observation information; and obtaining the historical meteorological observation information set corresponding to each visibility monitoring area according to the cluster analysis of the respective historical meteorological observation information according to the visibility monitoring area to which each historical meteorological observation information belongs.

[0009] As an implementation manner, determining each piece of auxiliary visibility classification information related to the historical system meteorological data according to each target auxiliary classifier includes: determining a target visibility monitoring area corresponding to the historical system meteorological data, determining a target auxiliary classifier corresponding to the target visibility monitoring area, and obtaining other target auxiliary classifiers corresponding to each other visibility monitoring area except the target visibility monitoring area; and respectively determining each piece of auxiliary visibility classification information related in the historical system meteorological data according to each other target auxiliary classifier.

[0010] As an implementation manner, the auxiliary visibility classification information includes an auxiliary visibility mark and a predicted support coefficient predicted for the auxiliary visibility mark; after determining each piece of auxiliary visibility classification information related to the historical system meteorological data, it further includes: obtaining an initial overlap mark set between the auxiliary visibility marks included in each piece of auxiliary visibility classification information and the original visibility prior marks related to the historical system meteorological data, and determining the predicted support coefficient corresponding to each visibility prior mark in the initial overlap mark set; removing the visibility prior marks in the original visibility prior marks related to the historical system meteorological data, for which the predicted support coefficient in the initial overlap mark set reaches the mark filtering index, to obtain the processed original visibility prior marks; and for the processed original visibility prior marks, entering the step of using the original visibility prior marks related to the historical system meteorological data and the augmented label sets of each piece of auxiliary visibility classification information as the template visibility prior marks related to the historical system meteorological data for execution.

[0011] As an implementation manner, when performing an endogenous learning on the meteorological data vector representation component once, the following steps are completed: extracting a predetermined number of system meteorological data knowledge templates from the basic training knowledge template set, and respectively extracting a predetermined number of template meteorological data record points associated with time stamps from each of the extracted system meteorological data knowledge templates through the meteorological data vector representation component; forming cohesive knowledge template pairs according to each template meteorological data record point belonging to the same system meteorological data knowledge template and forming discrete knowledge template pairs according to each template meteorological data record point belonging to different system meteorological data knowledge templates through the meteorological data vector representation component, where one belonging to different includes two template meteorological data record points; and performing multiple endogenous cyclic learning on the meteorological data vector representation component according to the cohesive knowledge template pairs and the discrete knowledge template pairs until a preset single optimization end judgment criterion is reached.

[0012] As an implementation manner, when extracting a predetermined number of template meteorological data recording points associated with timestamps from each of the extracted system meteorological data knowledge templates, one of the following steps is completed: for each extracted system meteorological data knowledge template, according to the time span corresponding to the system meteorological data knowledge template, set the target node corresponding to the system meteorological data knowledge template, and starting from the template meteorological data recording point corresponding to the target node on the system meteorological data knowledge template, extract a predetermined number of template meteorological data recording points associated with timestamps; for each extracted system meteorological data knowledge template, according to the time span corresponding to the system meteorological data knowledge template, divide the system meteorological data knowledge template into a predetermined number of knowledge template sub-meteorological data, and arbitrarily extract one template meteorological data recording point from each of the knowledge template sub-meteorological data to obtain a predetermined number of template meteorological data recording points associated with timestamps.

[0013] As an implementation manner, the initial multi-source data visibility recognition model includes an initial meteorological data vector characterization component, an initial station data vector characterization component, a feature integration component for integrating station data features and system meteorological data features, and a classifier; in one example-driven learning of the initial multi-source data visibility recognition model, it includes: inputting the system meteorological data knowledge template and the corresponding station observation data knowledge template in a batch into the initial multi-source data visibility recognition model, where the system meteorological data knowledge template is input into the initial meteorological data vector characterization component, and the station observation data knowledge template is input into the initial station data vector characterization component; loading the fusion feature of each system meteorological data feature output by the initial meteorological data vector characterization component and the station data feature output by the station data vector characterization component into the feature integration component to obtain the fused system meteorological data feature and station data feature; loading the feature fusion result output by the feature integration component into the classifier to obtain the predicted visibility classification result output for each sample meteorological observation information; combining the error between the visibility classification result and the corresponding template visibility prior label, obtaining the cost value according to the binary cross-entropy cost function, and optimizing the model parameters of the initial multi-source data visibility recognition model according to the cost value.

[0014] As an implementation, when at least the meteorological data vector characterization component after basic training, the site data vector characterization component after basic training, and the classifier after basic training are included in the multi-source data visibility recognition model after basic training, when it is determined that the visibility prior marker set during model refinement tuning does not match the visibility prior marker set in the classifier after basic training, the multi-source data visibility recognition model after basic training is refined and tuned, and the following steps are completed to obtain the target multi-source data visibility recognition model: Obtain a set of refinement tuning knowledge templates, where each refinement tuning knowledge template includes a refinement tuning system meteorological data knowledge template, a refinement tuning site observation data knowledge template related to the refinement tuning system meteorological data knowledge template, and a refinement tuning template visibility prior marker; Update the visibility prior marker set in the classifier after basic training according to each visibility prior marker during model refinement tuning; Based on the set of refinement tuning knowledge templates, perform multiple model refinement tunings on the updated multi-source data visibility recognition model to obtain the target multi-source data visibility recognition model after refinement tuning learning.

[0015] In a second aspect, the present application provides a computer system, including:

[0016] One or more processors;

[0017] A memory; one or more computer programs; wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the methods described above are implemented.

[0018] The beneficial effects of the present application at least include: The present application provides a method and system for meteorological quality data analysis based on multi-source data fusion and AI. In the basic training link of the model, first, endogenous learning is performed on the meteorological data vector characterization component and the site data vector characterization component respectively, and then, an initial multi-source data visibility recognition model that at least covers the initial meteorological data vector characterization component and the initial site data vector characterization component obtained by endogenous learning is subjected to example-driven learning, so that the basic training is divided into two links. The endogenous learning link can independently complete the training of a single type of system meteorological data and site observation data to obtain the model ability of feature mining. At the same time, combined with example-driven learning, continuously learn the features of system meteorological data and site observation data to complete the visibility level classification. In this way, the multi-source data visibility recognition model after basic training can fully acquire the knowledge of the basic training knowledge template set, simulate the prior markers included in the basic training knowledge template, and at the same time simulate the involvement effect between different types of data, improving the accuracy of visibility level classification.

[0019] When classifying visibility levels, after obtaining the meteorological data of the system to be analyzed and the observed data of the stations to be analyzed, based on the target multi-source data visibility recognition model obtained by refining and calibrating the multi-source data visibility recognition model obtained from basic training, according to the meteorological data of the system to be analyzed and the observed data of the stations to be analyzed, visibility analysis and recognition information is output. In this way, for the multi-source data visibility recognition model obtained after basic training according to the batch-based basic training mechanism, after refining and calibrating the multi-source data visibility recognition model obtained after basic training that fully simulates known knowledge during basic training, the obtained target multi-source data visibility recognition model can improve the visibility level classification process reflected by complex multi-source data and complete the high-quality analysis of meteorological observation information.

[0020] In the following description, some other features will be partly stated. When examining the following content and the drawings, those skilled in the art will partly discover these features, or may learn these features through production or application. Through practicing or using various aspects of the methods, tools, and combinations listed in the detailed examples described later, the features in the current application can be implemented and obtained. Brief Description of the Drawings

[0021] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application.

[0022] Figure 1 It is a flowchart of a method for meteorological quality data analysis based on multi-source data fusion and AI provided by an embodiment of the present application.

[0023] Figure 2 It is a schematic diagram of the composition of a computer system provided by an embodiment of the present application. Detailed Embodiments

[0024] The following describes the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation part of the embodiments of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.

[0025] In the embodiments of the present application, the execution subject of the meteorological quality data analysis method based on multi-source data fusion and AI is a computer system, including but not limited to servers, personal computers, laptops, tablets, smart phones, etc. The server includes but not limited to a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of computers or network servers in cloud computing. Among them, cloud computing is a type of distributed computing, which consists of a group of loosely coupled computer sets to form a super virtual computer. Among them, the computer system can run independently to implement the present application, or can be connected to the network and implement the present application through interaction with other computer systems in the network. Among them, the network where the computer system is located includes but not limited to the Internet, wide area network, metropolitan area network, local area network, VPN network, etc.

[0026] Embodiments of the present application provide a meteorological quality data analysis method based on multi-source data fusion and AI, as Figure 1 shown, the method includes:

[0027] Step S10: Obtain the meteorological data of the system to be analyzed, and determine the observed data of the site to be analyzed related to the meteorological data of the system to be analyzed.

[0028] In step S10, the meteorological data of the system to be analyzed is obtained and the observed data of the relevant site to be analyzed is determined. Regarding obtaining the meteorological data of the system to be analyzed, the computer system needs to extract the required data from a specific meteorological system. Taking the China Land Data Assimilation System as an example, this system is a large data processing system integrating various meteorological observation data and model data. The computer system can obtain detailed meteorological data about a specific target area (such as a city, region or country) by accessing this system. These data can include but not limited to various meteorological elements such as temperature, humidity, wind speed, wind direction, air pressure, etc. After obtaining the system meteorological data, the computer system also needs to find the actual observed data related to these data, and these observed data usually come from meteorological observation stations distributed at different geographical locations. Each observation station regularly collects and records various meteorological information, such as rainfall, sunshine duration, emission data (such as industrial or traffic emissions), air pollutant concentration data, etc. By comparing and analyzing with the system meteorological data, these site observed data can provide important references and verifications for subsequent data analysis.

[0029] For example, assume that a computer system obtains hourly temperature and humidity data of a certain city from the China Land Data Assimilation System. To gain a more comprehensive understanding of the meteorological conditions in this city, the computer system also needs to find the actual observational data recorded by meteorological observation stations in the city and its surrounding areas. These observational data may include information such as rainfall, wind direction and speed, emission data, air pollutant concentration data, etc. By combining the system meteorological data with the station observational data, the computer system can conduct a more in-depth analysis and prediction of the meteorological conditions in this city.

[0030] Step S20: Load the system meteorological data to be analyzed and the station observational data to be analyzed into the trained target multi-source data visibility recognition model to obtain the visibility analysis and recognition information output by the target multi-source data visibility recognition model, where the target multi-source data visibility recognition model is a machine learning model obtained by refining and tuning the multi-source data visibility recognition model after basic training.

[0031] In step S20, the previously obtained meteorological data and observational data are loaded into a pre-trained machine learning model to obtain the analysis and recognition information about visibility. First, the computer system integrates the system meteorological data to be analyzed and the station observational data to be analyzed into a format that the model can process. The system meteorological data contains information in multiple dimensions such as temperature, humidity, air pressure, etc., while the station observational data may contain real-time weather phenomenon records such as rainfall, wind speed, emission data, etc. These data need to be formatted into the form of feature vectors or arrays for input into the machine learning model. Then, the computer system loads these formatted data into the trained target multi-source data visibility recognition model. This model has learned how to extract key features from general meteorological data and station observational data and predict visibility during the basic training phase. However, in order to more accurately analyze the specific system meteorological data and station observational data, the model has also undergone refinement and tuning, that is, fine-tuning. Fine-tuning is to further optimize the model according to a specific task and dataset to improve its prediction performance in a specific scenario.

[0032] When implementing step S20, the target multi-source data visibility recognition model can be various types of machine learning models, such as deep neural networks (DNNs), convolutional neural networks (CNNs), or long short-term memory networks (LSTMs), etc. These models differ in structure and function, but all aim to make predictions by learning complex patterns in the data. Taking CNN as an example, it is particularly suitable for processing data with a grid structure, such as images or time series data. In this case, meteorological data and observation data can be regarded as two-dimensional or three-dimensional grid structures, and CNN can predict visibility by learning their spatial features. Finally, the target multi-source data visibility recognition model outputs visibility analysis and recognition information. This information may be the specific value, level, or classification label of visibility, etc., depending on the model design and task requirements. This information is of great value for understanding meteorological conditions, formulating weather forecasts, or making meteorology-related decisions.

[0033] For example, suppose a computer system obtains the system meteorological data and station observation data of a certain city within a week and loads these data into a fine-tuned target multi-source data visibility recognition model. This model may be a deep neural network that has been initially trained on a large amount of historical meteorological data and fine-tuned on the meteorological data of similar cities. The results output by the model may be the predicted visibility values for each day within a week in this city, in kilometers. These predicted values can be used to guide decision-making in fields such as aviation, transportation, or environmental monitoring.

[0034] It can be understood that the initial training of the multi-source data visibility recognition model, that is, the pre-training process, directly affects the recognition effect of the model. The focus of this application also lies in this initial training process. Specifically, the initial training process of the multi-source data visibility recognition model can include the following steps:

[0035] Step S110: Obtain a set of basic training knowledge templates. Each basic training knowledge template in the set of basic training knowledge templates includes a system meteorological data knowledge template, a station observation data knowledge template related to the system meteorological data knowledge template, and a prior visibility label for the template.

[0036] The set of knowledge templates for basic training contains various features that the model needs to learn and their corresponding labels. In this step, the computer system obtains a large amount of systematic meteorological data and related site observation data from various sources. These data are organized into individual basic training knowledge templates, each of which contains a systematic meteorological data knowledge template, a related site observation data knowledge template, and a prior visibility label for the template. The systematic meteorological data knowledge template is a data structure containing various meteorological elements, such as temperature, humidity, wind speed, wind direction, etc. This data can come from various meteorological data sources, such as meteorological satellites, meteorological radars, meteorological observation stations, etc. Each knowledge template is a snapshot of meteorological data at a specific time point and location.

[0037] The site observation data knowledge template is the actual observation data corresponding to the systematic meteorological data. It includes various meteorological phenomena observed at the site, such as rainfall, cloud cover, emission data, pollutant concentration, etc. This data usually comes from meteorological observation stations distributed at different geographical locations and is an important supplement and verification of the systematic meteorological data. The prior visibility label for the template is the true visibility label corresponding to each knowledge template. This label is known and represents the visibility value or level actually observed at that time point and location. This label is the target variable in the model training process, and the model needs to learn the features of the meteorological data and observation data to predict this label.

[0038] For example, suppose there is a set of basic training knowledge templates that contains 1000 knowledge templates. Each knowledge template contains the systematic meteorological data (such as temperature 25°C, humidity 60%, wind speed 5 m / s, etc.), site observation data (such as rainfall 0 mm, cloud cover 50%, Pm2.5 459, etc.) at a certain time point and location, and the corresponding prior visibility label (such as visibility level "good"). This set is used as the basic training data for the machine learning model, and the model learns the features and patterns in these data to predict the visibility label corresponding to new and unknown meteorological data.

[0039] In one implementation, in step S110, obtaining the set of basic training knowledge templates may specifically include:

[0040] Step S111: Generate a set of historical meteorological observation information corresponding to each visibility monitoring area according to the historical meteorological observation information corresponding to each visibility monitoring area; where each historical meteorological observation information includes a historical systematic meteorological data, historical site observation data related to the historical systematic meteorological data, and an original prior visibility label.

[0041] Specifically, the computer system accesses a database or data source storing historical meteorological observation information. This historical meteorological observation information typically contains meteorological data for each visibility monitoring area over a long time range (such as the past few years). Each visibility monitoring area may have its unique climate characteristics and meteorological patterns, so it is important to process the data for these areas separately. In step S111, the computer system processes this historical meteorological observation information to generate a set of historical meteorological observation information corresponding to each visibility monitoring area. These information sets contain a large amount of historical systematic meteorological data, related historical site observation data, and original visibility prior labels.

[0042] Historical systematic meteorological data refers to the meteorological data collected through meteorological systems (such as satellites, radars, weather stations, etc.), which reflect the meteorological conditions at a specific time and location, such as temperature, humidity, wind speed, etc.

[0043] Historical site observation data are data from actual meteorological observation sites, which provide real-time meteorological observation information about a specific location, such as rainfall, wind direction, emission data, air pollutant concentration, etc.

[0044] The original visibility prior labels are the true visibility values or grades corresponding to these historical data, which are known and will be used as target labels during the training of the machine learning model.

[0045] For example, assume there is a visibility monitoring area A, and its historical meteorological observation information includes daily meteorological data and corresponding visibility grades for the past year. The computer system will organize these data into a set of historical meteorological observation information, where each data point contains the daily systematic meteorological data (such as temperature, humidity, etc.), site observation data (such as rainfall, wind speed, etc.), and the corresponding visibility grade label. This set will be used as the basic data for subsequent machine learning model training.

[0046] Through the processing in step S111, the computer system can generate a series of sets of historical meteorological observation information specific to each visibility monitoring area, which provide rich and targeted data resources for subsequent model training.

[0047] Among them, as an implementation manner, in step S111, according to the historical meteorological observation information corresponding to each visibility monitoring area respectively, generating a set of historical meteorological observation information corresponding to each visibility monitoring area may specifically include:

[0048] Step S1111: According to each initial meteorological observation information within the preset time interval in each visibility monitoring area, an initial meteorological observation information set is formed. Each initial meteorological observation information includes an initial system meteorological data, historical site observation data related to the initial system meteorological data, and an original visibility prior label.

[0049] In step S1111, the computer system forms an initial meteorological observation information set based on the initial meteorological observation information in each visibility monitoring area within the preset time interval. These information sets are the basis for subsequent data processing and model training. Specifically, for each visibility monitoring area, the computer will collect all the initial meteorological observation information within the preset time interval (such as the past year, five years, or ten years) of the monitoring time in that area. These information usually includes initial system meteorological data (such as temperature, humidity, wind speed, etc.), historical site observation data related to the initial system meteorological data (such as rainfall, air pressure, etc. at a specific location), and an original visibility prior label (i.e., the actual visibility level or category at that time point).

[0050] For example, assume there is a visibility monitoring area A. The computer system will collect the initial meteorological observation information per hour in area A in the past year. These information may include the hourly temperature, humidity, and wind speed readings, as well as the rainfall and air pressure data at a specific site within that hour. In addition, each observation information will be attached with an original visibility prior label indicating the actual visibility level (such as "excellent", "good", "moderate", "poor", etc.) within that hour. After collecting these data, the computer will organize them into an initial meteorological observation information set. Each element in this set is a complete observation information, containing system meteorological data, site observation data, and visibility level label. Such a set facilitates subsequent data screening, feature extraction, and model training.

[0051] It should be noted that the "initial meteorological observation information" here refers to the original data before any screening or processing. In subsequent steps (such as step S1112), these initial information will be further processed and screened to remove redundant or irrelevant information, improving the data quality and the accuracy of model training.

[0052] Step S1112: Extract an initial meteorological observation information one by one from the initial meteorological observation information set, and complete the following single screening according to the extracted initial meteorological observation information until there is no unextracted initial meteorological observation information in the initial meteorological observation information set:

[0053] Step S11121: Determine the meteorological data characterization vector matching factors (i.e., the similarity between features) between the initial system meteorological data in the extracted initial meteorological observation information and the initial system meteorological data in other initial meteorological observation information in the initial meteorological observation information set respectively;

[0054] Step S11122: Remove from the initial meteorological observation information set other initial meteorological observation information whose meteorological data characterization vector matching factors with the extracted initial meteorological observation information reach the matching factor threshold;

[0055] Step S11123: Take the initial meteorological observation information in the initial meteorological observation information set after single screening as historical meteorological observation information;

[0056] Step S11124: According to the visibility monitoring area to which each historical meteorological observation information belongs, obtain the historical meteorological observation information set corresponding to each visibility monitoring area through cluster analysis (i.e., clustering) of each historical meteorological observation information.

[0057] The purpose of Step S1112 is to improve the data quality and the accuracy of subsequent model training by removing redundant or similar observation information. Specifically, the computer system processes each initial meteorological observation information in the initial meteorological observation information set one by one according to a preset process. For each extracted initial meteorological observation information, first determine the initial system meteorological data it contains, and calculate the meteorological data characterization vector matching factors between these data and the initial system meteorological data in other initial meteorological observation information. This matching factor is essentially a similarity measure used to quantify the similarity between two system meteorological data.

[0058] For example, assume there are two initial meteorological observation information A and B, which respectively contain system meteorological data of temperature, humidity, and wind speed. The computer system will calculate the similarity between A and B in these dimensions to obtain a meteorological data characterization vector matching factor. If this matching factor is higher than the preset matching factor threshold, it indicates that A and B are very similar in meteorological characteristics and may be redundant.

[0059] In this case, other initial meteorological observation information with too high similarity (i.e., the matching factor reaches or exceeds the threshold) to the currently extracted initial meteorological observation information is removed from the initial meteorological observation information set. The purpose of doing this is to reduce data redundancy and avoid overfitting or unnecessary computational overhead caused by overly similar data points during model training. After this round of screening, the remaining initial meteorological observation information is considered to be more representative, and they will be retained as historical meteorological observation information. These historical meteorological observation information not only contains the original system meteorological data and site observation data, but also removes redundant information through the screening process, making the data set more refined and efficient.

[0060] Finally, according to the visibility monitoring areas to which each historical meteorological observation information belongs, the computer will perform cluster analysis (such as clustering algorithms) on this information. The purpose of clustering is to group similar historical meteorological observation information into the same set, so as to construct more targeted models or conduct more in-depth data analysis for each visibility monitoring area in the follow-up. Through such a processing flow, step S1112 effectively improves the data quality and lays a solid foundation for subsequent data utilization.

[0061] Specifically, in step S11121, its core task is to calculate the meteorological data representation vector matching factor, that is, the feature similarity, between the initial system meteorological data in the extracted initial meteorological observation information and the meteorological data in other observation information.

[0062] Specifically, when the computer system processes this step, it first selects an initial meteorological observation information as a reference. This reference information contains specific initial system meteorological data, such as numerical values in multiple dimensions such as temperature, humidity, and wind speed. These data can be regarded as a multi-dimensional vector, where each dimension corresponds to a meteorological feature. Then, it traverses all other observation information in the initial meteorological observation information set and calculates the similarity between the initial system meteorological data contained in each of them and the meteorological data in the reference information one by one. The calculation of similarity is usually based on distance or correlation metrics between vectors, such as Euclidean distance, cosine similarity, etc. These metrics can quantify the closeness or direction consistency of two vectors in multi-dimensional space.

[0063] For example, assume that the initial system meteorological data of the reference information is a three-dimensional vector [25, 60, 5], representing the numerical values of temperature, humidity, and wind speed respectively. And the initial system meteorological data of another observation information in the set is [26, 58, 6]. The computer will calculate the similarity between these two vectors, such as using the Euclidean distance formula to calculate the distance between them. The closer the distance, the higher the similarity; conversely, the farther the distance, the lower the similarity.

[0064] In this process, the computer generates a similarity value for each observation information compared with the reference information, that is, the meteorological data characterization vector matching factor. This factor is essentially a numerical value used to quantify the similarity degree of two observation information in meteorological characteristics.

[0065] Through the processing of step S11121, the computer system can assign a similarity value to each observation information, and these values will be used to screen redundant information or perform other data processing operations in subsequent steps. This step is crucial for improving data quality and optimizing the model training process because it helps reduce the potential impact of similar or duplicate data on the model performance.

[0066] In step S11122, the computer system screens the initial meteorological observation information set according to a preset matching factor threshold to remove other observation information that is too similar to the currently extracted initial meteorological observation information in meteorological characteristics. Specifically, when the computer system completes step S11121, a set of meteorological data characterization vector matching factors is obtained, and these factors quantify the similarity between the currently extracted initial meteorological observation information and other observation information. Next, the computer compares these matching factors with the preset matching factor threshold. The matching factor threshold is a preset numerical value used to determine whether two observation information are similar enough to be regarded as redundant. If the matching factor between a certain observation information and the currently extracted observation information is higher than or equal to this threshold, then the computer will consider that these two observation information are too similar in meteorological characteristics and there is redundancy.

[0067] In this case, the computer removes other observation information in the initial meteorological observation information set that has too high a similarity (i.e., the matching factor reaches or exceeds the threshold) with the currently extracted observation information. The purpose of doing this is to reduce data redundancy and improve the data quality and the accuracy of subsequent model training.

[0068] For example, assume that the currently extracted initial meteorological observation information A contains data of temperature 25°C, humidity 60%, and wind speed 5 m / s. In the initial meteorological observation information set, there is another observation information B with a temperature of 26°C, humidity of 58%, and wind speed of 6 m / s. If the calculated meteorological data characterization vector matching factor shows that the similarity between A and B is very high and exceeds the preset matching factor threshold, then the computer will remove the observation information B from the set to avoid duplicate or redundant data in subsequent processing.

[0069] Through the processing of step S11122, the data in the initial meteorological observation information set is further refined and optimized, providing a more accurate and efficient data basis for subsequent data analysis and model training.

[0070] Step S11123 involves officially recognizing the information in the initial meteorological observation information set after singulation screening as historical meteorological observation information. This step is an important part of the data cleaning and preparation process, aiming to ensure that the data used is unique and representative, thereby improving the accuracy and reliability of subsequent analysis. Specifically, when the computer system executes step S11123, it will first review the completed singulation screening process. In this process, the computer has removed other observation information that is too similar in meteorological characteristics to the currently extracted initial meteorological observation information according to the preset matching factor threshold. In this way, the remaining observation information in the screened set is relatively unique and representative in terms of meteorological characteristics. Next, these screened initial meteorological observation information are officially marked as historical meteorological observation information. This means that these information not only contain the original meteorological data, such as temperature, humidity, wind speed, etc., but also have undergone a strict data cleaning and screening process and are considered high-quality and reliable data that can be used for subsequent data analysis, model training, or other related applications.

[0071] For example, assume that the initial meteorological observation information set contains 100 pieces of observation information. After singulation screening, 20 redundant pieces of information that are too similar in meteorological characteristics to the currently extracted observation information are removed. Then, in step S11123, the computer will officially recognize the remaining 80 unique and representative observation information as historical meteorological observation information and store them in an appropriate data structure for subsequent use. Through this step of processing, the computer system not only optimizes the data set, removes redundant information, but also provides a more accurate and reliable data basis for subsequent data analysis and model training.

[0072] Step S11124 involves performing cluster analysis based on the visibility monitoring area to which the historical meteorological observation information belongs to form a set of historical meteorological observation information corresponding to each area. The purpose of this step is to classify the observation data under similar meteorological conditions for subsequent regional analysis and model training.

[0073] Specifically, when the computer system executes step S11124, it will first identify the visibility monitoring area to which each historical meteorological observation information belongs. These information have been marked with corresponding area labels in the previous processing steps. Then, the computer will classify the historical meteorological observation information according to these area labels. Next, for each visibility monitoring area, the computer will perform cluster analysis, usually using clustering algorithms. Clustering algorithms are an unsupervised learning method that can group similar data points into the same cluster. In this scenario, similar data points refer to historical meteorological observation information that is similar in meteorological characteristics. Through cluster analysis, the computer can group these information according to the similarity of meteorological characteristics, forming a set of historical meteorological observation information corresponding to each visibility monitoring area.

[0074] For example, suppose there are two visibility monitoring areas A and B, corresponding to different regions of the city. After processing in step S11123, a set of historical meteorological observation information is obtained. Now, in step S11124, the computer will divide this information into two groups according to the area labels to which they belong. Then, for each group of information in area A and area B, the computer will respectively apply clustering algorithms, such as the K-means algorithm, to further divide them into several clusters, and the observation information within each cluster has high similarity in meteorological characteristics. Finally, through the processing of step S11124, a set of historical meteorological observation information corresponding to each visibility monitoring area is obtained. These sets are not only classified by area but also further subdivided according to the similarity of meteorological characteristics within each area. Such a data structure provides strong support for subsequent regional meteorological analysis, model training, and prediction.

[0075] Step S112: For each visibility monitoring area, based on the historical meteorological observation information corresponding to the visibility monitoring area, after example-driven learning, obtain the target auxiliary classifier corresponding to the visibility monitoring area.

[0076] Step S112 involves supervised learning of the historical meteorological observation information of a specific visibility monitoring area to generate the target auxiliary classifier corresponding to that area. This classifier will be used in subsequent steps to assist in generating more accurate visibility analysis and recognition information. Specifically, the computer system will perform example-driven learning for each visibility monitoring area using the set of historical meteorological observation information corresponding to it. This learning method is a supervised learning method that requires the training data to contain input features (here are system meteorological data and site observation data) and corresponding target labels (here are the original visibility prior markings).

[0077] In this process, the computer system will select a suitable machine learning algorithm or neural network structure to construct the target auxiliary classifier. This selection depends on the characteristics of the data, the complexity of the task, and the available computing resources. For example, if the data is linearly separable, a simple linear classifier (such as logistic regression) may be sufficient; if the data has complex non-linear relationships, then a more complex model (such as a deep neural network) may be required, which needs to be adaptively selected according to the actual data structure in the knowledge template. Once the appropriate model structure is selected, the computer system will use the data in the historical meteorological observation information set to train this model. The training process includes passing the input features to the model, comparing the output of the model with the actual target labels, and then adjusting the parameters of the model according to the comparison results to minimize the prediction error.

[0078] After sufficient iterative training, the model will gradually learn the mapping relationship from the input features to the target labels, thus becoming a target auxiliary classifier that can accurately predict the visibility level. Although this classifier is called a "pseudo-classifier" here, it is actually a trained and effective tool that can be used to assist in generating visibility analysis and identification information.

[0079] For example, suppose there is a visibility monitoring area A, and its historical meteorological observation information set contains 1000 data points. Each data point includes system meteorological data, site observation data, and the corresponding visibility level label. The computer system can choose to use a deep neural network as the model structure of the target auxiliary classifier and use these data points to train this network. After training, this network can predict the corresponding visibility level based on the new system meteorological data and site observation data. Although this prediction result may not be 100% accurate, it provides valuable reference information for subsequent visibility analysis.

[0080] Step S113: For each historical system meteorological data, according to each target auxiliary classifier, determine each auxiliary visibility classification information related to the historical system meteorological data, and use the original visibility prior label related to the historical system meteorological data and the augmented label set of each auxiliary visibility classification information as the template visibility prior label related to the historical system meteorological data.

[0081] Step S113 involves using the target auxiliary classifier generated in the previous steps to enhance the original data labels, thereby enriching the training data set and improving the generalization ability of the model. The core of this step is to expand the label set of each historical system meteorological data through the prediction results of multiple auxiliary classifiers.

[0082] Specifically, the computer system traverses each historical system meteorological data point. For each data point, it inputs it into each target auxiliary classifier trained in the previous step S112. These auxiliary classifiers generate corresponding auxiliary visibility classification information based on the input meteorological data features, that is, their prediction results for the visibility level to which the data point belongs.

[0083] It should be noted that the prediction results of these auxiliary classifiers may not be exactly the same because they may be based on different algorithms, model structures, or subsets of training data. Therefore, the auxiliary visibility classification information they provide can be regarded as a supplement to the original visibility prior label or an interpretation from another perspective.

[0084] Next, the computer system combines the original visibility prior label of each historical system meteorological data point with the auxiliary visibility classification information generated by each auxiliary classifier to form an augmented label set. This augmented label set not only contains the original true labels but also incorporates the pseudo-labels predicted by the model, thus providing richer supervision information for the model to learn. For example, suppose there is a historical system meteorological data point X with an original visibility prior label of "good". At the same time, there are two target auxiliary classifiers A and B. Classifier A predicts the visibility level of data point X as "good" based on its features, while classifier B predicts it as "average". In this case, the augmented label set of data point X will include the original label "good" and the two auxiliary labels "good" and "average".

[0085] In this way, step S113 generates a more comprehensive and diverse label set for each historical system meteorological data point, which helps the machine learning model learn a more complex and detailed mapping relationship from data features to labels during the training process. Ultimately, this will help improve the prediction accuracy and generalization ability of the model on unseen data.

[0086] In one implementation, step S113, to determine each auxiliary visibility classification information related to the historical system meteorological data according to each target auxiliary classifier, may specifically include:

[0087] Step S1131: Determine the target visibility monitoring area corresponding to the historical system meteorological data, determine the target auxiliary classifier corresponding to the target visibility monitoring area, and obtain the other target auxiliary classifiers corresponding to each other visibility monitoring area except the target visibility monitoring area.

[0088] The main task of step S1131 is to determine the target visibility monitoring area corresponding to the historical system meteorological data, find the corresponding target auxiliary classifier accordingly, and also obtain the auxiliary classifiers corresponding to other visibility monitoring areas except this target area. The purpose of this step is to prepare for using these classifiers to determine the relevant auxiliary visibility classification information in the historical system meteorological data in the subsequent steps.

[0089] Specifically, when the computer system executes step S1131, it first analyzes the metadata or tags of the historical system meteorological data to determine which specific visibility monitoring area these data belong to. This determination process may be based on the geographical location information in the data, the identifier of the monitoring station, or other relevant identifiers.

[0090] Once the target visibility monitoring area corresponding to the historical system meteorological data is determined, the computer will further search for and determine the target auxiliary classifier corresponding to this area. This target auxiliary classifier is pre-trained and specifically used to process the meteorological data of this specific area, and can output classification information related to the visibility of this area. In addition, the computer will also obtain the other target auxiliary classifiers corresponding to all other visibility monitoring areas except the target visibility monitoring area. These other target auxiliary classifiers are also pre-trained, corresponding to different monitoring areas respectively, and each has the ability to process the meteorological data of the corresponding area and output visibility classification information.

[0091] For example, assume there are three visibility monitoring areas A, B, and C, and each area has its own corresponding auxiliary classifier A', B', and C'. If the historical system meteorological data is determined to belong to area A, then A' is the target auxiliary classifier, while B' and C' are the other target auxiliary classifiers. The computer will obtain and use these classifiers to further analyze the historical system meteorological data to determine the auxiliary visibility classification information related to each area.

[0092] Through the processing of step S1131, the computer system lays a foundation for using each auxiliary classifier to determine the relevant auxiliary visibility classification information in the historical system meteorological data in the subsequent step S1132.

[0093] S1132: Determine the relevant auxiliary visibility classification information in the historical system meteorological data respectively according to each of the other target auxiliary classifiers.

[0094] Step S1132 involves using each other target auxiliary classifier to determine relevant auxiliary visibility classification information in the historical system meteorological data. The purpose of this step is to extract visibility-related classification information from meteorological data in multiple perspectives and regions, so as to more comprehensively understand the impact of meteorological conditions on visibility. Specifically, when the computer system executes step S1132, it first obtains each other target auxiliary classifier determined in step S1131. These classifiers are pre-trained machine learning models, and each model corresponds to a specific visibility monitoring area and has the ability to process meteorological data in that area. These models may be constructed based on algorithms such as decision trees, support vector machines, and neural networks, and they can output corresponding visibility classification information according to the input meteorological feature vectors.

[0095] Next, the computer will input the historical system meteorological data into these other target auxiliary classifiers. Each classifier will process and analyze the input meteorological data according to its own training data and algorithm logic. In this process, the classifier will extract features related to visibility, such as temperature, humidity, wind speed, etc., and judge the visibility category to which the meteorological data belongs based on the values of these features. For example, assume there is an auxiliary classifier based on a neural network that receives a meteorological data vector containing features such as temperature, humidity, and wind speed as input. The neural network will perform non-linear transformation and combination on these features according to its internal weights and biases, and finally output a label or probability distribution representing the visibility category. This label or probability distribution is the judgment result of the classifier on the relevant visibility classification information in the historical system meteorological data.

[0096] Through the processing of step S1132, the computer system can obtain auxiliary visibility classification information from multiple perspectives and regions. This information can be used in subsequent data fusion, model optimization, or decision support tasks to help people more accurately understand and predict the impact of meteorological conditions on visibility. At the same time, these classification information can also be used as one of the input features of other meteorological analysis or prediction models to improve the performance and accuracy of the models.

[0097] In one implementation, in step S113, the auxiliary visibility classification information includes an auxiliary visibility mark and a predicted support coefficient predicted for the auxiliary visibility mark; then, in step S113, after determining each auxiliary visibility classification information related to the historical system meteorological data, it further includes:

[0098] Step S113a: Obtain an initial overlapping mark set between the auxiliary visibility marks included in each auxiliary visibility classification information and the original visibility prior marks related to the historical system meteorological data, and determine the predicted support coefficients corresponding to each visibility prior mark in the initial overlapping mark set.

[0099] In step S113a, the computer system needs to process the relationship between the auxiliary visibility classification information and the original visibility prior label to determine the overlapping part between the two, and further analyze the prediction support coefficients of these overlapping labels.

[0100] Specifically, when the computer system executes step S113a, it first obtains the auxiliary visibility labels included in each piece of auxiliary visibility classification information. These auxiliary visibility labels are obtained based on the analysis results of different auxiliary classifiers on historical system meteorological data, and they represent the visibility conditions predicted according to specific algorithms or models. Then, the computer compares these auxiliary visibility labels with the original visibility prior labels related to the historical system meteorological data. The original visibility prior labels are visibility labels obtained based on meteorological observation data or other prior knowledge without using the auxiliary classifier. The purpose of the comparison is to find the overlapping part between the auxiliary visibility labels and the original visibility prior labels, that is, the labels that both consider to belong to the same visibility category.

[0101] After determining the overlapping labels, the computer further analyzes the prediction support coefficients of these overlapping labels. The prediction support coefficient represents the confidence level or probability of the prediction result of the auxiliary classifier for a certain auxiliary visibility label. Usually, this coefficient is a probability value between 0 and 1, and the closer it is to 1, the more confident the auxiliary classifier is in the prediction result of that label.

[0102] For example, assume that an auxiliary visibility classification information contains an auxiliary visibility label marked as "low visibility" with a prediction support coefficient of 0.9. At the same time, there is also an original visibility prior label of "low visibility" in the historical system meteorological data. In this case, these two labels form an overlapping label, and since the prediction support coefficient of the auxiliary visibility label is relatively high (0.9), it can be considered that this overlapping label is relatively reliable.

[0103] By executing step S113a, the computer system can determine the set of overlapping labels between the auxiliary visibility classification information and the original visibility prior labels, and understand the prediction support coefficients of these overlapping labels. This provides an important basis for further processing and applying these labels in subsequent steps (such as steps S113b and S113c).

[0104] Step S113b: Remove the visibility prior labels in the original visibility prior labels related to the historical system meteorological data whose prediction support coefficients reach the label filtering index in the initial overlapping label set, and obtain the processed original visibility prior labels.

[0105] Step S113b is responsible for removing those tags from the original visibility prior tags of historical system meteorological data that overlap with the auxiliary visibility classification information and whose prediction support coefficients reach or exceed a specific tag filtering criterion. The purpose of this step is to ensure that the finally used visibility tags are sufficiently accurate and reliable.

[0106] Specifically, when the computer system executes step S113b, it first reviews the initial set of overlapping tags determined in step S113a. This set contains those tags that appear in both the original visibility prior tags and the auxiliary visibility classification information, and their prediction support coefficients have also been calculated. Next, according to the preset tag filtering criterion, each visibility prior tag in the initial set of overlapping tags is screened. The tag filtering criterion is usually a threshold value used to determine whether the prediction support coefficient is high enough to ensure the reliability of the tag. For example, if the tag filtering criterion is set to 0.8, then only the visibility prior tags with a prediction support coefficient higher than or equal to 0.8 will be considered for retention.

[0107] During the screening process, the computer checks the prediction support coefficient of each visibility prior tag in the initial set of overlapping tags one by one. If the prediction support coefficient of a certain tag is lower than the tag filtering criterion, then this tag will be regarded as insufficiently reliable and removed from the original visibility prior tags.

[0108] For example, assume that there is a visibility prior tag labeled "medium visibility" in the initial set of overlapping tags, and its prediction support coefficient is 0.75. If the tag filtering criterion is set to 0.8, then the prediction support coefficient of this "medium visibility" tag is lower than the filtering criterion, so it will be removed from the original visibility prior tags. By executing step S113b, the computer system can filter out those visibility prior tags with lower prediction support coefficients and potentially inaccurate ones, thus ensuring higher accuracy and reliability of the subsequently used visibility tags.

[0109] Step S113c: For the processed original visibility prior tags, proceed to the step of taking the original visibility prior tags related to historical system meteorological data and the augmented tag sets of each auxiliary visibility classification information as the template visibility prior tags related to historical system meteorological data.

[0110] In step S113c, the computer system combines the processed original visibility prior tags with the augmented tag sets of the auxiliary visibility classification information to form a more comprehensive and accurate set of template visibility prior tags.

[0111] Specifically, when the computer system executes step S113c, it first obtains the original visibility prior markers processed in step S113b. These markers have been compared and screened with the auxiliary visibility classification information, and those markers with insufficient prediction support coefficients and possible inaccuracies have been removed, so they have high reliability and accuracy. Then, the computer system merges these processed original visibility prior markers with the augmented label sets of each auxiliary visibility classification information. The augmented label set may contain other relevant information or markers in addition to the auxiliary visibility markers in the auxiliary visibility classification information, which can provide additional visibility condition information. By merging these markers, the computer system can form a more rich and comprehensive set of template visibility prior markers. This set of template visibility prior markers will play an important role in subsequent meteorological data analysis. It can be used as one of the input features for training machine learning models (such as decision trees, support vector machines, neural networks, etc.) to train and optimize the model's prediction ability for visibility conditions. At the same time, it can also be directly used for the classification and annotation of meteorological data, providing more accurate and reliable visibility information for applications such as weather forecasting and climate research.

[0112] For example, to illustrate, assume that the processed original visibility prior markers contain a marker labeled "high visibility", and the augmented label set of a certain auxiliary visibility classification information contains a numerical feature representing the visibility condition (such as visibility distance). When executing step S113c, the computer system will merge these two pieces of information to form a template visibility prior marker that contains both the "high visibility" marker and the specific visibility value. Such a marker is both semantically clear and contains specific quantitative information, which has higher value for subsequent meteorological data analysis.

[0113] Step S114: Based on the historical system meteorological data, the template visibility prior markers related to the historical system meteorological data, and the historical site observation data, form a basic training knowledge template;

[0114] Step S115: Based on each basic training knowledge template, form a set of basic training knowledge templates.

[0115] Step S114 involves integrating different types of data (historical system meteorological data, template visibility prior markers, historical site observation data) into a unified format of basic training knowledge template. This template will provide structured input for subsequent model training. Specifically, the computer system will combine each corresponding set of historical system meteorological data, relevant template visibility prior markers, and historical site observation data according to a predetermined data structure and format to form a complete basic training knowledge template. This template is a multi-dimensional data structure that contains all the necessary information for machine learning model training.

[0116] For example, a basic training knowledge template may contain the following information: system meteorological data (such as temperature, humidity, wind speed, etc.) at a specific time point, which are represented in the form of numerical values or vectors; template visibility prior markers corresponding to that time point, which is an augmented set of labels including the original visibility level and the visibility level predicted by the auxiliary classifier; and historical site observation data at the same time point, such as rainfall, air pressure, etc. By integrating these different types of data into a unified template, step S114 ensures that the machine learning model can receive and process input data in a standardized manner, thereby improving the training efficiency and the accuracy of the model.

[0117] Step S115 is carried out based on step S114 and involves combining multiple individual basic training knowledge templates into a larger set, namely the basic training knowledge template set. This set will serve as the main data source for machine learning model training. Specifically, the computer system will traverse all the generated basic training knowledge templates and add them to the basic training knowledge template set one by one. This process may involve data storage, indexing, and management to ensure that each template in the set can be effectively accessed and used. The construction of the basic training knowledge template set is a key preparatory step before machine learning model training. Through this set, the model can be exposed to a large number of diverse and representative training samples, thereby learning the complex mapping relationship from input data to target labels. The learning of this mapping relationship is the core goal of machine learning model training, which determines the prediction ability and generalization performance of the model on future unseen data.

[0118] Step S120: Based on the system meteorological data knowledge templates included in each knowledge template in the basic training knowledge template set, perform multiple endogenous learning (i.e., self-supervised training) on the meteorological data vector representation component to obtain the initial meteorological data vector representation component after learning.

[0119] In step S120, the meteorological data vector representation component is endogenously learned using the basic training knowledge template set. The purpose of this stage is to optimize and enhance the feature extraction ability of the meteorological data vector representation component, so as to capture key information from the original meteorological data more accurately. Specifically, in step S120, the computer system first accesses the pre-constructed basic training knowledge template set. This set contains multiple knowledge templates, each of which is constructed based on the system meteorological data and contains key features and information related to specific meteorological phenomena or conditions. These knowledge templates provide learning standards and references for the vector representation component in the system meteorological data processing. Then, these knowledge templates are used to endogenously learn the meteorological data vector representation component, that is, self-supervised training. Self-supervised training is a training method that uses the structure or relationship of the data itself as a supervision signal. It does not require additional manually labeled data, but discovers information from within the data for learning. In this process, the vector representation component will try to extract features from the meteorological data and compare and adjust them with the features in the knowledge templates. Through continuous iteration and optimization, the accuracy and efficiency of its feature extraction are gradually improved.

[0120] For example, assume that there is a template in the knowledge template set that describes the characteristics of the system meteorological data under the "high temperature and dry" weather conditions, including the value ranges or patterns of key indicators such as temperature, humidity, and wind speed. During self-supervised training, the meteorological data vector representation component will try to extract these features from the input meteorological data and compare them with the features in the knowledge template. If the extracted features do not match or have significant differences from the features in the template, then the vector representation component will optimize the feature extraction method by adjusting its internal parameters and structure to more accurately capture the meteorological data features under the "high temperature and dry" weather conditions.

[0121] Through the self-supervised training in step S120, the meteorological data vector representation component can gradually learn how to extract key features and information from the original meteorological data, providing strong support for subsequent meteorological data analysis, prediction, and decision-making. This training process is automated and can be iteratively optimized on a large amount of meteorological data, thus continuously improving the performance and accuracy of the vector representation component. The finally obtained initial meteorological data vector representation component after learning will be used in subsequent meteorological data processing tasks.

[0122] In one implementation, in step S120, when performing one endogenous learning on the meteorological data vector representation component, the following steps are completed:

[0123] Step S121: Extract a predetermined number of system meteorological data knowledge templates from the basic training knowledge template set, and through the meteorological data vector characterization component, extract a predetermined number of template meteorological data recording points associated with timestamps from each of the extracted system meteorological data knowledge templates.

[0124] In step S121, the computer system extracts a predetermined number of system meteorological data knowledge templates from the basic training knowledge template set and performs further data extraction and processing on these templates. The purpose of this step is to prepare the necessary data and templates for the subsequent training process.

[0125] Specifically, first access the basic training knowledge template set stored in memory or a database. This set contains a large number of system meteorological data knowledge templates, each of which is constructed based on historical meteorological data and relevant domain knowledge, reflecting the data characteristics and patterns under different meteorological phenomena or conditions.

[0126] Next, according to a preset extraction rule or algorithm, select a predetermined number of knowledge templates from the set. This predetermined number can be determined according to specific training requirements, computing resources, or time constraints. For example, if the training goal is to quickly verify the effectiveness of a new meteorological data vector characterization component, then a smaller number of knowledge templates can be selected for preliminary training; while if the goal is to build a high-performance meteorological prediction model, more knowledge templates may need to be extracted to obtain more comprehensive training data.

[0127] Once a predetermined number of knowledge templates are extracted, the computer uses the meteorological data vector characterization component to perform further data extraction on these templates. This process involves converting the original meteorological data (such as temperature, humidity, wind speed, etc.) into a digital or vector form that can be processed by the computer. For each knowledge template, the meteorological data vector characterization component extracts a predetermined number of template meteorological data recording points associated with timestamps. These recording points not only contain the numerical information of the meteorological data but also reflect the changes and associations of the data in the time series through timestamps. For example, a template meteorological data recording point can represent a combination of temperature, humidity, and wind speed data collected at a specific time point (such as 14:00:00 on April 1, 2023).

[0128] In this way, step S121 provides the basic data and structured input for subsequent endogenous learning. These data will be used to construct cohesive knowledge template pairs (positive sample pairs) and discrete knowledge template pairs (negative sample pairs), and then optimize the performance of the meteorological data vector characterization component through contrastive learning.

[0129] Among them, as an implementation manner, in step S121, when extracting a predetermined number of template meteorological data record points associated with time stamps from each of the extracted system meteorological data knowledge templates, one of the following steps is completed:

[0130] Step S121a: For each system meteorological data knowledge template extracted, according to the time span (i.e., duration) corresponding to the system meteorological data knowledge template, set the target node (i.e., target time stamp) corresponding to the system meteorological data knowledge template, and starting from the template meteorological data record point corresponding to the target node on the system meteorological data knowledge template, extract a predetermined number of template meteorological data record points associated with time stamps.

[0131] Alternatively, step S121b: For each system meteorological data knowledge template extracted, according to the time span corresponding to the system meteorological data knowledge template, divide the system meteorological data knowledge template into a predetermined number of knowledge template sub-meteorological data, and arbitrarily extract one template meteorological data record point from each of the knowledge template sub-meteorological data to obtain a predetermined number of template meteorological data record points associated with time stamps.

[0132] In step S121a, the computer system further processes each system meteorological data knowledge template extracted from the basic training knowledge template set to obtain a specific number of template meteorological data record points associated with time stamps. This process is crucial for ensuring the data quality and consistency of subsequent machine learning model training. Specifically, the computer system first identifies the time span of each system meteorological data knowledge template, that is, the time range or duration covered by the template. This time span is important because it determines the richness and variation range of the meteorological data within the template. For example, a knowledge template covering several days or weeks may contain data on various weather conditions, while a knowledge template covering only a few hours may only reflect specific weather phenomena within a short period. Next, according to the time span of each knowledge template, the computer will set one or more target nodes. These target nodes are specific time points selected within the time span for extracting meteorological data record points from the template. The selection of target nodes can be based on equal time intervals or on certain specific data characteristics or events. For example, if the time span of the knowledge template is one week, the target nodes may be set at 12 noon every day, or according to important moments of weather changes (such as the start and end times of a storm). Once the target nodes are determined, meteorological data record points associated with these target nodes are extracted from each knowledge template. These record points not only contain the specific values of meteorological data (such as temperature, humidity, wind speed, etc.), but also contain the time stamp information associated with each data point. The time stamp information is crucial for subsequent data analysis and model training because it allows the model to understand the changes and associations of data over time.

[0133] For example, assume that a system meteorological data knowledge template covers data from April 1st to April 7th, 2023, and there is a data recording point associated with 12:00 noon every day. In step S121a, the computer can select this time point every day as the target node and extract the meteorological data recording points at 12:00 noon every day within these seven days. In this way, a set of seven template meteorological data recording points associated with timestamps is obtained, and these recording points can be used for subsequent machine learning model training.

[0134] Step S121a ensures that the data recording points extracted from each system meteorological data knowledge template not only meet the quantity requirements but also are representative and relevant in terms of time. This provides high-quality training data for the subsequent machine learning model, helping the model better learn and understand the internal laws and patterns of meteorological data.

[0135] In step S121b, the computer system adopts a strategy different from that in step S121a to extract the template meteorological data recording points associated with timestamps. The core of this step is to divide the system meteorological data knowledge template according to its time span and extract data from each divided subset.

[0136] Specifically, first, the computer system will identify and determine the time span of each system meteorological data knowledge template. The time span refers to the time range covered by the template, which can be several days, weeks, months, or even longer. Understanding the time span is crucial for subsequent data processing because it determines the data distribution and variation range. Then, according to this time span, the computer system will divide each system meteorological data knowledge template into a predetermined number of knowledge template sub-meteorological data. This process is similar to cutting a large time period into multiple small time periods. Each knowledge template sub-meteorological data contains a part of the data in the original template, and this data is continuous in time. For example, if a knowledge template covers data for an entire month, it can be divided into four sub-meteorological data, each containing data for one week. Then, the computer system will arbitrarily extract a template meteorological data recording point from each divided knowledge template sub-meteorological data. This recording point represents the meteorological conditions at a specific moment in the sub-meteorological data, including various meteorological parameters (such as temperature, humidity, wind speed, etc.) and the associated timestamp. Since each sub-meteorological data is divided from the original template, the extracted recording points are also ordered in time.

[0137] In this way, the computer system can extract a predetermined number of timestamp-associated template meteorological data record points from each system meteorological data knowledge template. These record points are not only representative but also capable of reflecting the temporal variations and meteorological characteristics in the original template. This extraction strategy helps ensure the diversity and balance of the data, providing high-quality training data for subsequent machine learning tasks.

[0138] For example, assume there is a system meteorological data knowledge template covering one month of data, and it is desired to extract 4 timestamp-associated record points from it. According to step S121b, the data for this month can first be divided into 4 sub-meteorological data, with each sub-meteorological data containing one week of data. Then, a record point is randomly selected for extraction from each sub-meteorological data. Eventually, 4 record points representing the meteorological conditions of different weeks will be obtained, with each record point containing detailed meteorological parameters and timestamp information.

[0139] Step S122: Through the meteorological data vector characterization component, cohesive knowledge template pairs are formed based on each template meteorological data record point belonging to the same system meteorological data knowledge template, and discrete knowledge template pairs are formed based on each template meteorological data record point belonging to different system meteorological data knowledge templates, where one belonging to different includes two template meteorological data record points.

[0140] In step S122, the computer system uses the meteorological data vector characterization component to further process the previously extracted template meteorological data record points to construct sample pairs for endogenous learning. This step is a crucial step in machine learning, especially when self-supervised learning methods such as contrastive learning are adopted.

[0141] First, the computer system will, based on the template meteorological data record points extracted in step S121, identify which record points belong to the same system meteorological data knowledge template. Since these record points are collected under the same or similar meteorological conditions, the correlation and consistency between them are relatively high. The computer system will pair these record points belonging to the same template two by two to form "cohesive knowledge template pairs" or "positive sample pairs". For example, if the meteorological data collected at two different time points both reflect clear weather conditions (high visibility), then these two data points may be formed into a positive sample pair.

[0142] The construction of cohesive knowledge template pairs is based on an assumption that data points belonging to the same template should have similar representations in the feature space. By enabling the model to learn this similarity, it can better capture the internal structure and patterns of meteorological data. Secondly, the computer system also constructs "discrete knowledge template pairs" or "negative sample pairs" from the recorded points of meteorological data knowledge templates belonging to different systems. These recorded points reflect different meteorological conditions or patterns, so their representations in the feature space should be far apart from each other. For example, a data point reflecting sunny weather and a data point reflecting heavy rain weather may be formed into a negative sample pair.

[0143] The construction of discrete knowledge template pairs is based on another assumption that data points belonging to different templates should have separate representations in the feature space. By enabling the model to learn this separability, it can better distinguish different meteorological conditions and patterns.

[0144] After constructing these sample pairs, the computer system can use them to perform endogenous learning (self-supervised training) on the meteorological data vector representation component. Specifically, it adjusts the parameters and structure of the component so that for the two data points in a positive sample pair, their feature representations are as similar as possible; while for the two data points in a negative sample pair, their feature representations are as far apart as possible. In this way, the meteorological data vector representation component can gradually learn how to extract meaningful features and patterns from the original meteorological data.

[0145] It should be noted that in practical applications, the construction methods of positive sample pairs and negative sample pairs can be more complex and diverse. For example, in addition to simple pairwise combinations, data augmentation techniques can also be considered to generate more sample pairs; or more complex sample pair construction strategies can be designed according to the temporal and spatial characteristics of meteorological data. However, no matter how it is designed, the core goal is to enable the model to better learn the internal structure and patterns of meteorological data.

[0146] Step S123: Perform multiple endogenous cyclic learning on the meteorological data vector representation component according to the cohesive knowledge template pairs and discrete knowledge template pairs until a preset single optimization end judgment criterion is reached.

[0147] In step S123, the computer system performs multiple endogenous cyclic learning on the meteorological data vector representation component using the previously constructed cohesive knowledge template pairs (positive sample pairs) and discrete knowledge template pairs (negative sample pairs). Endogenous cyclic learning, also known as self-supervised training, is a method of optimization through the supervision signals generated by the model itself. In this process, the model attempts to learn how to best represent the input data so as to produce similar outputs between positive sample pairs and different outputs between negative sample pairs.

[0148] Specifically, the cohesive knowledge template pairs and discrete knowledge template pairs are used as input data and processed by the meteorological data vector representation component. This component is a neural network or machine learning model, and its task is to extract the features of the input data and convert them into vector representations. These vector representations are then used to calculate the similarity or distance between sample pairs. During each cycle of the learning process, the computer adjusts the parameters of the meteorological data vector representation component according to the calculation results of the similarity or distance of the cohesive knowledge template pairs and discrete knowledge template pairs. If the vector representations of two data points belonging to the same template (positive sample pair) are far apart in the feature space, or the vector representations of two data points belonging to different templates (negative sample pair) are close in the feature space, then the computer will adjust the parameters of the model to reduce this inconsistency.

[0149] For example, assume there is a simple meteorological data vector representation component, which is a shallow neural network. In the first cycle of learning, this component can randomly initialize its parameters and extract features from the input cohesive knowledge template pairs and discrete knowledge template pairs. Then, the computer calculates the similarity or distance between sample pairs based on the extracted features and finds that the similarity of some positive sample pairs is low, while the similarity of some negative sample pairs is high. To correct this situation, the computer adjusts the weights and biases of the neural network so as to produce more accurate outputs in the next cycle of learning.

[0150] This process will continue for multiple times until a preset end judgment criterion for a single optimization is reached. This criterion can be a fixed number of iterations, a convergence threshold, a performance metric on a validation set, etc. Through multiple cycles of learning, the meteorological data vector representation component can gradually learn how to extract meaningful features and patterns from the original meteorological data and provide strong support for subsequent meteorological data analysis, prediction, and decision-making.

[0151] Step S123 improves the feature extraction ability and representation learning ability of the meteorological data vector representation component by performing self-supervised training on it using the cohesive knowledge template pairs and discrete knowledge template pairs. The improvement of this ability is crucial for subsequent meteorological data processing tasks because it directly affects the generalization ability and prediction accuracy of the model for unknown data.

[0152] Step S130: Based on the site observation data knowledge templates related to each system meteorological data knowledge template in the basic training knowledge template set, perform multiple endogenous learning on the site data vector representation component to obtain the initial site data vector representation component after learning.

[0153] In step S130, the computer system performs multiple endogenous learning processes on the site data vector representation component by using the site observation data knowledge template related to the system meteorological data knowledge template in the basic training knowledge template set. The purpose of this step is to optimize the site data vector representation component so that it can more effectively extract key features from site observation data and provide an accurate data basis for subsequent meteorological analysis and prediction.

[0154] Specifically, the site data vector representation component is a network component responsible for converting the original site observation data into a vector representation. Since the original data often contains a large amount of redundant information and noise, directly using it for model training may lead to poor results. Through vector representation, the data can be mapped to a low-dimensional space while retaining its key features, making it easier for the model to learn the internal laws of the data. Endogenous learning is a self-supervised learning method that uses the internal structure and relationships of the data itself to generate supervision signals without the need for additional labeled data. In this process, the computer system adjusts the parameters of the site data vector representation component according to the similarities and differences between the site observation data knowledge templates, enabling it to better distinguish different data patterns.

[0155] For example, assume there are two site observation data knowledge templates A and B, which represent two different weather conditions respectively. During the endogenous learning process, the computer system will attempt to adjust the parameters of the site data vector representation component so that the distance between A and B in the vector space is as far as possible. In this way, when faced with new site observation data, the component can determine the weather condition it belongs to based on its position in the vector space. This process will be iterated multiple times, and each iteration will adjust the parameters of the component according to the current representation effect. Eventually, an initial site data vector representation component after learning is obtained, which can effectively extract key features from site observation data and provide strong support for subsequent meteorological analysis and prediction. It should be noted that the site observation data knowledge template here can be a template generated based on historical data or a representative data template obtained through other means. The specific implementation method of the site data vector representation component can be selected according to actual needs. For example, a convolutional neural network (CNN) or a recurrent neural network (RNN) can be used to process it. The principle of performing multiple endogenous learning processes on the site data vector representation component to obtain the initial site data vector representation component after learning can refer to the process of performing endogenous learning on the meteorological data vector representation component in step S120.

[0156] Step S140: Based on the basic training knowledge template set, perform multiple instance-driven learning on the initial multi-source data visibility recognition model that at least covers the initial meteorological data vector representation component and the initial station data vector representation component, to obtain the multi-source data visibility recognition model after basic training.

[0157] In step S140, the computer system will use the basic training knowledge template set to perform supervised learning on the initial multi-source data visibility recognition model, which is also often referred to as instance-driven learning. This model at least includes the initial meteorological data vector representation component and the initial station data vector representation component. The core goal of this step is to enable the model to accurately identify relevant visibility information from multi-source data by learning a large number of labeled samples. Specifically, the initial multi-source data visibility recognition model is a complex model integrating various data processing and analysis functions. Among them, the initial meteorological data vector representation component is responsible for converting the original meteorological data into vector form for the model to perform mathematical operations and logical processing. Similarly, the initial station data vector representation component also undertakes the task of converting station observation data into vectors. These two components together constitute the data input layer of the model, providing the basis for subsequent data analysis and feature extraction.

[0158] When performing supervised learning, the computer system will select a large number of labeled samples from the basic training knowledge template set. These samples include various meteorological conditions and station observation data, as well as the corresponding visibility labels. By learning these samples, the model can gradually establish a mapping relationship from the input data to the output labels. During the learning process, the computer system will continuously adjust the parameters and structure of the model to minimize the difference between the predicted results and the true labels. This difference is usually measured by a loss function, such as mean square error, cross entropy, etc. As the number of iterations increases, the prediction ability of the model will gradually improve until it reaches the predetermined performance index or convergence condition.

[0159] Finally, the multi-source data visibility recognition model after basic training will have the ability to accurately identify visibility from multi-source data. It can output the corresponding visibility prediction results based on the input meteorological data and station observation data. These results have important application values in fields such as meteorological forecasting and traffic safety.

[0160] It should be noted that the supervised learning here is a learning method that depends on labeled data. Therefore, the quality and quantity of the basic training knowledge template set have a decisive impact on the training effect of the model. In practical applications, it is necessary to ensure that these templates are representative, accurate, and diverse to fully exert the learning ability and generalization performance of the model.

[0161] In one implementation, the initial multi-source data visibility recognition model includes an initial meteorological data vector characterization component, an initial site data vector characterization component, a feature integration component for integrating site data features and system meteorological data features, and a classifier.

[0162] Based on this, in step S140, in one example-driven learning of the initial multi-source data visibility recognition model, it may specifically include:

[0163] Step S141: Input the system meteorological data knowledge template and the corresponding site observation data knowledge template in a batch into the initial multi-source data visibility recognition model. Among them, input the system meteorological data knowledge template into the initial meteorological data vector characterization component, and input the site observation data knowledge template into the initial site data vector characterization component;

[0164] Step S142: Load the fusion features of each system meteorological data feature output by the initial meteorological data vector characterization component and the site data features output by the site data vector characterization component into the feature integration component to obtain the fused system meteorological data features and site data features;

[0165] Step S143: Load the feature fusion result output by the feature integration component into the classifier to obtain the predicted visibility classification result output for each sample meteorological observation information;

[0166] Step S144: Combine the error between the visibility classification result and the corresponding template visibility prior label, obtain the cost value according to the binary cross-entropy cost function, and optimize the model parameters of the initial multi-source data visibility recognition model based on the cost value.

[0167] For example, the initial multi-source data visibility recognition model includes an initial meteorological data vector characterization component, an initial site data vector characterization component, a feature integration component for integrating site data features and system meteorological data features, and a classifier. The feature integration component is, for example, a neural network or a cross-attention network that completes feature splicing. The classifier is, for example, a neural network architecture including an affine network and an activation layer. Specifically, for example, the initial meteorological data vector characterization component is a convolutional neural network layer; the initial site data vector characterization component is a convolutional neural network layer; the feature integration component for integrating site data features and system meteorological data features is a cross-attention network layer; the classifier includes an affine layer and an activation layer to complete multi-classification.

[0168] When performing example-driven learning on the initial multi-source data visibility recognition model, assume that a batch includes 6 basic training knowledge templates as input. Then, 6 systematic meteorological data knowledge templates and 6 station observation data knowledge templates are extracted from the set of basic training knowledge templates as the input to the initial multi-source data visibility recognition model. For example, the 6 systematic meteorological data knowledge templates are input into the initial meteorological data vector representation component; the 6 station observation data knowledge templates are input into the initial station data vector representation component; the fused features of each systematic meteorological data feature output by the initial meteorological data vector representation component and the station data features output by the station data vector representation component are loaded into the feature integration component to obtain the fused systematic meteorological data features and station data features; the feature integration result output by the feature integration component is loaded into the classifier to obtain the predicted visibility classification results for the output of each sample meteorological observation information; the cost value is obtained according to the binary cross-entropy cost function through the error between the visibility classification result and the prior label of the corresponding template visibility, and backpropagation is performed based on the cost value to optimize the model parameters of the initial multi-source data visibility recognition model.

[0169] Specifically, in step S141, the computer system inputs the system meteorological data knowledge templates and the corresponding station observation data knowledge templates in a batch into the initial multi-source data visibility recognition model. This process aims to enable the model to learn the mapping relationship from the input data to the target output through a large amount of data input. The computer system selects a batch of system meteorological data knowledge templates, which are extracted from a large amount of meteorological data and are representative and diverse. At the same time, the computer system also obtains the corresponding station observation data knowledge templates for these system meteorological data knowledge templates. These station observation data knowledge templates contain various meteorological element information collected by ground observation stations, such as temperature, humidity, wind speed, etc. Then, the computer system inputs the system meteorological data knowledge templates into the initial meteorological data vector representation component. This component is a neural network model responsible for converting the original meteorological data into a vector representation for subsequent processing and analysis. Through vector representation, the key features and patterns in the original data are retained and extracted. At the same time, the computer system also inputs the station observation data knowledge templates into the initial station data vector representation component. This component is also a neural network model, and its role is to convert the station observation data into a vector representation for easy fusion and analysis with the system meteorological data. In this step, the initial multi-source data visibility recognition model receives two types of data inputs: system meteorological data and station observation data. These two types of data will be further processed and integrated in the subsequent steps to extract the features and information useful for visibility recognition. Here, "batch" is an important concept. In machine learning, since the dataset is usually very large and it is impossible to load all the data into the memory for training at one time. Therefore, the dataset is usually divided into several small batches, and each batch contains a certain number of samples. The model will process these batches in sequence during training and update the model's parameters by calculating the loss function values of each batch. This processing method can not only reduce memory occupancy but also improve the training efficiency of the model.

[0170] In step S142, the computer system loads the fusion features of each system meteorological data feature output by the initial meteorological data vector characterization component and the site data features output by the site data vector characterization component into the feature integration component. The purpose of this step is to fuse the features from different data sources to obtain a more comprehensive and representative feature set for use by subsequent classifiers. Specifically, the initial meteorological data vector characterization component has converted the original system meteorological data into vector form and extracted key meteorological features from it. These features may include multiple dimensions such as temperature, humidity, wind speed, and wind direction, and each dimension is represented by a vector. Similarly, the initial site data vector characterization component has also converted the site observation data into vector form and extracted key features related to visibility.

[0171] The role of the feature integration component is to fuse these two sets of features. The fusion method can be simple concatenation or more complex operations or transformations. For example, the meteorological feature vector and the site feature vector can be directly concatenated into a longer vector, or they can be combined into a new feature vector through a certain function. This new feature vector will contain information from both data sources and can more comprehensively describe the impact of meteorological conditions and site observation data on visibility. The fused feature vector will be passed to the subsequent classifier. The classifier is a machine learning model whose role is to predict the classification result of visibility based on the input feature vector. In step S142, the output of the feature integration component provides richer and more accurate input information for the classifier, which helps to improve the prediction performance of the classifier. It should be noted that the way and specific implementation of feature integration depend on the actual application scenario and data characteristics. In actual operation, it may be necessary to adjust the way and parameters of feature integration according to experimental results and performance evaluation to achieve the best prediction effect.

[0172] In step S143, the computer system loads the feature fusion result output by the feature integration component into the classifier to obtain the predicted visibility classification result for each sample meteorological observation information. This step is a key link in the machine learning model, where the classifier plays the role of making predictions based on the input features. Specifically, the classifier is a trained machine learning model that can receive the feature fusion result as input and classify and predict visibility based on these features. In this process, the computer system passes the feature vector output by the feature integration component to the classifier. These feature vectors contain the key information extracted from the system meteorological data and the site observation data and are the basis for the classifier to make accurate predictions.

[0173] The classifier can be implemented using various algorithms or models, depending on the complexity of the problem and the characteristics of the data. For example, the classifier can be a decision tree model, a support vector machine (SVM), a random forest model, or a deep neural network, etc. In the above example, the classifier includes a fully connected layer (affine layer) and an activation layer (such as sigmoid). During the training process, the neural network adjusts its internal parameters according to the input feature vectors and the corresponding labels to minimize the prediction error. Once the training is completed, the neural network can receive new feature vectors as input and output the predicted visibility classification results.

[0174] It should be noted that the prediction result of the classifier is a probability distribution or a class label, indicating the possibility of the sample belonging to each visibility category. For example, in a three-classification problem (such as low visibility, medium visibility, high visibility), the classifier can output a three-dimensional probability vector [0.1, 0.7, 0.2], indicating that the sample has the highest probability of belonging to medium visibility. In addition, the performance and accuracy of the classifier need to be measured and optimized through evaluation metrics. Common evaluation metrics include accuracy, recall rate, F1 score, etc. In practical applications, factors such as the generalization ability, robustness, and adaptability to new data of the model also need to be considered.

[0175] In step S144, the computer system combines the error between the visibility classification result and the corresponding template visibility prior label, obtains the cost value according to the binary cross-entropy cost function, and optimizes the model parameters of the initial multi-source data visibility recognition model based on this cost value. This step is a key part of the machine learning model training, which involves error calculation, application of the cost function, and optimization of model parameters. Specifically, the computer system first compares the difference between the visibility classification result output by the classifier and the known template visibility prior label. These prior labels are pre-labeled true results used to guide the training of the model. The error can be calculated in various ways, such as mean square error, cross-entropy, etc. Here, the binary cross-entropy cost function is used.

[0176] The binary cross-entropy cost function is a cost function commonly used in binary classification problems, which measures the difference between the probability distribution predicted by the model and the true probability distribution. In the visibility recognition problem, the visibility can be divided into two categories (such as low visibility and high visibility), or the multi-classification problem can be converted into a binary classification problem in a certain way for processing. The mathematical expression of the binary cross-entropy cost function is:

[0177]

[0178] where J is the cost function value, m is the number of samples, y (i) is the true label (0 or 1) of the (i)-th sample, a(i) is the predicted output of the model for the (i)-th sample (a probability value between 0 and 1).

[0179] By calculating the cost function value, the computer system can quantify the degree of inconsistency between the model prediction result and the true result. The smaller the cost function value, the more accurate the model's prediction. Next, the computer system will use this cost value to optimize the parameters of the model. The goal of optimization is to adjust the parameters of the model (such as the weights and biases of the neural network) so that the model can make more accurate predictions in the next iteration, that is, to reduce the value of the cost function. Optimization algorithms can adopt common optimization algorithms such as gradient descent, stochastic gradient descent, and Adam. These algorithms will update the parameters of the model according to the gradient information of the cost value to gradually approach the optimal solution. Through multiple iterations of training, the performance of the model will gradually improve, and finally a multi-source data visibility recognition model that can accurately identify visibility will be obtained.

[0180] In one implementation, when the multi-source data visibility recognition model after basic training covers at least the meteorological data vector characterization component after basic training, the site data vector characterization component after basic training, and the classifier after basic training, and it is determined that the visibility prior marker set during model refinement tuning does not match the visibility prior marker set in the classifier after basic training (i.e., not converged), the multi-source data visibility recognition model after basic training is refined and tuned, and the following steps are completed to obtain the target multi-source data visibility recognition model:

[0181] Step S1: Obtain a set of refinement tuning knowledge templates, where each refinement tuning knowledge template includes a refinement tuning system meteorological data knowledge template, a refinement tuning site observation data knowledge template related to the refinement tuning system meteorological data knowledge template, and a refinement tuning template visibility prior marker.

[0182] The purpose of Step S1 is to collect data templates specific to the model refinement tuning stage to further improve the performance of the trained model. These refinement tuning knowledge templates are the key for the multi-source data visibility recognition model to further adapt to specific scenarios or improve accuracy after basic training. Specifically, each refinement tuning knowledge template contains three main parts: a refinement tuning system meteorological data knowledge template, a refinement tuning site observation data knowledge template related to the system meteorological data, and the corresponding refinement tuning template visibility prior marker.

[0183] The refined calibration system meteorological data knowledge template refers to the system meteorological data that has been screened and processed, which may not have been covered in the basic training stage or is crucial for improving the performance of the model. For example, meteorological data under certain extreme weather conditions may be crucial for improving the recognition accuracy of the model in specific situations. The refined calibration site observation data knowledge template is the data collected by ground observation stations corresponding to the above system meteorological data. These data are also screened and processed to ensure consistency with the goals of model refinement calibration. For example, site observation data under specific terrain or climate conditions may be crucial for the performance of the model under these specific conditions. The refined calibration template visibility prior marker is the true visibility marker provided for the above two sets of data. These markers are obtained based on actual observations or expert judgments and are used to guide the optimization direction of the model during the model refinement calibration process. For example, for a specific set of meteorological data and site observation data, the corresponding visibility prior marker may be "low visibility" or "high visibility".

[0184] After obtaining these refined calibration knowledge template sets, the computer system will use this data for further training and optimization of the model to improve the performance and accuracy of the model in specific scenarios. This usually involves operations such as fine-tuning the model parameters and optimizing the model structure. Through this process, a more accurate and reliable target multi-source data visibility recognition model can be obtained.

[0185] Step S2: Update the visibility prior markers set in the classifier after basic training according to each visibility prior marker during model refinement calibration.

[0186] Specifically, when entering the model refinement calibration stage and determining a new set of visibility prior markers, the computer system will first access the classifier component in the multi-source data visibility recognition model after basic training. This classifier component has been given a set of visibility prior markers during the previous basic training stage, and these markers are used to guide the learning and prediction process of the model. However, in the refinement calibration stage, in order to further improve the performance and accuracy of the model, new data sets can be introduced or the data processing method can be adjusted, which may cause the original visibility prior markers to be no longer applicable. Therefore, the goal of step S2 is to update the markers in the classifier according to the new prior markers determined during refinement calibration.

[0187] The update process involves replacing or modifying the original prior tag values in the classifier. For example, if the original prior tag for visibility was based on simple weather conditions (such as "sunny", "cloudy"), and more complex weather classifications (such as "hazy weather", "sandstorm weather") were introduced during the refinement and calibration stage, then the prior tags in the classifier need to be updated according to the new weather classifications. This step can be achieved by loading the model after basic training, finding the prior tag parameters in the classifier component, and replacing these parameters with new values. This process needs to ensure that the new prior tags exactly correspond to the tags in the refinement and calibration knowledge template set to guarantee the accuracy and consistency of the model in subsequent training.

[0188] Step S3: According to the refinement and calibration knowledge template set, perform multiple model refinement and calibration operations on the updated multi-source data visibility recognition model to obtain the target multi-source data visibility recognition model after refinement and calibration learning.

[0189] In step S3, the computer system will use the previously obtained refinement and calibration knowledge template set to perform targeted training and optimization on the model. Specifically, the computer system extracts the corresponding system meteorological data, site observation data, and the corresponding visibility prior tags according to each template in the refinement and calibration knowledge template set. These data will be used as inputs and fed into the classifier whose prior tags have been updated for training. During the training process, the model will adjust its internal parameters and structure according to the input data and the corresponding tags to make the prediction results of the model closer to the actual visibility situation.

[0190] For example, assume that the refinement and calibration knowledge template set contains a template for hazy weather. This template includes the system meteorological data (such as temperature, humidity, wind speed, etc.) in hazy weather, the site observation data (such as the readings of the visibility meter), and the corresponding visibility prior tag (such as "low visibility"). During the training process, the model will try to learn the associations and patterns between these data so that it can accurately predict the low visibility situation when encountering similar hazy weather in the future.

[0191] It should be noted that the model refinement and calibration in step S3 is an iterative process. The computer system will repeat the above training steps multiple times, and each training will use a part of the data in the refined calibration knowledge template set. Through continuous iterative training, the model can gradually learn more features and patterns, thereby improving its prediction performance in various scenarios. Finally, after sufficient iterative training, the computer system will obtain a target multi-source data visibility recognition model after refined calibration. This model not only inherits the knowledge and capabilities of the basic training model, but also is optimized and adjusted for specific tasks and scenarios, so it has higher performance and accuracy. In future practical applications, this model will be able to provide more accurate and reliable visibility prediction services. The model after basic learning at least includes a meteorological data vector representation component after basic learning, a station data vector representation component after basic learning, and a classifier after basic learning. Among them, the meteorological data vector representation component after basic learning is trained based on the initial meteorological data vector representation component during the example-driven learning in the basic training session, the station data vector representation component after basic learning is trained based on the initial station data vector representation component during the example-driven learning in the basic training session, and the classifier after basic learning is trained based on the initial classifier in the initial multi-source data visibility recognition model during the example-driven learning process in the basic training session.

[0192] As a way of model refinement and calibration, model refinement and calibration are carried out on the premise of not changing the architecture of the multi-source data visibility recognition model after basic learning.

[0193] When it is determined that the number and content of each visibility prior marker in the application are the same as those set in the basic training session, that is, when the visibility prior marker during model refinement and calibration matches the visibility prior marker in the classifier after basic learning, model refinement and calibration are carried out on the premise of not changing the architecture of the multi-source data visibility recognition model after basic learning.

[0194] The computer system obtains a refined calibration knowledge template set, where each refined calibration knowledge template includes a refined calibration system meteorological data knowledge template, a refined calibration station observation data knowledge template related to the refined calibration system meteorological data knowledge template, and a refined calibration template visibility prior marker; based on the refined calibration knowledge template set, the obtained multi-source data visibility recognition model after basic learning is subjected to multiple model refinement and calibrations to obtain a target multi-source data visibility recognition model after refined calibration learning.

[0195] As another way of model refinement and calibration, the target classifier in the multi-source data visibility recognition model after basic learning is specifically adjusted and then model refinement and calibration are carried out.

[0196] When the computer system determines that the number and content of the visibility prior markers set compared to those in the basic training session are inconsistent, that is, when the visibility prior markers during model refinement tuning are different from those in the classifier after basic learning, the computer system performs model refinement tuning on the premise of changing the model architecture in the multi-source data visibility recognition model after basic learning to obtain the target multi-source data visibility recognition model. The computer system obtains a set of refinement tuning knowledge templates, where each refinement tuning knowledge template includes a refinement tuning system meteorological data knowledge template, a refinement tuning station observation data knowledge template related to the refinement tuning system meteorological data knowledge template, and a refinement tuning template visibility prior marker; the computer system then updates the classifier in the multi-source data visibility recognition model after basic learning according to the respective visibility prior markers during model refinement tuning, and performs multiple model refinement tunings on the updated multi-source data visibility recognition model based on the set of refinement tuning knowledge templates to obtain the target multi-source data visibility recognition model after refinement tuning learning. After the computer system determines that the number and content of the visibility prior markers have changed, it constructs a refinement tuning knowledge template according to the number and content of the visibility prior markers; updates the classifier in the multi-source data visibility recognition model after basic learning according to the actual number and content of the visibility prior markers; and performs multiple model refinement tunings on the updated multi-source data visibility recognition model after basic learning based on the refinement tuning knowledge template to obtain the target multi-source data visibility recognition model after refinement tuning learning.

[0197] In this way, the multi-source data visibility recognition model after basic learning is variably updated and adapted. At the same time, since the extraction components of the system meteorological data features and station data features after basic learning have been basically trained during the basic training process; during model refinement tuning, only the normalization association of the new visibility prior markers needs to be trained, and the generalization of the multi-source data visibility recognition model after basic learning is high, and the method of model refinement tuning is simple and efficient. Further, after the computer system obtains the target multi-source data visibility recognition model for video classification during the actual processing process, it loads the system meteorological data to be analyzed and the station observation data to be analyzed into the target multi-source data visibility recognition model to obtain the visibility analysis and recognition information output by the target multi-source data visibility recognition model.

[0198] The embodiment of the present application also provides a computer system, such as Figure 2As shown, computer system 100 includes: a processor 101 and a memory 103. Among them, the processor 101 and the memory 103 are connected, such as through a bus 102. Optionally, the computer system 100 may further include a transceiver 104. It should be noted that in practical applications, the transceiver 104 is not limited to one, and the structure of the computer system 100 does not constitute a limitation to the embodiments of the present application. The processor 101 may be a CPU, a general-purpose processor, a GPU, a DSP, an ASIC, an FPGA or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logical blocks, modules and circuits described in connection with the disclosure of the present application. The processor 101 may also be a combination that implements computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0199] The bus 102 may include a path for transmitting information between the above components. The bus 102 may be a PCI bus or an EISA bus, etc. The bus 102 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 2 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0200] The memory 103 may be a ROM or other type of static storage device that can store static information and instructions, a RAM or other type of dynamic storage device that can store information and instructions, or an EEPROM, a CD-ROM or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0201] The memory 103 is used to store the application program code for implementing the solution of the present application, and is controlled by the processor 101 to execute. The processor 101 is used to execute the application program code stored in the memory 103 to implement the content shown in any of the foregoing method embodiments.

[0202] The embodiments of the present application provide a computer system. The computer system in the embodiments of the present application includes: one or more processors; a memory; one or more computer programs, where one or more computer programs are stored in the memory and are configured to be executed by one or more processors. When the one or more programs are executed by the processor, the above method is implemented.

[0203] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially according to the indication of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this text, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps. The above is only part of the embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A meteorological quality data analysis method based on multi-source data fusion and AI, characterized in that: include: Acquire meteorological data of the proposed analysis system, and determine observation data of the proposed analysis site related to the meteorological data of the proposed analysis system; The meteorological data of the proposed analysis system and the observation data of the proposed analysis site are loaded into the trained target multi-source data visibility recognition model to obtain visibility analysis recognition information output by the target multi-source data visibility recognition model, wherein the target multi-source data visibility recognition model is a machine learning model obtained by model refinement and adjustment of the multi-source data visibility recognition model after basic training; The basic training process of the multi-source data visibility recognition model includes the following steps: Acquire a basic training knowledge template set, wherein each basic training knowledge template in the basic training knowledge template set includes a system meteorological data knowledge template, a site observation data knowledge template related to the system meteorological data knowledge template, and a template visibility priori mark; According to the system meteorological data knowledge template included in each knowledge template in the basic training knowledge template set, the meteorological data vector representation component is endogenously learned multiple times to obtain the learned initial meteorological data vector representation component; According to the site observation data knowledge templates related to the meteorological data knowledge templates of each system in the basic training knowledge template set, the site data vector representation component is endogenously learned multiple times to obtain the learned initial site data vector representation component; According to the basic training knowledge template set, the initial multi-source data visibility recognition model covering at least the initial meteorological data vector representation component and the initial site data vector representation component is subjected to multiple example-driven learning to obtain a multi-source data visibility recognition model after basic training; Among them, when the meteorological data vector representation component is endogenously learned, the following steps are completed: Extracting a predetermined number of system meteorological data knowledge templates from the basic training knowledge template set, and extracting a predetermined number of template meteorological data record points associated with timestamps from each of the extracted system meteorological data knowledge templates through the meteorological data vector representation component; Through the meteorological data vector representation component, a cohesive knowledge template pair is formed according to each template meteorological data record point belonging to the same system meteorological data knowledge template, and a discrete knowledge template pair is formed according to each template meteorological data record point belonging to different system meteorological data knowledge templates, wherein one belongs to different systems and includes two template meteorological data record points, the cohesive knowledge template pair is a positive sample pair, and the discrete knowledge template pair is a negative sample pair; Performing multiple endogenous loop learning on the meteorological data vector representation component according to the cohesive knowledge template pair and the discrete knowledge template pair until a preset single optimization end judgment standard is reached; The initial multi-source data visibility recognition model includes an initial meteorological data vector representation component, an initial site data vector representation component, a feature integration component for fusing site data features and system meteorological data features, and a classifier; In an example-driven learning of the initial multi-source data visibility recognition model, it includes: Inputting the system meteorological data knowledge template and the corresponding site observation data knowledge template in a batch into the initial multi-source data visibility recognition model, wherein the system meteorological data knowledge template is input into the initial meteorological data vector representation component, and the site observation data knowledge template is input into the initial site data vector representation component; Loading the fusion features of the various system meteorological data features output by the initial meteorological data vector representation component and the site data features output by the site data vector representation component into the feature integration component to obtain the fused system meteorological data features and site data features; Loading the feature fusion result output by the feature integration component into the classifier to obtain the predicted visibility classification result output for each sample meteorological observation information; Combining the error between the visibility classification result and the corresponding template visibility prior mark, a cost value is obtained according to a binary cross entropy cost function, and model parameters of the initial multi-source data visibility recognition model are optimized according to the cost value.

2. The method according to claim 1, characterized in that The step of obtaining a basic training knowledge template set includes: According to the historical meteorological observation information corresponding to each visibility monitoring area, a set of historical meteorological observation information corresponding to each visibility monitoring area is generated; wherein each historical meteorological observation information includes a historical system meteorological data, historical site observation data related to the historical system meteorological data, and an original visibility priori mark; For each visibility monitoring area, according to the historical meteorological observation information corresponding to the visibility monitoring area, example-driven learning is performed to obtain a target auxiliary classifier corresponding to the visibility monitoring area; For each historical system meteorological data, according to each target auxiliary classifier, each auxiliary visibility classification information related to the historical system meteorological data is determined, and the original visibility prior label related to the historical system meteorological data and the augmented label set of each auxiliary visibility classification information are used as the template visibility prior label related to the historical system meteorological data; A basic training knowledge template is formed according to the historical system meteorological data, the template visibility priori mark related to the historical system meteorological data and the historical site observation data; The basic training knowledge template set is assembled according to each basic training knowledge template.

3. The method according to claim 2, characterized in that The generating of a set of historical meteorological observation information corresponding to each visibility monitoring area according to the historical meteorological observation information corresponding to each visibility monitoring area includes: According to each initial meteorological observation information in each visibility monitoring area whose monitoring time is within a preset time interval, an initial meteorological observation information set is formed, wherein each initial meteorological observation information includes an initial system meteorological data, historical site observation data related to the initial system meteorological data, and an original visibility priori mark; Extracting one piece of initial meteorological observation information from the initial meteorological observation information set one by one, and performing the following simplification screening according to the extracted initial meteorological observation information, until there is no unextracted initial meteorological observation information in the initial meteorological observation information set: Respectively determining meteorological data characterization vector matching factors between the initial system meteorological data in the extracted initial meteorological observation information and the initial system meteorological data in other initial meteorological observation information in the initial meteorological observation information set; Removing other initial meteorological observation information from the initial meteorological observation information set whose meteorological data characterization vector matching factor with the extracted initial meteorological observation information reaches a matching factor threshold; The initial meteorological observation information in the unified and filtered initial meteorological observation information set is used as the historical meteorological observation information; According to the visibility monitoring area to which each piece of historical meteorological observation information belongs, a set of historical meteorological observation information corresponding to each visibility monitoring area is obtained according to a cluster analysis of each of the historical meteorological observation information.

4. The method according to claim 2, characterized in that Determining various auxiliary visibility classification information related to the historical system meteorological data according to various target auxiliary classifiers includes: Determine the target visibility monitoring area corresponding to the historical system meteorological data, determine the target auxiliary classifier corresponding to the target visibility monitoring area, and obtain other target auxiliary classifiers corresponding to each other visibility monitoring area except the target visibility monitoring area; According to each other target auxiliary classifier, each related auxiliary visibility classification information in the historical system meteorological data is determined respectively.

5. The method according to any one of claims 2, 3 and 4, characterized in that: The auxiliary visibility classification information includes an auxiliary visibility mark and a prediction support coefficient predicted for the auxiliary visibility mark; After determining each auxiliary visibility classification information related to the historical system meteorological data, the method further includes: Acquire an initial overlapping mark set between the auxiliary visibility marks included in each auxiliary visibility classification information and the original visibility prior marks related to the historical system meteorological data, and determine the prediction support coefficient corresponding to each visibility prior mark in the initial overlapping mark set; Removing visibility prior marks whose prediction support coefficients reach the mark filtering index in the initial overlapping mark set from the original visibility prior marks related to the historical system meteorological data, to obtain processed original visibility prior marks; For the processed original visibility prior mark, the step of using the original visibility prior mark related to the historical system meteorological data and the augmented label set of each auxiliary visibility classification information as the template visibility prior mark related to the historical system meteorological data is executed.

6. The method according to claim 1, characterized in that When extracting a predetermined number of template meteorological data record points associated with time stamps from each extracted system meteorological data knowledge template, one of the following steps is completed: For each system meteorological data knowledge template extracted, a target node corresponding to the system meteorological data knowledge template is set according to the time span corresponding to the system meteorological data knowledge template, and starting from the template meteorological data recording point corresponding to the target node on the system meteorological data knowledge template, a predetermined number of template meteorological data recording points associated with a timestamp are extracted; For each system meteorological data knowledge template extracted, the system meteorological data knowledge template is divided into a predetermined number of knowledge template sub-meteorological data according to the time span corresponding to the system meteorological data knowledge template, and one template meteorological data record point is arbitrarily extracted from each knowledge template sub-meteorological data to obtain a predetermined number of template meteorological data record points associated with timestamps.

7. The method according to claim 1, characterized in that When the multi-source data visibility recognition model after basic training at least includes the meteorological data vector representation component after basic training, the site data vector representation component after basic training, and the classifier after basic training, when it is determined that the visibility priori mark set during model refinement and calibration does not match the visibility priori mark set in the classifier after basic training, the multi-source data visibility recognition model after basic training is refined and calibrated, and the following steps are completed to obtain the target multi-source data visibility recognition model: Acquire a set of refined adjustment knowledge templates, wherein each refined adjustment knowledge template includes a refined adjustment system meteorological data knowledge template, a refined adjustment site observation data knowledge template related to the refined adjustment system meteorological data knowledge template, and a visibility priori mark of the refined adjustment template; According to each visibility priori marker during model refinement and adjustment, updating the visibility priori marker set in the classifier after the basic training; According to the set of refined and calibrated knowledge templates, the updated multi-source data visibility recognition model is subjected to multiple model refinement and calibration to obtain a target multi-source data visibility recognition model after refined and calibrated learning.

8. A computer system, characterized in that: include: one or more processors; Memory; one or more computer programs; The one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method for predicting atmospheric visibility

    CN107942411A

  • Sea fog level intelligent forecasting method and system

    CN114280696A