A method, system and related device for generating weather forecast service text data set
By collecting, preprocessing and expanding data, combining large language models and quantitative evaluation algorithms, the semi-automated construction of weather forecast service text data sets is achieved, which solves the problems of inefficient data set quality and generation efficiency in the existing technology, and improves the capabilities of large language models in the field of weather forecast service.
Patent Information
- Application Number
- CN202510055944.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-14
AI Technical Summary
The existing large language models perform poorly in the professional field of weather forecasting services, mainly because the quality of training data cannot be guaranteed and the text data of weather forecasting services is relatively small. The process of building data sets is cumbersome and inefficient, and it cannot meet the data needs of large language models.
A method for generating weather forecast service text data sets is proposed. By collecting, preprocessing and expanding data, combining large language models and quantitative evaluation algorithms, the semi-automated construction of weather forecast service text data sets is realized, and the generation efficiency and quality of the data set are improved.
It effectively solves the problem of uneven quality of existing data sets. The generated data sets can meet the needs of different application scenarios, improves the ability of large language models in the field of weather forecasting services, and reduces the cost and time of manually constructing data sets.
Smart Images

Figure CN119476449B_ABST
Abstract
Description
Technical Field
[0001] The invention proposes a weather forecast service text data set generation method, system and related device, belonging to the technical field of meteorological data. Background Art
[0002] In recent years, companies such as OpenAI, Meta, and Baidu have successively launched large language models such as GPT, Llama, and Wenxinyiyan, which aim to understand, interpret, and generate texts similar to human languages, making breakthrough progress in application fields such as text generation, question-answering systems, and mathematical reasoning. Large language models are essentially a natural processing tool based on deep neural networks. They are usually based on the Transformer (an attention-based neural network) architecture and are large-parameter models obtained by training a large amount of text data. The complex scale and architecture have caused an "emergence effect" (specifically, the small model does not exist, but the large model has the capabilities and characteristics), so that it can be adapted to various vertical fields such as finance, education, transportation, and medical care through transfer learning, thereby improving business efficiency and obtaining economic benefits. Today, large language models in professional fields are trained based on professional data sets, so it is particularly important to build high-quality professional data sets.
[0003] At present, mainstream large language models perform poorly in the professional field of weather forecast services. This is mainly because the training data of the models mostly comes from online platforms such as websites, forums and social media. The data quality cannot be guaranteed, making it impossible to improve the ability of large language models in the meteorological field. At the same time, due to the particularity of the weather forecast industry, the weather forecast service products publicly released daily by official agencies such as the Central Meteorological Observatory have professional requirements in content and structure, and the quantity is limited. As a result, the text data for weather forecast services is generally small, and the process of building data sets is mostly manually written, which is cumbersome and inefficient, and cannot meet the data volume required for the current large language model application. Therefore, it is necessary to design a method for generating weather forecast service text data sets for the weather forecast service field to meet the needs of large, diverse, and professional data sets for large language model applications. Summary of the invention
[0004] In order to solve the problems existing in the prior art, the present invention proposes a method, system and related devices for generating a weather forecast service text dataset. The method can realize the semi-automatic construction of the weather forecast service text dataset, and effectively improve the generation efficiency of the dataset.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0006] In a first aspect, the present invention provides a method for generating a weather forecast service text data set, comprising:
[0007] Collect weather forecast service data;
[0008] Preprocessing the weather forecast service data to obtain weather forecast service text data;
[0009] Based on the large language model, the weather forecast service text data is expanded to obtain the weather forecast service text extension data set; the weather forecast service text extension data set is tested and evaluated by combining the large language model with the quantitative evaluation algorithm;
[0010] The weather forecast service text dataset is constructed by using the data in the weather forecast service text extension dataset that meets the verification and evaluation requirements.
[0011] As a further improvement of the present invention, the collecting of weather forecast service data includes:
[0012] By combining manual screening with web crawler technology, weather forecast service data for the public, media and meteorological practitioners are collected in batches from multiple online platforms.
[0013] As a further improvement of the present invention, the weather forecast service data is preprocessed to obtain weather forecast service text data, including:
[0014] Design the storage structure of the data set, which includes four parts: title, category, time, and content. The title represents the title of the original data, the category represents the meteorological elements contained in the original data, and needs to be extracted in a standardized manner by combining the large language model prompt word technology and data annotation algorithm. The time represents the release time of the data, and the content represents the main text of the original data.
[0015] The weather forecast service data is preprocessed based on the storage structure of the data set to obtain weather forecast service text data.
[0016] As a further improvement of the present invention, the weather forecast service text data is expanded to obtain a weather forecast service text extension data set, including:
[0017] If you need to generate weather forecast service text data for a single meteorological element, you only need to obtain the data containing the target element from the original data, build a few-shot prompt word project based on the data containing the target element, and use the large language model and thinking chain technology to enable the large language model to expand the weather forecast service text data according to the prompt word instructions of the few-shot prompt word project;
[0018] If you need to generate weather forecast service text data with multiple meteorological elements, you need to build a knowledge vector library, slice, store, retrieve and generate the weather forecast service text data in sequence, and build prompt word instructions based on the application intent, so that the large language model can generate new weather forecast service text data containing multiple meteorological elements.
[0019] After expansion, we get the weather forecast service text extended dataset.
[0020] As a further improvement of the present invention, if it is necessary to generate weather forecast service text data of multiple meteorological elements, it is necessary to construct a knowledge vector library, perform data slicing, data storage, data retrieval and data generation on the weather forecast service text data in sequence, and construct prompt word instructions according to the application intention, so that the large language model generates brand-new weather forecast service text data containing multiple meteorological elements, including:
[0021] Obtain data related to designated meteorological elements from preprocessed data by data title, category, and release time;
[0022] Building a knowledge vector library based on data related to specified meteorological elements to generate extended data is mainly divided into the following four steps:
[0023] The first step is data slicing: using content-based data slicing methods to analyze the chapter structure of the data, using the title hierarchy to determine the relationship between paragraphs, and then using the large language model to determine whether the content of adjacent paragraphs is similar. If it exceeds the threshold, the paragraphs are merged so that each piece of data contains complete meteorological information;
[0024] The second step is data storage: the sliced data is encoded using a fine-tuned embedding model, the text data is converted into a set of high-dimensional vectors, so that the high-dimensional vectors contain the deep semantic structure of the text, and are stored in the vector database;
[0025] The third step is data retrieval: according to the pre-generated data application requirements, a hybrid retrieval method is used to retrieve and recall the specified meteorological element content;
[0026] The fourth step is data generation: construct prompt words, combine them with the retrieved content, and let the large language model expand the data according to the logic of the prompt words. After the expansion is completed, the weather forecast service text expansion data set is obtained.
[0027] As a further improvement of the present invention, the method of combining a large language model with a quantitative evaluation algorithm to test and evaluate the weather forecast service text extension data set includes:
[0028] The quantitative evaluation algorithm uses the GLM model as the evaluation model to evaluate the extended data of the weather forecast service text extension dataset and determine whether the extended data can describe the complete content of the original data;
[0029] The GLM model was used to extract the description content corresponding to each level of meteorological elements in the original data and the extended data, and then the Jaccard similarity coefficient between the two was calculated;
[0030] The quantitative evaluation algorithm uses the Jaccard similarity coefficient, and the evaluation index is:
[0031] ,
[0032] in, Represents the description content in the original data, Represents the description content in the extended data. The Jaccard similarity coefficient represents the similarity between the two. It is used to judge the similarity between the two. The higher the result, the more similar the two are.
[0033] The evaluation of weather forecast service text data with multiple meteorological elements also includes:
[0034] The GLM model is used to evaluate the extended data as a whole to determine whether it contains the content of the specified meteorological elements and whether it conforms to the commonly used description logic;
[0035] The GLM model is used to extract the description content corresponding to each level of meteorological elements in the extended data and recalled data slices, and the evaluation index is calculated:
[0036]
[0037] in, Represents the type of meteorological element, Represents the number of meteorological elements. Representative recall data involving meteorological elements Description content, Represents meteorological elements involved in the extended data Description content, Represents the Jaccard similarity coefficient between the two.
[0038] In a second aspect, the present invention provides a weather forecast service text data set generation system, comprising:
[0039] A collection module, used to collect weather forecast service data;
[0040] A preprocessing module, used for preprocessing weather forecast service data to obtain weather forecast service text data;
[0041] The expansion and evaluation module is used to expand the weather forecast service text data based on the large language model to obtain the weather forecast service text expansion data set; the weather forecast service text expansion data set is tested and evaluated by combining the large language model with the quantitative evaluation algorithm;
[0042] The construction module is used to construct the weather forecast service text dataset from the data in the weather forecast service text extension dataset that meets the verification and evaluation requirements.
[0043] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method for generating a weather forecast service text data set when executing the computer program.
[0044] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the method for generating a weather forecast service text dataset.
[0045] In a fifth aspect, the present invention provides a computer program product, wherein the computer program product comprises computer instructions, wherein the computer instructions instruct a computer to execute the weather forecast service text data set generation method.
[0046] The beneficial effects of the present invention compared with the prior art are as follows:
[0047] The present invention provides a method for generating a weather forecast service text dataset containing single meteorological elements and multiple meteorological elements based on a large language model. The method can be oriented to different application scenarios in weather forecast services, and adopts prompt word engineering, enhanced retrieval and text generation technologies to generate weather forecast service text datasets containing different meteorological elements and text structures, effectively solving the problem that the quality of existing datasets is uneven and cannot meet current application requirements. At the same time, compared with the traditional manual construction method, the method can realize the semi-automatic construction of the weather forecast service text dataset, effectively improving the efficiency of dataset generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the embodiments of the present invention or the drawings of related technical solutions in the prior art are introduced below. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0049] Figure 1A flow chart of a method for generating a weather forecast service text data set provided by the present invention;
[0050] Figure 2 The overall architecture diagram of the method for generating text dataset for weather forecast service;
[0051] Figure 3 The sample result of the weather forecast service text data for a single meteorological element;
[0052] Figure 4 The sample results of text data of weather forecast service for multiple meteorological elements;
[0053] Figure 5 A weather forecast service text data set generation device provided by the present invention;
[0054] Figure 6 A schematic diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION
[0055] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0056] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.
[0057] Explanation of relevant terms:
[0058] The GLM model, full name General Language Model, is a large language model released by Zhipu AI.
[0059] Quantitative evaluation refers to a method of evaluation and analysis that establishes a mathematical model based on statistical data and uses the mathematical model to calculate the various indicators and their values of the analysis object.
[0060] A large language model (LLM) is a deep learning model trained with a large amount of text data that can generate natural language text or understand the meaning of language text.
[0061] Web crawler technology refers to the technology of automatically crawling the World Wide Web information according to certain rules. Web crawlers are also called web spiders, web robots, and more often called web chasers in the FOAF community; other less commonly used names include ants, automatic indexers, simulation programs, or worms.
[0062] Few-shot prompt engineering refers to guiding large language models to give higher quality results by providing a small number of examples (i.e., Few-shot) in prompt engineering.
[0063] Chain of Thought (CoT) technology is a technology that enables large models to perform better logical reasoning.
[0064] like Figure 1 As shown, the first object of the present invention is to provide a method for generating a weather forecast service text data set, comprising:
[0065] Collect weather forecast service data;
[0066] Preprocessing the weather forecast service data to obtain weather forecast service text data;
[0067] Based on the large language model, the weather forecast service text data is expanded to obtain the weather forecast service text extension data set; the weather forecast service text extension data set is tested and evaluated by combining the large language model with the quantitative evaluation algorithm;
[0068] The weather forecast service text dataset is constructed by using the data in the weather forecast service text extension dataset that meets the verification and evaluation requirements.
[0069] The method for generating a weather forecast service text data set provided by the present invention has the following principles: first, a large amount of weather forecast service data is collected from various weather forecast service sources (such as meteorological departments, news websites, social media, etc.). These data cover text descriptions of different times, places and weather conditions. The collected weather forecast service data is preprocessed, including steps such as cleaning, deduplication, word segmentation, and annotation to obtain structured weather forecast service text data. This step ensures the accuracy and consistency of the data and provides a basis for subsequent data expansion and evaluation. Based on a large language model, the preprocessed weather forecast service text data is expanded. The large language model has a strong text generation and understanding ability and can generate new text related to a given text. Through this step, more diverse weather forecast service text data can be generated, thereby constructing a richer data set. The generated weather forecast service text data set is tested and evaluated by combining a large language model with a quantitative evaluation algorithm. The large language model can evaluate the language consistency and accuracy of the data, while the quantitative evaluation algorithm can quantify the diversity and integrity of the data. Through this step, data that meets the test and evaluation criteria can be screened out to ensure the quality of the data set. The data that meets the inspection and evaluation criteria are constructed into a weather forecast service text dataset. This dataset can be used to train large language models, optimize and improve weather forecast services, etc.
[0070] Furthermore, the present invention can quickly generate a large amount of weather forecast service text data through the text generation capability of the large language model, thereby improving the generation efficiency of the data set. The large language model has a powerful natural language processing capability and can generate more accurate and natural weather forecast service text data. At the same time, with the assistance of the quantitative evaluation algorithm, the quality of the data set can be further ensured. Through the expansion capability of the large language model, more diverse weather forecast service text data can be generated, thereby enhancing the diversity of the data set. This helps to improve the recognition and understanding capabilities of the large language model for complex weather conditions. The traditional weather forecast service text data set generation method requires a lot of manual participation, including steps such as data collection, preprocessing and annotation. The method of the present invention reduces labor costs and improves work efficiency through automation and intelligent methods. The weather forecast service text data set generation method provided by the present invention has significant advantages and can be widely used in fields such as weather forecast services and large language model training.
[0071] The method of the present invention is described in detail below in conjunction with specific embodiments:
[0072] Example 1
[0073] The present invention provides a method for generating a weather forecast service text dataset. By utilizing the capabilities of a large language model in natural language processing, the generation efficiency and quality of the dataset are effectively improved through the techniques of prompt word engineering, enhanced retrieval and text generation. The specific steps are as follows:
[0074] Step 1. Data collection steps:
[0075] By combining manual screening with web crawler technology, weather forecast service data for the public, media and meteorological practitioners are collected in batches from multiple online platforms such as official meteorological websites, public accounts, columns, etc., to ensure the diversity, accuracy and professionalism of the original data content.
[0076] Step 2. Data preprocessing steps:
[0077] Due to the diversity of original data sources, the data format is complex and the content styles are numerous, so data preprocessing is required. First, the storage structure of the data set is designed, which includes four parts: title, category, time, and content. Among them, the title represents the title of the original data, the category represents which meteorological elements (such as wind, rain, haze, etc.) the original data contains, and it is necessary to combine the large language model prompt word technology and data annotation algorithm for standardized extraction, the time represents the release time of the data, and the content represents the main text of the original data.
[0078] The weather forecast service data is preprocessed based on the storage structure of the data set to obtain weather forecast service text data.
[0079] For example, subsequent preprocessing steps usually include one or more of the following aspects:
[0080] a) Data cleaning; Remove invalid data: Delete data entries that are irrelevant to weather forecast or have incomplete information. Process duplicate data: Identify and delete duplicate data records to ensure the uniqueness of the data set.
[0081] b) Data standardization; Title standardization: Ensure that the format and expression of all titles are consistent, such as unifying uppercase and lowercase letters, removing special characters, etc. Category standardization: Use large language model prompt word technology and data annotation algorithms to standardize the extraction of meteorological elements in the original data. Establish a standard vocabulary of meteorological elements, such as wind, rain, snow, fog, haze, etc., and match the extracted categories with the standard vocabulary. Time standardization: Convert the release time of the data into a unified format, such as "YYYY-MM-DD HH:MM:SS", for subsequent time series analysis and processing. Content standardization: Perform word segmentation, stop word removal, annotation and other processing on the main text to extract key information and reduce data dimensions.
[0082] Through these preprocessing steps, high-quality, standardized weather forecast service text data can be obtained, providing strong support for subsequent data analysis and model training.
[0083] Step 3. Data generation steps:
[0084] The overall architecture of the weather forecast service text dataset generation method is shown in the attached figure. Figure 2 shown.
[0085] This architecture is based on a large language model. It uses the language understanding, logical reasoning and instruction-following capabilities of the large language model to achieve rapid expansion of weather forecast service text data, thereby obtaining a high-quality weather forecast service text data set. First, according to the actual application scenario, if you need to generate weather forecast service text data for a single meteorological element (for example, select content related to rain), you only need to obtain the data containing the target element from the original data, build a few-shot prompt word project, and use a large language model including GLM, Llama, and Tongyi Qianwen. Apply the thinking chain technology so that the model can expand the data content according to the prompt word instructions. If you need to generate weather forecast service text data for multiple meteorological elements (for example, including wind and rain related content at the same time, and you can also select other specified meteorological elements), you need to build a knowledge vector library to generate the above content, which includes four parts:
[0086] 1) Data slicing: A content-based data slicing algorithm is used to slice the data containing the target elements in the original data to ensure that each data slice contains complete information, thereby reducing data loss in the recall stage.
[0087] 2) Data storage: Use the embedding model to vectorize the sliced data and store it in the knowledge vector library.
[0088] 3) Data retrieval: Based on data application requirements, a hybrid retrieval algorithm is used to recall data within a specified range from the knowledge vector library to ensure retrieval efficiency and quality.
[0089] 4) Data generation: Based on the data recalled in step 3) data retrieval, construct prompt word instructions according to the application intent, and let the large language model generate new weather forecast service text data containing multiple meteorological elements.
[0090] In this way, a large amount of new data about weather forecast services can be generated. In order to further ensure the quality of the extended data, this method uses a combination of a large language model and a quantitative evaluation algorithm to test and evaluate the extended data. The quantitative evaluation algorithm uses the Jaccard similarity coefficient, and the evaluation indicators are as follows:
[0091] ,
[0092] in Represents the description content in the original data, Represents the description content in the extended data. The Jaccard similarity coefficient represents the two and is used to judge the degree of similarity between the two. The higher the result, the more similar the two are.
[0093] Step 4: Build the dataset
[0094] Using the two methods in step 3, you can quickly build a weather forecast service text dataset.
[0095] Example 2
[0096] The present invention provides a method for generating a weather forecast service text data set, which utilizes the ability of a large language model in natural language processing and effectively improves the generation efficiency and quality of the data set through prompt word engineering, enhanced retrieval and text generation technologies.
[0097] It is necessary to determine the types of meteorological elements that the extended data needs to include according to the application scenario. If you want to generate weather forecast service text data of a single meteorological element, such as rain warnings, you can obtain the original data related to rain from the preprocessed data through the data title and category. Next, use large language models such as GLM, Llama, and Tongyi Qianwen to build Few-shot warning words. By adjusting the hybrid method of description and content reconstruction, combined with the thinking chain technology, the expansion of rain data can be gradually realized. The sample file is as follows: Figure 3 shown.
[0098] To ensure the quality of data generation, the extended data needs to be verified by the evaluation module. First, the GLM model is used as an evaluation model to evaluate the extended data of the weather forecast service text extended data set to determine whether the extended data can describe the complete content of the original data. Next, the GLM model is used to extract the description content corresponding to the meteorological elements at each level in the original data and the extended data, and then the Jaccard similarity coefficient between the two is calculated. The higher the value, the higher the quality of the generated data. If you want to generate multi-factor weather forecast service text data, for example: daily wind and rain weather tips.
[0099] First, obtain wind and rain related data from the preprocessed data through data title, category, and release time, such as weather bulletin, gale warning, and rainstorm warning (Note: weather forecast is related to time, so the release time of the data needs to be considered).
[0100] Next, we construct a knowledge vector library for the above three materials to generate extended data, which is mainly divided into the following four steps:
[0101] The first step is data slicing: using a content-based data slicing method, first analyze the chapter structure of the data, use the title hierarchy to determine the relationship between paragraphs, and then use models such as BERT to determine whether the content of adjacent paragraphs is similar. If it is above a certain threshold, the paragraphs are merged. In this way, each piece of data is guaranteed to contain complete meteorological information.
[0102] The second step is data storage: Use a fine-tuned embedding model (such as bge-Embedding, a word embedding structure) to encode the sliced data. It converts text data into a set of high-dimensional vectors that contain the deep semantic structure of the text and stores them in a vector database (Elasticsearch, Faiss, Milvus and other vector databases are all acceptable). In this way, efficient retrieval can be supported while ensuring data persistence.
[0103] Among them, Elasticsearch, Faiss, and Milvus are tools for processing large-scale data search and analysis, as vector databases.
[0104] The third step is data retrieval. According to the need to pre-generate daily wind and rain weather tips, a hybrid retrieval method is used to retrieve and recall wind and rain content, such as: BM25 (Best Matching 25) retriever, vector retriever, and the weight of each retriever can be set by yourself. This method makes full use of the advantages of multiple retrieval methods, has stronger semantic understanding capabilities, and increases fault tolerance, so that it can quickly and accurately recall data slices containing wind and rain weather from the above three materials.
[0105] The fourth step is data generation. Construct prompt words. Combined with the content recalled in step 3, let the large language model generate extended data according to the specified logic, so as to achieve the expansion of weather forecast service text data with multiple meteorological elements. The sample file is as follows: Figure 4 As shown, a complex weather forecast service text dataset is formed.
[0106] Next, we need to evaluate the extended multi-meteorological element weather forecast service text data. First, we use the GLM model to evaluate the extended data as a whole to determine whether it contains wind and rain elements and whether it conforms to the commonly used description logic. Secondly, we use the GLM model to extract the description content corresponding to each level of meteorological elements in the extended data and the recalled data slices, and calculate their evaluation indicators.
[0107] The GLM model is used to extract the description content corresponding to each level of meteorological elements in the extended data and recalled data slices, and the evaluation index is calculated as follows:
[0108]
[0109] in, Represents the type of meteorological element, Represents the number of meteorological elements. Representative recall data involving meteorological elements Description content, Represents meteorological elements involved in the extended data Description content, Represents the Jaccard similarity coefficient between the two.
[0110] The present invention takes data containing wind and rain elements as an example, and the calculation evaluation index is as follows:
[0111]
[0112] in, Represents the description of rain in the recalled data. Represents the description of rain in the extended data. Represents the Jaccard similarity coefficient between the two. Represents the description content related to wind in the recalled data. Represents the description of wind in the extended data. Represents the Jaccard similarity coefficient between the two. Since the generation results of weather forecast service text data with multiple meteorological elements are relatively complex, the above score can only be used as an auxiliary manual evaluation criterion. The higher the value, the more complete the generated content.
[0113] Through the above two methods, we can quickly build data sets for the field of weather forecast services. These data involve different meteorological elements and text structures, which can be used for model pre-training and fine-tuning. While helping the large language model learn rich meteorological knowledge, it also improves the large language model's ability to answer questions, generate weather forecast service text products of different specifications, and call meteorological algorithm interfaces, thereby facilitating the intelligent upgrade of weather forecast service business.
[0114] The method for generating a learning data set provided by the present invention can be applied to automatically generating weather forecast service materials, etc.
[0115] like Figure 5 As shown, the third object of the present invention is to provide a weather forecast service text data set generation system, comprising:
[0116] A collection module 100 is used to collect weather forecast service data;
[0117] The preprocessing module 200 is used to preprocess the weather forecast service data to obtain weather forecast service text data;
[0118] The expansion and evaluation module 300 is used to expand the weather forecast service text data based on the large language model to obtain the weather forecast service text expansion data set; the weather forecast service text expansion data set is tested and evaluated by combining the large language model with the quantitative evaluation algorithm;
[0119] The construction module 400 is used to construct the weather forecast service text data set from the data in the weather forecast service text extension data set that meets the verification evaluation.
[0120] The weather forecast service text data set generation system of the present invention is based on the weather forecast service text data set generation method mentioned above.
[0121] like Figure 6 As shown, the third purpose of the embodiment of the present invention is to provide an electronic device, including a memory 701, a processor 702, and a computer program stored in the memory 701 and executable on the processor, wherein the processor implements the above-mentioned weather forecast service text data set generation method when executing the computer program. It also includes a communication interface 703 and a bus 704.
[0122] A fourth objective of an embodiment of the present invention is to provide a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for generating a weather forecast service text data set is implemented.
[0123] A fifth objective of an embodiment of the present invention is to provide a computer program product, wherein the computer program product comprises computer instructions, and the computer instructions instruct a computer to execute the above-mentioned method for generating a weather forecast service text dataset.
[0124] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0125] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0126] The present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, readable storage media, optical storage, etc.) containing computer-usable program codes.
[0127] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0128] Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A method for generating a weather forecast service text dataset, characterized in that: The following steps are involved: Collect weather forecast service data; Preprocessing the weather forecast service data to obtain weather forecast service text data; It is necessary to determine the types of meteorological elements that the extended data needs to include according to the application scenario. For example, if you need to generate weather forecast service text data of a single meteorological element, you only need to obtain the data containing the target element from the original data, build a few-shot prompt word project based on the data containing the target element, and use the large language model and thinking chain technology to enable the large language model to expand the weather forecast service text data according to the prompt word instructions of the few-shot prompt word project. If it is necessary to generate weather forecast service text data of multiple meteorological elements, it is necessary to build a knowledge vector library, obtain data related to the specified meteorological elements from the preprocessed data through the data title, category, and release time; build a knowledge vector library based on the data related to the specified meteorological elements to generate extended data, which is divided into the following steps: Step 1: Use content-based data slicing method to analyze the chapter structure of the data, use the title level to determine the relationship between paragraphs, and then use the large language model to determine whether the content of adjacent paragraphs is similar. If it exceeds the threshold, merge the paragraphs so that each piece of data contains complete meteorological information; Step 2: Use the fine-tuned embedding model to encode the sliced data, convert the text data into a set of high-dimensional vectors, so that the high-dimensional vectors contain the deep semantic structure of the text, and store them in the vector database; Step 3: Based on the pre-generated data application requirements, a hybrid search method is used to search and recall the specified meteorological element content; Step 4: Construct prompt words, combine them with the retrieved content, and let the large language model expand the data according to the logic of the prompt words; generate new weather forecast service text data containing multiple meteorological elements, and obtain the weather forecast service text extension data set after expansion; The weather forecast service text extension dataset is tested and evaluated by combining a large language model with a quantitative evaluation algorithm. Specifically: The quantitative evaluation algorithm uses the GLM model as an evaluation model to evaluate the extended data of the weather forecast service text extended data set and determine whether the extended data can describe the complete content of the original data; the GLM model is used to extract the description content corresponding to the meteorological elements of each level in the original data and the extended data, and then the Jaccard similarity coefficient between the two is calculated; The quantitative evaluation algorithm uses the Jaccard similarity coefficient, and the evaluation index is: , in, Represents the description content in the original data, Represents the description content in the extended data. The Jaccard similarity coefficient represents the similarity between the two. It is used to judge the similarity between the two. The higher the result, the more similar the two are. The evaluation of weather forecast service text data with multiple meteorological elements also includes: First, the GLM model is used to evaluate the extended data as a whole to determine whether it contains the content of the specified meteorological elements and whether it conforms to the commonly used description logic; Then use the GLM model to extract the description content corresponding to each level of meteorological elements in the extended data and recalled data slices, and calculate the evaluation index: in, Represents the type of meteorological element, Represents the number of meteorological elements. Representative recall data involving meteorological elements Description content, Represents meteorological elements involved in extended data Description content, represents the Jaccard similarity coefficient between the two; The weather forecast service text dataset is constructed by using the data in the weather forecast service text extension dataset that meets the verification and evaluation requirements.
2. The method for generating a weather forecast service text data set according to claim 1, characterized in that: The collecting of weather forecast service data includes: By combining manual screening with web crawler technology, weather forecast service data for the public, media and meteorological practitioners are collected in batches from multiple online platforms.
3. The method for generating a weather forecast service text data set according to claim 1, characterized in that: The preprocessing of the weather forecast service data to obtain the weather forecast service text data includes: Design the storage structure of the data set, which includes four parts: title, category, time, and content. The title represents the title of the original data, the category represents the meteorological elements contained in the original data, and needs to be extracted in a standardized manner by combining the large language model prompt word technology and data annotation algorithm. The time represents the release time of the data, and the content represents the main text of the original data. The weather forecast service data is preprocessed based on the storage structure of the data set to obtain weather forecast service text data.
4. A weather forecast service text data set generation system, based on the weather forecast service text data set generation method according to any one of claims 1 to 3, characterized in that: include: A collection module, used to collect weather forecast service data; A preprocessing module, used for preprocessing weather forecast service data to obtain weather forecast service text data; The expansion and evaluation module is used to expand the weather forecast service text data based on the large language model to obtain the weather forecast service text expansion data set; the weather forecast service text expansion data set is tested and evaluated by combining the large language model with the quantitative evaluation algorithm; The construction module is used to construct the weather forecast service text dataset from the data in the weather forecast service text extension dataset that meets the verification and evaluation requirements.
5. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the method for generating a weather forecast service text data set as described in any one of claims 1 to 3 when executing the computer program.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for generating a weather forecast service text data set according to any one of claims 1 to 3 is implemented.
7. A computer program product, comprising computer instructions, characterized in that: The computer instructions instruct the computer to execute the method for generating a weather forecast service text data set as described in any one of claims 1-3.
Citation Information
Patent Citations
Meteorological disaster early warning scheme generation method based on deep learning
CN112435447A
Aviation safety report analysis and evaluation method based on text similarity model
CN115186660A
Text-image generation method, system and device and storage medium
CN117095083A
Knowledge graph construction method and device based on pre-trained large language model
CN117851610A