A Smart Question-Answering Method and System for Event Weather Based on a Large Language Model

By constructing a multi-source knowledge base and dedicated data interface tools, and combining large language models for intent recognition and data scheduling, the problem of information fragmentation in large-scale sports events has been solved, achieving efficient and accurate one-stop meteorological services.

CN122087066APending Publication Date: 2026-05-26广州市气象综合保障中心(广州市突发事件预警信息发布中心广州市气象数据中心) +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
广州市气象综合保障中心(广州市突发事件预警信息发布中心广州市气象数据中心)
Filing Date
2026-03-20
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies struggle to provide efficient and accurate one-stop weather services for large-scale sporting events. Users have to piece together information from multiple platforms themselves, resulting in incomplete and untimely information that fails to meet the event organizers' needs for efficient and accurate weather support.

Method used

We construct a multi-source knowledge base and dedicated data interface tools, and combine a large language model to perform intent recognition and data scheduling to generate multi-source information fusion responses that conform to user intent.

Benefits of technology

It enables users to obtain comprehensive responses such as accurate weather, event information, and travel advice in one go, improving the convenience and efficiency of information acquisition. The system has good scalability and maintainability, and can adapt to future changes in data sources and service needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122087066A_ABST
    Figure CN122087066A_ABST
Patent Text Reader

Abstract

This invention relates to the field of meteorological information service technology, and discloses a method and system for intelligent question answering in event meteorology based on a large language model. It aims to solve the technical problem that existing meteorological services cannot intelligently integrate multi-source heterogeneous information to respond to complex user needs. The method includes: constructing a multi-source knowledge base covering basic meteorological knowledge and event services; creating and encapsulating dedicated data interface tools; receiving user natural language questions, performing intent recognition based on a large language model, dynamically scheduling the interface tools and knowledge base to acquire multi-source data; and finally generating a natural language response that matches the user's intent through data fusion processing. This invention, through a combination of knowledge base construction, dedicated meteorological tools for events, intelligent scheduling, and multi-source fusion, realizes the transformation of event meteorological services from "fragmented" to "one-stop," improving the intelligence level of meteorological services and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of meteorological information service technology, and in particular to an intelligent question-and-answer method and system for sports event meteorology based on a large language model. Background Technology

[0002] Currently, public meteorological services for large-scale sporting events (such as the 15th National Games) mainly rely on traditional meteorological information dissemination channels and manual consultation services. These methods have revealed significant shortcomings in practice: when users raise comprehensive questions involving multiple dimensions of information such as weather, event schedules, venues, and travel, such as "Will the track and field events at Tianhe Stadium tomorrow afternoon be affected by rain, and what precautions should spectators take?", existing systems often struggle to provide satisfactory answers. The service experience is fragmented, forcing users to separately search for weather conditions, event schedules, and traffic guidance on different platforms, and then manually piece together and logically interpret the information. This method of information acquisition is not only inefficient, but also, in scenarios like sporting events where timeliness and accuracy are paramount, can lead to incomplete, untimely, or even inconsistent information for the public, severely impacting their viewing experience and travel arrangements, and failing to meet the urgent needs of event organizers for efficient and accurate meteorological support. Therefore, how to break through the limitations of existing service models in the complex scenario of large-scale events and provide the public with new meteorological services that can accurately understand the comprehensive intentions of the events and provide one-stop solutions in a timely manner has become a prominent challenge in the current situation. Summary of the Invention

[0003] The purpose of this invention is to overcome the shortcomings of existing service models and, addressing the urgent public demand for comprehensive meteorological services during major sporting events, provide a smart question-and-answer method and system for event meteorology based on a large language model. This invention aims to solve the core technical challenge of existing service models in providing accurate, efficient, and one-stop information services when faced with complex and comprehensive natural language queries from users. It offers meteorological services covering a wide range of scenarios, including weather inquiries, event information, disaster warnings, travel advice, and science Q&A, thereby improving the intelligence level of public meteorological services and user experience.

[0004] To achieve the above objectives, a first aspect of the present invention provides an intelligent question-answering method for sports weather based on a large language model, comprising the following steps: Construct a multi-source knowledge base related to event meteorological services, wherein the multi-source knowledge base includes at least a basic meteorological knowledge base and an event service knowledge base; Create and encapsulate several dedicated data interface tools, including at least a weather query interface tool for obtaining real-time or forecast weather data and a sports query interface tool for obtaining sports event information; Receive natural language questions from users; The intent of the natural language question is identified based on the large language model, and the corresponding dedicated data interface tool is dynamically scheduled and invoked to obtain the first data and / or the corresponding multi-source knowledge base is retrieved to obtain the second data based on the identified intent.

[0005] Input the first data and / or the second data into the large language model; The large language model fuses the first data and / or the second data to generate and output a natural language response based on multi-source information and conforming to the user's intent.

[0006] A second aspect of the present invention provides an intelligent question-and-answer system for sports weather based on a large language model, comprising: The knowledge base construction module is used to construct a multi-source knowledge base related to event meteorological services. The multi-source knowledge base includes at least a basic meteorological knowledge base and an event service knowledge base. An interface tool encapsulation module is used to create and encapsulate several dedicated data interface tools, including at least a weather query interface tool for obtaining real-time or forecast weather data and an event query interface tool for obtaining event information. The input receiving module is used to receive natural language questions input by the user; The intent recognition and scheduling module is communicatively connected to the input receiving module, the knowledge base construction module, and the interface tool encapsulation module, respectively. The intent recognition and scheduling module is used to: perform intent recognition on natural language questions from the input receiving module based on a large language model, and, based on the recognized intent, initiate a call request to the interface tool encapsulation module to obtain first data, and / or initiate a retrieval request to the knowledge base construction module to obtain second data. The data fusion processing module is communicatively connected to the intent recognition and scheduling module, and is used to receive the first data and / or the second data, and input the data into the large language model for fusion processing to generate a natural language response based on multi-source information and conforming to the user's intent. The output module, which is communicatively connected to the data fusion processing module, is used to output the generated natural language response to the user interface.

[0007] The intelligent question-answering method and system for sports weather based on a large language model proposed in this application has at least the following beneficial effects: Users no longer need to switch between different applications or platforms; they only need to ask a question in natural language to receive a comprehensive, contextualized response that integrates accurate weather, event information, travel advice, and science popularization, greatly improving the convenience and efficiency of information acquisition. The system can accurately analyze the complex needs implied in the user's natural language (such as the potential impact of weather on events, travel, and clothing), and intelligently schedule and integrate multi-source information accordingly, achieving a service upgrade from "mechanical response" to "intelligent understanding." Through standardized interface encapsulation and a continuous learning and optimization mechanism, the system not only possesses good scalability and maintainability, adapting to future data sources and service needs, but also continuously improves service quality based on actual user feedback, enhancing the system's flexibility and maintainability. Attached Figure Description

[0008] Figure 1 A schematic diagram of a sports event meteorological intelligent question-answering method based on a large language model is provided for an embodiment of the present invention; Figure 2 This invention provides an overall technical roadmap for a sports event meteorology intelligent question-answering method based on a large language model, as exemplified by this invention. Figure 3 A schematic diagram illustrating the process of knowledge base construction and RAG technology provided in this embodiment of the invention; Figure 4 A workflow design diagram for a weather query tool provided in an embodiment of the present invention; Figure 5 A workflow design diagram for scheduling the 15th National Games provided in this embodiment of the invention; Figure 6 This invention provides a schematic diagram illustrating the principle of a ReAct prompt method. Figure 7 This is a schematic diagram of a tool MCP-based process provided in an embodiment of the present invention; Figure 8 This is a schematic diagram illustrating a collaboration method between the Dify platform and the RAGFlow platform, provided in an embodiment of the present invention. Figure 9 A schematic diagram illustrating the principle structure of a sports event weather intelligent question-answering system based on a large language model, provided for an embodiment of the present invention; Figure label: Knowledge base construction module-100, interface tool encapsulation module-200, input receiving module-300, intent recognition and scheduling module-400, data fusion processing module-500, output module-600. Detailed Implementation

[0009] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0010] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0011] Example 1 This invention provides a preferred embodiment of an intelligent question-answering method for sports event weather based on a large language model. This embodiment uses the meteorological service support work for the 15th National Games as an application scenario example to illustrate the specific implementation process and technical effects of the invention. Those skilled in the art will understand that the following specific embodiments are only for explaining the invention and are not intended to limit the scope of protection of the invention. See also Figure 1 The following is a flowchart illustrating an intelligent question-answering method for sports weather based on a large language model, provided as an example of an embodiment of the present invention, including the following steps: S100. Construct a multi-source knowledge base: Construct a multi-source knowledge base related to the meteorological services for the event, wherein the multi-source knowledge base includes at least a basic meteorological knowledge base and an event service knowledge base.

[0012] Specifically, in this embodiment, the multi-source knowledge base refers to a knowledge storage system composed of multiple sets of knowledge from different sources and of different types, including at least a meteorological basic knowledge base and a sports event service knowledge base, which can provide comprehensive and rich knowledge support for intelligent Q&A on sports event meteorology; the meteorological basic knowledge base refers to a database specifically storing basic knowledge in the field of meteorology, covering the definition of meteorological elements, the formation principles of weather phenomena, the classification and characteristics of climate types, etc., which is an important basis for understanding, analyzing and forecasting meteorological phenomena; the sports event service knowledge base refers to a database used to store knowledge related to sports event services, including basic event information and venue information, which helps to provide comprehensive services to event-related personnel.

[0013] One possible approach involves first collecting materials for building the knowledge base. This collection should encompass authoritative sources in the meteorological field, such as professional textbooks, technical specifications, historical meteorological data, and early warning signal standards, as well as service documents related to the 15th National Games, including event schedules, venue introductions, spectator guides, and transportation information. After collection, the materials undergo systematic preprocessing, including format standardization (converting documents of different formats to a standard text format), content deduplication using the SimHash algorithm (a text deduplication algorithm) with a similarity threshold of 95%, and text cleaning to remove irrelevant characters and garbled text.

[0014] The knowledge base construction process involves the following two steps: First, based on the RAGFlow platform (an LLM rapid development platform that includes knowledge base construction functionality, whose integrated OCR algorithm can accurately identify the layout and content of fuzzy documents), the preprocessed documents are imported into the platform. Text slicing parameters are set: sentence-level slices are 128-256 characters long, paragraph-level slices are 512-1024 characters long, and document-level slicing is used for technical specification documents with high integrity requirements. Then, the bge-large-zh-v1.5 embedding model (a Chinese semantic vector embedding model) is used for vectorization to generate vector representations. Second, a dual-index structure is established—a BM25 keyword index (a keyword retrieval index based on the BM25 algorithm) and a vector index. When recalling knowledge base content based on query statements, the dual-index approach improves the accuracy of knowledge matching.

[0015] S200. Create and encapsulate dedicated data interface tools: Create and encapsulate several dedicated data interface tools, including at least a weather query interface tool for obtaining real-time or forecast weather data and a sports event query interface tool for obtaining sports event information.

[0016] Specifically, a dedicated data interface tool refers to a tool used to enable data interaction and transmission between different systems or data sources. In this embodiment, it includes at least a weather query interface tool and an event query interface tool, which can acquire and convert the required data according to specific protocols and rules. The weather query interface tool is a dedicated data interface tool specifically used to acquire real-time or forecast weather data from meteorological data sources. It connects with the meteorological department system, sends requests and receives meteorological element information according to preset interface protocols and data formats. The event query interface tool is a dedicated data interface tool mainly used to acquire event information from event-related data sources. It connects with the event organizer's management system, etc., and acquires various types of event data according to specific query conditions and interface specifications.

[0017] In one possible implementation, this embodiment of the invention employs multiple dedicated interfaces, including weather queries and event information queries. During the interface development phase, the following core interfaces are implemented: The weather query interface is developed based on a RESTful architecture (representational state transition, HTTP interface specification), receiving latitude and longitude coordinates (accuracy requirement of 4 decimal places) and time range parameters, and returning real-time weather, 24-hour forecast, and 7-day forecast data. The interface response time is optimized to 200-500 milliseconds. The event query interface supports multi-dimensional queries by venue name, match date, sport, etc., using a MySQL 8.0 database (a relational database management system, version 8.0), and improving query performance by establishing composite indexes. The interface encapsulation adopts the MCP protocol (model compatibility protocol, interface standardization encapsulation specification) for standardization processing. Specific implementation includes: first, defining tool calling specifications, clarifying the input and output parameters of each interface according to the JSON Schema format (JSON data schema, defining interface specifications); then, establishing an interface wrapper layer, registering each interface as a callable entity; and finally, configuring a model adaptation layer to generate corresponding calling parameters based on different types of large language models. This encapsulation method allows different models to call tools in a unified way, improving the system's compatibility and maintainability.

[0018] S300: Receives user input via natural question input.

[0019] Specifically, in the user input reception stage, users' questions or requests, expressed in natural language, are flexible and diverse, requiring the system to accurately understand their intent to provide appropriate answers. One possible implementation is to receive user natural language questions through a front-end interface, supporting both text and voice input. A robust input filtering mechanism should be established to reasonably limit the length of questions, while also implementing real-time detection and blocking of sensitive words to ensure the security and standardization of input content. Preferably, user input can be received through a WeChat mini-program front-end interface, with voice input implemented using the Tencent Cloud Speech Recognition API (a speech-to-text interface provided by Tencent Cloud), achieving a recognition accuracy greater than 95%. An input filtering mechanism should be established to limit the length of questions to 10-200 characters, while also implementing real-time detection and blocking of sensitive words to ensure the security and standardization of input content.

[0020] S400. Intent recognition and data retrieval based on a large model: The intent of the natural language question is recognized based on a large language model, and according to the recognized intent, the corresponding dedicated data interface tool is dynamically scheduled and called to obtain the first data, and / or the corresponding multi-source knowledge base is retrieved to obtain the second data.

[0021] Specifically, in this embodiment, the first data refers to the data obtained by dynamically scheduling and calling the corresponding dedicated data interface tool, such as weather data obtained by calling the weather query interface tool or event information obtained by calling the event query interface tool; the second data refers to the data obtained by searching the corresponding multi-source knowledge base, such as meteorological principle knowledge obtained from the meteorological basic knowledge base or event rules and venue information obtained from the event service knowledge base.

[0022] In one possible implementation, after the system receives a natural language question from the user, it enters the intent recognition and scheduling stage.

[0023] In the intent recognition stage, the system uses a continuously trained large language model as a foundation to construct an intent classification system, thereby accurately identifying the simple and complex intents contained in user questions. Among them, the "Fenghe" large model (a meteorological vertical domain large language model, which is a model that is specifically fine-tuned and trained using meteorological data based on the general large language model ChatGLM) is preferred.

[0024] For example, when faced with complex questions such as "Will the track and field competition at Tianhe Stadium tomorrow afternoon be affected by rain, and what should spectators pay attention to?", the system can not only identify the explicit need for weather inquiries, but also gain a deeper understanding of the implicit needs for event arrangement consultation and travel advice.

[0025] The dynamic scheduling process employs an execution plan-based scheduling mechanism. When a user poses a complex question, the large language model first generates a structured execution plan, specifying the interfaces and knowledge bases to be invoked and the execution order. For example, if a user asks, "Will it rain during the track and field competition at Tianhe Stadium tomorrow afternoon? Should I bring an umbrella?", the execution plan generated by the large language model would be: call the weather query interface, setting the parameters to the coordinates of Tianhe Stadium and the time range for tomorrow afternoon; retrieve the track and field competition schedule from the event knowledge base; and retrieve the meteorological science knowledge base to obtain rain protection measures. Subsequently, the system dynamically schedules the corresponding interfaces and knowledge bases according to this plan.

[0026] For complex intents requiring the invocation of multiple tools, the system meticulously plans the invocation order and parameter passing dependencies of each tool, and supports parallel invocation to improve processing efficiency. For example, when a user asks about the weather conditions of multiple venues, the system will call multiple weather query interfaces in parallel to quickly and efficiently meet the user's needs.

[0027] S500, Input the acquired data into the large language model: Input the first data and / or the second data into the large language model; S600, Large Language Model Fusion Data to Generate Natural Response: The large language model fuses the first data and / or the second data to generate and output a natural language response based on multi-source information and in accordance with the user's intent.

[0028] Steps S500 and S600 are the specific implementations of data input and fusion processing. In one possible implementation, during the data input stage, a data standardization pipeline is established to convert the JSON data (data in JavaScript object representation format) returned by the interface into natural language descriptions. The knowledge base retrieval results are ranked by relevance, and the top 3 (the top 3 results with a relevance score greater than 0.7) fragments are selected as input. The fusion processing stage employs a chain-of-thought-based prompt optimization technique (a prompting technique that guides the model's step-by-step reasoning). Specifically, the large language model is guided to reason according to the following logic: First, weather data is analyzed to identify key information such as precipitation probability, temperature, and wind force; combined with event information, influencing factors such as event type and venue conditions are considered; the knowledge base content is integrated, incorporating meteorological science knowledge and protective advice; a final response is generated, transforming professional information into easily understandable natural language. Simultaneously, the role of the large language model is predefined as a "event meteorological service consultant" in the prompt words, and output format requirements and compliance check rules are set to ensure that the generated content is both professional and accurate, and meets regulatory requirements. Through the above specific implementation, this embodiment successfully solves the technical problem of the inability to intelligently integrate multi-source heterogeneous information. Actual testing shows that the system has a high accuracy rate in intent recognition when handling complex user queries, the average response time per request is controlled within a reasonable range, and it supports a large number of concurrent user accesses. Compared with traditional service models, this invention achieves a qualitative leap from "fragmented service" to "one-stop intelligent service," improving the quality and efficiency of meteorological services for major sporting events.

[0029] This embodiment effectively solves the technical problem of the inability to intelligently integrate multi-source information in existing technologies through the above-described technical solution. Specifically, by constructing a unified multi-source knowledge base and a standardized interface tool layer, it achieves the organic integration of heterogeneous data; through precise intent recognition and intelligent scheduling mechanisms, it achieves a deep understanding of complex user needs; and through this content fusion generation method, it achieves the popularization of professional information. The organic combination of these technical means enables the system to provide users with one-stop, scenario-based intelligent weather services, solving the problem of information fragmentation in traditional service models in existing technologies.

[0030] Example 2 Based on Embodiment 1, this embodiment of the present invention also provides another embodiment of an intelligent question-answering method for sports event meteorology based on a large language model. This embodiment aims to provide multiple preferred implementation schemes to fully demonstrate the technical diversity and feasibility of the present invention. It integrates the core needs of meteorological services and sports event support, achieving intelligent fusion and precise response of multi-source heterogeneous information. This embodiment uses the meteorological service support work for the 15th National Games as an application scenario example to illustrate the specific implementation process and technical effects of the present invention in detail. Those skilled in the art will understand that the following specific embodiments are only for explaining the present invention and are not intended to limit the scope of protection of the present invention.

[0031] See Figure 2 This invention provides an example of an intelligent question-answering method for sports weather based on a large language model, outlining a general technical roadmap comprising five core stages: "knowledge base construction → base model selection and testing → agent construction → agent application and testing → continuous optimization." These stages are progressively related. The agent construction stage includes sub-modules such as interface development, workflow design, tool encapsulation, and prompt word optimization. The agent application and testing stage includes sub-steps such as interface integration, internal testing, and online testing. This framework achieves a closed-loop process from knowledge accumulation, model adaptation, and function construction to application deployment, effectively solving the core technical problem of the inability to intelligently integrate and respond to complex sports weather needs from users across multiple heterogeneous sources.

[0032] In the overall implementation process, the first step is to construct a multi-source knowledge base related to event meteorological services. This multi-source knowledge base includes at least a basic meteorological knowledge base and an event service knowledge base. Combined with relevant technical solutions, the construction and optimization of the knowledge base are achieved through the following steps: One possible approach is to extensively collect authoritative meteorological data and documents related to the 15th National Games. The collection scope covers basic meteorological knowledge (such as professional textbooks like "Principles of Meteorology" and "Handbook of Atmospheric Detection Technology," technical documents such as warning signal standards and meteorological observation specifications issued by the National Meteorological Administration, and easy-to-understand meteorological science manuals) as well as event services (the official event schedule of the 15th National Games, detailed introductions of the 12 venues in the Guangdong-Hong Kong-Macao Greater Bay Area, spectator guides, traffic guidance, and event meteorological support plans, etc.), accumulating more than 5,000 documents to ensure comprehensive knowledge coverage.

[0033] In one possible approach, a systematic preprocessing operation is performed on the collected documents. Scanned PDF documents (such as scanned copies of some old meteorological observation standards) undergo OCR recognition using the OCR algorithm integrated into the RAGFlow platform. This algorithm possesses excellent fuzzy document recognition capabilities, accurately identifying the layout structure of the scanned documents (such as column text, tables, and formula positions) and extracting the text content, effectively solving the problem of low recognition rate for fuzzy documents using traditional OCR. Format unification processing is performed on documents of different formats (Word, Excel, PPT, PDF), converting them to standard UTF-8 encoded text format. The SimHash algorithm is used for content deduplication, setting a similarity threshold of 95% to automatically remove duplicate or highly similar document segments (such as repeatedly issued meteorological warning standards). Regular expression matching is used to remove irrelevant characters (such as special symbols, page numbers, watermark text) and garbled characters from the documents, ensuring the quality and consistency of the preprocessed data.

[0034] In one possible implementation, when performing text slicing processing using the RAGFlow platform, a hybrid slicing strategy is employed, using sentence-level, paragraph-level, and document-level methods simultaneously for slicing based on the document's logical structure, to adapt to the knowledge storage needs of different document types. In a preferred embodiment, text slicing processing using the RAGFlow platform includes: employing a hybrid slicing strategy, using sentence-level, paragraph-level, and document-level methods simultaneously for slicing based on the document's logical structure. Figure 3 This invention provides a flowchart illustrating the knowledge base construction and RAG (Retrieval Enhanced Generation) technology, which is a core technical path for achieving intelligent fusion of multi-source heterogeneous information. Specifically, it includes a complete closed loop from raw data processing to intelligent response generation: For simple documents such as short meteorological science articles and introductions to sports venues, sentence-level segmentation is used, with a segment length of 128-256 characters, ensuring that each knowledge fragment focuses on a single knowledge point (e.g., "Typhoon warning signals are divided into four levels: blue, yellow, orange, and red"). For medium-length structured documents such as meteorological technical specifications and sports event support plans, segmentation is used, with a segment length preferably of 512-1024 characters, preserving logical connections within paragraphs (e.g., description of the meteorological observation equipment layout and data acquisition frequency of a venue). For documents with high integrity requirements (e.g., the overall meteorological support plan for the 15th National Games and chapters in meteorological textbooks), document-level segmentation is used to ensure the coherence of the knowledge system. In addition, for documents containing tables (such as event schedules and meteorological statistics), a hybrid slicing method of "extracting the whole table + splitting the cell content" is adopted, which not only preserves the table structure information, but also facilitates the retrieval of individual data items.

[0035] In one possible implementation, vectorization processing employs the bge-large-zh-v1.5 or bge-m3 embedding model to transform the sliced ​​knowledge fragments into vector representations. In a preferred embodiment, the vectorization processing includes: using the bge-large-zh-v1.5 or bge-m3 embedding model to transform the sliced ​​knowledge fragments into vector representations. For knowledge fragments dense with meteorological terminology (such as "correlation analysis of relative humidity, dew point temperature, and pressure field"), the bge-large-zh-v1.5 embedding model is preferred, as this model has better semantic capture capabilities for Chinese professional texts, generating a high-dimensional vector representation of 1024 dimensions, which can accurately characterize the semantic features of professional knowledge. For event service texts (such as venue addresses and spectator information), the bge-m3 embedding model can be selected to generate a 768-dimensional vector representation, reducing storage and computational overhead while ensuring semantic accuracy. During the vectorization process, a batch processing method is adopted, with 128 knowledge fragments processed in each batch. GPU acceleration (such as NVIDIA A100 graphics card) is used to improve processing efficiency. Finally, all vectors are stored in the Milvus vector database to build a structured semantic knowledge network.

[0036] In a preferred embodiment, constructing a multi-source knowledge base related to event weather services further includes optimizing the knowledge base, specifically including the following steps: employing a hybrid retrieval enhancement method, combining vector retrieval and BM25 keyword retrieval for knowledge retrieval; receiving user feedback on the answers; and performing add, delete, or modify operations on the knowledge base content based on the feedback. For example, using a hybrid retrieval enhancement method, combining vector retrieval and BM25 keyword retrieval for knowledge retrieval, where vector retrieval ensures semantic-level association matching by calculating the cosine similarity between the user's query vector and the knowledge base vector, and BM25 keyword retrieval achieves accurate keyword matching based on word frequency statistics. The synergistic effect of both improves the comprehensiveness and accuracy of knowledge retrieval. For example, when a user asks "high-temperature protection recommendations for the 15th National Games track and field competition," vector retrieval retrieves semantic fragments related to "sports protection under high-temperature weather conditions," and BM25 retrieval retrieves fragments containing keywords such as "15th National Games track and field competition" and "high-temperature warning"; and building a user feedback mechanism. The data collection module uses the "15th National Games Meteorological Service" mini-program interface to set up "Answer Satisfaction Rating" (1-5 points) and "Feedback" input boxes to receive user feedback on the answers. Based on the feedback, it performs add, delete, or modify operations on the knowledge base content. For example, when multiple users report that "the answer to the rain protection measures of a certain venue is inaccurate," technicians check the corresponding segment in the knowledge base and supplement the specific drainage facility parameters and temporary rain shelter information for that venue. When users report that a knowledge segment is repeated, the module automatically marks and deletes the duplicate content. When the knowledge point asked by a user is not covered, relevant authoritative information is added to the knowledge base and sliced ​​and vectorized updates are completed.

[0037] After completing the construction of the multi-source knowledge base, several dedicated data interface tools are created and encapsulated. These dedicated data interface tools include at least a weather query interface tool and a sports event query interface tool. Combining relevant technical solutions, the development, encapsulation, and workflow design of the interfaces are implemented, along with the standardization of the tool's MCP (Multi-Channel Programming Protocol). The specific encapsulation steps are as follows: First, organize the tool list, listing all business interfaces and their input parameters and output structures. The dedicated data interfaces developed in this embodiment include weather interfaces for obtaining real-time and forecast weather data, effective warning signal interfaces, radar image interfaces, typhoon status and forecast interfaces, event information interfaces, and venue information interfaces, totaling 6 types of core interfaces and 2 types of auxiliary interfaces (such as time conversion interfaces and search engine interfaces).

[0038] The development details of the reference interface are as follows: For example, the weather query interface is developed based on a RESTful architecture, which receives latitude and longitude coordinates (precision requirement: 4 decimal places) and returns data in three time dimensions: "real-time", "24 hours", and "7 days", including temperature (precision ±0.1℃), humidity (precision ±1%), wind direction (azimuth 0-360°), wind speed (unit: m / s, precision ±0.1m / s), precipitation probability (0-100%), and weather phenomena (such as sunny, cloudy, and rainstorm). The interface response time is optimized to 200-500 milliseconds. By optimizing the database query statement (such as creating a latitude and longitude index) and using Redis to cache hot data (such as real-time weather of popular venues), the response efficiency in high-concurrency scenarios is ensured.

[0039] Secondly, most interfaces cannot be used directly and need to be encapsulated into tools that can be directly called by the model through workflow design. For example, the weather interface requires latitude and longitude as input parameters, but users generally do not provide latitude and longitude information when asking questions; instead, they provide the location. Therefore, it is necessary to first extract the location from the user's question, convert the location into latitude and longitude, and then use it as a parameter to send a request to the weather interface. Furthermore, it is necessary to design a system to handle the case where the user's question does not contain location information.

[0040] See Figure 4This document provides a workflow design diagram for a weather query tool, exemplified by an embodiment of the present invention. The process starts at the "Start Node," with the input variable being the address parameter extracted from the user's query by the large language model. Subsequent nodes connect to the "Conditional Branch Node," which checks if the address is empty. If the address is not empty, the process proceeds to step 1, sequentially passing through the "Code Execution Node" (calling the Gaode API to convert the address to a standard address and then obtaining the latitude and longitude of this address), the "HTTP Request Node" (requesting data from the weather interface using the latitude and longitude as input parameters), and the "Code Execution Node" (processing the data returned by the interface and filtering irrelevant fields such as historical meteorological data). Finally, the processed weather data is returned via the "End Node." If the address is empty (i.e., the user's query does not contain a location), the process proceeds to step 2, with the default location being Guangzhou Tianhe Stadium (the main venue for the 15th National Games). The weather interface data is obtained using its latitude and longitude (longitude: 113.3215, latitude: 23.1357) as parameters. After processing by the "Code Execution Node," the result is returned via the "End Node." Each HTTP request node in the process is configured to retry three times upon failure to ensure the reliability of data acquisition.

[0041] See Figure 5 This example illustrates a workflow design for arranging events for the 15th National Games. The process begins at the "Start Node," requiring no input parameters. It then directly connects to the "HTTP Request Node" to request data from the event query interface, whose parameters by default include all data from the 15th National Games. Next, it connects to the "Code Execution Node" to format the raw data returned by the interface (e.g., extracting key information such as event details, time, and venue, and removing redundant fields). Finally, it returns structured event information via the "End Node," such as "August 20th, 9:00-11:00, Track and Field Events, Guangzhou Tianhe Stadium, Men's 100m Preliminaries."

[0042] Then, the tool is MCP-ized to make it more standardized and improve the accuracy of model calling tools.

[0043] Define the MCP tool protocol specification and write the tool call specification according to the JSON Schema format. Taking the weather query interface tool (tool name "getWeatherForecast") as an example, its specification definition is as follows: json { "name":"getWeatherForecast", "description":"Query real-time or forecast weather data for the event area, supporting all venues and surrounding areas of the 15th National Games". "parameters":{ "type":"object", "properties":{ "longitude":{ "type":"string", "description": "Longitude of the target area, with a precision of 4 decimal places, for example: 113.3215" }, "latitude":{ "type":"string", "description": "The latitude of the target area, with a precision of 4 decimal places, for example: 23.1357" } }, "required":["longitude","latitude"] } } Similarly, the JSON Schema specification for event query interfaces, venue information interfaces, etc., should be defined in accordance with the above format, clearly specifying the name, function description, input parameters (type, constraints, description) and required fields of each interface, to ensure that the large language model can accurately identify the interface call requirements.

[0044] The MCP tool is registered as a callable entity through an intermediate layer. A wrapper layer is preferably developed using the Python Flask framework to encapsulate a unified call entry point for each interface, implementing functions such as receiving interface requests, parameter validation, format conversion, and response return. For example, when the large language model initiates a weather query interface call request, the wrapper layer first validates whether the input parameters conform to the JSON Schema specification (e.g., whether the latitude and longitude format is correct, and whether the time range is within the enumerated values). If the parameters are abnormal, an error message is returned; if the parameters are valid, the request format is converted into an HTTP request format that the interface can recognize, forwarded to the weather query interface, and then converted into a structured data format that the large language model can process after receiving the returned data.

[0045] Based on different types of large language models, corresponding prompt words and function call parameter structures are generated. This embodiment adapts to the following large language models: DeepSeek-R1, Qwen3, and the "Fenghe" large model. For models that natively support FunctionCalling (such as Qwen3), the model adaptation layer directly generates parameter structures conforming to their function call format (such as JSON objects containing "name" and "parameters" fields). For models that do not natively support FunctionCalling (such as DeepSeek-R1), prompt words are generated through ReAct (inference-action prompts) to guide the model to initiate interface calls according to the logic of "inference needs → calling tools → obtaining results." In a preferred embodiment, the encapsulation using the MCP protocol includes the following steps: organizing tools, listing all business interfaces and their input parameters and output structures; defining the MCPSchema specification, writing tool call specifications according to the JSON Schema format; constructing an interface wrapper layer, registering the MCP tool as a callable entity through an intermediate layer; and configuring the model adaptation layer to generate corresponding prompt words and function call parameter structures based on different types of large language models.

[0046] In a preferred embodiment, dynamically scheduling and invoking the corresponding dedicated data interface tool to obtain first data, and / or retrieving the corresponding multi-source knowledge base to obtain second data includes: dynamically scheduling and invoking the corresponding dedicated data interface tool to obtain first data, wherein the execution plan includes a tool invocation strategy and / or a knowledge retrieval strategy; and executing the tool invocation operation and / or data retrieval according to the execution plan. In a preferred embodiment, generating the execution plan includes: for complex intents that require invoking multiple dedicated data interface tools, planning the invocation order of each tool in the plan; and defining the parameter passing dependencies between tool invocations in the plan.

[0047] Next, the system receives natural language questions input by the user. This step, combined with relevant technologies, is implemented in the following way: A smart question-and-answer entry is set up through the WeChat mini-program front-end interface, supporting both text and voice input methods. For text input, an input box is provided (supporting 10-200 characters), allowing users to directly input natural language questions (such as "Will it rain during the sailing competition at Shenzhen Bay Sports Center tomorrow afternoon?" "What are the high-temperature protection measures for the track and field competitions at the 15th National Games?"). For voice input, the Tencent Cloud speech recognition API is integrated to achieve speech-to-text functionality. This API has a recognition accuracy rate of over 95% and supports Cantonese and Mandarin (adapting to the needs of users in the Guangdong-Hong Kong-Macau Greater Bay Area). Users can long-press the voice input button to record their question, and upon releasing, the system automatically completes the recognition and converts it into text.

[0048] An input filtering mechanism is established to perform dual verification of the question content: First, length verification. If the text input exceeds 200 characters or the text after speech recognition exceeds 200 characters, a prompt will pop up saying "The question content is too long, please shorten it to within 200 characters"; Second, real-time sensitive word detection. A third-party sensitive word library (covering politically sensitive, violent, and abusive words, etc.) is called up, and a string matching algorithm is used for real-time detection. If a sensitive word is detected, the input is blocked and a prompt will appear saying "The question content contains inappropriate information, please modify and re-enter" to ensure the security and standardization of the input content.

[0049] Next, intent recognition and dynamic scheduling are performed. This step uses a large language model to recognize the intent of natural language questions and dynamically schedules them based on the intent using interface tools and / or knowledge base retrieval, combined with relevant technical solutions. The specific implementation is as follows: Employing a high-quality general-purpose large-scale language model or a large-scale meteorological vertical-domain model as a foundation, when a user's question involves multiple locations or events, the system calls multiple dedicated data interface tools in parallel to obtain data. For example, in the user's question involving weather queries for two venues, the system uses multi-threading technology to call the weather query interface twice in parallel, requesting weather data for Shenzhen Bay Sports Center and Guangzhou Tianhe Stadium respectively. Compared to serial calls, the response time is reduced by more than 40%, improving data acquisition efficiency.

[0050] According to the execution plan, the system dynamically schedules the corresponding dedicated data interface tools to obtain the first data and / or retrieves the second data from a multi-source knowledge base. During the tool invocation process, the standardized execution of the interface invocation is achieved through the MCP encapsulation layer and model adaptation layer constructed above; during the knowledge retrieval process, a hybrid retrieval enhancement method (vector retrieval + BM25 retrieval) is used to retrieve relevant knowledge fragments from the multi-source knowledge base, and the top 3 fragments with a relevance score greater than 0.7 are selected as the second data. For example, if a user asks "What are the wave weather conditions required for the surfing competition of the 15th National Games?", the system retrieves 3 relevant fragments after searching the knowledge base: "The suitable wave height for surfing competitions is 1.5-2.5 meters", "Wave periods greater than 8 seconds are more suitable for surfing competitions", and "Waves in the outer impact area of ​​a typhoon are unstable and not suitable for holding surfing competitions", which are used as the second data for subsequent fusion processing. In a preferred embodiment, dynamically scheduling and invoking the corresponding dedicated data interface tools includes: when it is identified that the user's question involves multiple locations or events, multiple dedicated data interface tools are invoked in parallel to obtain data.

[0051] Then, the acquired first data (data returned by the interface) and / or second data (knowledge base retrieval fragments) are input into the large language model. Before input, data standardization processing is performed: For the first data (such as JSON format data returned by the weather interface), it is converted into a natural language description through the data standardization pipeline. For example, "temperature: 32.5℃, humidity: 65%, precipitation probability: 20%" is converted into "real-time temperature 32.5℃, relative humidity 65%, precipitation probability 20%"; For the second data (knowledge base retrieval fragments), they are sorted from high to low according to relevance scores and input in the format of "knowledge fragment 1: XXX; knowledge fragment 2: XXX; knowledge fragment 3: XXX" to ensure that the large language model can clearly identify and utilize multi-source data.

[0052] Finally, the data is fused and processed, and a natural language response is generated. In this step, the large language model fuses the data and generates a natural language response. The specific implementation, combined with relevant technical solutions, is as follows: The fusion processing of the first data and / or the second data includes performing a prompt word optimization process, which includes: using Chain-of-Thought prompts to guide the large language model in reasoning; and / or using a Prompt Slot Filling structured template to ensure output stability. In a preferred embodiment, the fusion processing of the first data and / or the second data includes performing a prompt word optimization process, which includes the following steps: using Chain-of-Thought prompts to guide the large language model in reasoning; and / or using a Prompt Slot Filling structured template to ensure output stability.

[0053] The fusion processing of the first and / or second data further includes: predefining the role of the large language model in the prompt words, wherein the role is the event meteorological service consultant; and setting the output format and compliance requirements in the prompt words. In a preferred embodiment, the fusion processing of the first and / or second data further includes: predefining the role of the large language model in the prompt words, wherein the role is the event meteorological service consultant; and setting the output format and compliance requirements in the prompt words. The prompt content for the predefined role is: "You are a professional event meteorological service consultant for the 15th National Games. Your answers should be both professional and easy to understand, avoiding the use of overly obscure meteorological terms; for meteorological warning-related content, you must strictly follow the 'Measures for the Issuance and Dissemination of Meteorological Disaster Warning Signals,' accurately and standardly express yourself, and not exaggerate or downplay the risks"; the output format requirements are: "First give a clear conclusion, then explain the basis in points, and finally provide specific suggestions"; the compliance requirements include: "It is forbidden to release unverified meteorological information, it is forbidden to use expressions that may cause public panic, and the source of meteorological data must be indicated (e.g., 'Guangdong Provincial Meteorological Bureau Forecast for X Month X Day X Hour 202X')."

[0054] The method, following step S6, further includes: performing an audit step, using a combined human and AI audit mechanism to check the compliance of the natural language response; ensuring that the response content complies with meteorological regulations and event specifications. The intelligent audit module detects the response content (e.g., whether it contains illegal statements, whether the meteorological data is accurate, and whether the warning level description is standardized) through keyword matching and a rule engine. If the audit passes, the response is directly output to the user; if suspected illegal content is detected (e.g., "There will be a severe rainstorm tomorrow, the event will definitely be canceled"), it is pushed to the human audit queue for review by meteorological service specialists, corrected, and then output to ensure the compliance and authority of the response. In a preferred embodiment, the method, following step S600, further includes: performing an audit step, using a combined human and AI audit mechanism to check the compliance of the natural language response; ensuring that the response content complies with meteorological regulations and event specifications.

[0055] For example, a user asks, "What will the weather be like for the sailing regatta at Shenzhen Bay Sports Center tomorrow afternoon? Will it be able to proceed as scheduled?" The resulting reply after processing is: "

Conclusion

Basis

Recommendations

[0056] Regarding model selection, the method further includes selecting a large language model, with the following specific steps: simultaneously testing the DeepSeek, Qwen, and "Fenghe" large models; evaluating model performance from multiple dimensions including comprehension ability, generation quality, response speed, and compliance of answer content; and selecting the base model most suitable for the meteorological service needs of the event based on the test results. In a preferred embodiment, the method further includes selecting a large language model, including: simultaneously testing the DeepSeek, Qwen, and "Fenghe" large models; evaluating model performance from multiple dimensions including comprehension ability, generation quality, response speed, and compliance of answer content; and selecting the base model most suitable for the meteorological service needs of the event based on the test results.

[0057] Regarding testing and evaluation, the method further includes performing test evaluation, specifically with the following steps: simulating high-concurrency access to test interface performance and stability before agent deployment; setting up test cases related to weather services and events; and evaluating the model's response accuracy and response time in different scenarios. In a preferred embodiment, the method further includes performing test evaluation, comprising the following steps: simulating high-concurrency access to test interface performance and stability before agent deployment; setting up test cases related to weather services and events; and evaluating the model's response accuracy and response time in different scenarios.

[0058] High-concurrency testing employed JMeter to simulate high-concurrency access scenarios. Test subjects included the weather query interface, the sports event query interface, and the overall agent service. Concurrent user counts were set at 100, 200, 500, and 1000, with each concurrency level tested for 30 minutes. Interface response time, error rate, and server resource utilization were recorded. Test results showed that when the number of concurrent users was ≤100, all interface response times were ≤500ms, the error rate was 0%, and server CPU utilization was ≤40% and memory utilization was ≤50%. When the number of concurrent users was 200, the interface response time was ≤800ms, the error rate was ≤0.5%, and server resource utilization was ≤60%. When the number of concurrent users exceeded 500, the interface response time exceeded 1500ms, and the error rate rose to over 3%. Therefore, it was determined that the agent in this embodiment supports at least 100 concurrency levels (estimated based on model response time, approximately 300-500 user accesses per minute), meeting the user access needs during the 15th National Games.

[0059] 1000 test cases were set up, covering all application scenarios, including queries with different intents, multi-tool calls, and multi-knowledge base retrieval. Functional test results show that the agent achieves 100% functionality in scenarios such as weather queries, sports information, and science Q&A, with no missing functions. Performance test results show that the single-user model's response speed is at least 30 tokens / s (1 token is approximately 1-2 Chinese characters), and the average response time per request is controlled within 1 second, meeting user experience requirements.

[0060] The test agent was tested on different devices (smartphones: iPhone 13, Huawei Mate 60, Xiaomi 14; tablets: iPad Pro, Huawei MatePad) and different browsers (WeChat built-in browser, Chrome, Safari, Firefox). The test included input functions, question and answer responses, and result display. The compatibility test pass rate was 100%, with no interface errors or functional abnormalities.

[0061] Regarding continuous optimization, the method further includes performing continuous optimization processing, the specific steps of which are as follows: Based on user feedback on the answers, the knowledge base content is added, deleted, or modified. For example, when a user reports that "the emergency shelter information of a certain venue is inaccurate", the venue information fragment in the knowledge base is checked and updated; when a user asks about "the weather conditions required for the equestrian competition of the 15th National Games" and is not covered by the knowledge base, relevant authoritative information is added and sliced ​​and vectorized. The design of prompt words was adjusted based on actual usage data. By analyzing the types of user questions and the model's response, the reasoning steps of the Chain-of-Thought prompts were optimized. For example, for high-frequency compound queries such as "event + weather + advice", the logical guidance order in the prompt words was optimized to improve the relevance of the response. Adjust tool invocation strategies based on actual usage data. For example, statistics show that users frequently query "weather in the next 2 hours," so optimize the tool invocation strategy by prioritizing the use of short-term forecast interfaces and caching the data to shorten response time. For interfaces with high failure rates (such as radar chart interfaces), add retry mechanisms and backup interfaces to improve the reliability of tool invocation.

[0062] In a preferred embodiment, the method further includes performing continuous optimization processing, including: adding, deleting, or modifying knowledge base content based on user feedback on answers; adjusting prompt word design based on actual usage data; and adjusting tool invocation strategies based on actual usage data.

[0063] The continuous optimization cycle is set to once a week. User feedback data and system operation data are collected through automated data analysis tools to generate optimization reports. The technical team then formulates and implements optimization plans to ensure that the service quality of the intelligent agent continues to improve.

[0064] The accompanying drawings involved in this embodiment also include Figure 6 , Figure 7 , Figure 8 The specific explanation is as follows: Figure 6 This is a schematic diagram illustrating the principle of a ReAct prompt method provided in an embodiment of the present invention. It shows the logical flow of the ReAct prompt guiding the DeepSeek-R1 model to implement tool calls, which is divided into an "inference phase" and an "action phase". In the inference phase, the model analyzes the user's question requirements (such as "query the weather of venue A") and determines the tool to be called (weather query interface). In the action phase, the model generates a tool call request, obtains data, organizes the results, and finally generates an answer. This clearly presents the tool call implementation logic of a non-natively supported Function Calling model. Figure 7This invention provides a schematic diagram of a tool MCP encapsulation process, illustrating the MCP encapsulation process of a dedicated data interface tool. This process is a key path to achieve standardized compatibility between the tool and large language models, and specifically includes the following progressive steps: First, define the MCP tool protocol specification (similar to JSON Schema), using a format similar to JSON Schema to clarify the tool's calling specifications, including input parameter types, output structure constraints, required fields, etc. (for example, a weather query tool needs to define latitude and longitude accuracy, time range enumeration values, etc.), laying the format foundation for multi-model compatibility; Second, establish an MCP registry center to include all MCP-encapsulated tools under unified management, forming a tool "catalog," enabling large language models to query tool functions, calling methods, and other information, achieving centralized tool discovery and scheduling; Next, build the MCP tool into a function_call format, adapting to the function calling mechanism of large language models (such as the OpenAI function call format), enabling the model to generate standardized calling instructions and trigger tool execution (for example, encapsulating a weather query tool into a function_call structure containing "name and parameters"); Finally, construct the tool chain. CI (Continuous Integration of Toolchains) ensures the stability of tools during iteration through automated testing and deployment processes, guaranteeing that tools can still be accurately invoked by models after updates. This process, through a closed loop of "standard definition - centralized management - model adaptation - continuous assurance," achieves efficient collaboration between tools and multiple large language models, providing technical support for cross-model invocation of tools such as weather and event queries in the sports meteorological intelligent agent. In a preferred embodiment, encapsulating several dedicated data interface tools includes the following steps: Obtain official meteorological data interfaces from the meteorological bureau and event data interfaces provided by authorized information providers; invoke the workflow design function of the Dify platform to perform processing operations on the data returned by the meteorological data interfaces and the event data interfaces; combine and process the processed data; encapsulate the processed data into standardized tools that can be directly called by the model; and perform unified and standardized encapsulation processing on the dedicated interface tools using the MCP protocol.

[0065] In one possible implementation, the system first obtains official meteorological data from the meteorological bureau and event data from authorized information providers. The meteorological data interface covers real-time weather observations, numerical weather prediction products, and radar satellite data, while the event data interface includes event schedules, venue information, and participant information. These interfaces are accessed via API key authentication, with a request frequency limit of 60 calls per minute and a timeout of 5 seconds. Subsequently, the system utilizes the workflow design functionality of the Dify platform, configuring a data processing flow that includes data parsing, cleaning, and transformation nodes: the parsing node extracts key meteorological elements and event information from the JSON data; the cleaning node filters outliers, removing temperature and humidity data that exceed reasonable ranges; and the transformation node unifies the data format, converting wind speed units to kilometers per hour and time to Beijing time.

[0066] The processed data is combined and processed using an association algorithm. Data from the nearest meteorological observation station is matched based on the latitude and longitude of the event venue, and data fusion rules are established, prioritizing official meteorological bureau data when discrepancies exist among multiple data sources. The processed data is encapsulated into a standard tool, defining unified input parameters and output structures. A wrapper class written in Python implements parameter validation and caching functions. Finally, the MCP protocol is used for unified and standardized encapsulation, tool calling specifications are defined according to JSON Schema, a tool registration mechanism is established, and a model adaptation layer is configured to ensure that mainstream large language models can correctly recognize and call the tool. This implementation scheme completes the unified encapsulation and standardized processing of multi-source data interfaces, providing a reliable data foundation for intelligent question answering services and effectively solving the complexity problem of multi-source heterogeneous data access and processing.

[0067] Figure 8 This diagram illustrates a collaborative mechanism between the Dify and RAGFlow platforms, as provided in an embodiment of the present invention. The RAGFlow framework includes modules for document management, document parsing algorithms, knowledge vector libraries, and knowledge retrieval algorithms, enabling the construction and retrieval of the knowledge base. The Dify framework on the right includes modules for agent management, model management, and agents, enabling the construction of agents and tool invocation. The two platforms collaborate through an external knowledge base access interface. Knowledge fragments retrieved by RAGFlow are transmitted to the Dify platform and, together with data obtained from tool invocation, serve as model input to ultimately generate an answer, thus clarifying the technical path for multi-platform collaboration.

[0068] This embodiment effectively solves the problems of rigid interaction, fragmented information, poor interface compatibility, and insufficient balance between professional and popular aspects in existing technologies for event meteorological services through the detailed implementation steps described above. In particular, it addresses the core pain point of the inability to intelligently integrate multi-source heterogeneous information. Through the construction of a multi-source knowledge base, standardized encapsulation of MCP, dynamic scheduling mechanism, and fusion processing of large language models, it achieves a qualitative leap from "fragmented service" to "one-stop intelligent service," providing efficient, accurate, and personalized meteorological service support for major events such as the 15th National Games.

[0069] Example 3 Based on the above-described method embodiments, the present invention also provides an intelligent question-and-answer system for sports event weather based on a large language model. The hardware structure and module collaborative operation mode of the intelligent question-and-answer system for sports event weather based on a large language model of the present invention will be described in detail below through several preferred embodiments. Figure 9 This embodiment provides a schematic diagram illustrating the principle structure of a sports event weather intelligent question-answering system based on a large language model, as exemplified in this embodiment. Figure 9 As shown, the system includes a knowledge base construction module 100, an interface tool encapsulation module 200, an input receiving module 300, an intent recognition and scheduling module 400, a data fusion processing module 500, and an output module. The modules communicate with each other via network interfaces, preferably using RESTful APIs based on the HTTP protocol for data exchange, ensuring efficient collaboration between the system components. Specifically, it includes: The knowledge base construction module 100 is used to construct a multi-source knowledge base related to event meteorological services. The multi-source knowledge base includes at least a basic meteorological knowledge base and an event service knowledge base. The interface tool encapsulation module 200 is used to create and encapsulate several dedicated data interface tools. The dedicated data interface tools include at least a weather query interface tool for obtaining real-time or forecast weather data and a sports query interface tool for obtaining sports event information. The input receiving module 300 is used to receive natural language questions input by the user; The intent recognition and scheduling module 400 is communicatively connected to the input receiving module 300, the knowledge base construction module 100, and the interface tool encapsulation module 200, respectively. The intent recognition and scheduling module 400 is used to: perform intent recognition on natural language questions from the input receiving module 300 based on a large language model, and, based on the recognized intent, initiate a call request to the interface tool encapsulation module 200 to obtain first data, and / or initiate a retrieval request to the knowledge base construction module 100 to obtain second data. The data fusion processing module 500 is communicatively connected to the intent recognition and scheduling module 400, and is used to receive the first data and / or the second data, and input the data into the large language model for fusion processing to generate a natural language response based on multi-source information and conforming to the user's intent. The output module 600 is communicatively connected to the data fusion processing module 500 and is used to output the generated natural language response to the user interface.

[0070] In one possible implementation, the knowledge base construction module 100 specifically includes a data collection unit, a preprocessing unit, a text processing unit, and an optimization unit. The data collection unit automatically collects document data from authoritative sources such as the meteorological bureau's official website and the event organizing committee's database using web crawler technology, with a collection frequency set to once daily to ensure the timeliness of the data. The preprocessing unit performs a multi-level processing flow on the collected documents. First, it uses a deduplication mechanism based on the SimHash algorithm, setting a similarity threshold of 95% to automatically remove duplicate content; then, it uses regular expression matching to remove special characters and garbled text; for scanned documents, it calls the OCR recognition engine integrated in the RAGFlow platform, which can accurately recognize the layout structure of blurry documents. The text processing unit uses the RAGFlow platform to intelligently slice the preprocessed documents, employing a multi-granularity slicing strategy, automatically selecting the slicing level according to the document type, and converting the text fragments into 1024-dimensional vector representations using the bge-large-zh-v1.5 embedding model. The optimization unit establishes a closed-loop mechanism for user feedback. By analyzing user satisfaction rating data, it dynamically adjusts the knowledge base content. For example, when multiple users' answers to a certain type of question score below a threshold, the knowledge base content update process is automatically triggered.

[0071] Preferably, the interface tool encapsulation module 200 consists of an interface definition unit, a specification generation unit, and a registration unit. The interface definition unit uses a YAML configuration file to explicitly define the input and output parameters of each interface; for example, a weather query interface requires two essential parameters: latitude and longitude coordinates and a query time range. The specification generation unit generates standardized interface descriptions based on the MCP protocol, preferably using JSON Schema format to define interface specifications, including parameter types, value ranges, and validation rules. The registration unit publishes the encapsulated interface tools as callable services through a service registry center, using Consul as the service discovery component to achieve dynamic registration and discovery of the interface tools.

[0072] Preferably, the input receiving module 300 is deployed on the front end of the WeChat mini-program, employing multiplexing technology to process both text and voice input simultaneously. For voice input, Tencent Cloud's speech recognition SDK is integrated, using dual verification through acoustic and language models to improve recognition accuracy. The input filtering mechanism employs a Trie tree-based sensitive word matching algorithm to detect and block inappropriate content in real time.

[0073] Preferably, the intent recognition and scheduling module 400 includes an intent recognition unit, a plan generation unit, and a scheduling and execution unit. The intent recognition unit deploys a selected basic large language model and runs on an NVIDIA A100 GPU to accurately capture the complex intent in user queries. The plan generation unit uses a graph-based execution plan representation method to decompose complex tasks into multiple subtasks that can be executed in parallel. The scheduling and execution unit uses thread pool technology to achieve concurrent execution of multiple tasks; for example, when a user queries the weather for multiple venues, multiple weather query requests are initiated simultaneously, improving system response speed.

[0074] Preferably, the data fusion processing module 500 consists of a prompt word optimization unit and a fusion generation unit. The prompt word optimization unit adopts a template-based prompt word generation method, using slot filling technology to fill structured data into a preset prompt word template. The fusion generation unit deploys the "Fenghe" big model, using a chain-of-thought reasoning mechanism to deeply integrate multi-source information and generate a natural language response that meets the user's needs.

[0075] Preferably, the output module processes responses asynchronously via a message queue, using RabbitMQ as the message middleware to ensure system stability under high concurrency scenarios. The output content is then reviewed by a compliance check engine and presented to the user through a WeChat mini-program interface.

[0076] This embodiment, through the precise coordination of the aforementioned modules, achieves the following technical effects: First, the system can intelligently integrate multi-source heterogeneous information such as meteorological data, event information, and popular science knowledge, providing users with a one-stop service solution; second, through standardized interface encapsulation and intelligent scheduling mechanisms, the system's compatibility and response efficiency are improved; finally, based on a continuously optimized knowledge base and advanced natural language processing technology, the accuracy and professionalism of the service content are ensured. The entire system performed excellently in the meteorological service support for the 15th National Games, effectively solving the information fragmentation problem of traditional service models and providing reliable technical support for meteorological services for large-scale events.

[0077] Based on the above module descriptions and connection methods, and combined with conventional software development and system integration techniques, those skilled in the art can implement this system. The specific implementation of each module can be done using programming languages ​​such as Python and Java. The deployment environment needs to be equipped with appropriate computing resources, such as GPU servers for model inference, distributed databases for knowledge storage, and load balancers for request distribution.

[0078] The above description is merely a preferred embodiment of this application and does not limit the patent scope of this invention. Any equivalent structural or procedural transformations made based on the description and drawings of this invention, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this invention.

Claims

1. A smart question-answering method for sports weather based on a large language model, characterized in that, Includes the following steps: Construct a multi-source knowledge base related to event meteorological services, wherein the multi-source knowledge base includes at least a basic meteorological knowledge base and an event service knowledge base; Create and encapsulate several dedicated data interface tools, including at least a weather query interface tool for obtaining real-time or forecast weather data and a sports query interface tool for obtaining sports event information; Receive natural language questions from users; The intent of the natural language question is identified based on the large language model, and the corresponding dedicated data interface tool is dynamically scheduled and invoked to obtain the first data and / or the corresponding multi-source knowledge base is retrieved to obtain the second data based on the identified intent. Input the first data and / or the second data into the large language model; The large language model fuses the first data and / or the second data to generate and output a natural language response based on multi-source information and conforming to the user's intent.

2. The method according to claim 1, characterized in that, The construction of a multi-source knowledge base related to event meteorological services includes: Collect and preprocess meteorological and event documents; The preprocessed meteorological and event documents were sliced ​​and vectorized using the RAGFlow platform to form a structured semantic knowledge network.

3. The method according to claim 2, characterized in that, Text slicing using the RAGFlow platform includes: A hybrid fragmentation strategy is adopted, which uses sentence-level, paragraph-level and document-level methods to fragment documents based on their logical structure.

4. The method according to claim 3, characterized in that, The vectorization process includes: The bge-large-zh-v1.5 or bge-m3 embedding model is used to transform the sliced ​​knowledge fragments into vector representations.

5. The method according to claim 2, characterized in that, Building a multi-source knowledge base related to event meteorological services also includes optimizing the knowledge base, specifically including the following steps: A hybrid retrieval enhancement method is adopted, combining vector retrieval and BM25 keyword retrieval for knowledge retrieval; Receive user feedback on the answers; Based on the feedback, perform add, delete, or modify operations on the knowledge base content.

6. The method according to claim 1, characterized in that, Several dedicated data interface tools are encapsulated, including: Access the official meteorological data interface from the meteorological bureau and the event data interface provided by the event's authorized information provider; The workflow design function of the Dify platform is invoked to perform processing operations on the data returned by the meteorological data interface and the event data interface. The processed data is then combined and processed. The processed data is packaged into standardized tools that can be directly used by the model; The dedicated interface tool is encapsulated using the MCP protocol in a unified and standardized manner.

7. The method according to claim 1, characterized in that, Dynamically scheduling and invoking the corresponding dedicated data interface tool to obtain the first data, and / or retrieving the corresponding multi-source knowledge base to obtain the second data, includes: The large language model is instructed to generate an execution plan, which includes a tool invocation strategy and / or a knowledge retrieval strategy. The execution plan is used to execute tool calls for operations and / or data retrieval.

8. The method according to claim 7, characterized in that, Generating the execution plan includes: For complex intents that require calling multiple dedicated data interface tools, the order in which each tool is called should be planned in the aforementioned plan; The plan defines the parameter passing dependencies between tool calls.

9. The method according to claim 1, characterized in that, The fusion processing of the first data and / or the second data includes the design of prompt words, specifically including the following steps: Chain-of-Thought prompts are used to guide reasoning in a large language model; And / or, use the Prompt Slot Filling structured template to ensure output stability.

10. A sports event weather intelligent question-answering system based on a large language model, characterized in that, include: The knowledge base construction module is used to construct a multi-source knowledge base related to event meteorological services. The multi-source knowledge base includes at least a basic meteorological knowledge base and an event service knowledge base. An interface tool encapsulation module is used to create and encapsulate several dedicated data interface tools, including at least a weather query interface tool for obtaining real-time or forecast weather data and an event query interface tool for obtaining event information. The input receiving module is used to receive natural language questions input by the user; The intent recognition and scheduling module is communicatively connected to the input receiving module, the knowledge base construction module, and the interface tool encapsulation module, respectively. The intent recognition and scheduling module is used to: perform intent recognition on natural language questions from the input receiving module based on a large language model, and, based on the recognized intent, initiate a call request to the interface tool encapsulation module to obtain first data, and / or initiate a retrieval request to the knowledge base construction module to obtain second data. The data fusion processing module is communicatively connected to the intent recognition and scheduling module, and is used to receive the first data and / or the second data, and input the data into the large language model for fusion processing to generate a natural language response based on multi-source information and conforming to the user's intent. The output module, which is communicatively connected to the data fusion processing module, is used to output the generated natural language response to the user interface.