Climate event intelligent retrieval and question-answering method based on multi-source data fusion

Through the data fusion and dimension adjustment of the intelligent search of climate events and the Q&A platform, the problems of low response efficiency and low accuracy caused by unreasonable data dimension division are solved, and the response efficiency and accuracy are improved, ensuring the stable and efficient operation of the system under different load conditions.

CN120277196AActive Publication Date: 2025-07-08STATE QIHOU CENT +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510764283.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The existing intelligent search and question-and-answer platform for climate events is unreasonable, resulting in low response efficiency and low accuracy rate.

Method used

By obtaining the climate data sets of each data source for data fusion processing, dividing dimensions according to preset dimension division rules, combining the system resource consumption judgment level and search Q&A verification judgment results, dynamically adjust the dimensions, and perform search Q&A secondary verification and cold data migration adjustment.

Benefits of technology

It improves the reply efficiency and accuracy of the Q&A system, ensures that the system operates stably and efficiently under different load conditions, optimizes the data storage structure, and reduces the impact of invalid data on the retrieval Q&A performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277196A_ABST
    Figure CN120277196A_ABST
Patent Text Reader

Abstract

The invention discloses a climate event intelligent retrieval and question-answering method based on multi-source data fusion, and belongs to the technical field of electric digital data processing, and the method comprises the following steps: S1, obtaining a fused climate data set; s2, acquiring each climate data information bar in the fused climate data set, and performing key feature analysis, thereby storing each climate data information bar to a preset data storage position; s3, obtaining each division dimension, and performing retrieval question and answer verification to obtain a retrieval question and answer verification judgment result; s4, a system resource consumption judgment level is obtained, a division dimension adjustment judgment result is obtained based on analysis of the system resource consumption judgment level and the retrieval question and answer verification judgment result, and corresponding division dimension adjustment is carried out; and S5, obtaining a secondary verification judgment result of the retrieval questions and answers, and performing corresponding cold data migration adjustment, thereby solving the problems of low reply efficiency and low accuracy caused by unreasonable data dimension division in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic digital data processing, and particularly to an intelligent retrieval and question answering method for climate events based on multi-source data fusion. Background Art

[0002] With the increasing attention of people to climate events, the requirements for intelligent retrieval and question answering platforms for climate data are also constantly increasing. The existing climate event question answering systems are realized by acquiring and analyzing multi-source data in a set scenario.

[0003] For example, a method and system for scenario-based multi-source data fusion analysis disclosed in the invention patent announcement with the publication number CN113609360B includes: acquiring samples of multi-source data in a set scenario, preprocessing the multi-source data, and the preprocessing includes feature extraction; using the probability density estimation method to fuse the features of the preprocessed multi-source data and then performing normalization processing to form a multi-source data feature set in the set scenario; using the association rule mining algorithm to extract the association rules between the features in the multi-source data feature set and the corresponding specific scenario services, and matching and inferring the business applications and solutions in the corresponding target scenario; extracting the business features and knowledge features in the specific scenario as the basic data for constructing the knowledge and business library in the specific scenario, and then performing the fusion of features and knowledge, so that people can better obtain more insights from the massive and complex multi-source data, thereby realizing the intelligent cognition and automatic analysis of multi-source data in the specific scenario.

[0004] For example, a power file question-and-answer intelligent retrieval method and system disclosed in the invention patent announcement with the publication number CN117171333B includes: a power file question-and-answer intelligent retrieval method, including: Step S1, user semantic analysis, including: realizing user semantic concept extraction; realizing user semantic expansion; Step S2, document retrieval and processing, including: establishing a file database; measuring document similarity; constructing an intention graph to represent the relationship between document data and query statements; Step S3, answer extraction, including: presenting retrieval results according to the user's search intention in combination with traditional relevance features.

[0005] However, in the process of implementing the technical solution of the present invention in the embodiments of the present application, it is found that the above technologies have at least the following technical problems: In the prior art, when a user asks a question on an existing intelligent retrieval and question answering platform for climate events, a large number of climate event data are retrieved and analyzed for feature similarity, and then the reply result is fed back to the user. However, due to the large amount of complex and disordered data information stored in the platform, there will be a phenomenon of unreasonable data dimension division, resulting in the reply effect not satisfying the user. Therefore, the prior art has problems of low reply efficiency and low accuracy rate caused by unreasonable data dimension division. Summary of the Invention

[0006] By providing an intelligent retrieval and question - answering method for climate events based on multi - source data fusion, the embodiments of the present application solve the problems of low reply efficiency and low accuracy rate caused by unreasonable data dimension division in the prior art, and achieve an improvement in reply efficiency and an improvement in reply accuracy rate.

[0007] The embodiments of the present application provide an intelligent retrieval and question - answering method for climate events based on multi - source data fusion, including the following steps: S1. Obtain climate data sets of each data source, and perform data fusion processing to obtain a fused climate data set; S2. Obtain each climate data information item in the fused climate data set, and perform key feature analysis, and thus store each climate data information item in a preset data storage location; S3. Perform dimension division processing on each climate data information item according to a preset dimension division rule to obtain each division dimension, and perform retrieval question - answering verification to obtain a retrieval question - answering verification determination result; S4. Obtain system resource consumption parameters, analyze to obtain a system resource consumption determination level, and based on the analysis of the system resource consumption determination level and the retrieval question - answering verification determination result, obtain a division dimension adjustment determination result, and thus perform corresponding division dimension adjustment; S5. After the division dimension adjustment, perform secondary retrieval question - answering verification to obtain a secondary retrieval question - answering verification determination result, and thus perform corresponding cold data migration adjustment.

[0008] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. The intelligent retrieval and question - answering method for climate events based on multi - source data fusion provided by the present invention, through the analysis based on the system resource consumption determination level and the retrieval question - answering verification determination result, performs corresponding division dimension adjustment, improves the accuracy and response speed of the question - answering system, and further realizes the improvement of reply efficiency and the improvement of reply accuracy rate, effectively solving the problems of low reply efficiency and low accuracy rate caused by unreasonable data dimension division in the prior art.

[0009] 2. The present invention obtains system resource consumption parameters and analyzes to obtain a system resource consumption determination level, and performs division dimension adjustment based on this level and the retrieval question - answering verification determination result, so as to dynamically adapt to changes in system resources and data characteristics, and ensure that the system can operate stably and efficiently under different load conditions.

[0010] 3. By performing secondary retrieval question - answering verification after the division dimension adjustment to obtain a secondary retrieval question - answering verification determination result, and if the verification is unqualified, performing cold data migration adjustment, it can further optimize the data storage structure, reduce the impact of invalid data on the retrieval question - answering performance, and improve the accuracy rate of question - answering replies. Brief Description of the Drawings

[0011] Figure 1This is the flowchart of the implementation method for the intelligent retrieval and question answering method of climate events based on multi-source data fusion provided by the embodiments of the present application; Figure 2 This is the macro flowchart of the intelligent retrieval and question answering method of climate events based on multi-source data fusion provided by the embodiments of the present application; Figure 3 This is the detailed step diagram of the intelligent retrieval and question answering method of climate events based on multi-source data fusion provided by the embodiments of the present application. Specific implementation manners

[0012] The embodiments of the present application provide an intelligent retrieval and question answering method of climate events based on multi-source data fusion, which solves the problems of low response efficiency and low accuracy rate caused by unreasonable data dimension division in the prior art, and achieves the improvement of response efficiency and response accuracy rate.

[0013] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0014] As Figure 1 shown, this is the flowchart of the implementation method for the intelligent retrieval and question answering method of climate events based on multi-source data fusion provided by the embodiments of the present application. The method includes the following steps: S1. Obtain the climate data sets of each data source and perform data fusion processing to obtain a fused climate data set; S2. Obtain each climate data information item in the fused climate data set and perform key feature analysis, and thus store each climate data information item in a preset data storage location; S3. Perform dimension division processing on each climate data information item according to a preset dimension division rule to obtain each division dimension, and perform retrieval question answering verification to obtain a retrieval question answering verification determination result; S4. Obtain the system resource consumption parameters, analyze to obtain a system resource consumption determination level, and based on the analysis of the system resource consumption determination level and the retrieval question answering verification determination result, obtain a division dimension adjustment determination result, and thus perform corresponding division dimension adjustment; S5. Perform secondary retrieval question answering verification after the division dimension adjustment to obtain a secondary retrieval question answering verification determination result, and thus perform corresponding cold data migration adjustment.

[0015] In this embodiment, as Figure 2 shown, Figure 2This is the macro flowchart of the intelligent retrieval and question answering method for climate events based on multi-source data fusion provided by the embodiments of this application. Climate datasets from various data sources are processed for data fusion to obtain a fused climate dataset. Each climate data information item in the fused climate dataset is retrieved and key feature analysis is performed. Thus, each climate data information item is stored at a preset data storage location. Each climate data information item is processed for dimension division according to a preset dimension division rule to obtain each division dimension, and retrieval question answering verification is performed to obtain a retrieval question answering verification determination result. Based on the retrieval question answering verification determination result and the analysis of the retrieval question answering verification determination result, a division dimension adjustment determination result is obtained. Thus, corresponding division dimension adjustment is performed, and corresponding cold data migration adjustment is performed after the division dimension adjustment, and finally the process ends.

[0016] As Figure 3 shown, Figure 3 This is the detailed step diagram of the intelligent retrieval and question answering method for climate events based on multi-source data fusion provided by the embodiments of this application. The division dimension adjustment determination result is divided into two cases: performing division dimension dimension increase processing and performing division dimension dimension decrease processing. When performing division dimension dimension increase processing, the division dimension corresponding to the maximum data volume is processed for dimension increase to obtain an updated value of system resource consumption increase due to dimension increase, and then the increase ratio due to dimension increase is calculated. It is judged whether the result of dimension increase processing is qualified. If it is qualified, the dimension increase processing is completed; otherwise, the dimension increase processing is continued. When performing division dimension dimension decrease processing, first the contribution values of each division dimension are obtained, the division dimension to be adjusted is determined, and then dimension decrease processing is performed, and the question answering response effect value after dimension decrease is obtained. The dimension decrease effect adjustment ratio is calculated. If it is greater than the dimension decrease effect adjustment ratio threshold, the dimension decrease processing is completed; otherwise, the dimension decrease processing is repeated. Finally, corresponding cold data migration adjustment is performed.

[0017] It should be noted that "multi-source data fusion" refers to normalizing and synergistically integrating heterogeneous data (including remote sensing image data, vector observation data, time series numerical simulation data, radar reflectivity image data, and text record data, etc.) obtained from different observation platforms (satellite remote sensing systems, meteorological radars, ground observation stations, weather forecasts) according to a unified structural format and semantic standard to obtain data of a unified dimension.

[0018] Further, store each climate data information bar at a preset data storage location. The specific method is as follows: Obtain each climate data information bar in the fused climate data set, and perform key feature extraction to obtain the key feature group of each climate data information bar; Obtain the reference key feature set corresponding to each storage location in the preset system in the database; Respectively match the key feature groups of each climate data information bar with the reference key feature sets corresponding to each storage location to obtain the reference key feature set with the maximum feature matching degree, and obtain the storage location corresponding to this reference key feature set as the preset data storage location of this climate data information bar, thereby storing each climate data information bar at the preset data storage location.

[0019] In this embodiment, by using the reference key feature set to accurately locate each data entry, it is avoided that all data are uniformly stored in a single database table, thereby greatly reducing the overhead of scanning irrelevant storage partitions during data retrieval, improving the I / O efficiency. The matching feature mapping aggregates data of the same type of climate event or similar attributes to the same storage location, and the retrieval engine can directly locate the highly relevant partition. Therefore, when a user queries, the target data can be hit faster, the response delay can be shortened, and the system throughput can be improved.

[0020] The mapping relationship between the preset key features and the storage locations can be dynamically adjusted through a configuration table. When adding or adjusting climate elements (such as adding an "extreme wind speed" indicator), only the "reference key feature set" and its storage mapping need to be updated to take effect quickly. By storing the partitions with high feature matching degrees together according to the matching feature analysis, abnormal or irrelevant entries can be effectively eliminated, ensuring that the data information in each storage partition is highly consistent, and it is helpful for subsequent verification, backup, and archiving management, reducing data redundancy and consistency risks.

[0021] It should be noted that the key feature extraction is specifically as follows: Clean and standardize the obtained climate data to remove noise and outliers; According to domain knowledge and business requirements, select features related to climate events, such as time, location, event type, intensity, etc.; Convert the selected features into a unified format or dimension, for example, convert time to the standard time format and convert location to geographical coordinates; Apply a feature extraction algorithm (such as TF-IDF) to extract key features.

[0022] Among them, the key feature group is, for example: the key feature group of the typhoon event. Time feature: 2023-10-01 14:30, location feature: longitude 113.5°E, latitude 22.3°N, event type feature: typhoon, intensity feature: central pressure 960 hPa, maximum wind speed 45 m / s.

[0023] Feature matching is performed between the key feature groups of each climate data information item and the reference key feature sets corresponding to each storage location. The specific steps are as follows: Extract the key feature group of each climate data information item from the fused climate dataset. For example, for a typhoon event data item, its key feature group includes: Time: 2023-10-01 14:30, Location: 113.5°E, 22.3°N, Event type: Typhoon, Intensity: Central pressure 960 hPa, Maximum wind speed 45 m / s. Reference key feature sets: Obtain the preset reference key feature sets from the database. This feature set contains multiple reference key feature groups, and each reference key feature group is associated with a specific storage location in the database. For example, the reference key feature set includes: Storage location A: Time range from 2023-10-01 00:00 to 2023-10-01 23:59, Location from 113.0°E to 114.0°, 22.0°N to 23.0°N, Event type Typhoon, Intensity range Central pressure 950 hPa to 970 hPa, Maximum wind speed 40 m / s to 50 m / s; Storage location B: Time range from 2023-09-01 00:00 to 2023-09-30 23:59, Location from 115.0°E to 116.0°, 23.0°N to 24.0°N, Event type Heavy rain, Intensity range Precipitation 100 mm / h to 150 mm / h. Standardize the extracted key feature groups and reference key feature sets. For example: Convert the time to a unified time format, convert the location to a unified geographic coordinate format (such as decimal degrees of longitude and latitude), and convert the intensity indicators to a unified unit and range. Then perform feature matching. For the key feature group of each climate data information item, match it with the reference key feature sets of each storage location in sequence. Calculate the matching degree of each feature. For example: Time matching degree: If the time of the climate data is within the time range of storage location A, the time matching degree is 100%; otherwise it is 0%. Similarly, calculate the matching degrees of other features, and calculate the overall matching degree by comprehensively considering the matching degrees of all features (weighted sum of the matching degrees of each feature). Finally, compare the overall matching degrees of the climate data information item with each storage location to determine the best-matching storage location, and store the climate data information item in the preset data storage location corresponding to the storage location with the highest overall matching degree. If the overall matching degrees of multiple storage locations are the same, further judgment can be made according to business rules (such as the highest priority of the storage location) to obtain the preset data storage location.

[0024] Further, to obtain the retrieval Q&A verification determination result, the specific method is as follows: Obtain the preset division dimension rules in the database; divide each climate data information item according to the preset division dimension rules to obtain each division dimension, and jointly mark each division dimension as the initial division dimension; perform retrieval Q&A verification based on the initial division dimension, obtain the Q&A response effect parameters under the initial division dimension, and analyze to obtain the Q&A response effect value; obtain the preset Q&A response effect threshold in the database, and compare it with the Q&A response effect value to obtain the retrieval Q&A verification determination result. If the Q&A response effect value is less than the Q&A response effect threshold, the retrieval Q&A verification determination result is unqualified. If the Q&A response effect value is above the Q&A response effect threshold, the retrieval Q&A verification determination result is qualified.

[0025] In this embodiment, by performing retrieval Q&A verification under the initial division dimension, the system response effect under different dimension combinations can be pre-detected, ensuring that the Q&A results meet the set accuracy and latency indicators, avoiding users encountering incorrect or incomplete answers in the actual Q&A environment, and improving the overall system credibility and user experience.

[0026] It should be noted that the preset dimension rules can be flexibly configured for different climate event types (such as typhoons, droughts, heavy rains), and the verification model can be quickly switched under different scenarios. The preset division dimension rules are the division rules preset and stored in the database. If the Q&A response effect value is less than the Q&A response effect threshold, the retrieval Q&A verification determination result is unqualified. When the verification is unqualified, the subsequent "dimension adjustment and secondary verification" process can be automatically triggered to achieve multi-round optimization and meet the diverse usage requirements of scientific research, early warning, science popularization, etc.

[0027] Examples of the preset dimension rules are as follows: Division by time dimension: Divide the climate data into 2020, 2021, 2022, etc. Division by location dimension: Divide into continents (Asia, Europe, Africa, etc.), countries, provinces, etc. Division by climate event type: Such as typhoons, hurricanes, heavy rains, droughts, high temperatures, cold snaps, etc.

[0028] The automated verification determination process replaces manual testing, greatly reducing manual intervention and testing costs. By discovering and locating dimension division problems in advance, large-scale data reconstruction or emergency repair can be avoided, and the operation and maintenance efficiency can be improved.

[0029] Further, to obtain the Q&A response effect value, the specific steps are as follows: Obtain the Q&A response effect parameters under the initial division dimension within a preset time period. The Q&A response effect parameters include the average Q&A response duration, the Q&A push accuracy rate, the system retrieval accuracy, and the maximum fluctuation ratio of the Q&A response duration. Obtain the preset Q&A response baseline set in the database, perform a ratio operation with the Q&A response effect parameters to obtain the response ratio analysis result, and introduce the corresponding weighting factor based on the response ratio analysis result for quantitative coupling processing to obtain the Q&A response effect value. The Q&A response baseline set includes the average Q&A response duration baseline value, the Q&A push accuracy rate baseline value, the ideal system retrieval accuracy value, and the allowable fluctuation ratio of the Q&A response duration.

[0030] In this embodiment, the average Q&A response duration refers to the average time taken for the user to receive a complete answer from the system when testing within a preset time period, which can be obtained by analyzing the consumption logs. The Q&A push accuracy rate refers to the ratio of the number of times the system pushes the correct answer to the user (the content pushed by the system is liked and adopted by the user) to the total number of system pushes, which can be obtained by querying the records in the system logs. The system retrieval accuracy represents the average value of the matching degrees of each system push within a preset time period (the feature matching degree between the key features of the content pushed by the system and the key features of the user's question), and the feature matching degree can be analyzed by the NLTK tool. The maximum fluctuation ratio of the Q&A response duration can be obtained by calculating the difference between the maximum and minimum values of the system Q&A response duration to get the fluctuation difference, and then dividing the fluctuation difference by the average Q&A response duration.

[0031] Obtaining the Q&A response effect value by analyzing the Q&A response effect parameters (including the average Q&A response duration, the Q&A push accuracy rate, the system retrieval accuracy, and the maximum fluctuation ratio of the Q&A response duration) under the initial division dimension takes into account the mutual influence relationships among these parameters. For example, when the system retrieval accuracy (average matching degree) increases, it is necessary to scan candidate data more deeply or widely, which will cause the average Q&A response duration to increase without changing the information storage location, etc. If the average duration is too long, it will affect the user experience and cause the push satisfaction rate to decrease. If the response duration fluctuates greatly in each query (the maximum fluctuation ratio is high), even if the overall average response is fast and the retrieval is accurate, users will still reduce the number of likes and adoptions due to the unstable experience, thus dragging down the Q&A push accuracy rate.

[0032] The specific method for obtaining the Q&A response effect value is as follows: ; In the formula, represents the Q&A response effect value, represents the average Q&A response duration, represents the average Q&A response duration baseline value, Indicates the correct rate of Q&A push, Indicates the baseline value of the correct rate of Q&A push, Indicates the system retrieval accuracy, Indicates the ideal value of the system retrieval accuracy, Indicates the maximum fluctuation ratio of the Q&A response time, Indicates the allowable fluctuation ratio of the Q&A response time, Indicates the weighted factor of the average Q&A response time, Indicates the weighted factor of the correct rate of Q&A push, Indicates the weighted factor of the system retrieval accuracy, Indicates the weighted factor of the maximum fluctuation ratio of the Q&A response time.

[0033] The weighted factor of the average Q&A response time, the weighted factor of the correct rate of Q&A push, the weighted factor of the system retrieval accuracy, and the weighted factor of the maximum fluctuation ratio of the Q&A response time can be obtained from the database. For example, the weighted factor of the average Q&A response time can be obtained by obtaining the historical average Q&A response time stored in the database and the corresponding weighted factor of the average Q&A response time. Thus, a mapping set of the average Q&A response time is constructed, in which there is a one-to-one or many-to-one correspondence relationship. By inputting the data of the average Q&A response time to be used into the mapping set of the average Q&A response time, the weighted factor of the average Q&A response time can be obtained. The acquisition methods of other weighted factors are the same as that of the weighted factor of the average Q&A response time, and they can all be matched in the corresponding mapping sets. Among them, the weighted factor of the correct rate of Q&A push corresponds to the mapping set of the correct rate of Q&A push, the weighted factor of the system retrieval accuracy corresponds to the mapping set of the system retrieval accuracy, and the weighted factor of the maximum fluctuation ratio of the Q&A response time corresponds to the mapping set of the maximum fluctuation ratio of the Q&A response time.

[0034] Furthermore, system resource consumption parameters are obtained and analyzed to obtain a system resource consumption determination level. The specific method is as follows: obtain system resource consumption parameters, which include I / O resource utilization, memory page error rate, file descriptor usage, and task queue length; obtain a preset system resource consumption allowable set in a database, and perform a proportional operation on the system resource consumption parameters to obtain a resource consumption proportional operation analysis result; based on the resource consumption proportional operation analysis result, introduce a corresponding weighting factor for quantitative coupling processing to obtain a system resource consumption value; the system resource consumption allowable set includes an I / O resource utilization allowable value, a memory page error rate allowable value, a file descriptor usage upper limit value, and a task queue length upper limit value; obtain a system resource consumption threshold value preset in the database, and compare it with the system resource consumption value to obtain a system resource consumption determination level. If the system resource consumption value is below the system resource consumption threshold value, the system resource consumption determination result is a first-level resource consumption; if the system resource consumption value is greater than the system resource consumption threshold value, the system resource consumption determination result is a second-level resource consumption.

[0035] In this embodiment, the I / O resource utilization can be monitored by a performance monitor. The memory page error rate refers to the number of memory page errors that occur per unit time, reflecting the tension of memory resources, and can be obtained by using performance monitoring tools such as vmstat. The file descriptor usage refers to the number of file descriptors opened in the system, reflecting the file resource usage of the system, and can be monitored by using the lsof tool. The task queue length refers to the number of processes waiting for CPU execution in the system, reflecting the load of the system, and can be queried by using the top command line tool.

[0036] By analyzing the system resource consumption parameters (including I / O resource utilization, memory page fault rate, file descriptor usage, and task queue length), the mutual influence between these parameters is taken into account. For example, when the I / O resource utilization is high, the memory page fault rate will increase, because frequent data swapping in and out will lead to an increase in memory page faults, and memory page faults will trigger more I / O operations, increasing the I / O burden. Large file descriptor usage is usually accompanied by a large number of file operations, resulting in increased I / O resource utilization. At the same time, high I / O utilization is likely to increase the task queue length because the process is often blocked waiting for I / O operations to complete. Large file descriptor usage will also increase the task queue length due to frequent file operations. When the memory page fault rate is high, the process frequently waits for pages to be swapped in and out, and the task queue length increases accordingly.

[0037] Get the system resource consumption value. The specific method is: ; In the formula, Represents the system resource consumption value, Represents the I / O resource utilization rate, Represents the allowable value of the I / O resource utilization rate, Represents the memory page error rate, Represents the allowable value of the memory page error rate, Represents the file descriptor usage, Represents the upper limit value of the file descriptor usage, Represents the task queue length, Represents the upper limit value of the task queue length, Represents the weighting factor of the I / O resource utilization rate, Represents the weighting factor of the memory page error rate, Represents the weighting factor of the file descriptor usage, Represents the weighting factor of the task queue length.

[0038] The weighting factors of the I / O resource utilization rate, the memory page error rate, the file descriptor usage, and the task queue length can be obtained from the database. For example, the weighting factor of the I / O resource utilization rate can be obtained by obtaining the historical I / O resource utilization rate stored in the database and the corresponding weighting factor of the I / O resource utilization rate, thereby constructing an I / O resource utilization rate mapping set, in which there is a one-to-one or many-to-one correspondence relationship. By inputting the I / O resource utilization rate data to be used into the I / O resource utilization rate mapping set, the weighting factor of the I / O resource utilization rate can be obtained. The acquisition methods of other weighting factors are the same as that of the weighting factor of the I / O resource utilization rate, and they can all be obtained by matching in the corresponding mapping sets. Among them, the weighting factor of the memory page error rate corresponds to the memory page error rate mapping set, the weighting factor of the file descriptor usage corresponds to the file descriptor usage mapping set, and the weighting factor of the task queue length corresponds to the task queue length mapping set.

[0039] Furthermore, based on the analysis of the system resource consumption determination level and the retrieval question-and-answer verification determination result, the partition dimension adjustment determination result is obtained. The specific method is as follows: Based on the analysis of the system resource consumption determination level and the retrieval question-and-answer verification determination result, if the retrieval question-and-answer verification determination result is unqualified, the partition dimension adjustment determination result is to execute the partition dimension adjustment. At the same time, if the system resource consumption determination level is the first-level resource consumption, the partition dimension adjustment determination result is to execute the partition dimension augmentation process. If the system resource consumption determination level is the second-level resource consumption, the partition dimension adjustment determination result is to execute the partition dimension reduction process; if the retrieval question-and-answer verification determination result is qualified, the partition dimension adjustment determination result is not to execute the partition dimension adjustment.

[0040] In this embodiment, when the retrieval Q&A verification determination result is unqualified and the system resource consumption determination level is first-level resource consumption (i.e., the system resource consumption value is lower than the threshold), it indicates that the current resource consumption of the system is within an acceptable range, but the retrieval Q&A effect is not good. At this time, performing dimensionality augmentation processing on the division dimension can make the data division more detailed, thereby improving the data organization and retrieval efficiency. Since the system resources are sufficient, the dimensionality augmentation operation will not impose too much burden on the system performance. At the same time, it can improve the accuracy and response effect of the retrieval Q&A, which is beneficial to enhancing the data expression ability and retrieval accuracy, enabling the system to better meet the user's query needs, and making full use of the existing resource capabilities of the system.

[0041] If, in the case where the retrieval Q&A verification determination result is unqualified, the system resource consumption determination level is second-level resource consumption (i.e., the system resource consumption value exceeds the threshold), it indicates that the system resources are tense. At this time, performing dimensionality reduction processing on the division dimension can reduce the complexity of data processing and relieve the system resource pressure. Dimensionality reduction can remove redundant or low-contribution dimensions, enabling the system to concentrate resources on processing key data and improving the overall operation efficiency of the system. In the case of limited system resources, the dimensionality reduction operation helps to optimize resource allocation, ensure the stable operation of the system, improve the response effect of the retrieval Q&A, enhance the resource utilization efficiency of the system, and avoid system performance degradation or service interruption caused by excessive resource consumption, thus ensuring the continuous and effective operation of the system.

[0042] When the retrieval Q&A verification determination result is qualified, it means that the current division dimension can meet the requirements of the retrieval Q&A. The system does not need to perform dimension adjustment, which can maintain the stability and consistency of the system and avoid potential impacts on the system caused by unnecessary adjustments. At the same time, the stable division dimension helps to maintain the maintainability and understandability of the system, facilitating subsequent management and optimization work. The benefit of not performing division dimension adjustment is to save the cost and time of system adjustment, ensure the efficient operation of the system under a stable division dimension, and avoid the increase in system burden and complexity caused by frequent adjustments.

[0043] Further, perform dimensionality augmentation processing on the partitioning dimension. The specific method is as follows: Obtain the data volume of each partitioning dimension under the initial partitioning dimension, perform dimensionality augmentation processing on the partitioning dimension corresponding to the maximum data volume, obtain the system resource consumption value after dimensionality augmentation processing on the partitioning dimension, and mark it as the system resource consumption dimensionality augmentation update value; perform a differential ratio analysis on the system resource consumption dimensionality augmentation update value and the system resource consumption threshold to obtain the dimensionality augmentation update ratio; obtain the preset dimensionality augmentation update ratio threshold in the database; based on the dimensionality augmentation update ratio, the dimensionality augmentation update ratio threshold, the system resource consumption dimensionality augmentation update value, and the system resource consumption threshold analysis, obtain the dimensionality augmentation processing result; if the system resource consumption dimensionality augmentation update value is less than the system resource consumption threshold and the dimensionality augmentation update ratio is less than the dimensionality augmentation update ratio threshold, the dimensionality augmentation processing result is that the dimensionality augmentation processing is unqualified, and continue to perform dimensionality augmentation processing on the partitioning dimension until the system resource consumption dimensionality augmentation update value is above the system resource consumption threshold or the dimensionality augmentation update ratio is above the dimensionality augmentation update ratio threshold, and complete the dimensionality augmentation processing on the partitioning dimension; if the system resource consumption dimensionality augmentation update value is above the system resource consumption threshold or the dimensionality augmentation update ratio is above the dimensionality augmentation update ratio threshold, the dimensionality augmentation processing result is that the dimensionality augmentation processing is qualified, and directly complete the dimensionality augmentation processing on the partitioning dimension.

[0044] In this embodiment, by performing dimensionality augmentation on the partitioning dimension corresponding to the maximum data volume and analyzing the system resource consumption dimensionality augmentation update value, and by comparing the system resource consumption dimensionality augmentation update value with the system resource consumption threshold and comparing the dimensionality augmentation update ratio with the dimensionality augmentation update ratio threshold, it is possible to accurately judge whether the dimensionality augmentation processing has achieved the expected effect, avoid misjudgment of the dimensionality augmentation processing caused by a single judgment rule, ensure the accuracy and effectiveness of the dimensionality augmentation processing, and avoid unnecessary repeated operations.

[0045] Comparing the system resource consumption dimensionality augmentation update value with the system resource consumption threshold ensures that the system resource consumption during the dimensionality augmentation processing will not exceed the allowable range, helps to maintain the stable operation of the system, and avoids performance degradation or service interruption caused by excessive resource consumption. If the expected effect is still not achieved after the dimensionality augmentation processing, the system will continue to perform the dimensionality augmentation processing until the conditions are met, enabling the system to adapt to different data conditions and resource limitations, and improving the adaptability and robustness of the system. Through the dimensionality augmentation processing, the organization and structure of the data are optimized, enabling the data to better reflect the actual situation, improving the availability and retrieval efficiency of the data, being able to more accurately meet the query needs of users, providing more accurate answers, and by setting the dimensionality augmentation update ratio threshold, it is possible to avoid unlimited dimensionality augmentation operations, prevent waste of system resources, ensure that the dimensionality augmentation processing stops in a timely manner after achieving the expected effect, and maintain the high efficiency of the system.

[0046] It should be noted that the dimension increase processing of the division dimension continues. The specific method is as follows: repeatedly obtain the data volume of each division dimension after the current execution of the division dimension increase processing, perform the division dimension increase processing on the division dimension corresponding to the maximum data volume, obtain the system resource consumption value at this time, that is, obtain the system resource consumption increase update value for this round. From this, analyze the increase update ratio, increase update ratio threshold, system resource consumption increase update value, and system resource consumption threshold for this round of execution of the division dimension increase processing, and judge the increase processing result for this round. If the increase processing result for this round is still unqualified for the increase processing, continue the next round of execution of the division dimension increase processing. The steps of the next round of execution of the division dimension increase processing are exactly the same as those of this round of execution of the division dimension increase processing, until the system resource consumption increase update value is above the system resource consumption threshold or the increase update ratio is above the increase update ratio threshold, and the division dimension increase processing is completed.

[0047] It should be noted that the difference ratio analysis is performed on the system resource consumption increase update value and the system resource consumption threshold to obtain the increase update ratio. The specific steps are as follows: perform a difference processing on the system resource consumption increase update value and the system resource consumption threshold to obtain the system resource consumption difference, and divide the system resource consumption difference by the system resource consumption threshold to obtain the increase update ratio.

[0048] Further, perform dimensionality reduction processing on the division dimensions. The specific method is as follows: Obtain the contribution degree parameters of each division dimension under the initial division dimensions. The contribution degree parameters include the retrieval hit rate, the number of retrievable data instances, the memory occupancy, and the user adoption hit rate; Obtain the preset contribution degree baseline set in the database. The contribution degree baseline set includes the retrieval hit rate baseline value, the retrievable data instance number baseline value, the memory occupancy allowance value, and the user adoption hit rate baseline value; Perform proportional operations on the retrieval hit rate, the memory occupancy, and the user adoption hit rate with the retrieval hit rate baseline value, the memory occupancy allowance value, and the user adoption hit rate baseline value to obtain the contribution ratio operation result. Perform difference degree analysis on the number of retrievable data instances and the retrievable data instance number baseline value to obtain the contribution difference analysis result. Based on the contribution ratio operation result and the contribution difference analysis result and introduce the corresponding weighting factors for quantitative coupling processing to obtain the contribution value of each division dimension; Obtain the preset contribution degree threshold in the database and compare it with the contribution value of each division dimension. If the contribution value of a certain division dimension is less than the contribution degree threshold, mark this division dimension as the division dimension to be adjusted, and thus obtain each division dimension to be adjusted; Incorporate each division dimension to be adjusted into the corresponding target division dimension respectively to complete the dimensionality reduction processing of the division dimensions; Obtain the Q&A response effect value after the dimensionality reduction processing of the division dimensions and determine whether to perform repeated dimensionality reduction processing; Incorporate each division dimension to be adjusted into the corresponding target division dimension respectively. The specific steps are as follows: Obtain each division dimension in the database that is of the same dimension type as the division dimension to be adjusted and mark them as each reference target division dimension; Perform feature matching analysis on the division dimension to be adjusted and each reference target division dimension, and select the reference target division dimension with the highest feature matching degree with the division dimension to be adjusted as the target division dimension of this division dimension to be adjusted, and thus analyze and obtain the target division dimension corresponding to each division dimension to be adjusted; Incorporate each division dimension to be adjusted into the corresponding target division dimension respectively.

[0049] In this embodiment, by performing dimensionality reduction processing on the division dimensions, the complexity of data processing and storage requirements can be reduced, thereby reducing the consumption of system resources, which helps to improve the overall performance and response speed of the system. Dimensionality reduction can remove redundant or low-contribution dimensions, making the data more focused on key features, which helps to improve the accuracy and efficiency of retrieval and Q&A, enabling the system to respond to user queries faster and provide more relevant results. By obtaining the Q&A response effect value after the dimensionality reduction processing of the division dimensions and comparing it with the preset adjustment ratio threshold, it can be judged whether the dimensionality reduction processing has achieved the expected effect. If the adjustment ratio of the Q&A response effect after dimensionality reduction is lower than the threshold, it means that the dimensionality reduction processing is not completed and the dimensionality reduction operation needs to be continued to further optimize the data organization and retrieval efficiency. At the same time, it can also prevent over-dimensionality reduction and ensure the integrity and availability of the data, because over-dimensionality reduction will affect the accuracy of retrieval and Q&A.

[0050] By obtaining other partitioning dimensions of the same dimension type as the partitioning dimension to be adjusted, the consistency and relevance of data during the dimensionality reduction process can be ensured, which helps to maintain the semantic integrity and logical relationships of the data, enabling the dimensionality-reduced data to still accurately reflect the actual situation. Marking these partitioning dimensions of the same dimension type as reference target partitioning dimensions can provide a more appropriate merging target for the partitioning dimension to be adjusted. Through feature matching analysis, selecting the reference target partitioning dimension with the highest feature matching degree with the partitioning dimension to be adjusted for merging can ensure the rationality and effectiveness of the dimensionality reduction operation, avoid forcibly merging irrelevant data resulting in information chaos, and help improve the quality of the dimensionality-reduced data and the accuracy of retrieval and question answering.

[0051] The retrieval hit rate refers to the frequency at which data in a certain dimension is successfully hit during the question and answer retrieval process, that is, the ratio of the number of times the system uses the data in this dimension to provide answers to the total number of retrievals when answering user questions. Its acquisition method can be statistically obtained using data analysis tools (such as the Pandas library in Python). The number of retrievable data instances refers to the amount of data available for retrieval under a certain partitioning dimension, that is, the number of data records included in this dimension. Its acquisition method is obtained through database queries. The memory occupancy refers to the storage space size occupied by the data in a certain partitioning dimension in memory, which can be queried through the task manager. The user adoption hit rate refers to the acceptance degree of the answers provided by the question and answer system by users, that is, the ratio of the number of times users adopt the answers generated by the data in this dimension to the total number of times the data in this dimension is used for answering, which can be statistically obtained through user feedback logs.

[0052] It should be noted that by analyzing contribution parameters including the retrieval hit rate, the number of retrievable data instances, the memory occupancy, and the user adoption hit rate, and obtaining the contribution values of each partitioning dimension, the mutual influence relationships between these parameters are considered. For example: a high retrieval hit rate indicates that the system can effectively utilize the number of retrievable data instances. A large number of retrievable data instances can increase the hit rate, but if it exceeds the memory capacity and the memory occupancy soars, it will instead drag down the system efficiency. The user adoption hit rate is affected by the retrieval hit rate and the quality of data instances. A high hit rate and high-quality data can increase the user adoption rate. At the same time, too many instances may dilute attention and reduce the adoption hit rate. The memory occupancy is affected by the number of retrievable data instances. More instances result in a larger memory occupation.

[0053] The contribution values of each partitioning dimension are obtained, and the specific method is as follows: ; In the formula, represents the contribution value of the i-th partitioning dimension, i represents the number of the partitioning dimension, , represents the total number of partitioning dimensions, Denotes the retrieval hit rate of the $i$-th partitioning dimension, Denotes the baseline value of the retrieval hit rate, Denotes the baseline value of the number of retrievable data instances, Denotes the number of retrievable data instances of the $i$-th partitioning dimension, Denotes the memory occupancy of the $i$-th partitioning dimension, Denotes the allowable value of the memory occupancy, Denotes the user adoption hit rate of the $i$-th partitioning dimension, Denotes the baseline value of the user adoption hit rate, Denotes a constant, Denotes the weighting factor of the retrieval hit rate, Denotes the weighting factor of the number of retrievable data instances, Denotes the weighting factor of the memory occupancy, Denotes the weighting factor of the user adoption hit rate.

[0054] It should be noted that, Denotes that the constant is a number with an extremely small value, which is used to prevent calculation errors caused when the number of retrievable data instances is the same as the baseline value of the number of retrievable data instances, and to ensure the accuracy of the calculation results.

[0055] The weighting factor of the retrieval hit rate, the weighting factor of the number of retrievable data instances, the weighting factor of the memory occupancy, and the weighting factor of the user adoption hit rate can be obtained from the database. For example: the weighting factor of the retrieval hit rate can be obtained by acquiring the historical retrieval hit rate stored in the database and the corresponding weighting factor of the retrieval hit rate, and thus constructing a retrieval hit rate mapping set, where there is a one-to-one or many-to-one correspondence in this mapping set. By inputting the retrieval hit rate data to be used into the retrieval hit rate mapping set, the weighting factor of the retrieval hit rate can be obtained. The acquisition methods of the other weighting factors are the same as that of the weighting factor of the retrieval hit rate, and they can all be matched and obtained in the corresponding mapping sets. Among them, the weighting factor of the number of retrievable data instances corresponds to the mapping set of the number of retrievable data instances, the weighting factor of the memory occupancy corresponds to the mapping set of the memory occupancy, and the weighting factor of the user adoption hit rate corresponds to the mapping set of the user adoption hit rate.

[0056] Obtain each partitioning dimension in the database that has the same dimension type as the partitioning dimension to be adjusted, and mark them as each reference target partitioning dimension. For example, for the same dimension type: if the partitioning dimension to be adjusted is "Typhoon in May 2023" under the "Time Dimension", then the other partitioning dimensions under the "Time Dimension" (such as "Typhoon in the first half of 2023", etc.) retrieved will be marked as reference target partitioning dimensions.

[0057] Further, obtain the Q&A response effect value after dimensionality reduction processing of the partitioning dimension, and determine whether to perform repeated dimensionality reduction processing. The specific method is as follows: Obtain the Q&A response effect value after dimensionality reduction processing of the partitioning dimension, and mark it as the dimensionality reduction processed Q&A response effect value; Based on the dimensionality reduction processed Q&A response effect value, perform adjustment ratio analysis with the Q&A response effect value to obtain the dimensionality reduction effect adjustment ratio; Obtain the preset dimensionality reduction effect adjustment ratio threshold in the database, and analyze it with the dimensionality reduction effect adjustment ratio. If the dimensionality reduction effect adjustment ratio is below the dimensionality reduction effect adjustment ratio threshold, do not perform repeated dimensionality reduction processing. If the dimensionality reduction effect adjustment ratio is greater than the dimensionality reduction effect adjustment ratio, continue to perform repeated dimensionality reduction processing.

[0058] In this embodiment, if the dimensionality reduction effect adjustment ratio is greater than the preset threshold, it indicates that although the current dimensionality reduction processing has achieved certain effects, there is still room for improvement. Continuing dimensionality reduction can further optimize data organization, reduce redundant dimensions, and improve data compactness and relevance. Through repeated dimensionality reduction processing, the data structure can be adjusted more finely, making the response effect of the Q&A system more accurate and efficient, which helps to improve user satisfaction and the overall performance of the system. When the adjustment ratio of the Q&A response effect value after dimensionality reduction processing to the original effect value exceeds the preset threshold, it indicates that the current dimensionality reduction effect has not reached the expected goal. At this time, it is necessary to continue to perform dimensionality reduction processing to further improve the Q&A response effect.

[0059] Further, perform corresponding cold data migration adjustment. The specific method is as follows: After the partitioning dimension is adjusted, perform secondary verification of the retrieval Q&A to obtain the secondary verification determination result of the retrieval Q&A. If the secondary verification determination result of the retrieval Q&A is unqualified, perform cold data migration adjustment. If the secondary verification determination result of the retrieval Q&A is qualified, do not perform cold data migration adjustment; If cold data migration adjustment is performed, obtain the access parameters of each cold data information item within the preset time period. The access parameters include the number of accesses, the number of times adopted, and the adoption ratio; Obtain the preset access benchmark set in the database, and perform ratio calculation with the access parameters to obtain the access difference analysis result. Based on the access difference analysis result, introduce the corresponding weighting factor for quantitative coupling processing to obtain the access comprehensive adoption value of each cold data information item; Obtain the preset access comprehensive adoption threshold in the database and compare it with the access comprehensive adoption value of each cold data information item. If there is a cold data information item whose access comprehensive adoption value is above the access comprehensive adoption threshold, perform cold data migration adjustment on this cold data information item and migrate it to the hot data storage area, otherwise do not perform cold data migration adjustment.

[0060] In this embodiment, when increasing the dimension, if the retrieval Q&A verification fails due to the increase in data volume, migrating cold data to hot storage can improve the access speed and retrieval efficiency, and optimize the Q&A effect. When reducing the dimension, if data redundancy or low-quality data affects the verification result, migrating eligible cold data to the hot storage area, focusing on high-quality data, can improve the Q&A accuracy and relevance. Migrating cold data that meets specific conditions to the hot data storage area can improve data management efficiency, reduce the occupation of storage resources by infrequently used data, and optimize the storage structure.

[0061] After performing the dimensionality increase processing of the partition dimension, if the retrieval Q&A verification fails, this indicates that the best optimization has been achieved from the perspective of data dimension, because increasing the data dimension will cause a greater occupation of numerical resources. Therefore, excessive data dimensionality increase processing cannot be performed. At this time, the data organizational structure needs to be adjusted, and cold data needs to be migrated to the hot data area to make the cold data more accessible and retrievable, thereby improving the performance of the Q&A system and avoiding excessive occupation of system resources caused by excessive data dimensionality increase processing. After performing the dimensionality reduction processing of the partition dimension, if the retrieval Q&A verification fails, the corresponding cold data migration adjustment is considered because the data dimensionality reduction processing will not affect the increase in system resource occupation. At this time, if the reduction effect adjustment ratio is below the reduction effect adjustment ratio threshold, performing the data dimensionality reduction processing can no longer improve the Q&A response effect. Therefore, the corresponding cold data migration adjustment needs to be performed to change the data organizational structure, thereby increasing the Q&A response effect value.

[0062] In summary, in this embodiment, through the analysis of the system resource consumption determination level and the retrieval Q&A verification determination result, the corresponding partition dimension adjustment is performed, improving the accuracy and response speed of the Q&A system, and thus achieving the improvement of the reply efficiency and the reply correct rate, effectively solving the problems of low reply efficiency and low correct rate caused by unreasonable data dimension division in the prior art.

[0063] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0064] The present invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block of the flowchart illustrations and / or block diagrams, and combinations of flows and / or blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing apparatus create means for implementing the functions specified in the flowchart flow or flows and / or block or blocks. Figure 1 one flow or flows and / or block or blocks Figure 1 means for implementing the functions specified in one block or blocks or more.

[0065] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in one flow or flows and / or block or blocks. Figure 1 one flow or flows and / or block or blocks Figure 1 means for implementing the functions specified in one block or blocks or more.

[0066] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one flow or flows and / or block or blocks. Figure 1 one flow or flows and / or block or blocks Figure 1 means for implementing the functions specified in one block or blocks or more.

[0067] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0068] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A method for intelligent retrieval and question answering of climate events based on multi-source data fusion, characterized in that It includes the following steps: S1. Obtain the climate datasets of each data source, and perform data fusion processing to obtain a fused climate dataset; S2. Obtain each climate data information item in the fused climate dataset, and perform key feature analysis, and thus store each climate data information item at a preset data storage location; S3. Perform dimension division processing on each climate data information item according to a preset dimension division rule to obtain each division dimension, and perform retrieval question-and-answer verification to obtain a retrieval question-and-answer verification determination result; S4. Obtain the system resource consumption parameters, analyze to obtain the system resource consumption determination level, and based on the analysis of the system resource consumption determination level and the retrieval question-and-answer verification determination result, obtain the division dimension adjustment determination result, and thus perform corresponding division dimension adjustment; S5. After the division dimension is adjusted, perform secondary retrieval question-and-answer verification to obtain a secondary retrieval question-and-answer verification determination result, and thus perform corresponding cold data migration adjustment.

2. The intelligent retrieval and question answering method for climate events based on multi-source data fusion according to claim 1, wherein: The method of storing each climate data information item at a preset data storage location is specifically as follows: Obtain each climate data information item in the fused climate dataset, and perform key feature extraction to obtain the key feature group of each climate data information item; Obtain the reference key feature set corresponding to each storage location in the preset system in the database; Match the key feature group of each climate data information item with the reference key feature set corresponding to each storage location respectively to obtain the reference key feature set with the maximum feature matching degree, and obtain the storage location corresponding to the reference key feature set as the preset data storage location of this climate data information item, and thus store each climate data information item at the preset data storage location.

3. The intelligent retrieval and question answering method for climate events based on multi-source data fusion according to claim 1, characterized in that: The method of obtaining the retrieval question-and-answer verification determination result is specifically as follows: Obtain the preset dimension division rule in the database; Perform dimension division on each climate data information item according to the preset dimension division rule to obtain each division dimension, and jointly mark each division dimension as the initial division dimension; Perform retrieval question-and-answer verification based on the initial division dimension, obtain the question-and-answer response effect parameter under the initial division dimension, and analyze to obtain the question-and-answer response effect value; Obtain the preset question-and-answer response effect threshold in the database, and compare it with the question-and-answer response effect value to obtain the retrieval question-and-answer verification determination result. If the question-and-answer response effect value is less than the question-and-answer response effect threshold, the retrieval question-and-answer verification determination result is unqualified verification. If the question-and-answer response effect value is above the question-and-answer response effect threshold, the retrieval question-and-answer verification determination result is qualified verification.

4. The climate event intelligent retrieval and question answering method based on multi-source data fusion according to claim 3, characterized in that: The specific steps of obtaining the question-and-answer response effect value are as follows: Obtain the question-and-answer response effect parameter under the initial division dimension within a preset time period, and the question-and-answer response effect parameter includes the average question-and-answer response duration, the correct rate of question-and-answer push, the system retrieval accuracy, and the maximum fluctuation ratio of the question-and-answer response duration; Obtain the preset question-and-answer response baseline set in the database, and perform proportional operation with the question-and-answer response effect parameter to obtain the response ratio analysis result. Based on the response ratio analysis result, introduce the corresponding weighting factor for quantitative coupling processing to obtain the question-and-answer response effect value; The Q&A response baseline set includes the baseline value of the average Q&A response duration, the baseline value of the Q&A push accuracy rate, the ideal value of the system retrieval accuracy, and the allowable fluctuation ratio of the Q&A response duration.

5. The intelligent retrieval and question answering method for climate events based on multi-source data fusion according to claim 1, characterized in that: The method for obtaining the system resource consumption parameters and analyzing to obtain the system resource consumption determination level is as follows: Obtain the system resource consumption parameters, where the system resource consumption parameters include the I / O resource utilization rate, the memory page error rate, the file descriptor usage amount, and the task queue length; Obtain the preset system resource consumption allowable set in the database, perform a ratio operation with the system resource consumption parameters to obtain the analysis result of the resource consumption ratio operation, and introduce the corresponding weighting factor for quantization coupling processing based on the analysis result of the resource consumption ratio operation to obtain the system resource consumption value; The system resource consumption allowable set includes the allowable value of the I / O resource utilization rate, the allowable value of the memory page error rate, the upper limit value of the file descriptor usage amount, and the upper limit value of the task queue length; Obtain the preset system resource consumption threshold in the database, compare it with the system resource consumption value to obtain the system resource consumption determination level. If the system resource consumption value is below the system resource consumption threshold, the system resource consumption determination result is level-one resource consumption. If the system resource consumption value is greater than the system resource consumption threshold, the system resource consumption determination result is level-two resource consumption.

6. The intelligent retrieval and question answering method for climate events based on multi-source data fusion according to claim 1, characterized in that: The method for analyzing based on the system resource consumption determination level and the retrieval Q&A verification determination result to obtain the division dimension adjustment determination result is as follows: Based on the analysis of the system resource consumption determination level and the retrieval Q&A verification determination result, if the retrieval Q&A verification determination result is verification failed, the division dimension adjustment determination result is to execute the division dimension adjustment. At the same time, if the system resource consumption determination level is level-one resource consumption, the division dimension adjustment determination result is to execute the division dimension dimension-increasing processing. If the system resource consumption determination level is level-two resource consumption, the division dimension adjustment determination result is to execute the division dimension dimension-decreasing processing; If the retrieval Q&A verification determination result is verification passed, the division dimension adjustment determination result is not to execute the division dimension adjustment.

7. The intelligent retrieval and question answering method for climate events based on multi-source data fusion according to claim 6, characterized in that: The method for executing the division dimension dimension-increasing processing is as follows: Obtain the data volume of each division dimension under the initial division dimension, perform the division dimension dimension-increasing processing on the division dimension corresponding to the maximum data volume, obtain the system resource consumption value after the division dimension dimension-increasing processing, and mark it as the system resource consumption dimension-increasing updated value; Perform a difference ratio analysis on the system resource consumption dimension-increasing updated value and the system resource consumption threshold to obtain the dimension-increasing updated ratio; Obtain the preset dimension-increasing updated ratio threshold in the database; Based on the analysis of the dimension-increasing updated ratio, the dimension-increasing updated ratio threshold, the system resource consumption dimension-increasing updated value, and the system resource consumption threshold, obtain the dimension-increasing processing result; If the system resource consumption dimension-increasing updated value is less than the system resource consumption threshold and the dimension-increasing updated ratio is less than the dimension-increasing updated ratio threshold, the dimension-increasing processing result is that the dimension-increasing processing is unqualified, and continue to execute the division dimension dimension-increasing processing until the system resource consumption dimension-increasing updated value is above the system resource consumption threshold or the dimension-increasing updated ratio is above the dimension-increasing updated ratio threshold, and complete the division dimension dimension-increasing processing; If the dimensionality-increasing update value of system resource consumption is above the system resource consumption threshold or the dimensionality-increasing update ratio is above the dimensionality-increasing update ratio threshold, the dimensionality-increasing processing result is qualified for dimensionality-increasing processing, and the dimensionality-increasing processing of the divided dimension is directly completed.

8. The intelligent retrieval and question answering method for climate events based on multi-source data fusion according to claim 6, characterized in that: The execution of the dimensionality-reduction processing of the divided dimension is specifically as follows: Obtain the contribution degree parameters of each divided dimension under the initial divided dimension, where the contribution degree parameters include the retrieval hit rate, the number of retrievable data instances, the memory occupancy, and the user adoption hit rate; Obtain the preset contribution degree baseline set in the database, where the contribution degree baseline set includes the retrieval hit rate baseline value, the baseline value of the number of retrievable data instances, the allowable value of the memory occupancy, and the user adoption hit rate baseline value; Perform proportional operations on the retrieval hit rate, the memory occupancy, and the user adoption hit rate with the retrieval hit rate baseline value, the allowable value of the memory occupancy, and the user adoption hit rate baseline value to obtain the contribution ratio operation result, perform difference degree analysis on the number of retrievable data instances and the baseline value of the number of retrievable data instances to obtain the contribution difference analysis result, and perform quantitative coupling processing on the contribution ratio operation result and the contribution difference analysis result by introducing the corresponding weighting factors to obtain the contribution value of each divided dimension; Obtain the preset contribution degree threshold in the database and compare it with the contribution value of each divided dimension. If the contribution value of a certain divided dimension is less than the contribution degree threshold, mark this divided dimension as the divided dimension to be adjusted, and thus obtain each divided dimension to be adjusted; Merge each divided dimension to be adjusted into the corresponding target divided dimension respectively to complete the dimensionality-reduction processing of the divided dimension; Obtain the question-and-answer response effect value after the dimensionality-reduction processing of the divided dimension, and judge whether to perform repeated dimensionality-reduction processing; The specific steps of merging each divided dimension to be adjusted into the corresponding target divided dimension respectively are as follows: Obtain each divided dimension in the database that has the same dimension type as the divided dimension to be adjusted, and mark them as each reference target divided dimension; Perform feature matching analysis on the divided dimension to be adjusted and each reference target divided dimension, and select the reference target divided dimension with the highest feature matching degree with the divided dimension to be adjusted as the target divided dimension of this divided dimension to be adjusted, and thus analyze the target divided dimension corresponding to each divided dimension to be adjusted; Merge each divided dimension to be adjusted into the corresponding target divided dimension respectively.

9. The climate event intelligent retrieval and question answering method based on multi-source data fusion according to claim 8, characterized in that: The specific method of obtaining the question-and-answer response effect value after the dimensionality-reduction processing of the divided dimension and judging whether to perform repeated dimensionality-reduction processing is as follows: Obtain the question-and-answer response effect value after the dimensionality-reduction processing of the divided dimension, and mark it as the question-and-answer response effect value after the dimensionality-reduction processing; Perform adjustment ratio analysis on the basis of the question-and-answer response effect value after the dimensionality-reduction processing and the question-and-answer response effect value to obtain the dimensionality-reduction effect adjustment ratio; Obtain the preset dimensionality-reduction effect adjustment ratio threshold in the database and analyze it with the dimensionality-reduction effect adjustment ratio. If the dimensionality-reduction effect adjustment ratio is below the dimensionality-reduction effect adjustment ratio threshold, do not perform repeated dimensionality-reduction processing. If the dimensionality-reduction effect adjustment ratio is greater than the dimensionality-reduction effect adjustment ratio, continue to perform repeated dimensionality-reduction processing.

10. The climate event intelligent retrieval and question answering method based on multi-source data fusion according to claim 1, wherein: The corresponding cold data migration adjustment is specifically as follows: After the division dimension is adjusted, perform retrieval and question-and-answer secondary verification to obtain the determination result of the retrieval and question-and-answer secondary verification. If the determination result of the retrieval and question-and-answer secondary verification is unqualified, perform cold data migration adjustment. If the determination result of the retrieval and question-and-answer secondary verification is qualified, do not perform cold data migration adjustment; If cold data migration adjustment is performed, obtain the access parameters of each cold data information item within a preset time period, where the access parameters include the number of accesses, the number of times adopted, and the adoption ratio; Obtain the preset access benchmark set in the database, perform a ratio operation with the access parameters to obtain the access difference analysis result, and introduce the corresponding weighting factor based on the access difference analysis result for quantization coupling processing to obtain the access comprehensive adoption value of each cold data information item; Obtain the preset access comprehensive adoption threshold in the database and compare it with the access comprehensive adoption value of each cold data information item. If the access comprehensive adoption value of a certain cold data information item is above the access comprehensive adoption threshold, perform cold data migration adjustment on this cold data information item and migrate it to the hot data storage area, otherwise do not perform cold data migration adjustment.

Citation Information

Patent Citations

  • A method and system for scenario-based multi-source data fusion analysis

    CN113609360B

  • A question-answering intelligent retrieval method and system for power documents

    CN117171333B

  • Multi-mode intelligent question and answer reasoning method and device based on knowledge distillation

    CN119623643A

  • Factory equipment operation data analysis and anomaly detection system based on machine learning

    CN119720054A

  • Intelligent data splitting and optimizing method and system based on multiple dimensions

    CN120045613A