Intelligent retrieval and question-answering method for climate events based on multi-source data fusion

By integrating data and adjusting dynamic dimensions of climate event intelligent search and question-and-answer platform, the problems of low response efficiency and low accuracy caused by unreasonable data dimension division are solved, and the response efficiency and accuracy are improved, ensuring the stable and efficient operation of the system under different load conditions.

CN120277196BActive Publication Date: 2025-08-22STATE QIHOU CENT +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510764283.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-08-22
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The existing intelligent search and question-and-answer platform for climate events is unreasonable, resulting in low response efficiency and low accuracy rate.

Method used

By obtaining the climate data sets of each data source for data fusion processing, key feature analysis and preset dimension division are performed, combined with the system resource consumption judgment level and search Q&A verification and judgment results, the dimension division is dynamically adjusted, and search Q&A secondary verification and cold data migration adjustment are performed after adjustment.

Benefits of technology

It improves the reply efficiency and accuracy of the Q&A system, ensures that the system operates stably and efficiently under different load conditions, optimizes the data storage structure, and reduces the impact of invalid data on the retrieval Q&A performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277196B_ABST
    Figure CN120277196B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent retrieval and question-answering method for climate events based on multi-source data fusion, which belongs to the technical field of electronic digital data processing and includes the following steps: S1, obtaining a fused climate data set; S2, obtaining each climate data information item in the fused climate data set, and performing key feature analysis, thereby storing each climate data information item in a preset data storage location; S3, obtaining each division dimension, and performing retrieval question and answer verification to obtain a retrieval question and answer verification judgment result; S4, obtaining a system resource consumption judgment level, and obtaining a division dimension adjustment judgment result based on the system resource consumption judgment level and the retrieval question and answer verification judgment result analysis, thereby performing corresponding division dimension adjustment; S5, obtaining a retrieval question and answer secondary verification judgment result, thereby performing corresponding cold data migration adjustment, thereby solving the problem of low response efficiency and low accuracy caused by unreasonable data dimension division in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic digital data processing, and in particular to a climate event intelligent retrieval and question-answering method based on multi-source data fusion. Background Art

[0002] As people pay more and more attention to climate events, the requirements for intelligent retrieval and question-answering platforms for climate data are also increasing. The existing climate event question-answering system is implemented by obtaining and analyzing multi-source data of set scenarios.

[0003] For example, the invention patent announcement with announcement number CN113609360B discloses a method and system based on scenario-based multi-source data fusion analysis, including: obtaining samples of multi-source data of a set scenario, preprocessing the multi-source data, and the preprocessing includes feature extraction; using a probability density estimation method to fuse the preprocessed multi-source data and normalize it to form a multi-source data feature set for the set scenario; using an association rule mining algorithm to extract association rules between features in the multi-source data feature set and corresponding specific scenario businesses, matching and inferring the corresponding business applications and solutions in the target scenario; extracting business features and knowledge features in a specific scenario as the basic data for constructing a knowledge and business library in a specific scenario, and then fusing features and knowledge, so that people can better gain more insights from massive and complex multi-source data, thereby realizing intelligent cognition and automatic analysis of multi-source data in specific scenarios.

[0004] For example, the invention patent announcement with announcement number CN117171333B discloses a question-and-answer intelligent retrieval method and system for power documents, including: a question-and-answer intelligent retrieval method for power documents, including: step S1, user semantic analysis, including: realizing user semantic concept extraction; realizing user semantic expansion; step S2, document retrieval and processing, including: establishing a file database; measuring document similarity; constructing an intention graph to represent the relationship between document data and query statements; step S3, answer extraction, including: presenting retrieval results based on the user's search intention in combination with traditional correlation features.

[0005] However, in the process of implementing the technical solutions of the invention in the embodiments of the present application, the present application found that the above technology has at least the following technical problems:

[0006] In the existing technology, when a user asks a question, the existing climate event intelligent retrieval and question-answering platform obtains and retrieves a large amount of climate event data and performs feature similarity analysis, and then feeds back the answer result to the user. However, since the platform stores a large amount of complex and disordered data information, there will be unreasonable data dimension division, resulting in the answer effect not satisfying the user. Therefore, the existing technology has the problem of low answer efficiency and low accuracy due to unreasonable data dimension division. Summary of the Invention

[0007] The embodiments of the present application solve the problems of low response efficiency and low accuracy caused by unreasonable data dimension division in the prior art by providing an intelligent retrieval and question-answering method for climate events based on multi-source data fusion, thereby achieving improved response efficiency and improved response accuracy.

[0008] An embodiment of the present application provides a climate event intelligent retrieval and question-answering method based on multi-source data fusion, comprising the following steps: S1, obtaining climate data sets from various data sources, and performing data fusion processing to obtain a fused climate data set; S2, obtaining each climate data information piece in the fused climate data set, and performing key feature analysis, thereby storing each climate data information piece in a preset data storage location; S3, performing dimension division processing on each climate data information piece according to a preset dimension division rule to obtain each division dimension, and performing retrieval question-answering verification to obtain a retrieval question-answering verification judgment result; S4, obtaining system resource consumption parameters, analyzing to obtain a system resource consumption judgment level, and obtaining a division dimension adjustment judgment result based on the system resource consumption judgment level and the retrieval question-answering verification judgment result, thereby performing corresponding division dimension adjustment; S5, performing retrieval question-answering secondary verification after the division dimension adjustment to obtain a retrieval question-answering secondary verification judgment result, thereby performing corresponding cold data migration adjustment.

[0009] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0010] 1. The intelligent retrieval and question-answering method for climate events based on multi-source data fusion provided by the present invention makes corresponding adjustments to the division dimensions based on the system resource consumption judgment level and the retrieval question-answering verification judgment result analysis, thereby improving the accuracy and response speed of the question-answering system, thereby achieving improved response efficiency and response accuracy, and effectively solving the problems of low response efficiency and low accuracy caused by unreasonable data dimension division in the existing technology.

[0011] 2. The present invention obtains system resource consumption parameters and analyzes them to obtain the system resource consumption judgment level, and adjusts the division dimensions based on the level and the retrieval question and answer verification judgment results, so as to dynamically adapt to changes in system resources and data characteristics, and ensure that the system can operate stably and efficiently under different load conditions.

[0012] 3. By performing secondary verification of retrieval questions and answers after adjusting the division dimensions, the secondary verification judgment results of the retrieval questions and answers are obtained. If the verification fails, cold data migration adjustment is performed, which can further optimize the data storage structure, reduce the impact of invalid data on retrieval question and answer performance, and improve the accuracy of question and answer responses. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 A flowchart of an implementation method for the intelligent retrieval and question-answering method for climate events based on multi-source data fusion provided in an embodiment of the present application;

[0014] Figure 2 A macroscopic flowchart of the climate event intelligent retrieval and question-answering method based on multi-source data fusion provided in an embodiment of the present application;

[0015] Figure 3 A detailed step diagram of the climate event intelligent retrieval and question-answering method based on multi-source data fusion provided in an embodiment of the present application. DETAILED DESCRIPTION

[0016] The embodiments of the present application solve the problems of low response efficiency and low accuracy caused by unreasonable data dimension division in the prior art by providing an intelligent retrieval and question-answering method for climate events based on multi-source data fusion, thereby achieving improved response efficiency and improved response accuracy.

[0017] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0018] like Figure 1 As shown, there is a flow chart of an implementation method of the intelligent retrieval and question-answering method for climate events based on multi-source data fusion provided in an embodiment of the present application, the method comprising the following steps: S1, obtaining climate data sets from various data sources, and performing data fusion processing to obtain a fused climate data set; S2, obtaining each climate data information piece in the fused climate data set, and performing key feature analysis, thereby storing each climate data information piece in a preset data storage location; S3, performing dimension division processing on each climate data information piece according to a preset dimension division rule to obtain each division dimension, and performing retrieval question-answering verification to obtain a retrieval question-answering verification judgment result; S4, obtaining system resource consumption parameters, analyzing to obtain a system resource consumption judgment level, and obtaining a division dimension adjustment judgment result based on the system resource consumption judgment level and the retrieval question-answering verification judgment result, thereby performing corresponding division dimension adjustment; S5, performing retrieval question-answering secondary verification after the division dimension adjustment to obtain a retrieval question-answering secondary verification judgment result, thereby performing corresponding cold data migration adjustment.

[0019] In this embodiment, if Figure 2 As shown, Figure 2 The macroscopic flow chart of the intelligent retrieval and question-answering method for climate events based on multi-source data fusion provided in the embodiment of the present application includes climate data sets from various data sources, and data fusion processing is performed to obtain a fused climate data set. Each climate data information item in the fused climate data set is obtained, and key feature analysis is performed, thereby storing each climate data information item in a preset data storage location. Each climate data information item is dimensionally divided according to a preset dimension division rule to obtain each division dimension, and retrieval question-answering verification is performed to obtain a retrieval question-answering verification judgment result. Based on the retrieval question-answering verification judgment result and the retrieval question-answering verification judgment result analysis, a division dimension adjustment judgment result is obtained, thereby performing corresponding division dimension adjustments, and performing corresponding cold data migration adjustments after the division dimension adjustments, and finally terminating the process.

[0020] like Figure 3 As shown, Figure 3 The detailed step diagram of the intelligent retrieval and question-answering method for climate events based on multi-source data fusion provided in the embodiment of the present application, the division dimension adjustment judgment result is divided into two cases: executing the division dimension dimensionality increase processing and executing the division dimension dimensionality reduction processing. When executing the division dimension dimensionality increase processing, the division dimension corresponding to the maximum data volume is subjected to dimensionality increase processing, and the system resource consumption dimensionality increase update value is obtained, and then the dimensionality increase update ratio is calculated to determine whether the dimensionality increase processing result is qualified. If qualified, the dimensionality increase processing is completed, otherwise the dimensionality increase processing is continued. When executing the division dimension dimensionality reduction processing, first obtain the contribution value of each division dimension, determine the division dimension to be adjusted, and then perform dimensionality reduction processing, and obtain the question and answer response effect value after dimensionality reduction, calculate the dimensionality reduction effect adjustment ratio, and if it is greater than the dimensionality reduction effect adjustment ratio threshold, complete the dimensionality reduction processing; otherwise repeat the dimensionality reduction processing. Finally, perform the corresponding cold data migration adjustment.

[0021] It should be noted that "multi-source data fusion" refers to the normalization and collaborative integration of heterogeneous data (including remote sensing image data, vector observation data, time series numerical simulation data, radar reflectivity image data, and text record data, etc.) obtained from different observation platforms (satellite remote sensing systems, meteorological radars, ground observation stations, weather forecasts) according to a unified structural format and semantic standards to obtain data of unified dimensions.

[0022] Furthermore, each climate data information piece is stored in a preset data storage location. The specific method is: obtain each climate data information piece in the fused climate data set, and perform key feature extraction to obtain a key feature group of each climate data information piece; obtain a reference key feature set corresponding to each storage location in the system preset in the database; perform feature matching on the key feature group of each climate data information piece with the reference key feature set corresponding to each storage location, obtain a reference key feature set corresponding to the maximum feature matching degree, and obtain the storage location corresponding to the reference key feature set as the preset data storage location of the climate data information piece, thereby storing each climate data information piece in the preset data storage location.

[0023] In this embodiment, each data entry is accurately located by utilizing a reference key feature set, thereby avoiding storing all data in a single database table, thereby significantly reducing the overhead of scanning irrelevant storage partitions during data retrieval and improving I / O efficiency. Matching feature mapping aggregates data of similar climate events or similar attributes into the same storage location, and the retrieval engine can directly locate highly relevant partitions. Therefore, users can hit target data faster when querying, shorten response latency, and improve system throughput.

[0024] The mapping between preset key features and storage locations can be dynamically adjusted through a configuration table. When adding or adjusting climate factors (such as adding the "Extreme Wind Speed" metric), simply update the "Reference Key Feature Set" and its storage mapping for immediate effect. By grouping partitions with high feature matching based on feature analysis, anomalous or irrelevant entries are effectively eliminated, ensuring high data consistency within each storage partition. This also facilitates subsequent verification, backup, and archiving management, reducing data redundancy and consistency risks.

[0025] It should be noted that key features are extracted. The specific methods are: cleaning and standardizing the acquired climate data to remove noise and outliers; selecting features related to climate events, such as time, location, event type, intensity, etc., based on domain knowledge and business needs; converting the selected features into a unified format or dimension, such as converting time into a standard time format and converting location into geographic coordinates; applying feature extraction algorithms (such as TF-IDF) to extract key features.

[0026] Among them, the key feature group is, for example: the typhoon event key feature group time feature: 2023-10-01 14:30, location feature: longitude 113.5°E, latitude 22.3°N, event type feature: typhoon, intensity feature: central pressure 960hPa, maximum wind speed 45m / s.

[0027] The key feature groups for each climate data item are matched against the reference key feature sets corresponding to each storage location. Specifically, the key feature groups for each climate data item are extracted from the fused climate dataset. For example, for a typhoon event data item, its key feature groups include: Time: 2023-10-01 14:30, Location: 113.5°E, 22.3°N, Event Type: Typhoon, Intensity: Central Pressure 960hPa, Maximum Wind Speed ​​45m / s. Reference Key Feature Set: A preset reference key feature set is obtained from the database. This feature set contains multiple reference key feature groups, each of which is associated with a specific storage location in the database. For example, the reference key feature set includes: storage location A: time range 2023-10-01 00:00 to 2023-10-01 23:59, location 113.0° to 114.0° east longitude, 22.0° to 23.0° north latitude, event type typhoon, intensity range central pressure 950hPa to 970hPa, maximum wind speed 40m / s to 50m / s; storage location B: time range 2023-09-01 00:00 to 2023-09-30 23:59, location 115.0° to 116.0° east longitude, 23.0° to 24.0° north latitude, event type heavy rain, intensity range precipitation 100mm / h to 150mm / h. The extracted key feature groups and reference key feature sets are standardized. For example, time is converted to a unified time format, location is converted to a unified geographic coordinate format (such as decimal degrees for longitude and latitude), and intensity metrics are converted to a unified unit and range. Feature matching is then performed. For each climate data item's key feature group, it is matched against the reference key feature set for each storage location. The matching degree for each feature is calculated. For example, for time matching, if the climate data item's time falls within the time range of storage location A, the temporal matching degree is 100%; otherwise, it is 0%. The matching degrees for other features are calculated similarly. The overall matching degree is calculated by combining all the matching degrees of all features (taking a weighted sum of the matching degrees of each feature). Finally, the overall matching degrees of the climate data item and each storage location are compared to determine the best matching location. The climate data item is then stored in the pre-set data storage location corresponding to the storage location with the highest overall matching degree. If the overall matching degrees for multiple storage locations are the same, further determination can be made based on business rules (e.g., the storage location with the highest priority) to determine the pre-set data storage location.

[0028] Furthermore, the retrieval question and answer verification result is obtained, and the specific method is: obtain the preset division dimension rules in the database; divide each climate data information piece into dimensions according to the preset division dimension rules to obtain each division dimension, and jointly mark each division dimension as the initial division dimension; perform retrieval question and answer verification based on the initial division dimension, obtain the question and answer response effect parameters under the initial division dimension, and analyze to obtain the question and answer response effect value; obtain the question and answer response effect threshold preset in the database, and compare it with the question and answer response effect value to obtain the retrieval question and answer verification result. If the question and answer response effect value is less than the question and answer response effect threshold, the retrieval question and answer verification result is verification failure. If the question and answer response effect value is above the question and answer response effect threshold, the retrieval question and answer verification result is verification qualification.

[0029] In this embodiment, by performing retrieval question and answer verification under the initial division dimensions, the system response effect under different dimension combinations can be pre-detected to ensure that the question and answer results meet the set accuracy and latency indicators, thereby avoiding users encountering incorrect or incomplete answers in the actual question and answer environment, and improving the overall system credibility and user experience.

[0030] It should be noted that the preset dimension rules can be flexibly configured for different climate event types (such as typhoons, droughts, and rainstorms), allowing for rapid switching of verification models in different scenarios. The preset dimension partitioning rules are pre-stored in the database. If the Q&A response effect value is less than the Q&A response effect threshold, the retrieval Q&A verification result is deemed unqualified. If the verification fails, the subsequent "dimensional adjustment and secondary verification" process is automatically triggered, achieving multiple rounds of optimization to meet diverse usage needs such as scientific research, early warning, and popular science.

[0031] Preset dimension rules, for example: division by time dimension: divide climate data into 2020, 2021, 2022, etc.; division by location dimension: divide into continents (Asia, Europe, Africa, etc.), countries, provinces, etc.; division by climate event type: such as typhoons, hurricanes, heavy rains, droughts, high temperatures, cold waves, etc.

[0032] The automated verification and judgment process replaces manual testing, significantly reducing manual intervention and testing costs. By discovering and locating dimension division problems in advance, large-scale data reconstruction or emergency repairs can be avoided, thereby improving operation and maintenance efficiency.

[0033] Furthermore, the question and answer response effect value is obtained. The specific steps are: obtaining the question and answer response effect parameters under the initial division dimension within the preset time period, and the question and answer response effect parameters include the average question and answer response time, the question and answer push accuracy, the system retrieval accuracy and the maximum fluctuation ratio of the question and answer response time; obtaining the preset question and answer response baseline set in the database, and performing proportional operation with the question and answer response effect parameters to obtain the response proportion analysis result, introducing the corresponding empowerment factor based on the response proportion analysis result for quantitative coupling processing to obtain the question and answer response effect value; the question and answer response baseline set includes the baseline value of the average question and answer response time, the baseline value of the question and answer push accuracy, the ideal value of the system retrieval accuracy and the allowed fluctuation ratio of the question and answer response time.

[0034] In this embodiment, the average duration of a question-and-answer response refers to the average time it takes for a user to initiate a query request and for the system to return a complete answer when testing within a preset time period. This can be obtained by analyzing the time-consuming log. The accuracy of question-and-answer push refers to the ratio of the number of correct answers pushed to users by the system (the number of times the system-pushed content is liked and adopted by the user) to the total number of system pushes, which can be obtained by querying records in the system log. The system retrieval accuracy represents the average value of the matching degree of each system push within a preset time period (the feature matching degree of the key features of the system-pushed content and the key features of the user's question content). The feature matching degree can be analyzed using the NLTK tool. The maximum fluctuation ratio of the question-and-answer response time can be obtained by calculating the difference between the maximum and minimum values ​​of the system question-and-answer response time to obtain the fluctuation difference, and then dividing the fluctuation difference by the average question-and-answer response time.

[0035] The question-and-answer response effect value is obtained by analyzing the question-and-answer response effect parameters under the initial division dimensions (including the average question-and-answer response time, the question-and-answer push accuracy, the system retrieval accuracy, and the maximum fluctuation ratio of the question-and-answer response time). This takes into account the mutual influence between these parameters. For example, when the system retrieval accuracy (average matching degree) increases, it is necessary to scan the candidate data more deeply or widely. Without changing the information storage location, etc., the average question-and-answer response time will increase; if the average time is too long, it will affect the user experience and cause the push satisfaction rate to decrease; if the response time fluctuates greatly in each query (the maximum fluctuation ratio is high), even if the overall average response is fast and the retrieval is accurate, users will still reduce likes and adoptions due to the unstable experience, thereby lowering the question-and-answer push accuracy rate.

[0036] Get the question-answer response effect value. The specific method is:

[0037] ;

[0038] Where, Indicates the question-answer response effect value, represents the average time of question-answer response, represents the baseline value of the average duration of question-answer response, Indicates the accuracy rate of question and answer push. Indicates the baseline value of the correct rate of question and answer push. Indicates the system retrieval accuracy, Indicates the ideal value of system retrieval accuracy, Indicates the maximum fluctuation ratio of the question-answer response time, Indicates the allowed fluctuation ratio of the question and answer response time. represents the weighting factor of the average duration of question-answer response, It represents the weighting factor of the correct rate of question and answer push, represents the system retrieval precision weighting factor, Represents the weighting factor of the maximum fluctuation ratio of question and answer response time.

[0039] The weighting factor for the average duration of question and answer responses, the weighting factor for the accuracy of question and answer push, the weighting factor for the system retrieval precision and the weighting factor for the maximum fluctuation ratio of question and answer response duration can be obtained from the database. For example, the weighting factor for the average duration of question and answer responses can be obtained by obtaining the historical average duration of question and answer responses stored in the database, and the average duration weighting factor for question and answer responses corresponding to the historical average duration of question and answer responses, thereby constructing a mapping set for the average duration of question and answer responses, wherein there is a one-to-one or many-to-one correspondence in the mapping set. The weighting factor for the average duration of question and answer responses can be obtained by inputting the required average duration of question and answer responses into the mapping set for the average duration of question and answer responses. The other weighting factors are obtained in the same way as the weighting factor for the average duration of question and answer responses, and can all be matched in the corresponding mapping sets. Among them, the weighting factor for the accuracy of question and answer push corresponds to the mapping set for the accuracy of question and answer push, the weighting factor for the system retrieval precision corresponds to the mapping set for the system retrieval precision, and the weighting factor for the maximum fluctuation ratio of question and answer response duration corresponds to the mapping set for the maximum fluctuation ratio of question and answer response duration.

[0040] Furthermore, system resource consumption parameters are obtained and analyzed to obtain a system resource consumption determination level. The specific method is as follows: obtain system resource consumption parameters, which include I / O resource utilization, memory page error rate, file descriptor usage, and task queue length; obtain a preset system resource consumption allowable set in the database, and perform a proportional operation on the system resource consumption parameters to obtain a resource consumption proportional operation analysis result; based on the resource consumption proportional operation analysis result, introduce a corresponding weighting factor for quantitative coupling processing to obtain a system resource consumption value; the system resource consumption allowable set includes an I / O resource utilization allowable value, a memory page error rate allowable value, a file descriptor usage upper limit value, and a task queue length upper limit value; obtain a preset system resource consumption threshold value in the database, and compare it with the system resource consumption value to obtain a system resource consumption determination level. If the system resource consumption value is below the system resource consumption threshold value, the system resource consumption determination result is a first-level resource consumption; if the system resource consumption value is greater than the system resource consumption threshold value, the system resource consumption determination result is a second-level resource consumption.

[0041] In this embodiment, I / O resource utilization can be monitored using a performance monitor. The memory page fault rate refers to the number of memory page faults per unit time, reflecting the level of memory resource shortage. This can be monitored using performance monitoring tools such as vmstat. File descriptor usage refers to the number of open file descriptors in the system, reflecting the system's file resource usage. This can be monitored using the lsof tool. The task queue length refers to the number of processes in the system waiting for CPU execution, reflecting the system load. This can be obtained by querying the top command-line tool.

[0042] Analyzing system resource consumption parameters (including I / O resource utilization, memory page fault rate, file descriptor usage, and task queue length) takes into account the interplay between these parameters. For example, when I / O resource utilization is high, the memory page fault rate increases. Frequent data swapping leads to increased memory page faults, which in turn trigger more I / O operations, increasing the I / O burden. Excessive file descriptor usage, often accompanied by numerous file operations, increases I / O resource utilization. Furthermore, high I / O utilization can increase task queue lengths, as processes are often blocked waiting for I / O operations to complete. Excessive file descriptor usage can also lead to increased task queue lengths due to frequent file operations. When the memory page fault rate is high, processes frequently wait for pages to be swapped in and out, increasing the task queue length.

[0043] Get the system resource consumption value. The specific method is:

[0044] ;

[0045] Where, Indicates the system resource consumption value. Indicates I / O resource utilization, Indicates the allowed value of I / O resource utilization. Indicates the memory page fault rate, Indicates the allowed value of memory page fault rate. Indicates the file descriptor usage, Indicates the upper limit of file descriptor usage. Indicates the length of the task queue, Indicates the upper limit of the task queue length. Indicates the I / O resource utilization weighting factor, Represents the memory page fault rate weighting factor, Indicates the file descriptor usage weighting factor, Indicates the task queue length weighting factor.

[0046] The I / O resource utilization weighting factor, memory page fault rate weighting factor, file descriptor usage weighting factor and task queue length weighting factor can be obtained from the database. For example, the I / O resource utilization weighting factor can be obtained by obtaining the historical I / O resource utilization stored in the database, and the I / O resource utilization weighting factor corresponding to the historical I / O resource utilization, thereby constructing an I / O resource utilization mapping set, wherein there is a one-to-one or many-to-one correspondence in the mapping set. The I / O resource utilization weighting factor can be obtained by inputting the I / O resource utilization data to be used into the I / O resource utilization mapping set. The other weighting factors are obtained in the same way as the I / O resource utilization weighting factor, and can all be matched in the corresponding mapping set. Among them, the memory page fault rate weighting factor corresponds to the memory page fault rate mapping set, the file descriptor usage weighting factor corresponds to the file descriptor usage mapping set, and the task queue length weighting factor corresponds to the task queue length mapping set.

[0047] Furthermore, based on the analysis of the system resource consumption judgment level and the retrieval question and answer verification judgment results, the division dimension adjustment judgment result is obtained. The specific method is: based on the analysis of the system resource consumption judgment level and the retrieval question and answer verification judgment results, if the retrieval question and answer verification judgment result is that the verification is unqualified, then the division dimension adjustment judgment result is to execute the division dimension adjustment. At the same time, if the system resource consumption judgment level is the first-level resource consumption, then the division dimension adjustment judgment result is to execute the division dimension increase processing. If the system resource consumption judgment level is the second-level resource consumption, then the division dimension adjustment judgment result is to execute the division dimension reduction processing; if the retrieval question and answer verification judgment result is that the verification is qualified, then the division dimension adjustment judgment result is not to execute the division dimension adjustment.

[0048] In this embodiment, when the retrieval question and answer verification result is a failure and the system resource consumption is determined to be level one (i.e., the system resource consumption value is below the threshold), this indicates that the system's current resource consumption is within an acceptable range, but the retrieval question and answer performance is poor. At this point, performing dimensionality increase processing can further refine the data segmentation, thereby improving data organization and retrieval efficiency. Due to sufficient system resources, dimensionality increase does not impose an excessive burden on system performance. It can also improve the accuracy and responsiveness of retrieval questions and answers, thereby enhancing data representation and retrieval precision, enabling the system to better meet user query requirements while fully utilizing the system's existing resource capabilities.

[0049] If the query verification result is a failure, and the system resource consumption level is determined to be Level 2 (i.e., the system resource consumption value exceeds the threshold), this indicates that system resources are limited. In this case, performing dimensionality reduction can reduce the complexity of data processing and alleviate system resource pressure. Dimensionality reduction removes redundant or low-contribution dimensions, allowing the system to focus resources on processing critical data and improving overall system efficiency. When system resources are limited, dimensionality reduction helps optimize resource allocation and ensure stable system operation. It also improves the responsiveness of query and answer queries, enhances system resource utilization efficiency, avoids system performance degradation or service interruptions caused by excessive resource consumption, and ensures the continuous and effective operation of the system.

[0050] When the query-answer verification result is qualified, the current partitioning dimension meets the query-answer requirements, and the system does not need to adjust the dimension. This maintains system stability and consistency, and avoids the potential impact of unnecessary adjustments on the system. Furthermore, a stable partitioning dimension helps maintain the maintainability and understandability of the system, facilitating subsequent management and optimization. The benefit of not performing partitioning dimension adjustments is that it saves the cost and time of system adjustments, ensures efficient system operation under a stable partitioning dimension, and avoids the increased system burden and complexity caused by frequent adjustments.

[0051] Furthermore, dimension multiplication processing is performed, and the specific method is as follows: obtaining the data volume of each dimension under the initial dimension, performing dimension multiplication processing on the dimension corresponding to the maximum data volume, obtaining the system resource consumption value after dimension multiplication processing, and marking it as the system resource consumption dimension multiplication update value; performing a difference ratio analysis on the system resource consumption dimension multiplication update value and the system resource consumption threshold to obtain the dimension multiplication update ratio; obtaining the dimension multiplication update ratio threshold preset in the database; based on the dimension multiplication update ratio, the dimension multiplication update ratio threshold, the system resource consumption dimension multiplication update value and the system resource consumption threshold analysis, The result of the dimensionality increase processing is obtained; if the dimensionality increase update value of the system resource consumption is less than the system resource consumption threshold and the dimensionality increase update ratio is less than the dimensionality increase update ratio threshold, the result of the dimensionality increase processing is that the dimensionality increase processing is unqualified, and the dimension increase processing of the divided dimension is continued until the dimensionality increase update value of the system resource consumption is above the system resource consumption threshold or the dimensionality increase update ratio is above the dimensionality increase update ratio threshold, and the dimension increase processing of the divided dimension is completed; if the dimensionality increase update value of the system resource consumption is above the system resource consumption threshold or the dimensionality increase update ratio is above the dimensionality increase update ratio threshold, the result of the dimensionality increase processing is that the dimensionality increase processing is qualified, and the dimension increase processing of the divided dimension is directly completed.

[0052] In this embodiment, by performing dimensionality increase processing on the division dimension corresponding to the maximum data volume and analyzing the resource consumption dimensionality increase update value, by comparing the resource consumption dimensionality increase update value with the system resource consumption threshold, and comparing the dimensionality increase update ratio with the dimensionality increase update ratio threshold, it is possible to accurately judge whether the dimensionality increase processing has achieved the expected effect, avoiding the misjudgment of the dimensionality increase processing caused by a single judgment rule, ensuring the accuracy and effectiveness of the dimensionality increase processing, and avoiding unnecessary repeated operations.

[0053] Comparing the system resource consumption dimensionality increase update value with the system resource consumption threshold ensures that the system resource consumption during the dimensionality increase process does not exceed the allowed range, which helps maintain the stable operation of the system and avoids performance degradation or service interruption due to excessive resource consumption. If the expected effect is still not achieved after the dimensionality increase process, the system will continue to perform the dimensionality increase process until the conditions are met, so that the system can adapt to different data conditions and resource limitations, and improve the adaptability and robustness of the system. Through dimensionality increase processing, the organization and structure of the data are optimized, so that the data can better reflect the actual situation, improve the availability and retrieval efficiency of the data, and more accurately meet the user's query needs and provide more accurate answers. By setting the dimensionality increase update ratio threshold, unlimited dimensionality increase operations can be avoided, preventing waste of system resources, ensuring that the dimensionality increase process stops in time after achieving the expected effect, and maintaining the efficient operation of the system.

[0054] It should be noted that the specific method for continuing to execute the dimension increasing process is: repeatedly obtaining the data volume of each dimension after the current dimension increasing process is executed, performing dimension increasing process on the dimension corresponding to the maximum data volume, and obtaining the system resource consumption value at this time, that is, obtaining the system resource consumption dimension increasing update value of this round, thereby analyzing and obtaining the dimension increasing update ratio, dimension increasing update ratio threshold, system resource consumption dimension increasing update value and system resource consumption threshold of this round of dimension increasing process, and judging the result of this round of dimension increasing process, if the result of this round of dimension increasing process is still unqualified, and continuing the next round of dimension increasing process, the next round of dimension increasing process is exactly the same as the steps of this round of dimension increasing process, until the system resource consumption dimension increasing update value is above the system resource consumption threshold or the dimension increasing update ratio is above the dimension increasing update ratio threshold, and the dimension increasing process is completed.

[0055] It should be noted that the difference ratio analysis of the system resource consumption dimensionality update value and the system resource consumption threshold is performed to obtain the dimensionality update ratio. The specific steps are: performing difference processing on the system resource consumption dimensionality update value and the system resource consumption threshold to obtain the system resource consumption difference, and dividing the system resource consumption difference by the system resource consumption threshold to obtain the dimensionality update ratio.

[0056] Furthermore, dimension reduction processing of the partitioning dimensions is performed, and the specific method is as follows: obtaining the contribution parameters of each partitioning dimension under the initial partitioning dimension, the contribution parameters include retrieval hit rate, number of retrievable data instances, memory occupancy and user adoption hit rate; obtaining the contribution baseline set preset in the database, the contribution baseline set includes the retrieval hit rate baseline value, the retrievable data instance number baseline value, the memory occupancy allowable value and the user adoption hit rate baseline value; performing proportional operation based on the retrieval hit rate, memory occupancy and user adoption hit rate and the retrieval hit rate baseline value, the memory occupancy allowable value and the user adoption hit rate baseline value to obtain the contribution proportion operation result, performing difference analysis based on the number of retrievable data instances and the retrievable data instance number baseline value to obtain the contribution difference analysis result, and introducing the corresponding weighting factor based on the contribution proportion operation result and the contribution difference analysis result to perform quantitative coupling processing to obtain the contribution value of each partitioning dimension; obtaining the contribution threshold value preset in the database, and comparing it with the contribution threshold of each partitioning dimension The contribution values ​​of the division dimensions are compared. If the contribution value of a certain division dimension is less than the contribution threshold, the division dimension is marked as the division dimension to be adjusted, thereby obtaining the division dimensions to be adjusted; each division dimension to be adjusted is respectively incorporated into the corresponding target division dimension to complete the division dimension dimensionality reduction processing; the question-answer response effect value after the division dimension dimensionality reduction processing is obtained, and it is determined whether to perform repeated dimensionality reduction processing; each division dimension to be adjusted is respectively incorporated into the corresponding target division dimension, and the specific steps are: obtain each division dimension of the same dimension type as the division dimension to be adjusted in the database, and mark it as each reference target division dimension; perform feature matching analysis on the division dimension to be adjusted and each reference target division dimension, and select the reference target division dimension with the highest feature matching degree with the division dimension to be adjusted as the target division dimension of the division dimension to be adjusted, thereby obtaining the target division dimension corresponding to each division dimension to be adjusted; each division dimension to be adjusted is respectively incorporated into the corresponding target division dimension.

[0057] In this embodiment, by performing dimensionality reduction processing, the complexity of data processing and storage requirements can be reduced, thereby reducing the consumption of system resources and helping to improve the overall performance and response speed of the system. Dimensionality reduction can remove redundant or low-contribution dimensions, allowing data to focus more on key features, helping to improve the accuracy and efficiency of retrieval and question-answering, enabling the system to respond to user queries more quickly and provide more relevant results.

[0058] By obtaining the Q&A response performance value after dimensionality reduction and comparing it with the preset adjustment ratio threshold, we can determine whether the dimensionality reduction process has achieved the desired effect. If the adjustment ratio of the Q&A response effect after dimensionality reduction is lower than the threshold, it means that the dimensionality reduction process is not complete and further dimensionality reduction operations are required to further optimize data organization and retrieval efficiency. At the same time, it can also prevent excessive dimensionality reduction, which can affect retrieval and Q&A accuracy, and ensure data integrity and availability.

[0059] By obtaining other partitioning dimensions of the same type as the partitioning dimension to be adjusted, we can ensure the consistency and relevance of the data during the dimensionality reduction process, helping to maintain the semantic integrity and logical relationships of the data, so that the reduced data can still accurately reflect the actual situation. Marking these partitioning dimensions of the same type as the reference target partitioning dimensions can provide more suitable merging targets for the partitioning dimension to be adjusted. Through feature matching analysis, selecting the reference target partitioning dimension with the highest feature match with the partitioning dimension to be adjusted for merging ensures the rationality and effectiveness of the dimensionality reduction operation, avoids the information confusion caused by the forced merging of unrelated data, and helps improve the quality of the data after dimensionality reduction and the accuracy of retrieval questions and answers.

[0060] The retrieval hit rate refers to the frequency with which data from a particular dimension is successfully hit during the question-and-answer retrieval process. Specifically, it's the ratio of the number of times the system uses data from that dimension to the total number of searches when answering user questions. This can be obtained using data analysis tools (such as Python's Pandas library). The number of retrievable data instances refers to the amount of data available for retrieval under a particular dimension, specifically the number of data records contained in that dimension. This is obtained through database queries. Memory usage refers to the amount of memory space occupied by data from a particular dimension, which can be obtained through querying the task manager. The user adoption hit rate refers to the user's acceptance of the answers provided by the question-and-answer system. This is the ratio of the number of times users adopt answers generated by data from that dimension to the total number of times data from that dimension is used in answers. This can be obtained through user feedback logs.

[0061] It should be noted that the contribution values ​​of each dimension are derived by analyzing contribution parameters including retrieval hit rate, number of retrievable data instances, memory usage, and user adoption hit rate. This takes into account the interplay between these parameters. For example, a high retrieval hit rate indicates that the system can effectively utilize the number of retrievable data instances. A large number of retrievable data instances can improve the hit rate, but if the memory capacity is exceeded, the memory usage will soar, which will in turn hinder system efficiency. The user adoption hit rate is affected by the retrieval hit rate and the quality of the data instances. A high hit rate and high-quality data can increase the user adoption rate. At the same time, too many instances may dilute attention and reduce the adoption hit rate. Memory usage is affected by the number of retrievable data instances; a large number of instances consumes more memory.

[0062] The contribution value of each division dimension is obtained by:

[0063] ;

[0064] Where, Indicates the contribution value of the i-th partition dimension, i represents the number of the partition dimension, , Indicates the total number of partition dimensions, represents the retrieval hit rate of the i-th partition dimension, Indicates the retrieval hit rate baseline value, Indicates the baseline value of the number of retrievable data instances. Indicates the number of retrievable data instances of the i-th partition dimension, Indicates the memory usage of the i-th partition dimension, Indicates the allowed value of memory usage. represents the user adoption hit rate of the i-th partition dimension, Indicates the user adoption hit rate baseline value, represents a constant, Represents the retrieval hit rate weighting factor, represents the weighting factor of the number of retrievable data instances, Represents the memory usage weighting factor, Indicates the user adoption hit rate weighting factor.

[0065] It should be noted that The representation constant is an extremely small number used to prevent calculation errors caused when the number of retrievable data instances is the same as the baseline value of the number of retrievable data instances, thereby ensuring the accuracy of the calculation results.

[0066] The retrieval hit rate weighting factor, the retrievable data instance number weighting factor, the memory usage weighting factor and the user adoption hit rate weighting factor can be obtained from the database. For example, the retrieval hit rate weighting factor can be obtained by obtaining the historical retrieval hit rates stored in the database, and the retrieval hit rate weighting factors corresponding to the historical retrieval hit rates, thereby constructing a retrieval hit rate mapping set, wherein there is a one-to-one or many-to-one correspondence in the mapping set. The retrieval hit rate weighting factor can be obtained by inputting the retrieval hit rate data to be used into the retrieval hit rate mapping set. The other weighting factors are obtained in the same way as the retrieval hit rate weighting factor, and can all be matched in the corresponding mapping set. Among them, the retrievable data instance number weighting factor corresponds to the retrievable data instance number mapping set, the memory usage weighting factor corresponds to the memory usage mapping set, and the user adoption hit rate weighting factor corresponds to the user adoption hit rate mapping set.

[0067] Obtain all partition dimensions in the database that are of the same dimension type as the partition dimension to be adjusted, and mark them as reference target partition dimensions. For example, if the partition dimension to be adjusted is "Typhoons in May 2023" under the "Time Dimension", then mark other partition dimensions found under the "Time Dimension" (such as "Typhoons in the first half of 2023" etc.) as reference target partition dimensions.

[0068] Furthermore, the question and answer response effect value after the dimension reduction processing of the divided dimension is obtained to determine whether to perform repeated dimensionality reduction processing. The specific method is: obtain the question and answer response effect value after the dimension reduction processing of the divided dimension, and mark it as the dimension reduction processing question and answer response effect value; perform an adjustment ratio analysis based on the dimension reduction processing question and answer response effect value and the question and answer response effect value to obtain the dimensionality reduction effect adjustment ratio; obtain the dimensionality reduction effect adjustment ratio threshold preset in the database, and analyze it with the dimensionality reduction effect adjustment ratio. If the dimensionality reduction effect adjustment ratio is below the dimensionality reduction effect adjustment ratio threshold, repeated dimensionality reduction processing will not be performed. If the dimensionality reduction effect adjustment ratio is greater than the dimensionality reduction effect adjustment ratio, repeated dimensionality reduction processing will continue to be performed.

[0069] In this embodiment, if the adjustment ratio of the dimensionality reduction effect is greater than the preset threshold, it means that although the current dimensionality reduction process has achieved certain results, there is still room for improvement. Continuing dimensionality reduction can further optimize data organization, reduce redundant dimensions, and improve the compactness and relevance of data. By repeating the dimensionality reduction process, the data structure can be adjusted more finely, making the response effect of the question-and-answer system more accurate and efficient, which helps to improve user satisfaction and the overall performance of the system. When the adjustment ratio of the question-and-answer response effect value after the dimensionality reduction process to the original effect value exceeds the preset threshold, it indicates that the current dimensionality reduction effect has not achieved the expected goal. At this time, it is necessary to continue the dimensionality reduction process to further improve the question-and-answer response effect.

[0070] Furthermore, corresponding cold data migration adjustment is performed. The specific method is as follows: after the dimension adjustment, a secondary verification of the search question and answer is performed to obtain a secondary verification result of the search question and answer. If the secondary verification result of the search question and answer is a verification failure, the cold data migration adjustment is performed; if the secondary verification result of the search question and answer is a verification pass, the cold data migration adjustment is not performed; if the cold data migration adjustment is performed, the access parameters of each cold data information item within a preset time period are obtained, and the access parameters include the number of accesses, the number of adoptions, and the adoption ratio; the access benchmark set preset in the database is obtained and proportional operation is performed on the access parameters to obtain an access difference analysis result; based on the access difference analysis result, a corresponding empowerment factor is introduced to perform quantitative coupling processing to obtain an access comprehensive adoption value of each cold data information item; the access comprehensive adoption threshold preset in the database is obtained and compared with the access comprehensive adoption value of each cold data information item; if the access comprehensive adoption value of a cold data information item is above the access comprehensive adoption threshold, the cold data information item is subjected to cold data migration adjustment and migrated to the hot data storage area; otherwise, the cold data migration adjustment is not performed.

[0071] In this embodiment, when increasing the dimension, if the increased data volume causes the retrieval question and answer verification to fail, migrating cold data to hot storage can improve access speed and retrieval efficiency, thereby optimizing the question and answer results. When reducing the dimension, if data redundancy or low-quality data affects the verification results, migrating qualified cold data to the hot storage area can focus on high-quality data and improve the accuracy and relevance of the question and answer. Migrating cold data that meets specific conditions to the hot data storage area improves data management efficiency, reduces the storage resource usage of infrequently used data, and optimizes the storage structure.

[0072] After performing the dimension division and dimension increase processing, if the retrieval question and answer verification fails, this indicates that optimal optimization has been achieved from the data dimension perspective. Because increasing the data dimension will cause a greater occupation of numerical resources, excessive data dimension increase processing cannot be performed. Therefore, it is necessary to adjust the data organization structure at this time. It is necessary to migrate cold data to the hot data area to make cold data easier to access and retrieve, thereby improving the performance of the question and answer system and avoiding excessive system resource occupation caused by excessive data dimension increase processing. After performing the dimension division and dimension reduction processing, if the retrieval question and answer verification fails, the corresponding cold data migration adjustment is performed at this time to consider that the data dimension reduction processing will not affect the increase in system resource occupation. At this time, if the dimension reduction effect adjustment ratio is below the dimension reduction effect adjustment ratio threshold, Best's data dimension reduction processing can no longer improve the question and answer response effect. Therefore, it is necessary to perform corresponding cold data migration adjustments to change the data organization structure and thereby improve the question and answer response effect value.

[0073] To sum up, this embodiment improves the accuracy and response speed of the question-and-answer system by analyzing the system resource consumption judgment level and the retrieval question-and-answer verification judgment results, and thus makes corresponding adjustments to the division dimensions, thereby achieving improved response efficiency and improved response accuracy, and effectively solving the problems of low response efficiency and low accuracy caused by unreasonable data dimension division in the existing technology.

[0074] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0075] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0076] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0077] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0078] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0079] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. An intelligent retrieval and question-answering method for climate events based on multi-source data fusion, characterized by: The following steps are involved: S1. Obtain climate datasets from various data sources and perform data fusion processing to obtain a fused climate dataset; S2. Obtain each climate data information piece in the fused climate data set, perform key feature analysis, and store each climate data information piece in a preset data storage location; S3. Divide each climate data information piece into dimensions according to a preset dimension division rule to obtain each division dimension, and perform search question and answer verification to obtain a search question and answer verification judgment result; S4. Obtain system resource consumption parameters, analyze and obtain a system resource consumption determination level, analyze the system resource consumption determination level and the search question and answer verification determination results, obtain a partition dimension adjustment determination result, and perform corresponding partition dimension adjustments accordingly; S5. After the division dimension is adjusted, perform secondary verification of the search question and answer to obtain the secondary verification result, and make corresponding cold data migration adjustments accordingly. The method of obtaining the system resource consumption parameters and analyzing the system resource consumption determination level is as follows: Obtaining system resource consumption parameters, including I / O resource utilization, memory page fault rate, file descriptor usage, and task queue length; Obtain the preset system resource consumption allowable set in the database, and perform proportional operation on it with the system resource consumption parameter to obtain the resource consumption proportional operation analysis result. Based on the resource consumption proportional operation analysis result, introduce the corresponding empowerment factor to perform quantitative coupling processing to obtain the system resource consumption value; The system resource consumption allowable set includes an I / O resource utilization allowable value, a memory page error rate allowable value, a file descriptor usage upper limit value, and a task queue length upper limit value; Obtain a preset system resource consumption threshold in the database and compare it with the system resource consumption value to obtain a system resource consumption determination level. If the system resource consumption value is below the system resource consumption threshold, the system resource consumption determination result is level one resource consumption; if the system resource consumption value is greater than the system resource consumption threshold, the system resource consumption determination result is level two resource consumption; The system resource consumption judgment level and the retrieval question and answer verification judgment result analysis are used to obtain the division dimension adjustment judgment result. The specific method is as follows: Based on the analysis of the system resource consumption judgment level and the search question and answer verification judgment result, if the search question and answer verification judgment result is verification failure, the partition dimension adjustment judgment result is to perform partition dimension adjustment. At the same time, if the system resource consumption judgment level is first-level resource consumption, the partition dimension adjustment judgment result is to perform partition dimension increase processing. If the system resource consumption judgment level is second-level resource consumption, the partition dimension adjustment judgment result is to perform partition dimension reduction processing. If the result of the retrieval question and answer verification is that the verification is qualified, the result of the division dimension adjustment is that the division dimension adjustment will not be performed.

2. The method for intelligent retrieval and question-answering of climate events based on multi-source data fusion according to claim 1, characterized in that: The specific method of storing each climate data information piece in a preset data storage location is as follows: Obtain each climate data information piece in the fused climate data set, and perform key feature extraction to obtain a key feature group for each climate data information piece; Obtaining a reference key feature set corresponding to each storage location in the system preset in the database; The key feature groups of each climate data information piece are respectively matched with the reference key feature sets corresponding to each storage location to obtain the reference key feature set corresponding to the maximum feature matching degree, and the storage location corresponding to the reference key feature set is obtained as the preset data storage location of the climate data information piece, thereby storing each climate data information piece in the preset data storage location.

3. The climate event intelligent retrieval and question-answering method based on multi-source data fusion according to claim 1 is characterized by: The specific method for obtaining the search question and answer verification result is as follows: Obtain the preset partitioning dimension rules in the database; Divide each climate data information piece into dimensions according to a preset division dimension rule to obtain each division dimension, and mark each division dimension as an initial division dimension; Perform retrieval question and answer verification based on the initial division dimension, obtain the question and answer response effect parameters under the initial division dimension, and analyze to obtain the question and answer response effect value; Obtain the preset question and answer response effect threshold in the database and compare it with the question and answer response effect value to obtain the retrieval question and answer verification judgment result. If the question and answer response effect value is less than the question and answer response effect threshold, the retrieval question and answer verification judgment result is verification failure. If the question and answer response effect value is above the question and answer response effect threshold, the retrieval question and answer verification judgment result is verification qualification.

4. The method for intelligent retrieval and question-answering of climate events based on multi-source data fusion according to claim 3 is characterized by: The specific steps of obtaining the question-answer response effect value are as follows: Obtaining question-and-answer response effect parameters under the initial division dimensions within a preset time period, wherein the question-and-answer response effect parameters include the average question-and-answer response time, the correct rate of question-and-answer push, the system retrieval accuracy, and the maximum fluctuation ratio of the question-and-answer response time; Obtain the preset question-answer response baseline set in the database, and perform proportional calculations on it with the question-answer response effect parameters to obtain the response ratio analysis results. Based on the response ratio analysis results, introduce the corresponding weighting factors for quantitative coupling processing to obtain the question-answer response effect value; The question-and-answer response baseline set includes a baseline value for the average duration of question-and-answer responses, a baseline value for the accuracy of question-and-answer push, an ideal value for system retrieval accuracy, and an allowable fluctuation ratio for question-and-answer response duration.

5. The climate event intelligent retrieval and question-answering method based on multi-source data fusion according to claim 1 is characterized by: The specific method for performing the dimension division and dimension increase processing is as follows: Obtain the data volume of each partition dimension under the initial partition dimension, perform dimension-increasing processing on the partition dimension corresponding to the maximum data volume, obtain the system resource consumption value after the dimension-increasing processing, and mark it as the system resource consumption dimension-increasing update value; Perform a difference ratio analysis on the system resource consumption incremental update value and the system resource consumption threshold to obtain the incremental update ratio; Obtain the preset dimension increase update ratio threshold in the database; Based on the analysis of the dimension increase update ratio, dimension increase update ratio threshold, system resource consumption dimension increase update value and system resource consumption threshold, the dimension increase processing result is obtained; If the system resource consumption dimensionality increase update value is less than the system resource consumption threshold and the dimensionality increase update ratio is less than the dimensionality increase update ratio threshold, the dimensionality increase processing result is unqualified, and the dimension increase processing of the divided dimension is continued until the system resource consumption dimensionality increase update value is above the system resource consumption threshold or the dimensionality increase update ratio is above the dimensionality increase update ratio threshold, and the dimension increase processing of the divided dimension is completed; If the system resource consumption dimensionality increase update value is above the system resource consumption threshold or the dimensionality increase update ratio is above the dimensionality increase update ratio threshold, the dimensionality increase processing result is qualified, and the dimension increase processing of the divided dimension is completed directly.

6. The climate event intelligent retrieval and question-answering method based on multi-source data fusion according to claim 1, characterized in that: The specific method of performing the dimension division and dimensionality reduction processing is as follows: Obtaining contribution parameters of each division dimension under the initial division dimension, wherein the contribution parameters include retrieval hit rate, number of retrievable data instances, memory usage, and user adoption hit rate; Obtaining a contribution baseline set preset in a database, wherein the contribution baseline set includes a retrieval hit rate baseline value, a retrieval data instance number baseline value, a memory usage allowance value, and a user adoption hit rate baseline value; Based on the retrieval hit rate, memory usage, and user adoption hit rate, a proportional operation is performed with the retrieval hit rate baseline value, the memory usage allowable value, and the user adoption hit rate baseline value to obtain the contribution ratio operation result. Based on the difference degree analysis between the number of retrievable data instances and the baseline value of the number of retrievable data instances, a contribution difference analysis result is obtained. Based on the contribution ratio operation result and the contribution difference analysis result, corresponding weighting factors are introduced to perform quantitative coupling processing to obtain the contribution value of each division dimension. Obtain a contribution threshold preset in the database and compare it with the contribution value of each division dimension. If the contribution value of a division dimension is less than the contribution threshold, mark the division dimension as a division dimension to be adjusted, thereby obtaining the division dimensions to be adjusted; Merge each partition dimension to be adjusted into the corresponding target partition dimension to complete the partition dimension reduction process; Obtain the question-answer response effect value after the dimension reduction process, and determine whether to perform repeated dimension reduction processing; The specific steps of merging each to-be-adjusted partition dimension into the corresponding target partition dimension are as follows: Obtain each partition dimension of the same dimension type as the partition dimension to be adjusted in the database, and mark it as each reference target partition dimension; Perform feature matching analysis on the to-be-adjusted partitioning dimension and each reference target partitioning dimension, and select the reference target partitioning dimension with the highest feature matching degree with the to-be-adjusted partitioning dimension as the target partitioning dimension of the to-be-adjusted partitioning dimension, thereby obtaining the target partitioning dimension corresponding to each to-be-adjusted partitioning dimension; Each partitioning dimension to be adjusted is incorporated into the corresponding target partitioning dimension.

7. The climate event intelligent retrieval and question-answering method based on multi-source data fusion according to claim 6, characterized in that: The method of obtaining the question-answer response effect value after the dimension reduction process and determining whether to perform repeated dimension reduction process is as follows: Obtain the question-answer response effect value after the dimension reduction processing of the divided dimension, and mark it as the question-answer response effect value after the dimension reduction processing; Based on the dimension reduction processing of the question and answer response effect value and the adjustment ratio analysis of the question and answer response effect value, the dimension reduction effect adjustment ratio is obtained; Obtain the preset dimensionality reduction effect adjustment ratio threshold in the database and analyze it with the dimensionality reduction effect adjustment ratio. If the dimensionality reduction effect adjustment ratio is below the dimensionality reduction effect adjustment ratio threshold, repeated dimensionality reduction processing will not be performed. If the dimensionality reduction effect adjustment ratio is greater than the dimensionality reduction effect adjustment ratio, repeated dimensionality reduction processing will continue to be performed.

8. The method for intelligent retrieval and question-answering of climate events based on multi-source data fusion according to claim 1, characterized in that: The corresponding cold data migration adjustment is performed as follows: After the dimension adjustment, perform secondary verification of the search question and answer to obtain the secondary verification result. If the secondary verification result of the search question and answer is unqualified, perform cold data migration adjustment. If the secondary verification result of the search question and answer is qualified, do not perform cold data migration adjustment. If cold data migration adjustment is performed, access parameters of each cold data information piece within a preset period are obtained, wherein the access parameters include the number of accesses, the number of adoptions, and the adoption ratio; Obtain the preset access benchmark set in the database, and perform proportional calculation with the access parameter to obtain the access difference analysis result. Based on the access difference analysis result, introduce the corresponding empowerment factor for quantitative coupling processing to obtain the access comprehensive adoption value of each cold data information item; Obtain the preset accessed comprehensive adoption threshold in the database and compare it with the accessed comprehensive adoption value of each cold data information piece. If the accessed comprehensive adoption value of a certain cold data information piece is above the accessed comprehensive adoption threshold, the cold data information piece will be subjected to cold data migration adjustment and migrated to the hot data storage area. Otherwise, the cold data migration adjustment will not be performed.

Citation Information

Patent Citations

  • A method and system for scenario-based multi-source data fusion analysis

    CN113609360B

  • A question-answering intelligent retrieval method and system for power documents

    CN117171333B

  • Multi-mode intelligent question and answer reasoning method and device based on knowledge distillation

    CN119623643A

  • Factory equipment operation data analysis and anomaly detection system based on machine learning

    CN119720054A