A data question and answer method based on structured and unstructured data fusion
Patent Information
- Application Number
- CN202611289675.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-25
- Publication Date
- 2026-09-25
AI Technical Summary
后端缺少规范的表关联匹配规则与负载管控机制,极易构建多余的表关联逻辑,带来不必要的数据库资源开销;这种浪费软硬件资源的数据问答业务,会降低用户的查询效率与使用体验
1、本发明通过多维度日志特征加权计算结合密度聚类构建高频结构化数据池,摒弃人工配置频次阈值筛选热点数据的方式。依照接口实际负载配置对应存储分区容量,结合时序折算、执行状态量化精准量化查询热度,自动剔除无效访问数据并回收系统资源,解决传统热点数据筛选失真、无效数据堆积、存储资源配置失衡的弊端,为解析用户表层查询诉求、识别底层数据表关联关系提供精准可靠的热点数据源支撑。
Smart Images

Figure CN122817528A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electronic digital data processing technology, specifically a data question-answering method based on the fusion of structured and unstructured data. Background Technology
[0002] Existing data-driven question answering solutions are prone to generating semantic illusions when processing unstructured natural language queries, resulting in poor matching accuracy between data and query requests. Most business users lack knowledge of the underlying database table structure; when users initiate queries based on business needs, they can only describe the required data content at the table level, often implying inherent relationships between multiple tables. The lack of standardized table join matching rules and load management mechanisms on the backend easily leads to the construction of redundant table join logic, resulting in unnecessary database resource overhead. This waste of hardware and software resources in data-driven question answering services reduces user query efficiency and user experience. Summary of the Invention
[0003] The purpose of this invention is to provide a data question-answering method based on the fusion of structured and unstructured data to solve the problems raised in the prior art.
[0004] To achieve the above objectives, the present invention provides the following technical solution: a data question answering method based on the fusion of structured and unstructured data, the data question answering method comprising the following steps:
[0005] Step S1: Set the parameters of the high-frequency structured data pool according to the load characteristics of the historical data Q&A interface, extract the business log features of historical data Q&A, extract and store them in the high-frequency structured data pool through density clustering combined with weighted fusion calculation; Collect all historical data and Q&A interface operation monitoring data, extract the interface peak load value and interface steady state load value recorded in the monitoring data, and match and delineate the fixed storage capacity of each physical storage partition of the high-frequency structured data pool according to the two types of measured load values. Retrieve all original business logs archived within the persistent storage medium, and disassemble each log into three types of fixed characteristic fields embedded within it. These three types of fixed characteristic fields are, in order, the real-time load occupancy value of the interface, the request call timing marker, and the instruction execution status identifier of the database SQL. The entire domain business logs are aggregated to form a complete log sample set. The total number of valid samples corresponding to the load occupancy dimension, call time sequence dimension, and database SQL execution dimension are counted separately. Based on the proportion of the total number of valid samples in the log sample set, the weights of load occupancy value, time decay coefficient, and SQL normal execution pass rate are fixedly configured. Read the difference between the call timing marker of a single business log and the current system timing marker, match and bind historical logs whose difference exceeds the preset archiving duration threshold with a predetermined decay coefficient, and complete the original feature field value conversion processing; business logs whose timing difference does not reach the archiving duration threshold retain the original feature values without conversion processing; Based on the converted feature values, multidimensional popularity feature vectors are generated for each single search term and single query paradigm. These multidimensional popularity feature vectors are then aggregated from all business elements to form the input sample set for the density clustering algorithm. The calculation formula for the multidimensional popularity feature vectors is as follows: ; In the formula, S represents the multidimensional heat feature vector; Represented as load occupancy value weight; This is represented as the weight value of the time decay coefficient; This is represented by the weight of the SQL execution success rate; This is represented as the real-time load occupancy value of the interface; It is represented as a time-series characteristic value after attenuation reduction; This represents the quantized value of the database SQL command execution status identifier. The database SQL command execution status identifier is subjected to binary quantization mapping: when the database SQL command is executed normally, the corresponding quantized value is assigned to 1; when the database SQL command encounters various abnormal situations such as errors, interruptions, timeouts, and execution failures, the corresponding quantized value is assigned to 0. The original log only retains the text status marker field. After the program reads the text marker, it converts it into a quantified value that can participate in numerical calculation according to the above-mentioned preset fixed mapping rules, and obtains the quantified value of the database SQL command execution status marker. The region partitioning process is completed by using the local density distribution boundary of the input sample set in the feature vector space. The partitioning result is further divided into a high-density sample cluster set and a discrete isolated sample point set. All business elements within the high-density sample cluster set are written into a high-frequency structured data pool for persistent storage. Access records, debugging and testing requests, and invalid query records corresponding to the discrete isolated sample points are removed. The deletion operation includes GC reclamation. During the density clustering execution phase, cluster boundary processing is completed based on the inherent density distribution of the samples, without inputting manually configured frequency judgment threshold parameters; The existing methods for collecting hot query data for data-driven question-and-answer services suffer from a fundamental flaw: relying on manually setting frequency thresholds to filter hot data. This leads to several related drawbacks: data pool storage partitions are configured with fixed specifications, which don't match the actual load conditions of the interfaces, easily resulting in wasted storage resources or restricted hot data writes; the evaluation of query popularity only considers the number of accesses, failing to consider interface load, statement execution success or failure, and log data timeliness for a comprehensive assessment. Historically expired log data cannot be naturally weighted, and the fact that the SQL execution status in the original logs is in text format and cannot be directly used in numerical calculations distorts the results of popularity quantification; manual filtering struggles to accurately distinguish between normal business access and noisy data such as instantaneous traffic, debugging requests, and invalid queries, leading to a continuous accumulation of useless data in the data pool, consuming hardware and software resources, and lacking a standardized data recycling and disposal mechanism. By measuring peak and steady-state loads through interfaces, the storage capacity of each partition in the data pool is determined, ensuring that the storage configuration closely matches the actual operating load of the system. Multi-dimensional feature weights are assigned based on the proportion of log samples, and the computational proportion of long-term logs is reduced using time-series conversion. At the same time, the text identifiers of SQL executions are standardized and mapped into quantified values. Multi-dimensional popularity feature vectors of each group of search terms and query paradigms are objectively calculated using weighted formulas. Density clustering is then used to automatically separate hot clusters from discrete noise data based on the density distribution of the samples themselves. The entire process eliminates the intervention of manual thresholds, and stable and effective business hot data is persistently stored in a high-frequency structured data pool. Various invalid noise data are removed with GC collection, which not only ensures the objectivity and authenticity of the hot business element screening results, but also cleans up redundant data in the system, reducing the resource consumption caused by invalid data to the subsequent database statement assembly process.
[0006] Step S2: Process the high-frequency structured data pool and unstructured data by combining business log features with semantic analysis, and map them to obtain the mapped unstructured data pool; Extract fixed semantic identifiers carried in business logs, select retrieval terms and query paradigms stored in the high-frequency structured data pool as semantic benchmark ontology, and perform word segmentation and entity semantic matching verification on the original unstructured question and answer text data of the entire domain. Based on the semantic matching and corresponding binding relationship, a matching structured baseline ontology index identifier is attached to each group of unstructured raw data; a unified storage structure orchestration process is performed on all unstructured data that has completed index attachment, and all unstructured data entities with attached index identifiers are collected through a predetermined storage organization structure to generate a mapped unstructured data pool. The existing structured data and unstructured question-and-answer data are stored independently. There is no fixed semantic binding relationship between the two types of data. Unstructured natural language queries cannot be directly matched with the standard query paradigm of structured data. The data cannot form a fixed mapping relationship, which makes it difficult to assemble database statements based on unstructured content. The overall data link cannot operate in a coherent manner. Based on the standard query content fixed in the high-frequency structured data pool as a semantic reference, the entity matching and index binding of unstructured data are completed by using the semantic identifiers of business logs. The unstructured data is organized and collected through a unified orchestration method, and a fixed mapping relationship is established between structured hot data and original unstructured query data. The standardized construction of the mapping unstructured data pool is completed, providing a fixed data mapping support relationship for the assembly process of backend database execution statements.
[0007] Step S3: Perform real-time data trend prediction analysis on the high-frequency structured data pool based on time series fitting and Markov state transition, and use the results of the prediction analysis to adjust and update the mapping of the unstructured data pool. The full multidimensional heat feature vectors stored in the high-frequency structured data pool are extracted to construct a continuous time-series sample sequence. Time-series fitting is then performed based on this sample sequence to divide the discrete operating state intervals corresponding to the heat values. The calculation formula for time-series fitting is as follows: ; In the formula, E represents the time series fitting bias; T represents the total number of time series nodes contained in the entire time series sample sequence. This represents the fitted heat value of the t-th node in the time series fitting operation output. The Markov state transition matrix is generated by statistically analyzing the state transition records of adjacent time-series nodes. The Markov state transition calculation formula is as follows: ; In the formula, This represents the state transition probability when the heat state jumps from the i-th interval to the j-th interval; This represents the statistical number of times the heat status jumps from i to j in all time series samples; This represents the total number of jumps from heat state i to all state intervals k; m represents the total number of all discrete running state intervals obtained from the heat division. Using the current time series real heat index to the state interval as the initial state, and combining the heat index change baseline value obtained by time series fitting and the jump probability distribution recorded in the state transition matrix, the most probable heat index state at the next moment is calculated, and the state information of continuous moments is summarized to generate the final heat index trend prediction. The Markov state transition matrix and the time series fitting output are input together into the prediction unit to complete the real-time trend prediction and analysis of the heat changes of business elements within the high-frequency structured data pool. Read the status change records generated by trend prediction analysis, and locate the matching unstructured data entities within the unstructured data pool according to the preset index binding relationship between structured and unstructured data. Based on the predicted state change instructions of structured data, perform targeted rewriting and incremental supplementation operations on the index association fields and storage layout of the corresponding data entities in the mapped unstructured data pool, and complete the synchronous control and update of the mapped unstructured data pool. The existing two sets of data pools rely solely on a single semantic binding to complete static association configuration. The popularity status of hot data within the high-frequency structured data pool will continuously change with business access behavior. The status changes of structured data cannot be synchronously transmitted to the mapped unstructured data pool. The mapping index relationship between the two types of data pools remains fixed in the long term, and the data evolution pattern at the time-series level cannot be captured and identified. The index identifier attached to the unstructured data has a fixed deviation from the real-time structured hot data, causing the index matching content to be inconsistent with the actual business hot status when assembling database statements based on the mapping relationship. By leveraging time-series fitting, the temporal patterns of structured heat data are extracted. Markov state transition matrices characterize the transition rules between different heat states. The results of two types of operations are used to perform real-time inference of the state trends of structured data. The state changes obtained from the inference are used to synchronously correct the data index and storage layout structure within the mapping unstructured data pool, maintain the real-time consistency of the mapping relationship between the high-frequency structured data pool and the mapping unstructured data pool, and ensure that the data association relationship between the two-level data pools changes synchronously with the business operation status.
[0008] Step S4: Based on the data element model, build a database execution statement assembly factory to preload database metadata into the unstructured data pool; Extract the structure, field constraints, field data types, and data table relationships of the entire business data table to construct a standardized data meta-model. Based on the operational behavior dimension of the data meta-model, four independent lower-level assembly plant entities are obtained: new operation factory, delete operation factory, modify operation factory, and query operation factory. The four types of lower-level assembly plants are uniformly assigned to the top-level scheduling architecture of the database execution statement assembly plant. Read the request input parameter fields, input parameter quantity, input parameter data type, and parameter combination arrangement form recorded in the full historical business log. According to the business operation behavior corresponding to each type of lower-level assembly plant, collect all parameter samples matching the corresponding behavior, and solidify the parameter receiving interface specifications, parameter verification rules, and parameter mapping rules within each type of assembly plant according to the samples. The four types of assembly plants with fixed parameter configurations are connected to the preset mapping execution nodes inside the data access layer. Based on the database execution statement samples stored in the historical business logs, the static preloading and storage of standard database statements are completed inside the mapping execution nodes. Receive real-time question-and-answer parsing data output from the mapped unstructured data pool. The real-time question-and-answer parsing data carries an operation identifier field and structured input parameter data. The top-level database executes statements, and the assembly plant reads the operation identifier field to complete the targeted matching and scheduling of the lower-level assembly plant, and passes the corresponding structured input parameter data to the lower-level assembly plant. The lower-level assembly plant calls the pre-loaded database statement template, binds the incoming structured input parameter data to the preset placeholder position inside the database statement, assembles and generates a standard execution statement that can be directly sent to the database execution unit, and pushes the statement to the underlying database through the mapping execution node to complete the data read and write operations. The existing database execution statements rely on real-time string concatenation during runtime to generate execution statements. The statement generation logic is deeply coupled with the business code. Database statements corresponding to historically high-frequency questions and answers cannot be pre-stored. Each question and answer request requires real-time parsing of fields and concatenation of statement structure. The parameter receiving logic of different business operations is mixed in the same processing unit. When the number or type of parameters changes, the underlying execution code needs to be modified synchronously. At the same time, the real-time concatenated statement structure has fixed statement format differences, which will generate repeated syntax parsing overhead when the statement is sent to the database. Based on the data meta-model, four independent CRUD-specific assembly plants are decomposed. By combining historical business logs, the interface specifications, parameter rules, and underlying Mapper statements of each plant are pre-configured and statically pre-loaded. The real-time question-and-answer process only performs plant matching, parameter input, and parameter binding operations to output usable database execution statements. The statement structure splicing logic and real-time business request process are separated, and the writing format of database statements and parameter input specifications are uniformly controlled. Based on the pre-deployed assembly architecture, the standardized conversion of unstructured queries into database execution statements is completed.
[0009] Step S5: Based on historical business logs, distinguish the load characteristics of structured and unstructured data Q&A. By analyzing the load characteristics when transforming data structures in historical business logs, obtain the baseline load characteristics and control the database access for real-time data Q&A. The system log parsing and monitoring node is deployed to collect and parse the original running records of all historical business logs, and extract the database running load parameters corresponding to structured data question and answer scenarios and unstructured data question and answer scenarios. The load parameters include the number of data table associations, statement execution time, resource consumption, cross-table call frequency and data structure conversion time. The parsed historical business log load parameters are classified and grouped according to scenarios, and a dedicated historical load sample set for structured data Q&A and a dedicated historical load sample set for unstructured data Q&A are constructed respectively. Traverse the two types of historical load sample sets, and analyze the load characteristics of each of the two cross-scenario switching processes, namely, the conversion from structured data to unstructured data and the conversion from unstructured data to structured data, to extract the load fluctuation parameters, resource change parameters, and table association scheduling parameters corresponding to the two types of data structure conversion processes. Collect all scene load parameters and cross-scene transition load parameters under all steady-state operating conditions, remove abnormal fluctuation sample data, and generate a benchmark load feature set corresponding to the data question and answer service. The benchmark load feature set includes feature data of single-scene steady-state load features and cross-scene transition load features. Read the updated mapped unstructured data pool, parse the data table association combination structure corresponding to each data entity in the mapped unstructured data pool, traverse the historical load operation records corresponding to different data table association combination methods, match the benchmark load feature set to complete the data table association load feature calibration, and construct the data table association load matching set. The real-time load characteristics of the current data Q&A business are collected in real time by the log parsing monitoring node. The real-time load characteristics include the number of real-time data table associations, real-time statement execution time, real-time resource consumption, and real-time data structure transformation time. The collected real-time load features are compared and matched one by one with the benchmark load feature set and the data table associated load matching set. The data table associated combination structure with the highest matching degree with the real-time load features and conforming to the benchmark load feature range is selected, and the data table associated combination structure is determined as the first execution table associated structure of the current data question and answer business. Using the load operating range and load fluctuation range defined by the benchmark load feature set as constraints, the database access execution process of real-time data question and answer is controlled according to the scheduling rules corresponding to the first execution table association structure. The database access execution process includes the data table association method, resource call level and statement execution sequence during the database access process. The existing database access execution process of the data question-and-answer service lacks a fixed load constraint mechanism. The load fluctuations generated during the conversion between structured and unstructured data have no historical data reference standard. Moreover, the real-time data table association and combination methods are randomly called, making it difficult to determine the pattern of load operation. The real-time load fluctuations and cross-structure data conversion load changes have no boundary constraints, which can easily lead to abnormal resource consumption and disordered multi-table calls, making it impossible to uniformly manage the stability of database access execution.
[0010] Furthermore, this invention collects and classifies the dual-scenario load characteristics and cross-scenario conversion load characteristics of historical business logs by logging monitoring nodes, constructs a standardized benchmark load characteristic set and a data table association load matching set, selects historical steady-state operation data as the basis for load constraints of real-time business, determines the compliant priority table association execution structure by matching and comparing real-time load characteristics with benchmark characteristics, and applies standardized load constraints and scheduling constraints to the database access process of real-time data question and answer, so that the load status of real-time database access and table association scheduling method fully conform to the historical steady-state operation specifications, and realizes the standardized and controllable execution of the data question and answer database access process; To avoid the problem of invalid computing power consumption caused by redundant table joins, reduce repeated query behavior caused by response distortion from the source, and break the bad operation loop of redundant calls, load increase and frequent retries.
[0011] Step S6: Parse the database metadata and execute the returned result dataset. After data processing, send the data back to the front end to complete the visualization rendering output. Receive the raw result dataset returned by the underlying database execution unit after executing the standard database execution statement, perform field parsing on the raw result dataset, and extract the business data fields, field return status, number of data returned rows, field mapping identifier, and statement execution receipt information from the result dataset; Based on the field mapping specifications and data encapsulation format preset by the data meta-model, the parsed raw result data is subjected to structured regularization processing to unify the field naming format, data storage format and field hierarchy structure of the entire data, and remove redundant fields from the underlying database receipts, system-reserved blank fields and redundant error messages contained in the result dataset. According to the preset data interaction protocol of the front-end visualization rendering, the structured result data after normalization is encapsulated and the fields are reorganized in a hierarchical manner, and the corresponding search terms, query paradigms, data access time sequence markers and database execution status markers are bound to generate a standardized front-end interactive dataset. The standardized front-end interactive dataset, once encapsulated, is pushed to the front-end visualization module via a preset data transmission interface or front-end access path. The front-end visualization module reads the field structure and data content of the standardized front-end interactive dataset, matches it with a preset visualization component template to complete the data mapping and rendering, and outputs a visual Q&A result interface.
[0012] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention constructs a high-frequency structured data pool by combining multi-dimensional log feature weighted calculation with density clustering, eliminating the need for manually configuring frequency thresholds to filter hot data. It configures corresponding storage partition capacity according to the actual interface load, and accurately quantifies query popularity by combining time-series conversion and execution status quantification. It automatically removes invalid access data and reclaims system resources, solving the drawbacks of traditional hot data filtering such as distortion, invalid data accumulation, and unbalanced storage resource allocation. This provides accurate and reliable hot data source support for analyzing user surface query requests and identifying underlying data table relationships.
[0013] 2. This invention is based on building a dedicated semantic mapping system for structured data and unstructured query data. It relies on a standardized data meta-model to split four types of statement assembly units, pre-stores high-frequency query statements and assembles execution statements using template filling, abandons the real-time string concatenation processing method at runtime, decouples statement generation logic from business code, unifies database statement format standards, reduces database syntax parsing overhead, accurately matches users' natural language requirements with data table query logic, and improves the accuracy of matching query content with users' actual needs.
[0014] 3. This invention uses a time-series fitting combined with Markov state deduction mechanism to dynamically maintain the mapping relationship between two-level data pools. Based on historical operating data, a standardized load management system is built, and the optimal data table association combination structure is adaptively matched. This solves the defects of the original data mapping being fixed, the inaccurate identification of implicit associations in data tables, and the inability to manage database access load. It eliminates resource consumption caused by redundant table associations, reduces system resource consumption, and improves query response efficiency and user experience. Attached Figure Description
[0015] Figure 1 This is a flowchart illustrating a data question-answering method based on the fusion of structured and unstructured data according to the present invention. Figure 2 This is a schematic diagram illustrating the process of building a benchmark based on historical load and controlling database access in a data question-answering method based on the fusion of structured and unstructured data according to the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] Example 1: As Figure 1As shown, this invention provides a technical solution: a data question-answering method based on the fusion of structured and unstructured data. The data question-answering method includes the following steps: Set the parameters of the high-frequency structured data pool based on the load characteristics of the historical data Q&A interface, extract the business log characteristics of historical data Q&A, process them, and store them in the high-frequency structured data pool; The fixed storage capacity of each physical storage partition of the high-frequency structured data pool is determined based on the peak load and steady-state load values of the historical data Q&A interface. Extract three types of fixed feature fields from the original business logs inside the persistent storage medium: real-time interface load usage, request call timing markers, and database SQL command execution status identifiers. Configure the load usage value weight, time decay coefficient weight, and SQL normal execution pass rate weight based on the effective sample ratio of each feature in the full set of log samples. Historical log feature values are converted by matching the decay coefficient based on the log time-series difference; a multi-dimensional popularity feature vector corresponding to the search terms and query paradigms is constructed based on the converted feature values, and the input sample set for density clustering is generated. The high-density sample cluster set and the discrete isolated sample point set are divided according to the local density boundary of the feature vector space, and the business elements corresponding to the high-density sample clusters are written into the high-frequency structured data pool. In practical implementation, taking online retail operation data analysis as an example, the system relies on the full volume of operational query logs to complete the unified capture and summary of three types of feature fields. The weight allocation is automatically generated according to the actual distribution ratio of various log samples. The time-series conversion mechanism automatically weakens the calculation weight of long-standing historical query records. Density clustering relies on the data's own aggregation characteristics to divide effective query behavior into test and invalid spam access data. The entire process eliminates subjective biases caused by manual threshold intervention. During the deployment phase, the system ensures that the log collection link fully covers all business query interfaces to prevent the overall offset of the heat vector calculation caused by the absence of local logs.
[0018] By combining business log features with semantic analysis, high-frequency structured data pools and unstructured data are processed and mapped to obtain a mapped unstructured data pool. Extract fixed semantic identifiers carried in business logs, select retrieval terms and query paradigms stored in the high-frequency structured data pool as semantic benchmark ontology, and perform word segmentation and entity semantic matching verification on the original unstructured question and answer text data of the entire domain. Based on the semantic matching and corresponding binding relationship, a matching structured baseline ontology index identifier is attached to each group of unstructured raw data; a unified storage structure orchestration process is performed on all unstructured data that has completed index attachment, and all unstructured data entities with attached index identifiers are collected through a predetermined storage organization structure to generate a mapped unstructured data pool. In practical implementation, taking online retail operation data analysis as an example, the text of sales, inventory, and user profile queries raised by operators in natural language is segmented and decomposed. Entities are aligned and bound with exclusive index identifiers based on the standardized query sentences within the high-frequency structured data pool. The standardized unstructured query data is centrally stored according to a unified architecture, thereby establishing a fixed semantic mapping link between the two data pools. During the implementation process, it is necessary to standardize the semantic annotation rules of business-specific terms to avoid mismatches of synonymous business words, which could damage the binding and correspondence between the two data layers.
[0019] Predict data trends in the high-frequency structured data pool and adjust and update the mapping of the unstructured data pool based on the prediction results; Extract the full multidimensional heat feature vectors stored in the high-frequency structured data pool to construct a continuous time series sample sequence. Based on the time series sample sequence, complete the time series fitting operation, divide the discrete running state intervals corresponding to the heat values, and generate the Markov state transition matrix by statistically analyzing the state jump records of adjacent time series nodes. The Markov state transition matrix and the time series fitting output are input together into the prediction unit to complete the real-time trend prediction and analysis of the heat changes of business elements within the high-frequency structured data pool. Read the status change records generated by trend prediction analysis, and locate the matching unstructured data entities within the unstructured data pool according to the preset index binding relationship between structured and unstructured data. Based on the predicted state change instructions of structured data, perform targeted rewriting and incremental supplementation operations on the index association fields and storage layout of the corresponding data entities in the mapped unstructured data pool, and complete the synchronous control and update of the mapped unstructured data pool. In practical implementation, taking online retail operation data analysis as an example, the popularity vector sequence of various retail business queries is organized according to daily time-series slices. The fitting operation is combined with the state transition matrix to predict the changing trend of popular analysis needs. After the popularity status fluctuates, the index information and storage arrangement of the corresponding query data in the unstructured data pool are modified synchronously to maintain the real-time synchronization of the mapping relationship between the two sets of data. During the operation phase, it is necessary to fix the unified division standard of time-series slices to prevent the statistical distortion of state transition probability caused by different slice scales, which would affect the accuracy of data synchronization updates.
[0020] A database execution statement assembly plant is built based on the data metadata model to preload database metadata into the unstructured data pool; Extract the structure, field constraints, field data types, and data table relationships of the entire business data table to construct a standardized data meta-model. Based on the operational behavior dimension of the data meta-model, four independent lower-level assembly plant entities are obtained: add operation factory, delete operation factory, modify operation factory, and query operation factory. The four types of lower-level assembly plants are uniformly assigned to the top-level scheduling architecture of the database execution statement assembly plant. Read the request input parameter fields, input parameter quantity, input parameter data type and parameter combination arrangement of the full historical business log records. Based on the business operation behavior corresponding to each type of lower-level assembly plant, collect all parameter samples matching the corresponding behavior, and solidify the parameter receiving interface specifications, parameter verification rules and parameter mapping rules of each type of assembly plant according to the samples. The four types of assembly plants with fixed parameter configurations are connected to the preset mapping execution nodes inside the data access layer. Based on the database execution statement samples stored in the historical business logs, the standard database statements are statically preloaded and stored inside the mapping execution nodes. Receive real-time question-and-answer parsing data output from the mapped unstructured data pool. The real-time question-and-answer parsing data carries an operation identifier field and structured input parameter data. The top-level database executes statements, and the assembly plant reads the operation identifier field to complete the targeted matching and scheduling of the lower-level assembly plant, and passes the corresponding structured input parameter data to the lower-level assembly plant. The lower-level assembly plant calls the pre-loaded database statement template, binds the incoming structured input parameter data to the preset placeholder position inside the database statement, assembles and generates a standard execution statement that can be directly sent to the database execution unit, and pushes the statement to the underlying database through the mapping execution node to complete the data read and write operations. In practical implementation, taking online retail operation data analysis as an example, a unified data meta-model is built by sorting out the structural relationships of multiple data tables of retail system products, orders, members, and logistics. Independent assembly units are divided according to the four types of data behaviors: adding, deleting, modifying, and querying. Templates are pre-stored and parameter verification rules are solidified based on historical retail query statements. After receiving operational queries and parsing parameters, the template is directly filled to generate executable statements. During the deployment phase, the operation logic of the four types of assembly plants needs to be isolated to avoid mutual interference of parameter verification rules of different operation types, which may cause format errors in the assembly of database statements.
[0021] like Figure 2 As shown, the load characteristics of structured and unstructured data Q&A in historical business logs are obtained and analyzed to obtain baseline load characteristics that control the execution of database metadata for real-time data Q&A. The system log parsing and monitoring node is deployed to collect and parse the original running records of all historical business logs, and extract the database running load parameters corresponding to structured data question and answer scenarios and unstructured data question and answer scenarios. The load parameters include the number of data table associations, statement execution time, resource consumption, cross-table call frequency and data structure conversion time. The parsed historical business log load parameters are classified and grouped according to scenarios, and a dedicated historical load sample set for structured data Q&A and a dedicated historical load sample set for unstructured data Q&A are constructed respectively. Traverse the two types of historical load sample sets, and analyze the load characteristics of each of the two cross-scenario switching processes, namely, the conversion from structured data to unstructured data and the conversion from unstructured data to structured data, to extract the load fluctuation parameters, resource change parameters, and table association scheduling parameters corresponding to the two types of data structure conversion processes. Collect all scene load parameters and cross-scene transition load parameters under all steady-state operating conditions, remove abnormal fluctuation sample data, and generate a benchmark load feature set corresponding to the data question and answer service. The benchmark load feature set includes feature data of single-scene steady-state load features and cross-scene transition load features. Read the updated mapped unstructured data pool, parse the data table association combination structure corresponding to each data entity in the mapped unstructured data pool, traverse the historical load operation records corresponding to different data table association combination methods, match the benchmark load feature set to complete the data table association load feature calibration, and construct the data table association load matching set. The real-time load characteristics of the current data Q&A business are collected in real time by the log parsing monitoring node. The real-time load characteristics include the number of real-time data table associations, real-time statement execution time, real-time resource consumption, and real-time data structure transformation time. The collected real-time load features are compared and matched one by one with the benchmark load feature set and the data table associated load matching set. The data table associated combination structure with the highest matching degree with the real-time load features and conforming to the benchmark load feature range is selected, and the data table associated combination structure is determined as the first execution table associated structure of the current data question and answer business. Using the load operating range and load fluctuation range defined by the benchmark load feature set as constraints, the database access execution process of real-time data question and answer is controlled according to the scheduling rules corresponding to the first execution table association structure. The database access execution process includes the data table association method, resource call level and statement execution sequence during the database access process. In practical implementation, taking online retail operation data analysis as an example, the monitoring node continuously collects various load indicators of the two types of question-and-answer modes in the retail business and the entire process of data format conversion. After removing abnormal data caused by the instantaneous impact of business peaks, a standard benchmark load system is compiled. The multi-table association combination scheme is matched and adapted to the real-time operating load indicators of the system. The compliant table association logic is locked based on the load matching results, and the overall database access behavior is constrained. During implementation, the benchmark load sample library needs to be continuously updated, and the load range standard is revised according to the retail business volume iteration to prevent the old benchmark data from being unable to adapt to the current system operation status.
[0022] The result dataset returned by parsing database metadata is processed and then sent back to the front end for visualization rendering output. The system receives the raw result dataset from the database execution unit, extracts valid information such as business fields, execution status, number of data rows, and execution receipts from the dataset, and removes redundant fields, blank placeholder data, and redundant error messages from the underlying system. It then unifies and standardizes the data according to the field format and hierarchical structure specified by the data meta-model, and encapsulates it into standardized interactive data by combining query terms, access sequence, and execution status markers. This standardized interactive data is then sent to the front-end visualization module via a predetermined transmission interface, and the front-end calls the corresponding visualization component template to complete the data mapping and rendering. In practical implementation, taking online retail operation data analysis as an example, the system receives raw data on product sales, user consumption, and order flow from the database, removes invalid content such as system receipts and null fields returned from the underlying database, organizes the data format according to unified standards, binds corresponding query identifiers and execution records to generate standard interactive data packages, and pushes the data to the front-end retail data dashboard to complete the automatic rendering and display of line charts, rankings, and summary reports.
[0023] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, it is intended that all variations falling within the meaning and scope of equivalents of the claims be included within the present invention.
Claims
1. A data question-answering method based on the fusion of structured and unstructured data, characterized in that: The data question-answering method includes: Set the parameters of the high-frequency structured data pool based on the load characteristics of the historical data Q&A interface, extract the business log characteristics of historical data Q&A, process them, and store them in the high-frequency structured data pool; By combining business log features with semantic analysis, high-frequency structured data pools and unstructured data are processed and mapped to obtain a mapped unstructured data pool. Predict data trends in the high-frequency structured data pool and adjust and update the mapping of the unstructured data pool based on the prediction results; A database execution statement assembly plant is built based on the data metadata model to preload database metadata into the unstructured data pool; By analyzing the load characteristics of structured and unstructured data questions and answers obtained from historical business logs, a baseline load characteristic is obtained to control the execution of database metadata for real-time data questions and answers. The database metadata is parsed, and the returned dataset is processed and sent back to the front end for visualization rendering output.
2. The data question-answering method based on the fusion of structured and unstructured data according to claim 1, characterized in that: The process of extracting historical data question-and-answer business log features, processing them, and storing them in a high-frequency structured data pool includes: The fixed storage capacity of each physical storage partition of the high-frequency structured data pool is determined based on the peak load and steady-state load values of the historical data Q&A interface. Extract three types of fixed feature fields from the original business logs inside the persistent storage medium: real-time interface load usage, request call timing markers, and database SQL command execution status identifiers. Configure the load usage value weight, time decay coefficient weight, and SQL normal execution pass rate weight based on the effective sample ratio of each feature in the full set of log samples. Historical log feature values are converted by matching the decay coefficient based on the log time-series difference; a multi-dimensional popularity feature vector corresponding to the search terms and query paradigms is constructed based on the converted feature values, and the input sample set for density clustering is generated. The high-density sample cluster set and the discrete isolated sample point set are divided according to the local density boundary of the feature vector space, and the business elements corresponding to the high-density sample clusters are written into the high-frequency structured data pool.
3. The data question answering method based on the fusion of structured and unstructured data according to claim 1, characterized in that: The process of combining business log features with semantic analysis to process high-frequency structured data pools and unstructured data, and mapping them to obtain a mapped unstructured data pool, includes: Extract fixed semantic identifiers carried in business logs, select retrieval terms and query paradigms stored in the high-frequency structured data pool as semantic benchmark ontology, and perform word segmentation and entity semantic matching verification on the original unstructured question and answer text data of the entire domain. Based on the semantic matching and corresponding binding relationship, a matching structured baseline ontology index identifier is attached to each group of unstructured raw data; a unified storage structure orchestration process is performed on all unstructured data that has completed index attachment, and all unstructured data entities with attached index identifiers are collected through a predetermined storage organization structure to generate a mapped unstructured data pool.
4. The data question answering method based on the fusion of structured and unstructured data according to claim 1, characterized in that: The process of adjusting and updating the unstructured data pool based on the prediction results includes: Extract the full multidimensional heat feature vectors stored in the high-frequency structured data pool to construct a continuous time series sample sequence. Based on the time series sample sequence, complete the time series fitting operation, divide the discrete running state intervals corresponding to the heat values, and generate the Markov state transition matrix by statistically analyzing the state jump records of adjacent time series nodes. The Markov state transition matrix and the time series fitting output are input together into the prediction unit to complete the real-time trend prediction and analysis of the heat changes of business elements within the high-frequency structured data pool.
5. A data question-answering method based on the fusion of structured and unstructured data according to claim 4, characterized in that: The method of adjusting and updating the unstructured data pool based on the prediction results also includes: Read the status change records generated by trend prediction analysis, and locate the matching unstructured data entities within the unstructured data pool according to the preset index binding relationship between structured and unstructured data. Based on the predicted state change instructions of the structured data, targeted rewriting and incremental supplementation operations are performed on the index association fields and storage layout positions of the corresponding data entities in the mapped unstructured data pool to complete the synchronous control and update of the mapped unstructured data pool.
6. The data question answering method based on the fusion of structured and unstructured data according to claim 1, characterized in that: The database execution statement assembly factory built based on the data metadata model preloads database metadata into the unstructured data pool, including: Extract the structure, field constraints, field data types, and data table relationships of the entire business data table to construct a standardized data meta-model. Based on the operational behavior dimension of the data meta-model, four independent lower-level assembly plant entities are obtained: add operation factory, delete operation factory, modify operation factory, and query operation factory. The four types of lower-level assembly plants are uniformly assigned to the top-level scheduling architecture of the database execution statement assembly plant. Read the request input parameter fields, input parameter quantity, input parameter data type, and parameter combination arrangement form recorded in the full historical business log. Based on the business operation behavior corresponding to each type of lower-level assembly plant, collect all parameter samples matching the corresponding behavior, and solidify the parameter receiving interface specifications, parameter verification rules, and parameter mapping rules within each type of assembly plant according to the samples.
7. A data question-answering method based on the fusion of structured and unstructured data as described in claim 6, characterized in that: The database execution statement assembly factory built based on the data metadata model preloads database metadata from the unstructured data pool, and also includes: The four types of assembly plants with fixed parameter configurations are connected to the preset mapping execution nodes inside the data access layer. Based on the database execution statement samples stored in the historical business logs, the standard database statements are statically preloaded and stored inside the mapping execution nodes. Receive real-time question-and-answer parsing data output from the mapped unstructured data pool. The real-time question-and-answer parsing data carries an operation identifier field and structured input parameter data. The top-level database executes statements, and the assembly plant reads the operation identifier field to complete the targeted matching and scheduling of the lower-level assembly plant, and passes the corresponding structured input parameter data to the lower-level assembly plant. The lower-level assembly plant calls the pre-loaded database statement template, binds the incoming structured input parameter data to the preset placeholder position inside the database statement, assembles and generates a standard execution statement that can be directly sent to the database execution unit, and pushes the statement to the underlying database through the mapped execution node to complete the data read and write operations.
8. The data question-answering method based on the fusion of structured and unstructured data according to claim 1, characterized in that: The execution of database metadata for real-time data question answering, based on obtained baseline load characteristics, includes: The system log parsing and monitoring node is deployed to collect and parse the original running records of all historical business logs, and extract the database running load parameters corresponding to structured data question and answer scenarios and unstructured data question and answer scenarios. The load parameters include the number of data table associations, statement execution time, resource consumption, cross-table call frequency and data structure conversion time. The parsed historical business log load parameters are classified and grouped according to scenarios, and a dedicated historical load sample set for structured data Q&A and a dedicated historical load sample set for unstructured data Q&A are constructed respectively.
9. A data question-answering method based on the fusion of structured and unstructured data according to claim 8, characterized in that: The process of obtaining baseline load characteristics to control real-time data question-and-answer database metadata execution also includes: Traverse the two types of historical load sample sets, and analyze the load characteristics of each of the two cross-scenario switching processes, namely, the conversion from structured data to unstructured data and the conversion from unstructured data to structured data, to extract the load fluctuation parameters, resource change parameters, and table association scheduling parameters corresponding to the two types of data structure conversion processes. Collect all scene load parameters and cross-scene transition load parameters under all steady-state operating conditions, remove abnormal fluctuation sample data, and generate a benchmark load feature set corresponding to the data question and answer service. The benchmark load feature set includes feature data of single-scene steady-state load features and cross-scene transition load features.
10. A data question-answering method based on the fusion of structured and unstructured data according to claim 9, characterized in that: The process of obtaining baseline load characteristics to control real-time data question-and-answer database metadata execution also includes: Read the updated mapped unstructured data pool, parse the data table association combination structure corresponding to each data entity in the mapped unstructured data pool, traverse the historical load operation records corresponding to different data table association combination methods, match the benchmark load feature set to complete the data table association load feature calibration, and construct the data table association load matching set. The real-time load characteristics of the current data Q&A business are collected in real time by the log parsing monitoring node. The real-time load characteristics include the number of real-time data table associations, real-time statement execution time, real-time resource consumption, and real-time data structure transformation time. The collected real-time load features are compared and matched one by one with the benchmark load feature set and the data table associated load matching set. The data table associated combination structure with the highest matching degree with the real-time load features and conforming to the benchmark load feature range is selected, and the data table associated combination structure is determined as the first execution table associated structure of the current data question and answer business. Using the load operating range and load fluctuation range defined by the baseline load characteristic set as constraints, and according to the scheduling rules corresponding to the first execution table association structure, the database access execution process of real-time data question and answer is controlled. The database access execution process includes the data table association method, resource call level and statement execution sequence during the database access process.