Artificial intelligence digital science and technology platform based on big data analysis
By modeling data uncertainty and building probability indexing on the artificial intelligence digital technology platform for big data analysis, the problem of data item uncertainty relationship and data partitioning in massive data resource management and retrieval is solved, and efficient and accurate data retrieval and query result sorting is achieved.
Patent Information
- Application Number
- CN202510529433.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the management and retrieval of massive data resources, the data items and their corresponding uncertainty characteristics are not clearly defined, resulting in a lack of accurate judgment of the confidence level of data items in the data retrieval process, which is prone to deviations in the search result, and the construction of probability indexes is not divided into data partitions, resulting in large-scale redundant data traversals often faced in the search process of probability data, affecting the efficiency of data retrieval.
An artificial intelligence digital technology platform based on big data analysis was designed, including data uncertainty modeling module, probability index building module, semantic query conversion module and search result sorting module. By establishing uncertain parameter structures for data items, structured probabilistic data records are generated, and data partitioning is carried out according to the preset spatial and numerical ranges, a multi-dimensional probability index structure is constructed, natural language query is optimized and probability query conditions are parsed to improve retrieval efficiency.
By closely binding data items with their uncertain parameters, structured probabilistic data records are generated to ensure the accuracy and reliability of data queries, reduce the search space and filtering difficulty of probabilistic data retrieval, improve the pertinence and response speed of data retrieval, and improve the accuracy of query and user interaction experience.
Smart Images

Figure CN120067166A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information retrieval, and in particular, to an artificial intelligence digital technology platform based on big data analysis. Background Art
[0002] An artificial intelligence digital technology platform based on big data analysis is a data management and retrieval platform constructed based on big data analysis methods and artificial intelligence technologies, mainly used for effective uncertainty modeling of massive data resources, construction of probabilistic indexes, semantic query conversion, and sorting of retrieval results.
[0003] When the prior art manages and retrieves massive data resources, the association relationship between data items and their corresponding uncertainty characteristics is not clearly defined, and data items are usually stored independently, which leads to a lack of accurate judgment of the confidence level of data items during the data retrieval process and is prone to retrieval result deviation; in addition, when constructing a probabilistic index, data partitions are not divided, resulting in frequent traversal of a large amount of redundant data during the probabilistic data retrieval process, affecting the data retrieval efficiency. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of the present invention is to solve the deficiencies existing in the prior art, and to propose an artificial intelligence digital technology platform based on big data analysis.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: An artificial intelligence digital technology platform based on big data analysis includes:
[0006] A data uncertainty modeling module, which receives a big data stream, determines the confidence interval boundary for each data item to represent the uncertainty of the data item, generates a set of data item uncertainty parameters, combines each data item with the corresponding set of data item uncertainty parameters and stores them to form a structured probabilistic data record;
[0007] A probabilistic index construction module, based on the structured probabilistic data record, partitions the data according to a preset spatial range and numerical range to form probabilistic attribute groups, and establishes index entries pointing to the storage location of the structured probabilistic data record according to the probabilistic attribute groups and entity link probabilities, and constructs a multi-dimensional probabilistic index structure;
[0008] A semantic query conversion module, which receives a natural language query, identifies the concepts and constraint conditions in the query through named entity recognition and relationship extraction, maps the concepts to domain ontology elements to form an ontology mapping of the query concepts, and combines the ontology mappings of the query concepts using the logical relationships defined by the domain ontology to convert and generate a logical query representation to obtain an optimized logical query plan;
[0009] A retrieval result sorting module, based on the optimized logical query plan, parses probabilistic query conditions, and uses the multi-dimensional probabilistic index structure for data filtering to obtain a candidate probabilistic record set.
[0010] Preferably, the step of obtaining the data item uncertainty parameter set is as follows:
[0011] Receive a single data item carried by each record in the large data stream, extract the acquisition time interval, measurement value jump amplitude, sensor accuracy, standard measurement fluctuation range, upper limit of sensor sampling error, and record span of the data item, and generate a data item measurement configuration group;
[0012] According to the data item measurement configuration group, calculate the confidence interval boundary estimate value;
[0013] Combine the confidence interval boundary estimate value with the measurement value jump amplitude, standard fluctuation range, and error upper limit to form an uncertainty parameter unit, generate a corresponding uncertainty parameter unit for each data item, and generate a data item uncertainty parameter set.
[0014] Preferably, the step of obtaining the structured probabilistic data record is as follows:
[0015] Bind the data item uncertainty parameter set to the corresponding data item, and create a unique identification index key for each data item to generate a data item combination structure with an identification index;
[0016] Based on the data item combination structure with an identification index, perform data serialization and compression, and write the compressed data item combination structure into the storage space one by one to form a structured probabilistic data record.
[0017] Preferably, the step of obtaining the probabilistic attribute grouping is as follows:
[0018] Extract the spatial coordinate points, data measurement values, confidence interval boundary estimate values, record time points, data sampling accuracy levels, and numerical fluctuation intensities of each record in the structured probabilistic data record to generate a partition reference attribute group;
[0019] According to the partition reference attribute group, calculate the partition adaptation intensity value;
[0020] Based on the partition adaptation intensity value, perform parallel clustering distribution judgment and critical similarity screening on all structured probabilistic data records to generate probabilistic attribute grouping.
[0021] Preferably, the step of obtaining the multi-dimensional probabilistic index structure is as follows:
[0022] Extract the structured probability data records in the probability attribute group, obtain the spatial distribution range, entity matching frequency, record data density, number of linked entities within the most recent week, average confidence interval boundary value, and entity hit rate, and form a link positioning basic information group;
[0023] Based on the link positioning basic information group, calculate the entity link confidence factor;
[0024] Based on the entity link confidence factor, construct a mapping structure with entity identifier, data record storage path, and grouped positioning hash index as combined items, and write the mapping structure into the index control table of the structured probability data record to generate a multi-dimensional probability index structure.
[0025] Preferably, the steps for obtaining the ontology mapping of the query concept are as follows:
[0026] Receive the natural language query content input by the user, extract the entity name, restricted range, time condition, and semantic association features in the natural language query content, and form a natural language query feature set;
[0027] Based on the natural language query feature set, through named entity recognition and relationship extraction logic, determine the association relationship and semantic roles between entities, and generate an entity relationship recognition structure with determined semantic role tags;
[0028] Based on the entity relationship recognition structure with determined semantic role tags, call the predefined domain ontology element set, perform entity concept matching according to the semantic roles, obtain the mapping relationship from entities to ontology elements, and form the ontology mapping of the query concept.
[0029] Preferably, the steps for obtaining the optimized logical query plan are as follows:
[0030] Receive the ontology mapping of the query concept, read the entity type, attribute name, constraint relationship, and class hierarchy path of each ontology mapping, call the predefined logical connection structure and attribute inheritance rules in the domain ontology, perform node connection, path combination, and constraint merging on the ontology mapping, and generate a logical query representation;
[0031] Based on the logical query representation, calculate the path selection priority score;
[0032] Based on the path selection priority score, select the one with the lowest path selection priority score from all candidate paths as the main query path, and call the field order, connection order, and execution priority rules of the main query path to combine and obtain the optimized logical query plan.
[0033] Preferably, the steps for obtaining the candidate probability record set are as follows:
[0034] Based on the optimized logical query plan, call the entity identifier, data record storage path, and grouped positioning hash index in the multi-dimensional probability index structure to perform combined matching and record path filtering, and generate a set of probability data records after preliminary filtering;
[0035] Based on the set of probability data records after preliminary filtering, compare the probability query conditions one by one, eliminate the data records that do not meet the probability query conditions, and retain all the data records that meet the probability query conditions to form a candidate probability record set.
[0036] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0037] In the present invention, by separately establishing an uncertainty parameter structure for each data item in the data stream, the data item is tightly bound to the corresponding uncertainty parameter to generate structured probability data records, ensuring the accuracy and reliability during the query and call process of the data, and avoiding the retrieval error caused by the isolated storage of data items; through the probability attribute grouping within the preset space and numerical range and the establishment of a multi-dimensional probability index structure, the search space and filtering difficulty of probability data are reduced, the pertinence and response speed of data retrieval are improved, and the redundant traversal caused by the lack of data partitioning during traditional probability data retrieval is solved; through the ontology element mapping and logical combination of natural language queries, the query path selection is optimized, enabling the logical query process to dynamically adapt to the user's query requirements and data structure changes, improving the accuracy of the query and the user's interaction experience, and avoiding the problem that traditional static query plans are difficult to adapt to changing requirements; by parsing the probability query conditions through the optimized logical query plan, the data records are secondarily matched and filtered to obtain a candidate probability record set, ensuring the accuracy and relevance of the probability retrieval results, reducing the frequency of invalid records, and improving the value of data calls. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is the platform flow chart of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0039] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0040] Please refer to Figure 1 , the present invention provides a technical solution: an artificial intelligence digital technology platform based on big data analysis includes:
[0041] A data uncertainty modeling module receives a large data stream, determines confidence interval boundaries for each data item to represent the uncertainty of the data item, generates a set of data item uncertainty parameters, combines each data item with the corresponding set of data item uncertainty parameters and stores them to form a structured probabilistic data record;
[0042] A probabilistic index construction module partitions the data based on a preset spatial range and numerical range based on the structured probabilistic data record to form probabilistic attribute groups, and establishes index entries pointing to the storage locations of the structured probabilistic data records according to the probabilistic attribute groups and entity link probabilities, constructing a multi-dimensional probabilistic index structure;
[0043] A semantic query transformation module receives a natural language query, identifies the concepts and constraints in the query through named entity recognition and relation extraction, maps the concepts to domain ontology elements to form an ontology mapping of the query concepts, and combines the ontology mappings of the query concepts using the logical relationships defined in the domain ontology to transform and generate a logical query representation, obtaining an optimized logical query plan;
[0044] A retrieval result sorting module parses the probabilistic query conditions based on the optimized logical query plan, filters the data using the multi-dimensional probabilistic index structure, and obtains a set of candidate probabilistic records.
[0045] The steps for obtaining the set of data item uncertainty parameters are as follows:
[0046] Receive a single data item carried by each record in the large data stream, extract the acquisition time interval, measurement value jump amplitude, sensor accuracy, standard measurement fluctuation range, upper limit of sensor sampling error, and record span of the data item to generate a data item measurement configuration group;
[0047] Calculate the estimated value of the confidence interval boundary according to the data item measurement configuration group, and the calculation formula is:
[0048] ;
[0049] where, is the estimated value of the confidence interval boundary of the th data item, is the acquisition time interval between two consecutive acquisition points in the th data item, is the jump amplitude within the current observation period of the th data item, is the accuracy level of the sensor of the th data item, is the standard fluctuation range of the th data item in the last 30 days, is the The maximum error value recorded by a data item sensor, is the time span length of the g-th data item in consecutive records;
[0050] The estimated confidence interval boundary, the jump amplitude of the measured value, the standard fluctuation range, and the upper error limit are jointly composed of an uncertainty parameter unit. An uncertainty parameter unit corresponding to each data item is generated to generate a data item uncertainty parameter set.
[0051] Specifically, receive the content of each record from a continuous large data stream, read and parse the independent data items contained therein one by one, extract the acquisition timestamp for each data item in the order of arrival and record it as a time series, use the time difference between adjacent data items to determine the data acquisition time interval, and compare it with a pre-established acquisition time interval range table. For example, the lower limit is set to 1s and the upper limit is set to 10s. When the time interval is less than 1s or greater than 10s, it is marked as an invalid record and the elimination action is performed. At the same time, parse the measured value information in each record and calculate the jump amplitude of the measured value through the absolute difference between the current record and the previous record. Compare the jump amplitude with the sensor accuracy extracted from the sensor specification book in advance. If the jump amplitude is higher than the accuracy level recorded by the sensor, the record is specially marked during statistics for subsequent inspection. The upper limit of the sensor sampling error and the record span are determined by comparing the original log and the index number respectively. The upper limit of the sensor sampling error is taken from the sensor specification marked by the manufacturer, and the record span is obtained by calculating the difference in the corresponding consecutive positions of adjacent records in the data stream and compared with a preset threshold of 500. When the span exceeds 500, subsequent records are numbered as a new batch. Finally, the data acquisition time interval, the jump amplitude of the measured value, the sensor accuracy, the standard measurement fluctuation range, the upper limit of the sensor sampling error, and the record span are combined to form a data item measurement configuration group.
[0052] Formula: , The benefit of the formula is that by incorporating the data acquisition time interval, jump amplitude, sensor accuracy, standard fluctuation range, the maximum error recorded by the sensor, and the time span of consecutive records into the same calculation framework, it is possible to comprehensively measure the influence of multiple factors on the final confidence interval boundary in one operation, providing a more flexible and accurate probability judgment basis for subsequent data analysis under the premise of relatively controllable computational complexity, enabling the system to quantify and control uncertainty.
[0053] The acquisition step of is to mark the arrival timestamps between the g-th data item and the (g - 1)-th data item under the same sensor, take the difference and record it as , to obtain accurate values, it is necessary to read the data timestamps one by one from the original sensor logs within a week and calculate the differences. Referring to the actual sampling period and data recording time of the device, all the differences are unified in seconds. After that, the stopwatch record is compared with the sensor logs, and finally is corrected. For example, in an industrial control process, if the sampling timestamps are 12.4s and 15.6s in sequence, then is 3.2s.
[0054] The obtaining steps of are as follows: for the data items within the same observation period, the jump amplitude is measured by comparing the numerical difference between this data item and the previous data item. The specific method is to subtract the two measurement values and take the absolute value, and then compare the obtained result with a device identification threshold to evaluate the significance of the jump. This device identification threshold is obtained by statistically analyzing 300 hours of industrial operation records. During this process, the operation mechanism is refined and the time points with intense numerical fluctuations are counted to calculate the number of fluctuations per hour, and this number is compared with the overall measurement value level of the device, so as to refine a benchmark value for the significance of fluctuations. For example, the benchmark value is refined to be 2.0 through the distribution of the differences measured multiple times. If the current difference is higher than 2.0, it is considered that the jump is obvious; otherwise, the ordinary change record is maintained. When the actually measured jump amplitude reaches 5.5, it is recorded as .
[0055] The obtaining steps of are as follows: referring to the accuracy nominal value provided by the sensor manufacturer and combining the data stability detection results in the past month, the two are uniformly measured. First, read the output accuracy index of this device from the hardware specification book provided by the manufacturer, and then perform superposition processing with the accuracy deviation statistical value measured continuously on-site for 30 days. The superposition processing uses the following formula: , where and represent the weight coefficients of the specification accuracy and the on-site accuracy deviation respectively. and are the specification accuracy and the on-site accuracy deviation respectively. The setting basis of these two weight coefficients is to evaluate the influence degree of the specification annotation and the on-site measured accuracy on the final measurement result. When it is statistically found that the average on-site accuracy deviation is 0.015 and the specification nominal value is 0.01, if and , then .
[0056] The acquisition steps are as follows: extract all records of this data item from the normal operating condition data within the past 30 days, and perform maximum and minimum value induction on the sampling values of this record from 0:00 to 24:00 every day, take the difference and record it as the daily fluctuation range, accumulate it for 30 days to form a fluctuation range sequence, and then perform weighted average on this sequence to obtain the standard fluctuation range of this data item. In order to quantify the weighted average process, the following formula can be used: , where and respectively represent the maximum and minimum values of this data item on the d-th day, represents the proportion of the data record volume of this day in the total record volume of the entire 30 days. When conducting statistics on a certain industrial equipment, after calculation, .
[0057] The acquisition steps are as follows: when the sensor is continuously operating, extract the maximum value from the maximum reading error of each calibration in its actual use environment and record it. Establish this value as the maximum error value recorded by the sensor. For example, during the calibration stage, the sensor is compared and detected multiple times, and the obtained error values are 0.03, 0.05, 0.01, 0.02, etc. Take the maximum value of 0.05 as , and at the same time compare it with the error limit marked by the manufacturer again. If the maximum value does not exceed the manufacturer's limit, take this maximum value; otherwise, perform the sensor replacement operation. In this example, .
[0058] The acquisition steps are as follows: calculate the index difference for the position sequence of this data item in the large data stream from the first occurrence to the current record, and exclude discontinuous record paragraphs during statistics. If the pause caused by network or equipment maintenance exceeds the established span threshold, start re-numbering from a new batch. For example, a certain machine has 1440 theoretical records in a day. If the difference between the effective record indexes corresponding to the data item g on that day is 12, then .
[0059] Calculation process:
[0060] Let , , , , , , and substitute these parameters into the formula in turn:
[0061] ;
[0062] First calculate , then , and then square this value to get , and Sum them up to get approximately , divide it by 2 to get , take the square root after taking the absolute value to get .
[0063] This result indicates that when the acquisition time interval, jump amplitude, sensor accuracy, standard fluctuation range, maximum error, and time span are all the above values, the estimated value of the confidence interval boundary of this data item is approximately 0.05748. If the same operation is performed on other data items under the same conditions, a set of values representing the uncertainty interval ranges of different data items can be formed, providing a quantitative basis for subsequent comprehensive screening at the probability level.
[0064] According to the estimated value of the confidence interval boundary obtained in the previous step, combine it with the jump amplitude, standard fluctuation range, and upper limit of the sensor sampling error obtained previously. When targeting the same data item, sequentially read the estimated value of the confidence interval boundary and the jump amplitude of the measured value of this data item and compare them. When it is found that the jump amplitude exceeds the benchmark reference value collected in advance, mark the fluctuation degree of this data in the comparison table. Subsequently, limit the interval of this data under normal working conditions according to the record of the standard fluctuation range of 2.6. If it exceeds 2.6, it will be highlighted in the record table. Compare with the value of the upper limit of the sensor sampling error of 0.05. When it exceeds 0.05, it will prompt that there is a significant deviation. In this way, collect the estimated value of the confidence interval boundary of each data item together with its jump amplitude, standard fluctuation range, and error upper limit, and then further pair them according to the data item number. Finally, complete the summary of the uncertainty parameter units for all data items and establish a list including the corresponding parameter units for each data item. This list is marked as the data item uncertainty parameter set.
[0065] The steps to obtain structured probability data records are as follows:
[0066] Bind the data item uncertainty parameter set to the corresponding data item, create a unique identification index key for each data item, and generate a data item combination structure with identification index;
[0067] Based on the data item combination structure with identification index, perform data serialization and compression, and write the compressed data item combination structure into the storage space item by item to form structured probability data records.
[0068] Specifically, bind the data item uncertainty parameter set to the corresponding data item. First, screen out the corresponding uncertainty parameter groups from each data item obtained previously, and associate them one by one according to the unique serial number of the data item. Pair items such as the estimated value of the confidence interval boundary, standard fluctuation range, jump amplitude, and upper limit of sampling error included in the uncertainty parameter group with this data item. During the pairing process, an internal marker for retrieval needs to be set. For example, an index prefix is formed by combining letters and numbers, such as "ITEM" followed by an incrementing numerical value to generate an initial unique identifier. Then, according to specific requirements, each uncertainty parameter is further segmented and confirmed. If there is a situation where the jump amplitude of the measured value in the uncertainty parameter is higher than the threshold of 3, an additional label "_HIGH" is appended to its unique identifier for distinction. The threshold of 3 is obtained from the long-term monitoring of the increment of measured values in the industrial scenario. By statistically analyzing the difference sequence of measured values of 120-day data, its average difference and standard deviation are obtained, and a boundary that can significantly reflect jumps is determined. The above identification logic aims to make subsequent retrieval more convenient. After the binding is completed, an indexable mapping table is generated in the order of data items. The mapping table includes the unique identifier index key of each data item and the corresponding uncertainty parameter group. When organizing the mapping table, it is necessary to ensure that the unique identifier index keys of each data item do not repeat. For the repeated parts, the numbers are corrected and checked again. When all data items are bound, a data item combination structure with identification indexes is constructed.
[0069] Based on the data item combination structure with identification index, first read the unique identification index key of each data item and the bound uncertainty parameter, and arrange the timestamp, numerical content and uncertainty parameter of the data item record corresponding to each index key in sequence, and judge the continuity of adjacent records. If the index is discontinuous or the interval is too large, it is regarded as cross-segment data and classified before serialization. Here, an additional symbol "GAP" can be set and marked after the index, and then the sequential serialization method is used to recombine these records into a long string or binary data stream in ascending order of the index, combined with the compression mode confirmed in advance, such as the fixed-length format and dictionary mapping are often used in industrial big data processing to deal with repeated occurrences. The index part is simplified, and all index records are packed and compressed one by one in this way. If a mark with a measured value jump amplitude higher than 3 is encountered during the compression, an additional mark field is reserved for such records during compression to record possible boundary deviations. Finally, when writing to the storage space in the order of the mapping table, it is necessary to check whether the number of bytes written each time exceeds the reference threshold of 2000. The reference threshold of 2000 comes from the test of the current machine storage read and write performance. By gradually increasing the amount of written data and recording the read and write time, a more suitable batch size is found. When the continuously written data block exceeds 2000 bytes, the next batch is automatically split. After all the combined data are compressed, a structured probabilistic data record is formed in the storage space.
[0070] The steps to obtain the probability attribute grouping are:
[0071] Extract the spatial coordinate points, data measurement values, confidence interval boundary estimates, recording time points, data sampling accuracy levels and numerical fluctuation intensity of each record in the structured probability data records to generate a partition reference attribute group;
[0072] According to the partition reference attribute group, the partition adaptation strength value is calculated using the following formula:
[0073] ;
[0074] in, For the The partition adaptation strength value of the structured probabilistic data record, and For the The spatial coordinate points extracted from the structured probability data records, For the The structured probability data records the cumulative value fluctuation intensity in the last 30 days. For the The confidence interval boundary estimates of the data items bounded by the structured probabilistic data records, For the Sampling accuracy level of structured probability data records is the measurement value of the th structured probability data record is the spatial jump amplitude of the
[0075] Based on the partition adaptation intensity value, parallel clustering distribution judgment and critical similarity screening are performed on all structured probability data records to generate probability attribute groups.
[0076] Specifically, read all available entries from the previously obtained structured probability data records. For each record, sequentially extract the spatial coordinate points contained therein and compare them with the available coordinate ranges. For example, compare whether the coordinate values are within the range of abscissa 0 to 500 and ordinate 0 to 500. If there are coordinate values outside this range, add an identifier to the list for subsequent separate comparison of this part. Then read the data measurement value and compare it with the pre-prepared measurement reference range. For example, in a pressure measurement scenario, compare the range of 0 MPa to 2 MPa, or in a temperature measurement scenario, compare the range of 0 °C to 90 °C. Note the deviation information for measurement values outside the reference range and record it in a deviation index table. Subsequently, continue to read the confidence interval boundary estimate value corresponding to the current record and combine it with the data measurement value for determination. When it is found that the confidence interval boundary estimate value is too small or too large, a mark can be made in the comparison table. At the same time, judge the sampling stability of the record according to the data sampling accuracy level. By comparing the sampling accuracy level with the device factory parameters, record the compliance with the device specifications and list it in a statistical list. Then, combined with the numerical fluctuation intensity information, check its measurement fluctuation in the recent period of time and determine whether the fluctuation breaks through the pre-set amplitude threshold. This threshold is obtained based on the historical fluctuation statistics of the industrial site. For example, classify and count the amplitude data during the operation of mechanical equipment, find the maximum and minimum fluctuation values and their distribution patterns over a period of time, and then determine a boundary value that can reflect obvious fluctuations. If it is observed that the fluctuation intensity of some data items reaches or exceeds this boundary value, a special label is marked on the corresponding entry. Finally, integrate the spatial coordinate points, data measurement values, confidence interval boundary estimate values, recording time points, data sampling accuracy levels, and numerical fluctuation intensities into partition reference attribute groups.
[0077] Formula: , the benefit of the formula is that by comprehensively considering the logarithmic relationship between the sum of the squares of the spatial coordinates, the cumulative numerical fluctuation intensity and the estimated value of the confidence interval boundary, the sampling accuracy level, and the coupling effect between the measured value and the spatial jump amplitude, multiple dimensions are unified and integrated into the same computable partition adaptation intensity index. This index can reflect the influence of multiple information such as spatial distribution, fluctuation amplitude, and measurement accuracy in a single operation. When subsequent clustering or similarity determination is performed, more targeted partition division and screening can be carried out according to this value.
[0078] The acquisition step of is as follows. First, read the abscissa value in the u-th record from the established structured probability data record. When the device is tested, the original abscissa is .
[0079] The acquisition step of is to read the original measured value of the ordinate of the u-th record. For example, in a statistical measurement, the original ordinate is .
[0080] The acquisition step of is to perform a cumulative operation on the numerical fluctuation intensity using the data values corresponding to this record within the most recent 30 days. First, sum up the fluctuation differences for each day to obtain the daily fluctuation amount, and then aggregate the daily fluctuation amounts for 30 days. The overall cumulative numerical fluctuation intensity is obtained through the following calculation formula: , where represents the fluctuation amplitude of the measured value on the d-th day, represents the weighting coefficient obtained according to the daily data volume distribution. This weighting coefficient is obtained by statistically calculating the working hours and load of the on-site device. When the working hours are longer and the load is higher, the corresponding daily weight value is larger. For example, during a period of full-load operation, by recording the fluctuation amounts of each time period of the device, we can obtain . If the working hours in a day are 12 hours, then will increase accordingly. Substitute these values into the formula and calculate item by item to obtain .
[0081] The acquisition step of is to find the estimated value of the confidence interval boundary obtained in the previous process for this structured probability data record.
[0082] The acquisition steps are as follows: read the sampling accuracy level information corresponding to this structured probability data record. This information includes both the theoretical accuracy provided by the sensor manufacturer and the accuracy compensation obtained from on-site actual data comparison tests. During on-site testing, compare the outputs of multiple sensors under the same measurement conditions, and count the measurement differences between all sensors and the standard measuring tool to form an accuracy error distribution. Then, match this distribution with the manufacturer's theoretical accuracy to obtain a comprehensive sampling accuracy level. For example, in a comparison test, the reading of the standard measuring tool is 50.00, and the outputs of each sensor are between 49.96 and 50.10. After statistics, the average difference is 0.02, and the standard deviation is 0.03. Combine with the manufacturer's data to confirm that this error level corresponds to .
[0083] The acquisition steps are as follows: directly read from the measurement value field of the current record and perform type verification. This measurement value is written into the structured probability data record after being generated by the on-site sensor. If it is a numerical type, it can be directly taken out. If it is detected that there is a non-numerical format, data elimination will be performed and relevant marks will be retained. Common measurement values in industry include pressure, temperature, current, etc. During this process, different numerical valid interval judgments can be made for different types. For example, perform an interval comparison for temperature from 0°C to 90°C and for pressure from 0 MPa to 2 MPa. Retain the measurement values that meet the rules uniformly, and mark the abnormal ones in the record table. For example, if a record corresponds to a pressure measurement value of 1.2 MPa, then .
[0084] The acquisition steps are as follows: by analyzing the spatial coordinate changes involved in this record within the last 5 days, use the coordinate difference method to calculate the spatial jump amplitude. Perform difference operations on the coordinate changes for each day and accumulate the difference results to obtain the overall spatial jump amplitude within five days. In the specific operation, a minimum jump detection threshold can be set, which is determined by combining the positioning error statistics of the actual operation trajectory of the device. For example, if it is found that the average positioning error reaches about 0.1 during operation, then this value is used as the minimum jump detection threshold. If the sum of the coordinate differences within a day is greater than this threshold, it will be accumulated. After five days, summarize the sum of all differences that exceed the threshold. For example, after calculating a certain record, we get .
[0085] Calculation process:
[0086] Let , , , , , , , and substitute these values into the formula in turn:
[0087] ;
[0088] First, calculate , and then , take , then , the numerator part is . When calculating the denominator , then . At the same time , add it to to get . The denominator is . Divide the numerator 44.975 by the denominator 3.8538 to get approximately 11.66, and finally take the square root to get .
[0089] This result indicates that when the horizontal and vertical coordinates, cumulative value fluctuation intensity, confidence interval boundary estimate value, sampling precision level, measured value, and spatial jump amplitude of this record are all the above values, the calculated partition adaptation intensity is approximately 3.413, which is in a relatively high adaptation intensity range. When enough are statistically obtained, they can be compared in subsequent parallel clustering and similarity screening, and records with similar or the same magnitude can be divided into the same partition or the same attribute group.
[0090] Based on the obtained partition adaptation intensity value, read the corresponding to different records one by one and classify and label them through pre-established numerical segments. For example, for records between 0 and 2, label them as low adaptation, for records between 2 and 5, label them as medium adaptation, and for records greater than 5, label them as high adaptation. These segmentation thresholds are determined based on the distribution rules of the statistical device sampling precision level and the on-site measured fluctuation intensity. By analyzing the distribution of the past 200 records, combining the data mean and variance and sorting out the intervals that can distinguish different adaptation degrees, and then judging the parallel clustering distribution of all records according to these labels, concentrating the records in the same adaptation interval for processing. Identify the critical similarity within each cluster by comparing the differences in horizontal and vertical coordinates and measured values. When both the coordinate deviation and the measured value deviation do not exceed the pre-established difference threshold, it is determined that these records have a high similarity. This difference threshold can be set according to the on-site data scale and precision requirements. For example, obtain a set of coordinate difference means based on the past measurement sequence and take twice of it as the segmentation boundary, and at the same time make the maximum and minimum range judgments on the measured values. If both the coordinate and measured value differences meet the requirements, then aggregate these similar records again. After multiple rounds of parallel calculations, multiple relatively compact clustering distribution groups can be obtained, and finally the critical similarity screening is completed, thereby generating the probability attribute grouping.
[0091] The steps for obtaining the multi-dimensional probability index structure are as follows:
[0092] Extract the structured probability data records in the probability attribute groups, and obtain the spatial distribution range, entity matching frequency, record data density, number of linked entities in the most recent week, average confidence interval boundary value, and entity hit rate, to form a link positioning basic information group;
[0093] Based on the link positioning basic information group, calculate the entity link confidence factor, and the calculation formula is:
[0094] ;
[0095] where, is the link confidence factor between the th group of probability attribute groups and the entity link, is the distribution range of the th group of structured probability data records in the current space, is the historical entity matching frequency corresponding to each record in the th group, is the number of linked entities in the th group in the most recent 7 days, is the average data density of the structured probability data records within the th group, is the average confidence interval boundary value of the th group, is the total number of hit linked entities in the th group;
[0096] Based on the entity link confidence factor, construct a mapping structure with entity identifier, data record storage path, and group positioning hash index as joint items, and write the mapping structure into the index control table of the structured probability data records to generate a multi-dimensional probability index structure.
[0097] Specifically, extract the structured probability data records contained in the probability attribute groups that have completed partitioning. Read the corresponding spatial coordinate information item by item and view the spatial coordinate range. Compare the coordinate values with the pre-established reference coordinate intervals. For example, compare the abscissa from 0 to 500 and the ordinate from 0 to 500. If there are records exceeding this interval, mark them as coordinate deviation in the comparison table. Then read the entity matching frequency of the corresponding record. This frequency can be accumulated and statistically calculated by whether the record is associated with specific entities in the past month, resulting in a matching count represented in integer form. List the records with too high or too low matching counts separately and mark them additionally. Then combine the record data density with the number of linked entities in the past seven days, view the cumulative distribution of data entries in each time period, and compare the number of linked entities with a pre-determined threshold. This threshold is set by the analysis team summarizing the data scale and entity association characteristics when reviewing all historical records. In this example, the value is 20. When the number of linked entities in the record reaches or exceeds 20, it is considered that the link scale is large, and these records are separately registered in a sub-comparison table. Subsequently, calculate the average confidence interval boundary value of this group based on the existing records. This confidence interval boundary value is obtained in the previous stage when comprehensively evaluating the uncertainty of measurement data. Here, directly call its value and calculate the average of several records. Finally, read the entity hit rate corresponding to each record. This entity hit rate represents the proportion of the frequency of successful matches for a specific entity within a week. This value can be obtained by retrieving the entity association entries in the database in the past seven days and comparing them with the structured probability data records of the current group. When the entity hit rate exceeds the benchmark percentage of 30, mark it additionally. Summarize and encapsulate the above information, including the spatial distribution range, entity matching frequency, record data density, number of linked entities in the most recent week, average confidence interval boundary value, and entity hit rate, into a link positioning basic information group.
[0098] Formula: , the benefit of the formula is that it incorporates the distribution range, historical entity matching frequency, number of linked entities, average data density, average confidence interval boundary value, and total number of hit linked entities into the same operation framework. Through a single calculation, it can comprehensively measure the credibility of the record group in entity linking. When making a determination on entity association, this value can directly reflect the degree of integration between each group and the entity.
[0099] The acquisition steps for are as follows. First, read the The spatial coordinate range of a group of structured probability data records is obtained by statistically calculating the minimum and maximum values of the horizontal and vertical coordinates of all records and then calculating the difference between the two. At the equipment site, the coordinates of each record will be checked against the reference coordinate system, and the complete coordinate coverage interval will be obtained after extracting the coordinate differences. Then, the sum of the squares of the horizontal and vertical lengths of this interval will be calculated and the square root will be taken to obtain the final distribution range value. If it is found through on-site statistics that the difference between the minimum and maximum coordinates is 50 and 60 respectively, the distribution range can be calculated. For the convenience of subsequent fusion, one decimal place can be taken and recorded as 78.1.
[0100] The acquisition steps of [are as follows. For each record in the group, the historical entity matching situations are summarized and the number of matches or matching scores are obtained. Since the entity matching frequency needs to consider whether each record has ever been associated with a certain entity and the association strength, the following formula can be defined to quantify the historical matching situation of the records: , where is the total number of records in this group, represents the flag indicating whether the th record matches the entity. If it matches, it is 1; otherwise, it is 0. represents the matching strength. When a single match involves multiple similar entities, a higher value can be obtained by accumulation. The method of obtaining these matching information on-site is to trace the historical logs and screen the part with the entity association ID, mark the records that meet the association conditions and accumulate them according to the strength values, and finally obtain . For example, in a period of historical data, there are a total of 30 records in the group, and 20 of them carry the entity matching flag. Their matching strength is given by the associated environmental data. When the accumulated sum is found to reach 8.0 after summarization, can be set as 8.0.
[0101] The acquisition steps of [are as follows. From last week until now, the entities linked by this group are counted and the number of unique entities is recorded. The specific method is to check the queries or operations for entity linking to this group in the database within a week, confirm the entity IDs one by one and remove duplicates for counting, and obtain the total quantity as . For example, if there are 25 entity queries and links to this group within a week, corresponding to 10 different entity IDs, can be obtained.
[0102] The acquisition steps of [are as follows. Among all the structured probability data records in the group, calculate the data volume size of each record and take the average. To quantify the data volume of each record, the following formula can be set: , where represents the actual byte or numerical capacity of the nth record in the group, including the space occupied by the record during storage. is the total data volume of all records in the group, and finally divided by the number of records to obtain the average data density. For example, when sorting sensor logs, it is found that the record size ranges from 300 to 400 bytes. If there are 50 records in the group and the total sum is 18000 bytes, then , indicating that the average data density reaches 360 bytes.
[0103] The steps to obtain are as follows: Extract the confidence interval boundary values obtained in the previous stage for this group and calculate the average value. Each record will obtain an estimated value of the confidence interval boundary based on parameters such as the acquisition time interval, measurement value jump amplitude, sensor accuracy, measurement fluctuation range, and error upper limit in the previous process. The multiple records included in this group can add up these confidence interval boundary values and divide by the number of records to obtain .
[0104] The steps to obtain are as follows: Count the total number of link entities hit by this group in all records. It is necessary to check the corresponding entity ID and link label for each record one by one. When a record that meets the link condition is detected, increment the count by 1. Finally, add up the entity hit marks of all records to obtain .
[0105] Calculation process:
[0106] Let , , , , , , and substitute these values into the formula in turn:
[0107] ;
[0108] First calculate , then , the numerator part is , in the denominator , , the denominator is , divide the numerator 188.3 by 363.355 to get approximately 0.518, and finally .
[0109] The result shows that the confidence factor of the current group on the entity link is about 0.518. If multiple groups are calculated under the same conditions, different link confidence factor distributions can be obtained. When the value is high, it means that the link between the group and the entity is relatively stable. When the value is low, it means that the link between the group and the entity is relatively loose. Combined with this value, groups with high matching degrees can be quickly screened out when the index structure is constructed and then processed later.
[0110] Based on the obtained entity link confidence factor, read the corresponding The value is segmented and annotated through a numerical comparison process. For example, if it is set in the range of 0 to 0.5, it is recorded as a low link trust, if it is in the range of 0.5 to 1.0, it is recorded as a moderate link trust, and if it exceeds 1.0, it is recorded as a high link trust. These segment boundaries are determined by comparing past entity association logs and on-site monitoring data. Several link records are summarized and their overall distribution in different numerical segments is counted. All records that meet the high link trust are gathered again and the entity identification is read. When reading, the group positioning hash index is first retrieved, and the corresponding table of partition positioning information and entity ID is found from it. Then the data storage path of each record is remapped to sort out a union that can be quickly queried. Item list, this list includes entity identifiers, data record storage paths and group positioning hash indexes. After the mapping structure is formed, a check is performed to check whether there are repeated entity identifiers. If it is found that the same entity identifier belongs to multiple different groups, their confidence factors are tracked separately and the higher one is retained, and the lower one is marked. At the same time, the data record storage path is verified for consistency to confirm that it does not conflict with the index control table generated in the previous stage. When encountering a storage path that has been registered in the index control table, the corresponding hash index is directly updated to the same mapping structure. Through such cross-comparison, the induction of all entity identifiers is completed, and finally the mapping structure is written into the index control table in batches to form a multi-dimensional probabilistic index structure.
[0111] The steps to obtain the ontology mapping of the query concept are:
[0112] Receive natural language query content input by the user, extract entity names, limited ranges, time conditions, and semantic association features in the natural language query content, and form a natural language query feature set;
[0113] Based on the natural language query feature set, the association relationship and semantic role between entities are determined through named entity recognition and relationship extraction logic, and an entity relationship recognition structure with determined semantic role tags is generated;
[0114] Based on the entity relationship recognition structure with definite semantic role tags, call the predefined set of domain ontology elements, perform entity concept matching according to the semantic roles, obtain the mapping relationship from entities to ontology elements, and form the ontology mapping of the query concept.
[0115] Specifically, receive the natural language query content input by the user. First, read the entire text and split it into words and phrases. During the splitting process, compare each punctuation mark, number, letter, and mixed character sequence with the internal dictionary table one by one. This dictionary table has previously been manually collected and organized with several common industry terms, place names, personal names, and professional nouns, and also includes some common quantifiers and range descriptions. Then, centrally identify the text units obtained from the splitting. When identifying, lock the potential entity names according to the entry categories in the dictionary table, and at the same time scan for adjectives or adverbs that appear in the text. If restrictive words such as "greater than", "less than", "between" are detected in the corresponding context, use them as clues for the limited range, and then use the preset numerical extraction rules to identify the subsequent numbers or units. For example, when "room temperature needs to be maintained between 10°C and 30°C" is detected, record 10°C and 30°C respectively as the lower and upper limits of the limited range. At the same time, search for various date and time keywords, such as "yesterday", "this week", "2025-04-09", etc., record all date and time forms item by item in the text, and set a start and end mark to distinguish specific time conditions. If an incomplete time segment is encountered, such as "around 2 pm", standardize it after comparing it with the time zone configuration table in the system. After extracting the entity names, limited ranges, and time conditions, continue to search for semantic association features. By locating verbs and associated conjunctions in the text, such as "cause", "correspond to", "associate", etc., perform a mapping comparison with the entities in the front and back texts. If a phrase that meets the structure is matched, register it as a potential semantic association feature. If there are ambiguous words during the process, mark them and append them to the entity. Finally, assemble the relevant information into a mapping table according to all entity names, limited ranges, time conditions, and semantic association features. This mapping table contains the optional ranges, date time periods, and associated verbs or conjunctions corresponding to each entity name, forming a set of natural language query features.
[0116] Based on the natural language query feature set obtained in the previous step, sequentially read the already extracted entity names, scoping, time conditions, and semantic association features, and determine the association relationships and semantic roles between entities through a pre-trained named entity recognition and relationship extraction process. During the specific training process, first construct a corpus containing industry texts and regular dialogue texts. The total number of words in this corpus is about 3 million. When annotating on this corpus, independent tags are respectively assigned to the entity names, time phrases, and quantifier phrases in each text. Then, the annotated corpus is segmented into a training set and a validation set. During actual training, a serialization method is used to encode the input text, and each word or character corresponds to a vector representation. A bidirectional long short-term memory network combined with a conditional random field structure is used for named entity recognition training. The input end includes a sequence of word vectors and an embedded vector at the character level. The context information is retained through a bidirectional gating mechanism in the middle layer of the network. Finally, the output layer uses a conditional random field to decode the annotation sequence. During the training process, the cross-entropy loss is continuously calculated by comparing the predicted sequence with the annotation sequence and the gradients are updated. The recognition accuracy is continuously monitored on the validation set. When the recognition accuracy and recall rate reach the set thresholds after multiple rounds of iteration, for example, the recognition accuracy threshold is set to 85% and the recall rate threshold is set to 80% in the industry scenario, after meeting the requirements, it is confirmed that the named entity recognition effect of the model meets the standard. Then, while the model annotates each input text with entities, the relationship extraction process is also embedded in it. After extracting the entities, specific verbs or prepositions between the entities will be detected. If a relationship pattern defined, such as "occurs in", "belongs to", "is associated with", etc., is detected, the relationship and semantic role will be output together, generating an entity relationship recognition structure with definite semantic role markings.
[0117] Based on the entity relationship recognition structure with definite semantic role tags, read item by item and compare with the predefined set of domain ontology elements within the system. This set of domain ontology elements was sorted out by business experts and the development team together at the start of the project, covering common concepts, attributes, and hierarchical relationships in the industry. Each ontology element has a unique identifier and descriptive information. When a certain entity is marked as the "subject" or "target" role in the recognition structure, compare the descriptions of the subject and target type concepts in the set of ontology elements. When a concept name is found to correspond to the character expression or synonyms of this entity, link this entity to the corresponding ontology element. At the same time, read the additional attributes existing in the recognition structure, such as the restricted scope or time condition associated with this entity, and when invoking the hierarchy of the ontology element, compare whether this attribute can match the corresponding ontology relationship. For example, if the current ontology has a sub-concept "constant temperature range" for the concept of "temperature", then the "10°C to 30°C" that appears in the statement can be matched to this sub-concept. If the match is successful, record this information in the entity mapping table. Repeat the above steps for all entities and their semantic roles. When no suitable ontology concept is found for some entities, record them as unknown entities and retain their semantic role tags. Finally, establish a mapping relationship between all successfully mapped entities and the corresponding ontology concepts to form the ontology mapping of the query concept.
[0118] The steps to obtain the optimized logical query plan are as follows:
[0119] Receive the ontology mapping of the query concept, read the entity type, attribute name, constraint relationship, and class hierarchy path of each ontology mapping, and call the defined logical connection structure and attribute inheritance rules in the domain ontology to perform node connection, path combination, and constraint merging on the ontology mapping to generate a logical query representation;
[0120] Based on the logical query representation, calculate the path selection priority score, and the calculation formula is:
[0121] ;
[0122] Among them, is the path selection priority score corresponding to the th logical query representation, is the hit frequency of the field in the database in this path, is the filtering degree of this path field, is the total execution jump count of this path, is the connection cost of the path field, is the amount of data row processing during the execution of this path, is the number of successful executions of this path in the past 30 days;
[0123] Based on the path selection priority score, select the path with the lowest path selection priority score from all candidate paths as the main query path, and call the field order, connection order, and execution priority rules of the main query path to combine and obtain an optimized logical query plan.
[0124] Specifically, receive the ontology mapping of the query concept, read the entity type, attribute name, constraint relationship, and class hierarchy path of each ontology mapping, and disassemble this information from a pre-compiled ontology definition file, which contains detailed records of entity types and their hierarchical relationships. For example, it records the main entity "device" and its subordinate sub-hierarchy "sensor", and lists the attribute names that each can contain and the available constraint methods. After reading, match it with the actual query concept. During the matching process, it is necessary to check whether the entities and constraint conditions of the query concept correspond to the same or synonymous elements in the file. If synonymous elements are found, include them in the associated reference table. If no matching item is found, append a prompt text in the record table to mark that there is no corresponding ontology element for this element. Subsequently, read these matched entities and their attribute fields item by item, and merge them to obtain a node connection list. In the node connection list, entities with a direct superior-subordinate relationship are sorted closely, entities with a common parent or similar constraints are grouped into the same branch, and constraint condition information is added below the same branch to indicate the constraint merging process. When it is detected that a certain attribute name corresponds to multiple constraint relationships, summarize them numerically. For example, merge the multiple qualification ranges of the same attribute into an interval. When it is detected in the attribute inheritance rule that recursive matching of the inheritance hierarchy is required, further search upward along the parent node to see if there are higher-level concepts that can merge the same or similar constraints. For example, take the stricter one of the temperature limits of "device" and the temperature limits of "sensor" as the final constraint. After node connection, path combination, and constraint merging, a complete logical query representation is formed, where each node has a corresponding entity or attribute type, and at the same time stores the specific connection order from the parent level to the sub-level. Retain this logical query representation in the data structure for subsequent path selection.
[0125] Formula: , The benefit of the formula is that it incorporates multiple factors such as the hit frequency of fields in the database, field screening degree, total execution jump count, path field connection cost, data row processing volume, and the number of successful executions in the past 30 days into numerical operations for a path, and uses a single score value to represent the priority of this path during query execution. In the query plan optimization phase, more appropriate execution routes can be selected by comparing the scores of each path.
[0126] The acquisition steps are as follows: it is necessary to traverse the fields involved in the current path in the database access log or system statistics, count the total number of times the field appears or is called in the past period of time, perform a proportional operation on this number and the total number of appearances of all fields in the same period to obtain the hit frequency value of the field in the database. If the field name in the scenario is "device_status" and it is requested 230 times in a week, while the total number of requests for all fields is 5000 times, then the hit frequency of this field can be calculated as , recorded as 0.046. When there are multiple fields in the path, extract their hit frequencies respectively and take the average value to form , for example, in a path containing 3 fields, the field hit frequencies are 0.046, 0.03, and 0.02 respectively, then .
[0127] The acquisition steps are as follows: perform a screening degree evaluation on the fields involved in this path, that is, judge how many data rows a certain field can filter out when being queried. This screening degree can be quantified through the following formula: , the number of rows filtered out comes from the comparison statistics of the past query execution situation on site. For example, when the field "device_status" performs conditional filtering in a week, 800 rows are filtered out, and in all queries, this field combines and processes 2000 rows. Thus, is obtained. If there are multiple fields in the path, the total screening degree of each field can be calculated and then comprehensively measured in the final combined conditions. For example, when there are 3 fields that respectively obtain 0.4, 0.2, and 0.6, the average value can be taken .
[0128] The acquisition steps are as follows: calculate the total number of hops in a complete query execution of this path. The number of hops can be defined as the number of operations for a connection or a shard retrieval in a relational database or a distributed storage environment, and it can be counted in the execution plan log. If this path involves 2 connections and 1 index jump, add them up to get .
[0129] The acquisition steps are as follows: here, the connection cost refers to the resource burden consumed when associating and connecting fields during query execution, including memory occupation and data transmission consumption, and it can be estimated through the following formula: , where is the total number of connections that occur under this path, represents the data throughput ratio during each connection, represents the time for this data throughput to be transmitted in memory and on the network. Add up the resource consumption values of each connection and record it as Add them together to get the final connection cost. For example, a certain path is tested , , , , then .
[0130] The acquisition steps of are as follows: count the number of data row processing during the execution of this path. It is necessary to record both the number of rows before screening and the number of rows finally passed to the next step. The sum of the two or the peak value at the key time point can be used as a measure. When performing multi-table joins, the number of intermediate result rows generated after combining Table A and Table B can be included in the total processing volume. For example, in a path query, the initial table contains 1000 rows, and 300 rows remain after connection screening. When associating with the third table, it becomes 500 rows, and finally 200 rows are returned. Then these processes can be added together .
[0131] The acquisition steps of are as follows: obtain it by accumulating the number of successful executions of the same path in the past 30 days in the query system log. It is necessary to find the execution records that match the path characteristics in the log and verify whether they end successfully. Successfully ending means not being terminated halfway or throwing an error. For example, a specific path is called 50 times within 30 days, and 45 of them are executed successfully. Then .
[0132] Calculation process:
[0133] Let , , , , , , substitute these values into the formula:
[0134] ;
[0135] First calculate the numerator: , then , the numerator is , for the denominator, first calculate , then , the denominator is , divide the numerator 1.961 by the denominator 2001.728 to approximately get .
[0136] The result shows that the current path selection priority score is approximately 0.000979. The smaller the value, the more likely the path is to be one of the main execution paths after comprehensively considering the hit frequency, screening degree, hop count, connection cost, data row processing volume, and execution records in the past 30 days. When compared with other paths, if this value is the lowest among all candidate paths, it can be set as the main query path in the query optimization stage and provide information such as field order and connection order for the subsequent execution process.
[0137] Based on the obtained path selection priority scores, compare each candidate path one by one and other parameters and calculate the corresponding scoring results. Sort the scoring results in ascending order of values. During the sorting process, parallelly mark the paths with the same scoring values, and retrieve the execution log records of each path in the past week to confirm whether it meets the latest system configuration thresholds. If the execution process of a path generates a data row processing volume or connection cost exceeding the set threshold, add a note in the comparison table. For example, set the threshold to 1500 data row processing volume. When it is detected that a certain path generates a high load of 1800 rows during actual operation, mark it as having a high load. After the comparison, select the path with the lowest score and not marked as having a high load as the main query path, read its field order, list the connection order one by one, and select the corresponding priority level in the predefined execution priority rule set. Finally, combine to obtain an optimized logical query plan.
[0138] The steps to obtain the candidate probability record set are as follows:
[0139] Based on the optimized logical query plan, call the entity identifier, data record storage path, and grouped location hash index in the multi-dimensional probability index structure, perform combined matching and record path screening, and generate a preliminarily filtered set of probability data records;
[0140] Based on the preliminarily filtered set of probability data records, compare each probability query condition one by one, eliminate the data records that do not meet the probability query conditions, and retain all the data records that meet the probability query conditions to form a candidate probability record set.
[0141] Specifically, based on the optimized logical query plan, read the corresponding entity identification information in the multi-dimensional probability index structure established in the early stage, extract the data record storage path and the grouped positioning hash index therein, pair the entity identification with the storage path, and check item by item whether it meets the path combination requirements listed in the query plan. When it is found that a storage path is repeated or incomplete, record its index information first, and then read the originally generated probability attribute grouped data item by item with the hash index as a clue. During this process, it is necessary to strictly compare the established hash value with the entity identification. If the hash value is inconsistent with the entity identification, this record will be temporarily excluded. If they are consistent, it will be marked as a matching item. Then, sort the data records of this part of the matching items in chronological order into a sequence for further screening. Subsequently, according to the access range threshold set by the technical personnel during the on-site operation and maintenance stage, for example, the situation where the entity identification contains more than 10 key symbols or the storage path byte count is greater than 200 is used as a high-load mark and grouped separately in the matching list. After grouping, check item by item whether there is a situation where the storage path jumps significantly or the centralized index fails by referring to the connection order specified in the query plan. If it is found during the comparison that the data record has a coordinate out-of-bounds or the measured value does not meet the expected range, it will be labeled with a deviation label. Keep the matching records that meet all connection and index rules in a new temporary list and calculate the number of data items in this temporary list. When the number of data items exceeds the system default capacity threshold of 300, perform proportional splitting, and further re-integrate each split segment according to the grouped positioning hash index, thus obtaining a preliminary filtering sequence that can be processed quickly. Finally, summarize all the probability data records in the preliminary filtering sequence into a preliminary filtered probability data record set.
[0142] Based on the preliminary filtered probability data record set, read the parameter content of each record therein in turn and compare it with the probability query conditions. When comparing, first check the basic fields such as the measured value, confidence interval boundary estimate value, sensor accuracy level, etc. If the measured value needs to be in the range of 0 to 90, check item by item against this range and mark the records that do not meet the requirements. If a preset threshold of 0.05 is set for the confidence interval boundary estimate value, check whether each record is greater than or less than 0.05. Mark the part higher than 0.05 as out of range. When cross-comparing the same sensor accuracy level, if this level is lower than a certain reference value of 1.5, it is considered low and an additional mark is made in the table. Then, count the jump amplitude in the record and the numerical fluctuation situation in the most recent day or week. When the fluctuation situation exceeds the threshold defined by the maintenance personnel in historical observations, place this record in a separate list. Keep all the records that meet the query conditions in the main list. Finally, after summarization, obtain a record set that meets the probability query conditions, forming a candidate probability record set.
Claims
1. An artificial intelligence digital technology platform based on big data analysis, characterized in that: The platform includes: A data uncertainty modeling module receives a large data stream, determines a confidence interval boundary for each data item to represent the uncertainty of the data item, generates a data item uncertainty parameter set, combines each data item with the corresponding data item uncertainty parameter set and stores them to form a structured probabilistic data record; A probabilistic index construction module, based on the structured probabilistic data record, performs data partitioning according to a preset spatial range and numerical range to form probabilistic attribute groups, and establishes index entries pointing to the storage location of the structured probabilistic data record according to the probabilistic attribute groups and entity link probabilities to construct a multidimensional probabilistic index structure; The semantic query conversion module receives natural language queries, identifies concepts and constraints in the query through named entity recognition and relationship extraction, and maps concepts to domain ontology elements to form an ontology mapping of query concepts. The ontology mapping of the query concepts is combined using the logical relationships defined in the domain ontology, and converted to generate a logical query representation to obtain an optimized logical query plan. The search result sorting module analyzes the probability query conditions based on the optimized logical query plan, and uses the multi-dimensional probability index structure to filter data to obtain a candidate probability record set.
2. The artificial intelligence digital technology platform based on big data analysis according to claim 1 is characterized in that: The steps for obtaining the data item uncertainty parameter set are: Receive a single data item carried by each record in a large data stream, extract the data item's collection time interval, measurement value jump amplitude, sensor accuracy, standard measurement fluctuation range, sensor sampling error upper limit and record span, and generate a data item measurement configuration group; Calculating confidence interval boundary estimates based on the data item measurement configuration group; The confidence interval boundary estimate value, the measurement value jump amplitude, the standard fluctuation range and the upper limit of the error are combined to form an uncertainty parameter unit, and a corresponding uncertainty parameter unit is generated for each data item to generate a data item uncertainty parameter set.
3. The artificial intelligence digital technology platform based on big data analysis according to claim 1 is characterized in that: The steps of obtaining the structured probability data record are: Binding the data item uncertainty parameter set to the corresponding data item, creating a unique identification index key for each data item, and generating a data item combination structure with an identification index; Based on the data item combination structure with identification index, data serialization and compression are performed, and the compressed data item combination structure is written into the storage space one by one to form a structured probabilistic data record.
4. The artificial intelligence digital technology platform based on big data analysis according to claim 1 is characterized in that: The steps of obtaining the probability attribute grouping are: Extracting the spatial coordinate point, data measurement value, confidence interval boundary estimate, recording time point, data sampling accuracy level and value fluctuation intensity of each record in the structured probabilistic data record to generate a partition reference attribute group; Calculating a partition adaptation strength value according to the partition reference attribute group; Based on the partition adaptation strength value, parallel cluster distribution judgment and critical similarity screening are performed on all structured probabilistic data records to generate probabilistic attribute groupings.
5. The artificial intelligence digital technology platform based on big data analysis according to claim 1 is characterized in that: The steps for obtaining the multidimensional probability index structure are: Extracting the structured probability data records in the probability attribute grouping, obtaining the spatial distribution range, entity matching frequency, record data density, number of linked entities in the last week, average confidence interval boundary value and entity hit rate, and forming a link positioning basic information group; Calculating the entity link confidence factor based on the link positioning basic information group; Based on the entity link confidence factor, a mapping structure with entity identification, data record storage path and group positioning hash index as joint items is constructed, and the mapping structure is written into the index control table of the structured probabilistic data record to generate a multidimensional probabilistic index structure.
6. The artificial intelligence digital technology platform based on big data analysis according to claim 1 is characterized in that: The steps for obtaining the ontology mapping of the query concept are: Receive natural language query content input by the user, extract entity names, limited ranges, time conditions, and semantic association features in the natural language query content, and form a natural language query feature set; Based on the natural language query feature set, the association relationship and semantic roles between entities are determined through named entity recognition and relationship extraction logic, and an entity relationship recognition structure with determined semantic role tags is generated; Based on the entity relationship recognition structure with certain semantic role tags, a set of predefined domain ontology elements is called, entity concepts are matched according to semantic roles, and a mapping relationship from entity to ontology elements is obtained to form an ontology mapping of the query concept.
7. The artificial intelligence digital technology platform based on big data analysis according to claim 1 is characterized in that: The steps for obtaining the optimized logical query plan are: Receive the ontology mapping of the query concept, read the entity type, attribute name, constraint relationship and class hierarchy path of each ontology mapping, call the logical connection structure and attribute inheritance rules defined in the domain ontology, perform node connection, path combination and constraint merging on the ontology mapping, and generate a logical query representation; Calculating a path selection priority score based on the logical query representation; Based on the path selection priority score, the path with the lowest path selection priority score is selected from all candidate paths as the main query path, and the field order, connection order and execution priority rules of the main query path are called to combine and obtain an optimized logical query plan.
8. The artificial intelligence digital technology platform based on big data analysis according to claim 1 is characterized in that: The steps for obtaining the candidate probability record set are: Based on the optimized logical query plan, the entity identifier, data record storage path and group positioning hash index in the multidimensional probabilistic index structure are called to perform combination matching and record path screening to generate a preliminary filtered probabilistic data record set; Based on the initially filtered probability data record set, the probability query conditions are compared one by one, data records that do not meet the probability query conditions are eliminated, and all data records that meet the probability query conditions are retained to form a candidate probability record set.
Citation Information
Cited By
Personnel information big data screening method and system based on person and certificate identification
CN120743907A
Personnel information big data screening method and system based on human and certificate identification
CN120743907B
Non-motor vehicle violation data processing method and system
CN121166677A
An off-chip storage structure setting method of an accurate matching flow table, a chip and a network card
CN122547706A
An off-chip storage structure setting method of an accurate matching flow table, a chip and a network card
CN122547706B