Data quick retrieval method and system based on active identification, terminal and medium

By assigning industrial Internet identification to equipment terminals on the industrial Internet platform and building a property index structure of two-layer B-class tree data structures, the problems of slow retrieval speed and low accuracy of traditional data retrieval methods in large-scale and high-complex industrial data processing are solved, and fast and accurate data retrieval is achieved, which improves data management level.

CN120067109APending Publication Date: 2025-05-30INSPUR YUNZHOU (SHANDONG) IND INTERNET CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510267287.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When traditional data retrieval methods process large-scale and high-complex industrial data, they show the disadvantages of slow retrieval speed and reduced accuracy, which is difficult to meet the requirements of modern industrial production for data immediacy and accuracy.

Method used

Using a fast data retrieval method based on active identification, we use the industrial Internet identification to each device terminal, parse the identification to obtain data records, and build the attribute index structure of the two-layer B-class tree data structure, including B-tree and B+tree, to achieve efficient data retrieval.

Benefits of technology

It realizes fast and accurate retrieval of industrial data, improves data utilization efficiency and management level, has the advantages of fast retrieval speed, high accuracy, strong security and good user experience, and is suitable for the industrial Internet field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067109A_ABST
    Figure CN120067109A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data retrieval, in particular to a quick data retrieval method and system based on active identification, a terminal and a medium, and the method comprises the following steps: endowing each equipment terminal with an industrial internet identification as a unique digital identity; the industrial internet identifier is analyzed, and a data record of the equipment terminal is obtained; an attribute index structure is constructed according to the data records, the attribute index structure comprises two layers of B-type tree data structures, the first layer adopts a B tree to group all the data records according to attributes, and the second layer independently establishes a B + tree for each of other attributes in the grouped data records of the first layer; and obtaining a retrieval condition, and performing data retrieval based on the attribute index structure to obtain an attribute retrieval result. According to the method, an efficient attribute index structure is actively identified and constructed through the industrial internet, so that the industrial data can be quickly and accurately retrieved, and the data utilization efficiency and the management level are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data retrieval, and particularly to a method, system, terminal and medium for fast data retrieval based on active identification. Background Art

[0002] With the deep promotion of the Industry 4.0 era, the core position of the industrial Internet platform in the intelligent manufacturing system has become increasingly prominent, and data management has become a key link to ensure the efficient operation of this system. This link is not just a concept, but specifically covers a series of refined operations from the initial data collection, secure storage, in-depth analysis, efficient sharing to strict protection. These operations jointly support the realization of intelligent manufacturing, help enterprises optimize production processes, and significantly improve production efficiency and innovation capabilities.

[0003] In the context of the booming development of the industrial Internet, industrial data has shown unprecedented complexity. The huge data scale, diverse sources, and heterogeneous structures have brought great challenges to the correlation and integration of data, and it is difficult to fully explore and utilize the data value. This not only limits the support of data for production decisions, but also affects the response speed of enterprises to market changes.

[0004] Facing such a large industrial data set, traditional data retrieval methods seem inadequate. When dealing with large-scale and highly complex industrial data, these methods often show the disadvantages of slow retrieval speed and decreased accuracy. This directly leads to low data acquisition efficiency and is difficult to meet the immediate and accurate requirements of modern industrial production for data. Summary of the Invention

[0005] To solve the technical problems of slow retrieval speed and low accuracy when existing data retrieval methods process large-scale and highly complex industrial data, the present invention provides a method for fast data retrieval based on active identification, and also provides a system, a terminal and a medium for fast data retrieval based on active identification.

[0006] To achieve the above object, in the first aspect, the technical solution adopted by a method for fast data retrieval based on active identification in the present invention is as follows: A method for fast data retrieval based on active identification includes the following steps: Assign an industrial Internet identifier to each device terminal as a unique digital identity; Parse the industrial Internet identifier to obtain the data records of the device terminal. Each data record contains various attribute information of the device terminal, each type of attribute information has corresponding data items, and each data item has a unique identification information; Construct an attribute index structure according to data records. The attribute index structure includes two - layer B - type tree data structures. The first layer uses a B - tree to group all data records according to attributes, and the second layer separately constructs a B + - tree for each other attribute in the data records grouped in the first layer; Obtain retrieval conditions, perform data retrieval based on the attribute index structure, and obtain an attribute retrieval result.

[0007] As a preferred implementation of the data fast retrieval method based on active identification, after constructing the attribute index structure, split the attribute index structure into multiple sub - attribute index structures, and store the multiple sub - attribute index structures dispersedly on several computing nodes. Each computing node can perform data retrieval based on its stored sub - attribute index structure and transmit the calculation result to the terminal node for aggregation.

[0008] As a preferred implementation of the data fast retrieval method based on active identification, store the sub - attribute index structures with an average access frequency exceeding a preset frequency value on the same computing node, adjacent computing nodes, or computing nodes closer to the terminal node.

[0009] As a preferred implementation of the data fast retrieval method based on active identification, in the second - layer B + - tree of the attribute index structure, optimize the leaf nodes of the B + - tree with a data volume exceeding a preset quantity value using a Bloom filter, and / or optimize the leaf nodes of the B + - tree for processing time - series data using a linear model.

[0010] As a preferred implementation of the data fast retrieval method based on active identification, after constructing the attribute index structure according to data records, when new data records are added, identify the attribute information in the new data records, determine the insertion position of the new data records in the already constructed attribute index structure based on the attribute information, and insert the corresponding index items; when data records are deleted or modified, obtain the identification information of the deleted or modified data records, locate and delete or modify the corresponding index items in the already constructed attribute index structure based on the identification information of the deleted or modified data records.

[0011] As a preferred implementation of the data fast retrieval method based on active identification, the data records also include text - type data. After obtaining the data records of the device terminal, the method further includes: Perform word - segmentation processing on the text - type data in the data records to split continuous text into individual keywords; Use the inverted index algorithm to establish an inverted index structure. When establishing the inverted index structure, traverse the text - type data in all data records, associate each keyword with the corresponding data record, and store them in the index; After obtaining the retrieval conditions, data retrieval is performed based on the inverted index structure to obtain the keyword retrieval results.

[0012] As a preferred implementation of the data fast retrieval method based on active identification, after obtaining the retrieval conditions, the retrieval conditions are parsed to extract the attribute retrieval conditions, keyword retrieval conditions, and logical operators in the retrieval conditions; For the attribute retrieval conditions, data retrieval is performed based on the attribute index structure to obtain the attribute retrieval results; For the keyword retrieval conditions, data retrieval is performed based on the inverted index structure to obtain the keyword retrieval results; According to the rules of the logical operators, the attribute retrieval conditions and the keyword retrieval conditions are merged to obtain the final retrieval results.

[0013] In a second aspect, the technical solution adopted by a data fast retrieval system based on active identification in the present invention is as follows: A data fast retrieval system based on active identification includes an active identification module, an identification parsing module, a data indexing module, and a data retrieval module. Among them, the active identification module is configured to assign an industrial Internet identification to each device terminal as a unique digital identity; the identification parsing module is configured to parse the industrial Internet identification to obtain the data records of the device terminal; the data indexing module is configured to construct an attribute index structure according to the data records. The attribute index structure includes a two-layer B-tree data structure. The first layer uses a B-tree to group all data records according to attributes, and the second layer separately constructs a B+ tree for each other attribute in the data records grouped in the first layer; the data retrieval module is configured to obtain the retrieval conditions and perform data retrieval based on the attribute index structure to obtain the attribute retrieval results.

[0014] In a third aspect, the technical solution adopted by a terminal in the present invention is as follows: A terminal includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of any one of the above-mentioned data fast retrieval methods based on active identification are implemented.

[0015] In a fourth aspect, the technical solution adopted by a medium in the present invention is as follows: A storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any one of the above-mentioned data fast retrieval methods based on active identification are implemented.

[0016] The beneficial effects of the present invention include: 1. The present invention actively identifies and constructs an efficient attribute index structure through the industrial Internet, enabling fast and accurate retrieval of industrial data, improving data utilization efficiency and management level. It has the advantages of fast retrieval speed, high accuracy, strong security, and good user experience, and has broad application prospects in the field of industrial Internet.

[0017] 2. In the attribute index structure of the present invention, in the first layer, all data records are grouped according to attributes, which can manage and analyze data more efficiently, better organize and manage large-scale data, and improve the efficiency of data retrieval and analysis. In the second layer, multiple B+ trees are used to process data records of different attribute categories respectively, which is an efficient index design method. Through layering and multi-attribute indexing, the efficiency and flexibility of data retrieval can be significantly improved. The attribute index structure of the present invention is particularly suitable for the fast retrieval requirements of multi-dimensional and large-scale data in the industrial Internet scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a schematic flow chart of a data fast retrieval method based on active identification in a specific embodiment of the present invention; Figure 2 It is a schematic structural diagram of a data fast retrieval system based on active identification in a specific embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] For easy understanding, some terms of the prior art that appear in the present invention are first explained: 1. The industrial Internet identification and coding technology is an important part of the industrial Internet network system. It assigns an industrial Internet identifier to each object through the network and, with the help of the industrial Internet identifier resolution system, realizes information query and sharing across regions, industries, and enterprises. This identifier is similar to the domain name resolution system (DNS) in the Internet field and is the nerve center supporting the interconnection and interoperability of the industrial Internet. Such as EUI-64, UUID, etc., which provide a unique globally identifiable identity for each device in the IPv6 environment, helping to build efficient data interaction in large-scale industrial data. It can realize the rapid registration, positioning, and maintenance of devices, enhancing the reliability and response speed of the system.

[0021] 2. A B-tree is a self-balancing multi-way search tree designed to efficiently handle the storage and retrieval of large amounts of data and is suitable for external storage systems such as disks. A B+ -tree is a variation and optimization of the B-tree. It is also a multi-way balanced search tree. It is improved based on the B-tree and is more suitable for range queries and sequential access. To make the objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the specific embodiments. Obviously, the embodiments described below are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.

[0022] Referring to Figure 1 , a data rapid retrieval method based on active identification proposed in this embodiment includes the following steps: S1. Assign an industrial Internet identifier to each device terminal as a unique digital identity. That is, at the source where data is generated, a unique identification code is assigned to each data item, so that it has unique identification information. For example, a standard identification form such as a Universally Unique Identifier (UUID) is used. This identifier accompanies the data throughout its life cycle. Whether in the storage, transmission, or retrieval stage, each data item can be accurately distinguished by this unique identifier.

[0023] S2. Analyze the industrial Internet identifier to obtain the data records of the device terminal. Each data record contains multiple attribute information of the device terminal. Each type of attribute information has corresponding data items, and each data item carries unique identification information.

[0024] An attribute is a characteristic or property of data. For example, device ID, device model, affiliated workshop, timestamp, temperature, pressure, operating status, etc.

[0025] A data item is the specific value of a certain attribute, such as temperature = 75°C, device ID = 001, etc. Each data item carries unique identification information, and each data item can be distinguished and its source traced based on the identification information. Suppose a device operation record contains the following information: device ID = 001, timestamp = 2023-10-01 12:00:00, temperature = 75°C, pressure = 100 kPa, operation status = normal. First, a combination including the device ID, timestamp, and data item name can be used as the basis for the unique identifier. Among them, "device ID = 001" itself is an identifier, and "timestamp = 2023-10-01 12:00:00" provides the time background of the data record. For each data item (such as temperature, pressure, operation status), a name or type identifier can be attached. Then, the unique identification information of "temperature 75°C" in this device operation record can be "001_2023-10-01_12:00:00_temperature_75°C", and for data simplicity, the value can be excluded, such as "001_2023-10-01_12:00:00_temperature"; the unique identification information of "pressure = 100 kPa" can be "001_2023-10-01_12:00:00_pressure_100 kPa", and similarly, the value can be omitted; the unique identification of "operation status = normal" can be "001_2023-10-01_12:00:00_operation status_normal", and similarly, the value can be omitted.

[0026] A data record is a complete data unit composed of multiple related data items and is a complete description of a certain entity or event. It can include various attribute information such as device model, production time, operation status, performance parameters, maintenance records, etc. For example, a device operation record may contain the following information: device ID = 001, timestamp = 2023-10-01 12:00:00, temperature = 75°C, pressure = 100 kPa, operation status = normal.

[0027] S3. Construct an attribute index structure based on the data records. The attribute index structure includes a two-layer B-tree data structure. The first layer uses a B-tree to group all data records according to attributes, and the second layer separately constructs a B+ tree for each other attribute in the data records grouped in the first layer.

[0028] Taking the operation data record table of the following device terminal as an example, the process of constructing the attribute index structure is explained:

[0029] Table 1 Operation Data Record Table of the Device Terminal When the first layer uses a B-tree to group all data records according to attributes, taking Table 1 as an example, that is, using device ID, timestamp, temperature, pressure, and operating status as classification attributes respectively to group the data records. Taking the grouping of data records according to the device ID as the classification attribute as an example, the data records are divided into the following groups: 1. Device ID = 001: 2023-10-01 12:00:00, Temperature = 75°C, Pressure = 100 kPa, Operating Status = Normal; 2023-10-01 12:05:00, Temperature = 78°C, Pressure = 105 kPa, Operating Status = Normal.

[0030] 2. Device ID = 002: 2023-10-01 12:00:00, Temperature = 80°C, Pressure = 110 kPa, Operating Status = Normal; 2023-10-01 12:05:00, Temperature = 85°C, Pressure = 115 kPa, Operating Status = Warning.

[0031] 3. Device ID = 003: 2023-10-01 12:00:00, Temperature = 70°C, Pressure = 95 kPa, Operating Status = Normal; 2023-10-01 12:05:00, Temperature = 72°C, Pressure = 98 kPa, Operating Status = Normal.

[0032] The principle of grouping the data records according to the timestamp, temperature, pressure, and operating status as classification attributes is the same as above and will not be elaborated.

[0033] In the prior art, the common index structures are relatively single, usually a whole B+-tree structure. In this embodiment, a two-layer B-tree data structure is adopted. The first layer uses a B-tree to group all data records according to attributes. In this embodiment, grouping all data records according to attributes in the first layer can manage and analyze data more efficiently, organize and manage large-scale data better, and improve the efficiency of data retrieval and analysis.

[0034] In the second layer, a B+-tree is separately established for each of the other attributes in the data records grouped in the first layer, that is, for each grouped data record, a B+-tree index is constructed according to the other attributes respectively. Taking the data records obtained by grouping according to the device ID as the classification attribute above and constructing B+-tree indexes according to temperature (°C) and pressure (kPa) respectively as an example, as follows: 1. Device ID = 001: Temperature B+ tree: 75°C → Record 1, 78°C → Record 2; Pressure B+ tree: 100 kPa → Record 1, 105 kPa → Record 2.

[0035] 2. Device ID = 002: Temperature B+ tree: 80°C → Record 1, 85°C → Record 2; Pressure B+ tree: 110 kPa → Record 1, 115 kPa → Record 2.

[0036] 3. Device ID = 003: Temperature B+ tree: 70°C → Record 1, 72°C → Record 2; Pressure B+ tree: 95 kPa → Record 1, 98 kPa → Record 2.

[0037] The second layer uses multiple B+ trees to separately process data records of different attribute categories, which is an efficient index design method. Through layering and multi-attribute indexing, the efficiency and flexibility of data retrieval can be significantly improved. This method is particularly suitable for the fast retrieval requirements of multi-dimensional and large-scale data in the industrial Internet scenario.

[0038] S4. Obtain the retrieval condition, perform data retrieval based on the attribute index structure, and obtain the attribute retrieval result. During retrieval, the retrieval condition input by the user can be accurately matched with the identification information to accurately judge the data that meets the retrieval condition, prevent confusion caused by reasons such as data similarity, fundamentally avoid misretrieval, and ensure that the retrieved data is the truly required data. Using the identification information of the retrieval structure can also achieve data traceability, trace back information such as the source and generation process of the data, further verify the accuracy and relevance of the data, and improve the quality of the overall retrieval result.

[0039] As the data scale continues to expand, if all the attribute index structures are centrally stored in one node, it will cause an excessive storage burden on that node, and a large amount of index data needs to be traversed during retrieval, resulting in extremely low retrieval efficiency. Therefore, in some embodiments, after constructing the attribute index structure, the attribute index structure is split into multiple sub-attribute index structures, and the multiple sub-attribute index structures are scattered and stored on several computing nodes. Each computing node can perform data retrieval based on its respective stored sub-attribute index structure and transmit the calculation results to the terminal node for aggregation. Thus, by adopting the distributed index method, the sub-attribute index structures are stored one-to-one or many-to-one on several computing nodes respectively. After obtaining the retrieval conditions, the system distributes the retrieval conditions to each computing node related to the retrieval conditions. Each computing node related to the retrieval conditions can perform data retrieval simultaneously based on the sub-attribute index structure and transmit the calculation results to the terminal node for aggregation to obtain the final attribute retrieval result, which can improve the retrieval efficiency, reduce the storage pressure on a single node, and also make the attribute index structure clearer.

[0040] In addition, the sub-attribute index structures with an average access frequency exceeding the preset frequency value can be stored in the same computing node, adjacent computing nodes, or computing nodes closer to the terminal node. This can increase the hit rate of data during retrieval, enabling more data that meets the retrieval conditions to be obtained with one or fewer traversals, thus accelerating the retrieval.

[0041] In order to optimize the attribute index structure and improve the retrieval performance in specific scenarios, in some embodiments, in the second-level B+ tree of the attribute index structure, Bloom filters are used to optimize the leaf nodes of the B+ tree with a data volume exceeding a preset value, and / or linear models are used to optimize the leaf nodes of the B+ tree for processing time series data. This adaptive optimization of the B+ tree leaf nodes for different application scenarios enables the data structure to better meet the requirements of specific scenarios. Specifically, when processing large-scale data retrieval, especially when there are a large number of queries for data that may not exist (i.e., determining whether a certain data exists in a set), the retrieval efficiency of the traditional B+ tree will be affected. A Bloom filter is a probabilistic data structure that can efficiently determine whether an element belongs to a certain set. Although there is a certain false positive rate, its space efficiency and query efficiency are extremely high. When retrieving data, first quickly determine whether the data may exist in the leaf node through the Bloom filter. If the Bloom filter determines that it does not exist, the leaf node is directly excluded without actual data comparison; if it is determined that it may exist, then further search in the leaf node data to reduce unnecessary data reading and comparison operations and improve the retrieval efficiency. In some application scenarios, the data has a strong linear relationship, such as time series data and numerical data distributed according to a certain linear rule. Taking time series data retrieval as an example, by constructing a linear model for time and the corresponding data values, when retrieving data within a certain time range, the data interval that meets the conditions in the leaf node can be quickly located based on the linear model, reducing the amount of data traversal and improving the retrieval efficiency.

[0042] In data management, an index is like a navigation map for data, which can quickly locate the required data and improve the retrieval efficiency. However, as the data changes continuously, the index also needs to be optimized and maintained regularly to continuously perform at its best. Therefore, in some embodiments, after constructing the attribute index structure based on the data records, when the data records increase, the attribute information in the newly added data records is identified, and based on the attribute information, the insertion position of the newly added data records in the already constructed attribute index structure is determined and the corresponding index entries are inserted. If the leaf node space is insufficient, a node splitting operation is performed to adjust the attribute index structure; when the data records are deleted or modified, the identification information of the deleted or modified data records is obtained, and based on the identification information of the deleted or modified data records, the corresponding index entries are located and deleted or modified in the already constructed attribute index structure. When modifying the corresponding index entries, the old data can be deleted first and then the new data can be inserted, or directly modified. If the leaf node does not meet the balance condition after deletion or modification, node merging or adjustment is performed to maintain consistency. The continuously updated attribute index structure can accurately reflect the latest state of the data, quickly locate the target data during retrieval, improve the overall performance of the system, and meet the requirements of the business for the high efficiency and accuracy of data retrieval.

[0043] During the index update process, according to the storage optimization strategy, factors such as CPU cache, multi-core processors, SIMD, and data prefetching are considered. The attribute index structure is optimized and adjusted, such as adjusting the filling degree of nodes, merging adjacent nodes, etc., to improve the performance of the index.

[0044] In terms of constructing an efficient index, this embodiment can also reasonably design the index structure according to factors such as the characteristics of the data and the usage frequency. For example, for data that is often retrieved according to a time range, an index based on timestamps can be constructed, and the data can be sorted in chronological order. In this way, when querying data within a specific time period, the corresponding data block can be quickly located, reducing the amount of data traversed, thereby significantly improving the retrieval speed.

[0045] In fact, data records not only include attribute information but also text data. For example, device operation logs, which detail various information during the operation of the device terminal, such as error messages during device failures, device parameter adjustment records, etc., contain a large amount of text descriptions, and these contents are all text data. Another example is product manuals, which detail the functions, usage methods, technical parameters, etc. of the product, and the text descriptions therein are also text data.

[0046] The hybrid B+ tree index structure is mainly suitable for processing structured data, especially those that require range queries and sorting operations, and is not suitable for processing text data. Therefore, for the text data in the data records, after obtaining the data records of the device terminal, the method further includes: Performing word segmentation on the text data in the data records to split the continuous text into individual keywords; Adopting an inverted index algorithm to establish an inverted index structure. When establishing the inverted index structure, traverse the text data in all data records, associate each keyword with the corresponding data record, and store it in the index; After obtaining the retrieval condition, perform data retrieval based on the inverted index structure to obtain the keyword retrieval result.

[0047] The inverted index algorithm uses keywords as index terms. By performing word segmentation on text data, it extracts the keywords and then records the data records where each keyword is located. During the construction process, a large amount of text data needs to be processed and analyzed to ensure the accurate extraction of keywords and their effective association with relevant data records. For example, when technicians need to retrieve specific fault information through keywords, the inverted index algorithm will come into play, quickly locating the position of the log records containing the keyword based on the keyword, facilitating technicians to analyze the cause of the fault. Or when users need to find the usage method of a specific function of a product, the inverted index algorithm can directly find the corresponding content position in the product manual according to the input keyword, helping users quickly obtain information.

[0048] Therefore, this application uses an attribute index structure employing a B-tree data structure and an inverted index algorithm to jointly process retrieval conditions. They can complement each other to meet the diverse retrieval needs of the system and can also be used jointly to process complex retrieval conditions, enabling various retrieval strategies such as keyword retrieval, attribute retrieval, and combined retrieval.

[0049] Based on the solution introducing the inverted index structure, after obtaining the retrieval conditions, this method needs to parse the retrieval conditions, extract the attribute retrieval conditions, keyword retrieval conditions, and logical operators in the retrieval conditions; For the attribute retrieval conditions, data retrieval is performed based on the attribute index structure to obtain the attribute retrieval results; For the keyword retrieval conditions, data retrieval is performed based on the inverted index structure to obtain the keyword retrieval results; According to the rules of logical operators, the attribute retrieval conditions and keyword retrieval conditions are merged to obtain the final retrieval result.

[0050] During the process of merging the results of each sub-query, for the "AND" logical operator, the intersection of the results of each sub-query is taken; for the "OR" logical operator, the union of the results of each sub-query is taken; for the "NOT" logical operator, the complement of the result of a certain sub-query is taken.

[0051] To further improve the adaptability and query efficiency of the index, this embodiment also uses a machine learning model (such as a recursive model index RMI) to replace the traditional index structure. The implementation steps of using a machine learning model (such as a recursive model index RMI) to replace traditional structures such as B+ trees are as follows: 1. Data feature extraction and preprocessing Data collection: Collect historical query logs and dataset distribution characteristics (such as key value ranges, query frequencies, data distributions, etc.).

[0052] Feature Engineering: Extract key features of data, including key-value distribution (such as uniform, skewed), query patterns (such as point query, range query), data update frequency, etc.

[0053] Data Normalization: Standardize key values (such as normalize to the interval [0, 1]) to ensure numerical stability of model input.

[0054] 2. Recursive Model Indexing (RMI) Architecture Design Hierarchical Model Construction: Top-level Model (Root Model): Use linear regression or shallow neural network to map input key values to the index interval of the second-level model.

[0055] Middle-level Models: Use multi-layer perceptron (MLP) or decision tree to further refine the mapping of key values to the third-level model.

[0056] Bottom-level Models (Leaf Models): Directly predict the physical location of data in storage (such as memory address or disk block) through lightweight models (such as linear regression).

[0057] Model Scale Configuration: Dynamically adjust the number of models in each layer according to data volume and query complexity (for example: 1 model in the top layer, 100 models in the middle layer, 10,000 models in the bottom layer).

[0058] 3. Model Training and Optimization Training Data Generation: Build a training set based on historical queries, with the input being key values (Key) and the output being the storage location (Position) of the corresponding data.

[0059] Loss Function Design: Adopt weighted mean squared error (WMSE), assign higher weights to high-frequency query key values, and optimize the prediction accuracy of the model for hot data.

[0060] Distributed Training: Use a GPU cluster to accelerate the model training for large-scale data sets, and improve the convergence speed through dynamic learning rate adjustment (such as Adam optimizer).

[0061] Model Compression: Quantize (INT8) or prune the bottom-level models to reduce inference latency and adapt to edge device deployment.

[0062] 4. Index Integration and Query Routing Model Embedding into Storage Engine: Compile the trained RMI model into a lightweight inference module (such as TensorRT engine) and integrate it into the database kernel.

[0063] Query Routing Mechanism: After the user submits a query key value, the RMI model predicts its location layer by layer: Key → Top - level model → Middle - level model → Bottom - level model → Predicted location Directly access the storage (such as memory pages or SSD blocks) according to the predicted location, reducing the disk seek or cache jump overhead of the traditional B + tree.

[0064] Fault - tolerance mechanism: If the deviation between the model - predicted location and the actual data exceeds a threshold (such as ±10%), trigger binary search for error correction and record the error samples for model iterative update.

[0065] 5. Performance optimization and hardware adaptation CPU / GPU heterogeneous acceleration: Use SIMD instructions (such as AVX - 512) to accelerate model inference, and control the single - prediction latency within 1 μs.

[0066] For large - scale batch queries, enable GPU parallel prediction (such as CUDA kernel functions), and the throughput is increased by more than 10 times.

[0067] Storage alignment optimization: According to the predicted location distribution, prefetch adjacent data blocks into the CPU cache (such as L1 / L2), and the cache hit rate is increased to more than 95%.

[0068] Power consumption control: Adopt model sparsification technology, turn off some model branches at low load, and reduce the power consumption by 30%.

[0069] Technical advantages and effects Query efficiency: Compared with the traditional B + tree, the point - query latency is reduced by 60% (from 100 μs to 40 μs), and the range - query throughput is increased by 3 times.

[0070] Storage overhead: The model volume is only 1 / 5 of the traditional index (for example: 100 GB of data corresponds to a 20 - MB model).

[0071] Dynamic adaptability: Support online update, and the model self - optimization time is shortened from the hour - level to the minute - level after the data distribution changes suddenly.

[0072] Through the above steps, machine - learning indexes (such as RMI) can significantly improve the index performance in high - concurrency and dynamic - data scenarios while maintaining transaction consistency, and are applicable to industrial Internet scenarios such as the Internet of Things and real - time analysis.

[0073] In terms of retrieval, in some embodiments, a hash algorithm can be used to generate a fixed-length hash value by performing a hash operation on the key features of the data. When searching, the hash value of the target data can be quickly compared and searched in the storage area, which is particularly suitable for accurate search scenarios and can quickly locate the target data in massive data. Pre-retrieval and post-retrieval methods can also be introduced, such as re-ordering retrieval results, using LLM to generate pseudo documents, etc., to improve the retrieval quality. The indexing process can also be optimized by sliding windows, fine-grained block division, metadata enhancement and other methods to improve the retrieval effect. Multi-stage retrieval can also be performed according to different types and requirements of questions to obtain more accurate retrieval results. The quality of indexed data can also be improved, including enhancing data granularity, optimizing index structure, adding metadata, alignment optimization and hybrid retrieval. It can also break the limitations of the traditional RAG framework, provide a more flexible organization method, allow modules and processes to be adjusted according to specific problems, and introduce new modules such as search modules, memory modules, and additional generation modules to expand the functions of RAG.

[0074] In order to enhance data security, in some embodiments, in the identity authentication link, first, a multi-factor identity authentication method is adopted. In addition to the conventional username and password login, fingerprint recognition, facial recognition, SMS verification code and other methods are combined to increase the accuracy and reliability of user identity confirmation, and prevent illegal users from stealing a single password and other means to impersonate others to access data. Second, an identity authentication system based on digital certificates is established, and digital certificates are issued to legitimate users. The certificates contain the user's identity information and public key, etc. During login verification, the user's identity is confirmed by verifying the legitimacy of the digital certificate and encrypted interaction with the server, effectively resisting the risks of network attacks and forged identities.

[0075] In terms of permission control, first, implement role-based permission access control (RBAC), define different roles such as administrators, ordinary employees, data analysts, etc. according to different positions and different business needs within the enterprise or system, and assign corresponding data access rights to each role. For example, administrators can add, delete, modify and query all data, while ordinary employees can only view data within a specific range. Through this role division, the scope of data access can be finely controlled. Second, perform dynamic permission management, and update user permissions in real time according to factors such as changes in business processes and adjustments to data importance. For example, when the sensitivity of data increases, the user group that can access the data can be promptly reduced to ensure that the data is always under safe and controllable permission management.

[0076] To enhance the user experience, in some embodiments, in terms of a user-friendly data retrieval interface, first, the interface design follows the principle of simplicity and intuitiveness, with a reasonable layout. The retrieval box is placed in a prominent position, and clear operation prompts are provided. For example, sample retrieval statements are default displayed in the retrieval box to guide users to correctly enter retrieval conditions, enabling users without much professional knowledge to easily start the retrieval operation. Second, visual interactive elements are adopted. For example, selectable retrieval fields are shown through a drop-down menu, and the classification statistics of retrieval results are shown through charts, etc., enabling users to more intuitively understand the data and retrieval functions, facilitating them to quickly find the data they want without having to grope through complex text descriptions and commands.

[0077] In terms of providing multiple retrieval methods, first, in addition to supporting conventional keyword retrieval, a fuzzy retrieval function is also provided. When users are not very sure about the exact retrieval terms, relevant data results can also be obtained by entering partial keywords or retrieval conditions with wildcards, expanding the flexibility and coverage of the retrieval. Second, combined retrieval is supported, allowing users to simultaneously enter multiple retrieval conditions and define the relationship between these conditions through logical operators (such as "AND", "OR", "NOT", etc.), meeting the complex and precise data screening needs of users, enabling users to retrieve data from multiple perspectives according to their personalized needs, and enhancing the accuracy and effectiveness of the retrieval.

[0078] As Figure 2 shown, the following is a data rapid retrieval system based on active identification provided by an embodiment of the present disclosure. A data rapid retrieval system based on active identification and a data rapid retrieval method based on active identification in the above embodiments belong to the same inventive concept. Details not described in detail in the embodiment of the data rapid retrieval system based on active identification can refer to the embodiment of the data rapid retrieval method based on active identification above.

[0079] A data rapid retrieval system based on active identification includes an active identification module, an identification parsing module, a data indexing module, and a data retrieval module. Among them, the active identification module is configured to assign an industrial Internet of Things identification to each device terminal as a unique digital identity; the identification parsing module is configured to parse the industrial Internet of Things identification to obtain the data record of the device terminal; the data indexing module is configured to construct an attribute index structure according to the data record. The attribute index structure includes a two-layer B-tree data structure. The first layer uses a B-tree to group all data records according to attributes, and the second layer separately constructs a B+ tree for each other attribute in the data records grouped in the first layer; the data retrieval module is configured to obtain retrieval conditions and perform data retrieval based on the attribute index structure to obtain an attribute retrieval result.

[0080] In some embodiments, the data indexing module is also capable of performing word segmentation on the text data in the data records, splitting the continuous text into individual keywords; using the inverted index algorithm to establish an inverted index structure. When establishing the inverted index structure, it traverses the text data in all data records, associates each keyword with the corresponding data record, and stores them in the index.

[0081] Correspondingly, after obtaining the retrieval condition, the data retrieval module can parse the retrieval condition, extract the attribute retrieval condition, keyword retrieval condition, and logical operator in the retrieval condition; for the attribute retrieval condition, perform data retrieval based on the attribute index structure to obtain the attribute retrieval result; for the keyword retrieval condition, perform data retrieval based on the inverted index structure to obtain the keyword retrieval result; according to the rules of the logical operator, merge the attribute retrieval condition and the keyword retrieval condition to obtain the final retrieval result.

[0082] An embodiment of the present application also proposes a terminal, including a memory, a processor, a communication unit, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of a data fast retrieval method based on active identification; the memory, the processor, and the communication unit communicate through one or more buses.

[0083] The processor may include one or more processing units. For example, the processor may include a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0084] Among them, the processor may be the nerve center and command center of the terminal. The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching instructions and executing instructions.

[0085] The memory is used to store the execution instructions of the processor. The memory can be implemented by any type of volatile or non-volatile storage terminal or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. When the execution instructions in the memory are executed by the processor, the terminal can execute some or all of the steps in the above-mentioned embodiments of the data fast retrieval method based on active identification.

[0086] The wireless communication function of the electronic device can be implemented by an antenna, a wireless communication module, a modulation and demodulation processor, a baseband processor, etc.

[0087] The wireless communication module can provide solutions for wireless communication including wireless local area network, Bluetooth, global navigation satellite system, frequency modulation, near field communication technology, infrared technology, etc. applied to the electronic device.

[0088] This embodiment also proposes a storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of any one of the above-mentioned data fast retrieval methods based on active identification.

[0089] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for rapid data retrieval based on active identification, characterized in that: The following steps are involved: Assign an industrial Internet ID to each device terminal as a unique digital identity; Parse the industrial Internet identifier and obtain the data record of the device terminal. Each data record contains multiple attribute information of the device terminal. Each attribute information has a corresponding data item, and each data item has unique identification information. Build an attribute index structure based on data records. The attribute index structure includes two layers of B-tree data structures. The first layer uses B-tree to group all data records according to attributes. The second layer builds a separate B+ tree for each other attribute in the data records after the first layer of grouping. Get the search conditions, perform data search based on the attribute index structure, and obtain the attribute search results.

2. The method for rapid data retrieval based on active identification according to claim 1, characterized in that: After constructing the attribute index structure, the attribute index structure is split into multiple sub-attribute index structures, and the multiple sub-attribute index structures are dispersedly stored on several computing nodes. Each computing node can retrieve data based on its own stored sub-attribute index structure and transmit the calculation results to the terminal node for aggregation.

3. The method for rapid data retrieval based on active identification according to claim 2 is characterized in that: The sub-attribute index structure whose average access frequency exceeds a preset frequency value is stored in the same computing node, a close computing node, or a computing node that is closer to the terminal node.

4. The method for rapid data retrieval based on active identification according to claim 1, characterized in that: In the second-level B+ tree of the attribute index structure, the leaf nodes of the B+ tree whose data volume exceeds a preset value are optimized using Bloom filters, and / or the leaf nodes of the B+ tree that processes time series data are optimized using linear models.

5. The method for rapid data retrieval based on active identification according to claim 1 is characterized in that: After the attribute index structure is constructed according to the data records, when the data records are added, the attribute information in the newly added data records is identified, and the insertion position of the newly added data records in the constructed attribute index structure is determined based on the attribute information, and the corresponding index items are inserted; When a data record is deleted or modified, identification information of the deleted or modified data record is obtained, and based on the identification information of the deleted or modified data record, the corresponding index item is located in the constructed attribute index structure and deleted or modified.

6. The method for rapid data retrieval based on active identification according to claim 1 is characterized in that: The data record also includes text data. After acquiring the data record of the device terminal, the method further includes: Perform word segmentation on the text data in the data records, dividing the continuous text into independent keywords; An inverted index algorithm is used to establish an inverted index structure. When establishing the inverted index structure, the text data in all data records are traversed, each keyword is associated with the corresponding data record, and stored in the index; After obtaining the search conditions, data is searched based on the inverted index structure to obtain keyword search results.

7. The method for rapid data retrieval based on active identification according to claim 6 is characterized in that: After obtaining the search conditions, the search conditions are parsed to extract the attribute search conditions, keyword search conditions and logical operators in the search conditions; For attribute retrieval conditions, data is retrieved based on the attribute index structure to obtain attribute retrieval results; For keyword search conditions, data is searched based on the inverted index structure to obtain keyword search results; According to the rules of logical operators, the attribute search conditions and keyword search conditions are combined to obtain the final search results.

8. A data rapid retrieval system based on active identification, characterized in that: It includes active identification module, identification resolution module, data indexing module and data retrieval module; The active identification module is configured to assign an industrial Internet identification as a unique digital identity to each device terminal; The identification resolution module is configured to resolve the industrial Internet identification and obtain the data record of the device terminal; The data index module is configured to construct an attribute index structure according to the data records. The attribute index structure includes two layers of B-tree data structures. The first layer uses the B-tree to group all data records according to the attributes. The second layer separately builds a B+ tree for each other attribute in the data records after the first layer grouping. The data retrieval module is configured to obtain retrieval conditions, perform data retrieval based on the attribute index structure, and obtain attribute retrieval results.

9. A terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of a method for rapid data retrieval based on active identification as described in any one of claims 1 to 7 are implemented.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a method for rapid data retrieval based on active identification as described in any one of claims 1 to 7 are implemented.