Data retrieval method and device, computer program product and readable storage medium

By generating entity type tags for data in a distributed storage system and updating them adaptively, the problems of low retrieval accuracy and efficiency in distributed storage systems are solved, and stable and efficient retrieval is achieved even when the logical location of data changes.

CN120821743APending Publication Date: 2025-10-21JINAN INSPUR DATA TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511156316.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

The lack of mature data retrieval methods in distributed storage systems results in poor retrieval accuracy and low efficiency, and makes it impossible to cope with the impact of data logical location migration.

Method used

Generate entity type labels for data in a distributed storage system, identify data types using a pre-trained entity recognition model, determine the target entity type when a query condition is received, retrieve data that meets the condition from the stored data of the target entity type, and adaptively update the labels based on dynamically changing features.

Benefits of technology

It improves the accuracy and efficiency of data retrieval, maintains retrieval stability when the logical location of data changes, and allows for flexible retrieval based on specific content under entity type.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120821743A_ABST
    Figure CN120821743A_ABST
Patent Text Reader

Abstract

The invention discloses a data retrieval method and device, a computer program product and a readable storage medium, belongs to the field of data retrieval, is used for optimizing a data retrieval process through entity type tags, and solves the problems of poor data retrieval precision and low efficiency. According to the method, the entity type label can be generated for any storage data in the distributed storage system, when the data query condition is received, the entity type to which the data query condition belongs can be determined to serve as the target entity type, and the data query condition is determined from the storage data of which the entity type label is the target entity type. The storage data meeting the data query condition is used as the target query data, on one hand, the data retrieval process is not influenced by the change of the logic position of the storage data, the data retrieval precision can be improved, on the other hand, the user can perform data retrieval through the specific content under the entity type, and the retrieval efficiency and flexibility can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data retrieval, and in particular to a data retrieval method, device, computer program product and readable storage medium. Background Art

[0002] The amount of data stored in a distributed storage system is huge, and the distributed storage system may migrate the logical location of the data. In this case, the relevant technology lacks a mature data retrieval method, resulting in poor retrieval accuracy and low retrieval efficiency when retrieving data in the distributed storage system.

[0003] Therefore, how to provide a solution to the above technical problems is a problem that those skilled in the art need to solve at present. Summary of the Invention

[0004] The purpose of the present invention is to provide a data retrieval method, device, computer program product and readable storage medium. For any stored data in a distributed storage system, an entity type label is generated. When a data query condition is received, the entity type to which the data query condition belongs can be determined as the target entity type, and the stored data that meets the data query condition from the stored data with the entity type label as the target entity type is used as the target query data. On the one hand, the data retrieval process is not affected by changes in the logical location of the stored data, which can improve the data retrieval accuracy. On the other hand, users can search for data through the specific content under the entity type, which is conducive to improving retrieval efficiency and flexibility.

[0005] To solve the above technical problems, the present invention provides a data retrieval method, comprising:

[0006] For any stored data in the distributed storage system, an entity type tag is added to the stored data; the entity type tag is used to represent the entity type to which the stored data belongs;

[0007] Upon receiving the data query condition, determining the entity type to which the data query condition belongs as the target entity type;

[0008] From the stored data whose entity type label is the target entity type, the stored data that meets the data query condition is determined as the target query data.

[0009] On the other hand, for any stored data in the distributed storage system, labeling the stored data with an entity type tag includes:

[0010] For any stored data in the distributed storage system, standardize the stored data according to the preset format requirements to obtain standardized data;

[0011] For any stored data in the distributed storage system, identifying the entity type of the stored data;

[0012] For any standardized data, the entity type of the storage data corresponding to the standardized data is bound to the standardized data as an entity type tag and stored.

[0013] On the other hand, for any stored data in the distributed storage system, identifying the entity type of the stored data includes:

[0014] Identifying the entity type of the stored data using a pre-trained entity recognition model;

[0015] The entity recognition model includes: an entity recognition model obtained by pre-training through a semantic alignment algorithm.

[0016] On the other hand, the data retrieval method further includes:

[0017] In response to an operation instruction for target data, determining whether the target data has a distributed lock; the target data is standardized data specified by the operation instruction;

[0018] If the distributed lock is not available, add a distributed lock to the target data and execute the operation instruction;

[0019] If a distributed lock is available, the sender of the operation instruction will be notified that the target data is locked.

[0020] On the other hand, executing the operation instruction includes:

[0021] Execute the operation instruction and start timing;

[0022] When the timing time reaches a preset time threshold, determining whether the target data has successfully executed the operation required by the operation instruction;

[0023] If successfully executed, the target operation applied during the timing period is made effective;

[0024] If the execution is unsuccessful, the operation applied to the target operation during the timing period is canceled.

[0025] On the other hand, after determining the stored data satisfying the data query condition from the stored data whose entity type label is the target entity type and using the stored data as the target query data, the data retrieval method further includes:

[0026] sorting the target query data;

[0027] Paging the target query data according to the sorting;

[0028] Push the identification information of the target query data on the homepage;

[0029] In response to the page change instruction, the identification information of the target query data in the target page is pushed;

[0030] In response to the identification information selection instruction, the target query data corresponding to the selected identification information is pushed.

[0031] On the other hand, when the data query condition is received, the entity type to which the data query condition belongs is determined, and the target entity type includes:

[0032] Upon receiving a data query condition, determining whether the data query condition is a mixed type data query condition; the mixed type data query condition is a data query condition composed of multiple sub-conditions;

[0033] If the data query condition is a mixed type, for any sub-condition in the data query condition, determine the entity type to which the sub-condition belongs as the candidate entity type;

[0034] If there is stored data with multiple entity type tags in the distributed storage system, then for mixed-type data query conditions, multiple target entity types are determined from each candidate entity type of the data query condition;

[0035] If there is no stored data with multiple entity type tags in the distributed storage system, then for the mixed-type data query condition, a target entity type is determined from each candidate entity type of the data query condition;

[0036] If the data query condition is not a mixed type, the entity type to which the data query condition belongs is used as the target entity type.

[0037] On the other hand, for mixed-type data query conditions, when multiple target entity types are determined, the stored data that satisfies the data query condition from the stored data whose entity type label is the target entity type is determined as the target query data, including:

[0038] Construct an association weight matrix between multiple target entity types, where the weight values ​​are determined based on the co-occurrence frequency of entity types in historical queries and the business logic relevance;

[0039] Prioritize multiple target entity types based on an association weight matrix;

[0040] Search the stored data that meets the corresponding sub-conditions under each entity type in order of priority, and calculate the comprehensive matching score between the stored data that meets the corresponding sub-conditions and all sub-conditions;

[0041] The stored data whose comprehensive matching score exceeds the preset score threshold is selected as the target query data.

[0042] On the other hand, determining the entity type to which the data query condition belongs includes:

[0043] Perform semantic analysis on data query conditions to extract core keywords and context-related information;

[0044] Perform fuzzy matching on the core keywords and the preset entity type vocabulary, and calculate the matching confidence level based on the contextual information.

[0045] If the matching confidence exceeds the preset confidence threshold, the corresponding entity type is used as the target entity type;

[0046] If the preset confidence threshold is not exceeded, a candidate list of entity types is pushed for the user to select, and the user's selection results are recorded to optimize the vocabulary matching rules.

[0047] On the other hand, the data retrieval method further includes:

[0048] For any stored data, obtaining dynamic change characteristics of the stored data; wherein the dynamic change characteristics include access frequency and content modification frequency of the stored data;

[0049] Based on the dynamic change characteristics of the stored data, the entity type label of the stored data is adaptively updated.

[0050] On the other hand, the adaptive updating of the entity type label of the stored data based on the dynamic change characteristics of the stored data includes:

[0051] Monitor the access frequency and content modification frequency of stored data in real time, and re-identify the entity type of the stored data when either the access frequency or the content modification frequency reaches a corresponding threshold;

[0052] If the re-identified entity type of the stored data is inconsistent with the original entity type label, a label update suggestion is generated and pushed to the administrator;

[0053] In response to a confirmation instruction from the administrator, the re-identified entity type is used as a new entity type label for the stored data.

[0054] On the other hand, the dynamic change feature also includes changes in the entity type of the associated data;

[0055] Adaptively updating the entity type label of the stored data based on the dynamic change characteristics of the stored data includes:

[0056] When the entity type of the associated data of the stored data changes, verify whether the entity type label of the stored data is accurate;

[0057] If it is inaccurate, the stored data will be marked as pending update and the administrator will be notified.

[0058] On the other hand, the dynamic change characteristics also include the life cycle stages of the stored data; the life cycle stages include the generation period, the active period and the archiving period;

[0059] Adaptively updating the entity type label of the stored data based on the dynamic change characteristics of the stored data includes:

[0060] Determining an update sensitivity coefficient of the entity type tag corresponding to the current life cycle stage of the stored data according to a preset first correspondence, wherein the first correspondence is a correspondence between the life cycle stage and the update sensitivity coefficient of the entity type tag;

[0061] Determine a tag update trigger value based on the product of the quantized value of the dynamic change characteristic of the stored data and the update sensitivity coefficient;

[0062] When the tag update trigger value reaches the preset trigger threshold, the entity type tag re-identification process is started;

[0063] Among them, the life cycle stage is comprehensively determined by storage duration, last access time and business association rules.

[0064] On the other hand, determining the tag update trigger value based on the product of the quantized value of the dynamic change characteristic of the stored data and the update sensitivity coefficient includes:

[0065] Determining the quantitative value of each sub-feature of the dynamic change feature of the stored data, wherein the sub-feature includes at least one of the access frequency of the stored data, the content modification frequency, the entity type change of the associated data, and the life cycle stage;

[0066] Normalize the quantized values ​​of each sub-feature;

[0067] Perform weighted summation on each quantized value after normalization to obtain the quantized value of the dynamic change feature;

[0068] The tag update trigger value is determined based on the product of the quantized value of the dynamic change characteristic of the stored data and the update sensitivity coefficient.

[0069] On the other hand, the determining of the quantized value of each sub-feature in the dynamic change feature of the stored data includes: quantizing the access frequency or the content modification frequency according to the corresponding preset quantization interval and the quantization rule within the preset quantization interval.

[0070] On the other hand, the determining of the quantized value of each sub-feature in the dynamic change feature of the stored data includes:

[0071] For entity type changes of associated data, quantification is performed based on the association level between the associated data of the entity type change and the stored data; the association level is divided according to the business dependency between the associated data and the stored data.

[0072] On the other hand, the determining of the quantized value of each sub-feature in the dynamic change feature of the stored data includes:

[0073] For the life cycle stage, if the life cycle stage is not an archiving period, the quantitative value is determined by the time decay relationship;

[0074] If the life cycle stage is the archiving stage, the product of the initial value and the first preset ratio is used as the quantitative value of the life cycle stage;

[0075] The time decay relationship includes:

[0076] Quantized value = initial value × e (-k×t) +Activity compensation value;

[0077] Wherein, k is the attenuation coefficient, t is the storage duration, and the activity compensation value is the product of the number of visits in the past preset period and the second preset ratio.

[0078] To solve the above technical problems, the present invention further provides a data retrieval device, comprising:

[0079] memory for storing computer programs;

[0080] A processor is configured to implement the steps of the data retrieval method described above when executing the computer program.

[0081] To solve the above technical problems, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above data retrieval method are implemented.

[0082] Beneficial effects: The present invention provides a data retrieval method. Taking into account that stored data has an entity type to which it belongs, and the entity type can be further subdivided into more specific content as a query condition, the present invention can generate an entity type label for any stored data in a distributed storage system. When a data query condition is received, the entity type to which the data query condition belongs can be determined as the target entity type, and the stored data that meets the data query condition from the stored data whose entity type label is the target entity type can be used as the target query data. On the one hand, the data retrieval process is not affected by changes in the logical position of the stored data, and the data retrieval accuracy can be improved. On the other hand, users can perform data retrieval through the specific content under the entity type, which is conducive to improving retrieval efficiency and flexibility.

[0083] The present invention also provides a data retrieval device, a computer program product, and a readable storage medium, which have the same beneficial effects as the above data retrieval method. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the relevant technologies and the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0085] Figure 1 A schematic flow chart of a data retrieval method provided by the present invention;

[0086] Figure 2 A schematic structural diagram of a data retrieval system provided by the present invention;

[0087] Figure 3 A schematic structural diagram of a data retrieval device provided by the present invention;

[0088] Figure 4 A schematic structural diagram of a computer-readable storage medium provided by the present invention. DETAILED DESCRIPTION

[0089] The core of the present invention is to provide a data retrieval method, device, computer program product and readable storage medium. For any stored data in a distributed storage system, an entity type label is generated. When a data query condition is received, the entity type to which the data query condition belongs can be determined as the target entity type, and the stored data that meets the data query condition from the stored data with the entity type label as the target entity type is used as the target query data. On the one hand, the data retrieval process is not affected by changes in the logical location of the stored data, which can improve the data retrieval accuracy. On the other hand, users can search for data through the specific content under the entity type, which is conducive to improving retrieval efficiency and flexibility.

[0090] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0091] Please refer to Figure 1 , Figure 1 A data retrieval method according to the present invention includes:

[0092] S101: For any stored data in the distributed storage system, an entity type tag is added to the stored data; the entity type tag is used to represent the entity type to which the stored data belongs;

[0093] Specifically, taking into account the technical problems in the above background technology, and considering that the stored data has an entity type to which it belongs, and the entity type can be further subdivided into more specific content as query conditions, the embodiment of the present invention intends to label the stored data in the distributed storage system with an entity type, so as to facilitate users to retrieve the stored data through the more specific content under the entity type, thereby improving the retrieval flexibility, retrieval accuracy and efficiency. Based on this consideration, in this step, any stored data in the distributed storage system can first be labeled with an entity type.

[0094] For example, the stored data may contain information such as name and age that belongs to a "person", so the entity type to which the stored data belongs may be a person. In addition, the entity type may also include a country or region (more specific content may include the country name or geographical location and population data, etc.) or a company (more specific content may include the company name, registration time, annual revenue, etc.), etc., and the embodiments of the present invention do not limit this.

[0095] Specifically, the distributed storage type may be of various types, such as a Ceph distributed sorting system, etc., which is not limited in the embodiment of the present invention.

[0096] S102: upon receiving the data query condition, determining the entity type to which the data query condition belongs as the target entity type;

[0097] Specifically, to better illustrate the embodiments of the present invention, please refer to Figure 2 , Figure 2 This is a structural diagram of a data retrieval system provided by the present invention. Users can send data query conditions through the client and retrieve data from the distributed storage system.

[0098] Specifically, since the data query conditions provided by the user are likely to include more specific content under the entity type, in order to match the entity type label of the stored data, in this step, when the data query conditions are received, the entity type to which the data query conditions belong can be determined as the target entity type, so that data can be retrieved from the stored data of the target entity type in subsequent steps.

[0099] S103: Determine, from the stored data whose entity type label is the target entity type, the stored data that meets the data query condition as the target query data.

[0100] Specifically, based on the target entity type determined in the above steps, the stored data with the entity type label of the target entity type can be first filtered out, and then the stored data that meets the data query conditions can be determined from the stored data with the entity type label of the target entity type as the target query data. Specifically, the "stored data with the entity type label of the target entity type" can be traversed, and during the traversal process, it is determined whether each stored data meets the specific content of the data query conditions. For example, the data query conditions can be: the name is John Smith, the age is 20 to 40 years old, etc., which is not limited in this embodiment of the present invention.

[0101] The present invention provides a data retrieval method. Considering that stored data has an entity type to which it belongs, and the entity type can be further subdivided into more specific content as a query condition, the present invention can generate an entity type tag for any stored data in a distributed storage system. When a data query condition is received, the entity type to which the data query condition belongs can be determined as the target entity type, and the stored data that meets the data query condition from the stored data whose entity type tag is the target entity type can be used as the target query data. On the one hand, the data retrieval process is not affected by changes in the logical position of the stored data, and the data retrieval accuracy can be improved. On the other hand, users can perform data retrieval through the specific content under the entity type, which is conducive to improving retrieval efficiency and flexibility.

[0102] Based on the above embodiment:

[0103] As an optional embodiment, for any stored data in the distributed storage system, labeling the stored data with an entity type tag includes:

[0104] For any stored data in the distributed storage system, the stored data is standardized according to the preset format requirements to obtain standardized data;

[0105] For any stored data in the distributed storage system, identify the entity type of the stored data;

[0106] For any standardized data, the entity type of the storage data corresponding to the standardized data is bound and stored as an entity type label with the standardized data.

[0107] Specifically, considering that the format of the original stored data may be non-standard or non-standard, if the stored data can be standardized according to the preset format requirements, it will be beneficial to improve the accuracy of the stored data. Therefore, in an embodiment of the present invention, any stored data in the distributed storage system can be standardized according to the preset format requirements to obtain standardized data. Then, for any stored data in the distributed storage system, the entity type of the stored data can be identified, and then the entity type of the stored data corresponding to the standardized data is bound and stored as an entity type tag with the standardized data, thereby obtaining standardized storage data associated with the entity type tag, which is convenient for subsequent efficient and accurate retrieval.

[0108] The preset format requirements may be of various types, for example, may include not allowing continuous spaces or not allowing characters in a blacklist, etc., which is not limited in the embodiment of the present invention.

[0109] For example, for the raw data {"name":"John Smith","age":30} and {"Name":"JaneDoe","Age":25} stored in a distributed storage system, we first normalize the data to lowercase the field names (for example, converting "Name" to "name"). Next, we use spaCy to identify the entity type of each stored data entry (for example, all are persons). Finally, we bind the normalized data (for example, {"name":"John Smith","age":30}) to its corresponding entity type label (person) using Django's Object-Relational Mapping (ORM) and store it in the database's Entity model.

[0110] Specifically, NLP (Natural Language Processing) libraries are a collection of tools used to implement natural language processing tasks, supporting operations such as entity recognition and semantic analysis of text data. In related solutions, NLP libraries (such as NLTK (Natural Language Toolkit) or spaCy) are primarily used to identify the entity types of stored data. For example, in the technical briefing example, by loading spaCy's pre-trained models (such as en_core_web_sm), the text information in the stored data (such as the name "John Smith") is processed to identify the entities contained therein and their types (for example, "John Smith" corresponds to the entity type "person"). This provides a basis for subsequently labeling the stored data with entity types, thereby supporting efficient data retrieval based on entity types.

[0111] Among them, spaCy is an industrial-grade NLP library, known for its efficient entity recognition, part-of-speech tagging and other functions. It supports a variety of pre-trained models and is suitable for quickly processing text data in practical applications.

[0112] ORM is a programming technology used to establish a mapping relationship between object-oriented programming languages ​​and relational databases, allowing developers to perform database operations by manipulating objects without having to directly write SQL statements. In the data retrieval solution of a distributed storage system, ORM plays a role in data storage and retrieval: by mapping standardized data (such as data processed by entity recognition and semantic alignment) to objects in the code (such as the Entity class in the technical briefing example), data and entity type labels are bound and stored. At the same time, with the help of the query interface provided by ORM (such as the filter method based on DjangoORM in the example), data retrieval operations can be performed efficiently, filtering out target data that meets the entity type labels and query conditions, simplifying the data management process and improving the convenience and efficiency of data operations in the distributed storage system.

[0113] As an optional embodiment, for any stored data in the distributed storage system, identifying the entity type of the stored data includes:

[0114] Identify the entity type of the stored data through the pre-trained entity recognition model;

[0115] The entity recognition model includes: an entity recognition model obtained by pre-training through a semantic alignment algorithm.

[0116] Specifically, considering that the pre-trained model combined with semantic alignment technology can improve the recognition accuracy of entity types, the embodiment of the present invention can identify the entity type of the stored data through a pre-trained entity recognition model, and the entity recognition model includes: an entity recognition model obtained by pre-training through a semantic alignment algorithm, thereby improving the accuracy of entity type recognition, ensuring the reliability of entity type labels, and thus improving retrieval accuracy.

[0117] Of course, in addition to this specific form, "for any stored data in the distributed storage system, identifying the entity type of the stored data" can also be implemented in other ways, and the embodiments of the present invention are not limited here.

[0118] As an optional embodiment, the data retrieval method further includes:

[0119] In response to an operation instruction for target data, determining whether the target data has a distributed lock; the target data is standardized data specified by the operation instruction;

[0120] If there is no distributed lock, add a distributed lock to the target data and execute the operation instruction;

[0121] If a distributed lock is available, the sender of the operation instruction will be notified that the target data is locked.

[0122] Specifically, considering that multiple users in a distributed system may operate the same data at the same time, which may cause data conflicts and inconsistencies, in order to improve data consistency, in an embodiment of the present invention, in response to an operation instruction for the target data, it can be determined whether the target data has a distributed lock. If it does, a distributed lock is added to the target data and the operation instruction is executed. If it has a distributed lock, the sender of the operation instruction is prompted that the target data is locked, thereby limiting the number of users who can operate on the target data at the same time. The exclusivity of data operations is controlled by distributed locks, data conflicts caused by concurrent operations are avoided, and consistency during data operations is guaranteed.

[0123] As an optional embodiment, executing the operation instruction includes:

[0124] Execute the operation instruction and start timing;

[0125] When the timing reaches a preset time threshold, it is determined whether the target data has successfully executed the operation required by the operation instruction;

[0126] If successfully executed, the target operation applied during the timing period will take effect;

[0127] If the execution is not successful, the operation applied by the target operation during the timing period is undone.

[0128] Specifically, considering that data operations may fail due to timeouts and other reasons, it is necessary to ensure the integrity and consistency of the operation results. Therefore, in an embodiment of the present invention, timing can be started when the operation instruction is executed, and when the timing time reaches a preset time threshold, it is determined whether the target data has successfully executed the operation required by the operation instruction. If it is successfully executed, the operation applied to the target operation during the timing period will take effect. If it is not successfully executed, the operation applied to the target operation during the timing period will be revoked, thereby monitoring the operation status through timing. If successful, the operation will take effect, and if failed, the operation will be revoked, thereby avoiding data inconsistency caused by partial execution and ensuring data integrity.

[0129] The preset time threshold can be set flexibly and autonomously, and the embodiment of the present invention does not limit this.

[0130] For example, when performing an update operation on the data of name="John Smith", a timer is started (for example, the preset time threshold is 30 seconds). If the age field is successfully updated from 30 to 35 within 30 seconds, the transaction is committed and the update takes effect. If the update is not completed within 30 seconds (for example, due to a network failure), the update operation during the timer is canceled and the age remains at 30.

[0131] Of course, in addition to this specific form, "executing the operation instruction" can also be implemented in other ways, which are not limited in the embodiment of the present invention.

[0132] As an optional embodiment, after determining stored data that meets the data query condition from the stored data whose entity type label is the target entity type and using the stored data as the target query data, the data retrieval method further includes:

[0133] Sort the target query data;

[0134] Paginate the target query data by sorting;

[0135] Push the identification information of the target query data on the home page;

[0136] In response to the page change instruction, the identification information of the target query data in the target page is pushed;

[0137] In response to the identification information selection instruction, the target query data corresponding to the selected identification information is pushed.

[0138] Specifically, considering that the number of search results may be huge, if all the search results are pushed directly, it will consume a lot of resources and reduce efficiency. Therefore, it is necessary to optimize the display method to improve the user experience. In the embodiment of the present invention, the target query data can be sorted; the target query data can be paged according to the sorting; the identification information of the target query data on the home page can be pushed; in response to the page change instruction, the identification information of the target query data in the target page is pushed; in response to the identification information selection instruction, the target query data corresponding to the selected identification information is pushed.

[0139] Among them, on the one hand, the embodiment of the present invention can push the target query data in pages, thereby improving the data retrieval efficiency and instantaneous resource consumption. On the other hand, the identification information of the target query data can be pushed to the user, and in response to the identification information selection instruction, the target query data corresponding to the selected identification information can be pushed, so that the identification information with a lower data volume can be initially pushed, and the target query data corresponding to the selected identification information can be pushed when needed, which is conducive to further improving the data retrieval efficiency and instantaneous resource consumption.

[0140] Specifically, for example, after retrieving multiple records that meet the criteria (target query data), they are sorted in ascending order by the age field and paginated (e.g., displaying 10 records per page). The identifiers for the first page (e.g., the id list [1, 2, ..., 10]) are pushed to the user. When the user triggers a page change (e.g., viewing page 2), the identifiers for page 2 ([11, 12, ..., 20]) are pushed. When the user selects identifier 1, the corresponding complete data set {"name":"John Smith","age":30} is pushed.

[0141] As an optional embodiment, when a data query condition is received, an entity type to which the data query condition belongs is determined, and the target entity type includes:

[0142] Upon receiving the data query condition, determining whether the data query condition is a mixed type data query condition; a mixed type data query condition is a data query condition composed of multiple sub-conditions;

[0143] If the data query condition is of mixed types, for any sub-condition in the data query condition, determine the entity type to which the sub-condition belongs as the candidate entity type;

[0144] If there is stored data with multiple entity type tags in the distributed storage system, then for mixed-type data query conditions, multiple target entity types are determined from each candidate entity type of the data query condition;

[0145] If there is no stored data with multiple entity type tags in the distributed storage system, then for mixed-type data query conditions, a target entity type is determined from each candidate entity type of the data query condition;

[0146] If the data query condition is not a mixed type, the entity type to which the data query condition belongs is used as the target entity type.

[0147] Specifically, considering that a query condition may contain multiple sub-conditions (mixed types), and in a distributed storage system, stored data may or may not have multiple entity type tags, if it has multiple entity type tags, multiple target entity types can be determined for data retrieval; if it does not have multiple entity type tags, a single target entity type can be determined for data retrieval. Therefore, in an embodiment of the present invention, when a data query condition is received, if it is a mixed type data query condition, for any sub-condition in the data query condition, the entity type to which the sub-condition belongs is determined as a candidate entity type; if there is stored data with multiple entity type tags in the distributed storage system, then for the mixed type data query condition, multiple target entity types are determined from each candidate entity type of the data query condition (for example, each candidate entity type is used as the target entity type); if there is no stored data with multiple entity type tags in the distributed storage system, then for the mixed type data query condition, one target entity type is determined from each candidate entity type of the data query condition; if it is not a mixed type data query condition, the entity type to which the data query condition belongs is used as the target entity type.

[0148] Specifically, based on the above operations, the required number of target entity types can be adaptively determined for retrieval according to the characteristic of "whether the data stored in the distributed system has multiple entity type labels". Even if the number of entity types involved in the data query conditions entered by the user does not match the "number of entity types involved" of the data stored in the distributed storage system, data retrieval can be carried out smoothly.

[0149] For example, when receiving a mixed-type query condition such as "name is John and organization is A," its sub-conditions are determined to correspond to two candidate entity types: "person" and "organization." If the distributed storage system contains data with multiple entity type tags, both "person" and "organization" are determined as target entity types. If the distributed storage system does not contain data with multiple entity type tags, a single candidate entity type can be selected as the target entity type, for example, "person" is selected as the target entity type.

[0150] As an optional embodiment, for mixed-type data query conditions, when multiple target entity types are determined, stored data that meets the data query conditions is determined from the stored data with the entity type label of the target entity type, and the target query data includes:

[0151] Construct an association weight matrix between multiple target entity types, where the weight values ​​are determined based on the co-occurrence frequency of entity types in historical queries and the business logic relevance;

[0152] Prioritize multiple target entity types based on an association weight matrix;

[0153] Search the stored data that meets the corresponding sub-conditions under each entity type in order of priority, and calculate the comprehensive matching score between the stored data that meets the corresponding sub-conditions and all sub-conditions;

[0154] The stored data whose comprehensive matching score exceeds the preset score threshold is selected as the target query data.

[0155] Specifically, considering that for mixed queries of multiple target entity types, the relevance of the results can be guaranteed by weighing the association relationships of each entity type, in an embodiment of the present invention, an association weight matrix between multiple target entity types is first constructed, and then the multiple target entity types are prioritized according to the association weight matrix. The stored data that meets the corresponding sub-conditions under each entity type are retrieved in order of priority, and the comprehensive matching score of the stored data that meets the corresponding sub-conditions and all sub-conditions is calculated. Finally, the stored data whose comprehensive matching score exceeds the preset score threshold is screened as the target query data, which can improve the accuracy and relevance of the query results.

[0156] A specific example involves constructing an association weight matrix with the target entity types of "user" and "item." The "user-item" weight is set to 0.8 based on the co-occurrence frequency (60%) and business logic relevance (40%) in historical queries. Prioritizing "user" > "item," the search begins with user data that meets the "city of Beijing" requirement, followed by product data that meets the "item A" requirement. For each piece of data, the combined matching score for the two sub-conditions is calculated (e.g., a user data matching score of 80 points + an item data matching score of 70 points yields a combined score of 75 points). Data with a score exceeding 70 is selected as the target query data.

[0157] As an optional embodiment, determining the entity type to which the data query condition belongs includes:

[0158] Perform semantic analysis on data query conditions to extract core keywords and context-related information;

[0159] Perform fuzzy matching on the core keywords and the preset entity type vocabulary, and calculate the matching confidence level based on the contextual information.

[0160] If the matching confidence exceeds the preset confidence threshold, the corresponding entity type is used as the target entity type;

[0161] If the preset confidence threshold is not exceeded, a candidate list of entity types is pushed for the user to select, and the user's selection results are recorded to optimize the vocabulary matching rules.

[0162] Specifically, considering that the query conditions may be expressed in a non-standard way (such as colloquial expressions), which may lead to difficulties in entity type recognition, semantic analysis, fuzzy matching and confidence judgment can be used to improve recognition accuracy and fault tolerance, and the vocabulary can be optimized in combination with user feedback to further improve subsequent recognition effects. Therefore, in an embodiment of the present invention, the data query conditions can be semantically analyzed, core keywords and context-related information can be extracted, the core keywords can be fuzzy matched with the preset entity type vocabulary, and the matching confidence can be calculated in combination with the context-related information; if the matching confidence exceeds the preset confidence threshold, the corresponding entity type is used as the target entity type; if it does not exceed the preset confidence threshold, a list of entity type candidates is pushed for user selection, and the user selection results are recorded to optimize the vocabulary matching rules.

[0163] As an optional embodiment, the data retrieval method further includes:

[0164] For any stored data, obtain dynamic change characteristics of the stored data; wherein the dynamic change characteristics include access frequency and content modification frequency of the stored data;

[0165] Based on the dynamic change characteristics of the stored data, the entity type label of the stored data is adaptively updated.

[0166] Specifically, considering that the access and modification of stored data may cause its entity type to change, static labels will reduce retrieval accuracy. Therefore, in order to further improve data retrieval accuracy, in an embodiment of the present invention, for any stored data, the dynamic change characteristics of the stored data (the access frequency and content modification frequency of the stored data) can be obtained, and then based on the dynamic change characteristics of the stored data, the entity type label of the stored data can be adaptively updated, thereby improving the accuracy of the entity type label of the stored data, thereby improving the accuracy of data retrieval.

[0167] As an optional embodiment, based on the dynamic change characteristics of the stored data, adaptively updating the entity type label of the stored data includes:

[0168] Monitor the access frequency and content modification frequency of stored data in real time, and re-identify the entity type of the stored data when either the access frequency or the content modification frequency reaches a corresponding threshold;

[0169] If the re-identified entity type of the stored data is inconsistent with the original entity type label, a label update suggestion is generated and pushed to the administrator;

[0170] In response to a confirmation instruction from the administrator, the re-identified entity type is used as a new entity type label for the stored data.

[0171] As an optional embodiment, the dynamic change feature also includes a change in the entity type of the associated data;

[0172] Based on the dynamic change characteristics of stored data, the entity type labels of the adaptively updated stored data include:

[0173] When the entity type of the associated data of the stored data changes, verify whether the entity type label of the stored data is accurate;

[0174] If it is inaccurate, the stored data will be marked as pending update and the administrator will be notified.

[0175] Specifically, on the one hand, considering that when either of the "access frequency and content modification frequency" is high, there may be a need to update the entity type label. On the other hand, when the entity type of the associated data of the stored data changes, there may also be a need to update the entity type label. Therefore, in an embodiment of the present invention, the access frequency and content modification frequency of the stored data can be monitored in real time. When either the access frequency or the content modification frequency reaches the corresponding threshold, the entity type of the stored data is re-identified, and after the administrator confirms the re-identified entity type, the re-identified entity type is used as the new entity type label of the stored data. On the other hand, when the entity type of the associated data of the stored data changes, the entity type label of the stored data can be verified to be accurate; if it is inaccurate, the stored data is marked as a label to be updated state (the entity type of the stored data can also be re-identified; if the re-identified entity type of the stored data is inconsistent with the original entity type label, a label update suggestion is generated and pushed to the administrator) and the administrator is prompted, thereby ensuring that the entity label type of the stored data can be updated in a timely and accurate manner.

[0176] For example, the system monitors the access frequency of "name="John Smith" in real time (reaching 100 visits per day, exceeding a threshold of 80 visits). A pre-trained model is then used to re-identify the entity type. If the result changes from "person" to "company," a label update suggestion is generated and pushed to the administrator. After the administrator confirms, the label is updated to "company." When the entity type of related data (e.g., stored data with the same entity type generated by the same user at the same time) changes, the data label is verified for accuracy. If inaccurate, it is marked as "label pending update" and the administrator is notified.

[0177] As an optional embodiment, the dynamic change feature further includes the life cycle stage of the stored data; the life cycle stage includes a generation period, an active period, and an archiving period;

[0178] Based on the dynamic change characteristics of stored data, the entity type labels of the adaptively updated stored data include:

[0179] Determining an update sensitivity coefficient of the entity type tag corresponding to the current life cycle stage of the stored data according to a preset first correspondence, where the first correspondence is a correspondence between the life cycle stage and the update sensitivity coefficient of the entity type tag;

[0180] Determine a tag update trigger value based on the product of the quantized value of the dynamic change characteristic of the stored data and the update sensitivity coefficient;

[0181] When the tag update trigger value reaches the preset trigger threshold, the entity type tag re-identification process is started;

[0182] Among them, the life cycle stage is comprehensively determined by storage duration, last access time and business association rules.

[0183] Specifically, considering that the label update requirements of stored data in different life cycle stages (generation period, active period, and archiving period) are different, the trigger value can be calculated through the correspondence between the life cycle stage and the update sensitivity coefficient, thereby achieving reasonable triggering updates and reducing unnecessary identification operations. Therefore, in an embodiment of the present invention, the update sensitivity coefficient of the entity type label corresponding to the current life cycle stage of the stored data can be determined according to the preset first correspondence, and the label update trigger value is determined based on the product of the quantized value of the dynamic change characteristics of the stored data and the update sensitivity coefficient; when the label update trigger value reaches the preset trigger threshold, the re-identification process of the entity type label is started.

[0184] For example, the lifecycle phase of stored data is determined to be "active" based on the storage duration (three months), the last access time (one day ago), and business association rules. Based on the first correspondence (the update sensitivity coefficient for the active period is 0.8), combined with the quantized value of 50 for the dynamic change feature, the tag update trigger value is calculated as 50 × 0.8 = 40. Because 40 exceeds the preset trigger threshold of 30, the system initiates the entity type tag re-identification process.

[0185] As an optional embodiment, determining the tag update trigger value based on the product of the quantized value of the dynamic change characteristic of the stored data and the update sensitivity coefficient includes:

[0186] Determining the quantitative value of each sub-feature of the dynamic change feature of the stored data, wherein the sub-feature includes at least one of the access frequency of the stored data, the content modification frequency, the entity type change of the associated data, and the life cycle stage;

[0187] Normalize the quantized values ​​of each sub-feature;

[0188] Perform weighted summation on each quantized value after normalization to obtain the quantized value of the dynamic change feature;

[0189] The tag update trigger value is determined based on the product of the quantized value of the dynamic change characteristic of the stored data and the update sensitivity coefficient.

[0190] Specifically, considering that the quantized values ​​of the dynamic change feature can be determined efficiently and accurately by normalizing and weighted summing the quantized values ​​of each sub-feature, and then the label update trigger value can be obtained, the embodiment of the present invention can respectively determine the quantized value of each sub-feature in the dynamic change feature of the stored data, and normalize the quantized value of each sub-feature; perform weighted summation on each quantized value after normalization to obtain the quantized value of the dynamic change feature.

[0191] As an optional embodiment, respectively determining the quantized value of each sub-feature in the dynamic change feature of the stored data includes:

[0192] The access frequency or content modification frequency is quantified according to the corresponding preset quantization interval and the quantization rule within the preset quantization interval.

[0193] Specifically, considering that the access frequency or content modification frequency can be quantized efficiently and accurately through the corresponding preset quantization interval and the quantization rules within the preset quantization interval, the embodiment of the present invention can quantize the access frequency or content modification frequency according to the corresponding preset quantization interval and the quantization rules within the preset quantization interval.

[0194] For example, the preset access frequency quantization intervals and rules are: 0-10 times per month corresponds to a quantization value of 10, 11-30 times corresponds to 30, 31-50 times corresponds to 50, and 50 times or more corresponds to 80. The access frequency of a certain stored data in the past 30 days was 35 times, and the quantization value according to the rule is 50.

[0195] As an optional embodiment, respectively determining the quantized value of each sub-feature in the dynamic change feature of the stored data includes:

[0196] For changes in the entity type of associated data, quantification is performed based on the association level between the associated data and the stored data; the association level is divided according to the business dependency between the associated data and the stored data.

[0197] Specifically, considering that changes in the type of associated data at different association levels have different impacts on stored data, the embodiments of the present invention divide the association levels according to business dependencies and quantify changes in the entity type of the associated data. This can accurately reflect the impact of the "entity type of the associated data" and improve the accuracy of quantification.

[0198] Among them, the business dependency between associated data and stored data may include the proportion of data shared fields. For example, if the proportion of data shared fields between a certain associated data and stored data is greater than the first preset ratio threshold, it can be determined as the first-level (highest) association level; if it is greater than the second preset ratio threshold and less than the first preset ratio threshold, it can be determined as the second-level association level; if it is greater than the third preset ratio threshold and less than the second preset ratio threshold, it can be determined as the third-level association level, etc. The embodiments of the present invention are not limited here.

[0199] As an optional embodiment, respectively determining the quantized value of each sub-feature in the dynamic change feature of the stored data includes:

[0200] For the life cycle stage, if the life cycle stage is not an archiving period, the quantitative value is determined by the time decay relationship;

[0201] If the life cycle stage is the archiving stage, the product of the initial value and the first preset ratio is used as the quantitative value of the life cycle stage;

[0202] The time decay relationships include:

[0203] Quantized value = initial value × e (-k×t) +Activity compensation value;

[0204] Wherein, k is the attenuation coefficient, t is the storage duration, and the activity compensation value is the product of the number of visits in the past preset period and the second preset ratio.

[0205] Specifically, considering that the probability of changing the entity type of storage data in the archiving period is low, it can be discussed separately. For storage data in the non-archiving period, the probability of changing the entity type is usually related to the storage duration and activity. Therefore, in the embodiment of the present invention, a time decay relationship is designed for storage data in the non-archiving period, so that the life cycle stage can be accurately quantified.

[0206] For example, if the lifecycle stage of stored data is the active stage (not the archival stage), the initial value is 100, the decay coefficient k is 0.01, the storage duration t is 30 days, the number of accesses in the past 30 days is 50, and the second preset ratio is 0.5, then the activity compensation value is 50 × 0.5 = 25, and the quantized value is 100 × e^(-0.01 × 30) + 25 ≈ 100 × 0.7408 + 25 ≈ 99.08. If the lifecycle stage is the archival stage and the first preset ratio is 0.2, then the quantized value is 100 × 0.2 = 20.

[0207] Please refer to Figure 3 , Figure 3 The present invention provides a structural diagram of a data retrieval device, which includes:

[0208] Memory 31, for storing computer programs;

[0209] The processor 32 is configured to implement the steps of the data retrieval method in the aforementioned embodiment when executing the computer program.

[0210] For an introduction to the data retrieval device provided by the embodiment of the present invention, please refer to the aforementioned embodiment of the data retrieval method, and the embodiment of the present invention will not be described in detail here.

[0211] Please refer to Figure 4 , Figure 4 This is a structural diagram of a computer-readable storage medium provided by the present invention. A computer program 42 is stored on the computer-readable storage medium 41. When the computer program 42 is executed by a processor, the steps of the data retrieval method in the aforementioned embodiment are implemented.

[0212] For an introduction to the computer-readable storage medium provided in an embodiment of the present invention, please refer to the aforementioned embodiment of the data retrieval method, and the embodiment of the present invention will not be described in detail here.

[0213] The present invention also provides a computer program product, comprising a computer program / instruction, which implements the steps of the data retrieval method in the aforementioned embodiment when executed by a processor.

[0214] For an introduction to the computer program product provided by the embodiment of the present invention, please refer to the aforementioned embodiment of the data retrieval method, and the embodiment of the present invention will not be described in detail here.

[0215] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from the other embodiments. Reference is made to the corresponding similar parts between the various embodiments. It should also be noted that, in this specification, the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or device comprising a set of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element specified by the phrase "comprises a..." does not preclude the presence of other identical elements in the process, method, article, or device comprising that element. The foregoing description of the disclosed embodiments will enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.

Claims

1. A data retrieval method, characterized in that: include: For any stored data in the distributed storage system, an entity type tag is added to the stored data; the entity type tag is used to represent the entity type to which the stored data belongs; Upon receiving the data query condition, determining the entity type to which the data query condition belongs as the target entity type; From the stored data whose entity type label is the target entity type, the stored data that meets the data query condition is determined as the target query data.

2. The data retrieval method according to claim 1, wherein: The step of labeling any stored data in the distributed storage system with an entity type tag includes: For any stored data in the distributed storage system, standardize the stored data according to the preset format requirements to obtain standardized data; For any stored data in the distributed storage system, identifying the entity type of the stored data; For any standardized data, the entity type of the storage data corresponding to the standardized data is bound to the standardized data as an entity type tag and stored.

3. The data retrieval method according to claim 2, characterized in that: For any stored data in the distributed storage system, identifying the entity type of the stored data includes: Identifying the entity type of the stored data using a pre-trained entity recognition model; The entity recognition model includes: an entity recognition model obtained by pre-training through a semantic alignment algorithm.

4. The data retrieval method according to claim 2, wherein: The data retrieval method further includes: In response to an operation instruction for target data, determining whether the target data has a distributed lock; the target data is standardized data specified by the operation instruction; If the distributed lock is not available, add a distributed lock to the target data and execute the operation instruction; If a distributed lock is available, the sender of the operation instruction will be notified that the target data is locked.

5. The data retrieval method according to claim 4, characterized in that: The executing the operation instruction includes: Execute the operation instruction and start timing; When the timing time reaches a preset time threshold, determining whether the target data has successfully executed the operation required by the operation instruction; If successfully executed, the target operation applied during the timing period is made effective; If the execution is unsuccessful, the operation applied to the target operation during the timing period is canceled.

6. The data retrieval method according to claim 1, wherein: After determining the stored data satisfying the data query condition from the stored data whose entity type label is the target entity type and using the stored data as the target query data, the data retrieval method further includes: sorting the target query data; Paging the target query data according to the sorting; Push the identification information of the target query data on the homepage; In response to the page change instruction, the identification information of the target query data in the target page is pushed; In response to the identification information selection instruction, the target query data corresponding to the selected identification information is pushed.

7. The data retrieval method according to claim 1, characterized in that: When receiving the data query condition, determining the entity type to which the data query condition belongs, as the target entity type includes: Upon receiving a data query condition, determining whether the data query condition is a mixed type data query condition; the mixed type data query condition is a data query condition composed of multiple sub-conditions; If the data query condition is a mixed type, for any sub-condition in the data query condition, determine the entity type to which the sub-condition belongs as the candidate entity type; If there is stored data with multiple entity type tags in the distributed storage system, then for mixed-type data query conditions, multiple target entity types are determined from each candidate entity type of the data query condition; If there is no stored data with multiple entity type tags in the distributed storage system, then for the mixed-type data query condition, a target entity type is determined from each candidate entity type of the data query condition; If the data query condition is not a mixed type, the entity type to which the data query condition belongs is used as the target entity type.

8. The data retrieval method according to claim 7, characterized in that: For mixed-type data query conditions, when multiple target entity types are determined, the stored data that meets the data query conditions from the stored data whose entity type label is the target entity type is determined as the target query data, including: Construct an association weight matrix between multiple target entity types, where the weight values ​​are determined based on the co-occurrence frequency of entity types in historical queries and the business logic relevance; Prioritize multiple target entity types based on an association weight matrix; Search the stored data that meets the corresponding sub-conditions under each entity type in order of priority, and calculate the comprehensive matching score between the stored data that meets the corresponding sub-conditions and all sub-conditions; The stored data whose comprehensive matching score exceeds the preset score threshold is selected as the target query data.

9. The data retrieval method according to claim 1, characterized in that: Determining the entity type to which the data query condition belongs includes: Perform semantic analysis on data query conditions to extract core keywords and context-related information; Perform fuzzy matching on the core keywords and the preset entity type vocabulary, and calculate the matching confidence level based on the contextual information. If the matching confidence exceeds the preset confidence threshold, the corresponding entity type is used as the target entity type; If the preset confidence threshold is not exceeded, a candidate list of entity types is pushed for the user to select, and the user's selection results are recorded to optimize the vocabulary matching rules.

10. The data retrieval method according to any one of claims 1 to 9, characterized in that: The data retrieval method further includes: For any stored data, obtaining dynamic change characteristics of the stored data; wherein the dynamic change characteristics include access frequency and content modification frequency of the stored data; Based on the dynamic change characteristics of the stored data, the entity type label of the stored data is adaptively updated.

11. The data retrieval method according to claim 10, characterized in that: Adaptively updating the entity type label of the stored data based on the dynamic change characteristics of the stored data includes: Monitor the access frequency and content modification frequency of stored data in real time, and re-identify the entity type of the stored data when either the access frequency or the content modification frequency reaches a corresponding threshold; If the re-identified entity type of the stored data is inconsistent with the original entity type label, a label update suggestion is generated and pushed to the administrator; In response to a confirmation instruction from the administrator, the re-identified entity type is used as a new entity type label for the stored data.

12. The data retrieval method according to claim 11, characterized in that: The dynamic change feature also includes changes in the entity type of the associated data; Adaptively updating the entity type label of the stored data based on the dynamic change characteristics of the stored data includes: When the entity type of the associated data of the stored data changes, verify whether the entity type label of the stored data is accurate; If it is inaccurate, the stored data will be marked as pending update and the administrator will be notified.

13. The data retrieval method according to claim 12, characterized in that: The dynamic change characteristics also include the life cycle stages of the stored data; the life cycle stages include the generation period, the active period and the archiving period; Adaptively updating the entity type label of the stored data based on the dynamic change characteristics of the stored data includes: Determining an update sensitivity coefficient of the entity type tag corresponding to the current life cycle stage of the stored data according to a preset first correspondence, wherein the first correspondence is a correspondence between the life cycle stage and the update sensitivity coefficient of the entity type tag; Determine a tag update trigger value based on the product of the quantized value of the dynamic change characteristic of the stored data and the update sensitivity coefficient; When the tag update trigger value reaches the preset trigger threshold, the entity type tag re-identification process is started; Among them, the life cycle stage is comprehensively determined by storage duration, last access time and business association rules.

14. The data retrieval method according to claim 13, wherein: The determining of the tag update trigger value based on the product of the quantized value of the dynamic change characteristic of the stored data and the update sensitivity coefficient includes: Determining the quantitative value of each sub-feature of the dynamic change feature of the stored data, wherein the sub-feature includes at least one of the access frequency of the stored data, the content modification frequency, the entity type change of the associated data, and the life cycle stage; Normalize the quantized values ​​of each sub-feature; Perform weighted summation on each quantized value after normalization to obtain the quantized value of the dynamic change feature; The tag update trigger value is determined based on the product of the quantized value of the dynamic change characteristic of the stored data and the update sensitivity coefficient.

15. The data retrieval method according to claim 14, characterized in that: The determining of the quantized value of each sub-feature in the dynamic change feature of the stored data includes: quantizing the access frequency or the content modification frequency according to the corresponding preset quantization interval and the quantization rule within the preset quantization interval.

16. The data retrieval method according to claim 14, characterized in that: The step of respectively determining the quantized value of each sub-feature in the dynamic change feature of the stored data includes: For entity type changes of associated data, quantification is performed based on the association level between the associated data of the entity type change and the stored data; the association level is divided according to the business dependency between the associated data and the stored data.

17. The data retrieval method according to claim 14, characterized in that: The step of respectively determining the quantized value of each sub-feature in the dynamic change feature of the stored data includes: For the life cycle stage, if the life cycle stage is not an archiving period, the quantitative value is determined by the time decay relationship; If the life cycle stage is the archiving stage, the product of the initial value and the first preset ratio is used as the quantitative value of the life cycle stage; The time decay relationship includes: Quantized value = initial value × e (-k×t) +Activity compensation value; Wherein, k is the attenuation coefficient, t is the storage duration, and the activity compensation value is the product of the number of visits in the past preset period and the second preset ratio.

18. A data retrieval device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the data retrieval method according to any one of claims 1 to 17 when executing the computer program.

19. A computer program product, characterized in that The method comprises a computer program / instruction, which, when executed by a processor, implements the steps of the data retrieval method according to any one of claims 1 to 17.

20. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the data retrieval method according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • Information query method and device, computer storage medium and terminal

    CN109753517A

  • Semantic-based data lake query system and method

    CN114218400A

  • Data processing method and device, electronic equipment and storage medium

    CN114647703A

  • Data query method and device, electronic equipment and storage medium

    CN115438150A

  • Entity label processing method and device based on elastic search engine

    CN119377467A