Data processing method and device of data map engine, equipment and storage medium

By identifying the clarity of the search intention and adopting corresponding search strategies, the problem of finding and using data in massive complex and diverse data is solved, and fast and accurate data retrieval is achieved, and the retrieval accuracy of data maps is improved.

CN119938803APending Publication Date: 2025-05-06BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411998431.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In a massive and complex and diverse data, it is difficult for the existing technology to quickly and accurately find and use data, resulting in challenges in data management.

Method used

Different search strategies are adopted by identifying the clarity of search intentions. When the search intention is unclear, search for keywords of the entity name; when the search intention is clear, search for keywords of multiple entity attributes to improve the search accuracy.

Benefits of technology

It realizes that the required data can be retrieved quickly and accurately in a massive and complex data storage system, improving the retrieval accuracy of data maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938803A_ABST
    Figure CN119938803A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device of a data map engine, equipment and a storage medium, and relates to a data map and intelligent search technology. The data processing method of the data map engine comprises the following steps: determining a retrieval intention of a retrieval request; if the retrieval intention is not clear, retrieving the entity name in a data storage system according to the retrieval keyword in the retrieval request to obtain a query result; if the retrieval intention is clear, a plurality of entity attributes are retrieved in the data storage system according to retrieval keywords in the retrieval request, a query result is obtained, and the entity attributes comprise at least two of an entity name, an entity description, an entity field name and an entity field description. Therefore, the data map engine can quickly and accurately retrieve the data from massive complex and diversified data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing, specifically to technologies such as data maps and intelligent search, and in particular to a data processing method, device, equipment and storage medium for a data map engine. Background Art

[0002] With the development of modern information technology, data in various industries are of various types and sources. At the same time, in the era of big data where the amount of data is growing exponentially, the diversity and complexity of data have led to certain challenges in data management.

[0003] Data maps can provide users with a comprehensive, intuitive and secure data management experience. Therefore, data management and data search are performed based on data maps, serving users and owners of data tables such as data analysis, data development, data mining, and data operations, so that users can quickly find, understand and use data.

[0004] Therefore, there is an urgent need to propose a data map engine for massive, complex and diverse data, so that data can be found and used more accurately and quickly in massive, complex and diverse data. Summary of the invention

[0005] The present disclosure provides a data processing method, device, equipment and storage medium for a data map engine.

[0006] According to a first aspect of the present disclosure, a data processing method for a data map engine is provided, comprising: determining the search intent of a search request; if the search intent is unclear, searching for entity names in a data storage system according to search keywords in the search request to obtain query results; if the search intent is clear, searching for multiple entity attributes in the data storage system according to search keywords in the search request to obtain query results, wherein the entity attributes include at least two of an entity name, an entity description, an entity field name, and an entity field description.

[0007] According to a second aspect of the present disclosure, a data processing device is provided, including: an analysis unit for analyzing the retrieval intent of a retrieval request; a retrieval unit for searching for entity names in a data storage system according to retrieval keywords in the retrieval request to obtain query results when the retrieval intent is unclear; the retrieval unit is also used for searching for multiple entity attributes in the data storage system according to retrieval keywords in the retrieval request to obtain query results when the retrieval intent is clear, the entity attributes including at least two of an entity name, an entity description, an entity field name, and an entity field description.

[0008] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data processing method of the data map engine described in the first aspect.

[0009] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the data processing method of the data map engine described in the first aspect.

[0010] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising: a computer program, the computer program being stored in a readable storage medium, at least one processor of an electronic device being able to read the computer program from the readable storage medium, and the at least one processor executing the computer program so that the electronic device executes the data processing method of the data map engine described in the first aspect.

[0011] According to the technical solution provided by the present invention, a search strategy corresponding to a matching search scenario is constructed according to different search scenarios. Specifically, whether the search intent is clear is identified. When the search intent is unclear, a keyword search is performed on the entity name to clarify the search location and improve the search accuracy. When the search intent is clear, a keyword search is performed on multiple entity attributes, and the search results of multiple entity attributes are combined to obtain the final query result. The search intent is clarified through multiple entity attributes to improve the search accuracy. This is suitable for massive, complex and diverse data map engines.

[0012] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0014] Figure 1 is a schematic diagram of an application scenario to which the embodiments of the present disclosure are applicable;

[0015] Figure 2 is a schematic diagram according to a first embodiment of the present disclosure;

[0016] Figure 3 is a schematic diagram according to a second embodiment of the present disclosure;

[0017] Figure 3ais a schematic diagram of using a word segmenter according to an entity retrieval process of the present disclosure;

[0018] Figure 4 is a schematic diagram according to a third embodiment of the present disclosure;

[0019] Figure 5 is a schematic diagram of a process written according to the disclosed data;

[0020] Figure 5a is a schematic diagram of an entity instantiation process according to the present disclosure;

[0021] Figure 6 is a schematic diagram according to a fourth embodiment of the present disclosure;

[0022] Figure 7 is a schematic diagram according to a fifth embodiment of the present disclosure;

[0023] Figure 8 FIG. 8 is a schematic block diagram of an example electronic device 800 that may be used to implement embodiments of the present disclosure. DETAILED DESCRIPTION

[0024] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0025] With the development of modern information technology, data in various industries are of various types and sources, so it is necessary to manage data. In the era of big data where the amount of data is growing exponentially, the diversity and complexity of data lead to certain challenges in data management.

[0026] The data map can provide users with a comprehensive, intuitive and secure data management experience. Therefore, the data map is built around data search, and metadata collection is enabled in the data map. The system will automatically collect existing metadata from multi-source heterogeneous data, while collecting incremental metadata every day and aggregating it to the data map to serve users and owners of data tables such as data analysis, data development, data mining, and data operations, so that they can quickly find, understand and use data.

[0027] However, there is still no data map engine that can accurately and quickly find and use data in response to massive amounts of complex and diverse data. Therefore, there is an urgent need to propose a data map engine that can quickly find and use data in massive amounts of complex and diverse data.

[0028] In order to solve the above problems, the present disclosure provides a data processing method, device, equipment and storage medium of a data map engine, which relates to the field of data processing, specifically to data maps, intelligent search and other technologies, and is applied to data management scenarios required for data segmentation, data development, data mining, data operations, etc. In the present disclosure, one side provides a management system for massive, complex and diverse data, and on the other side provides different retrieval strategies for different retrieval scenarios. Based on different retrieval scenarios, it is possible to quickly and accurately retrieve the required data in massive, complex and diverse data storage systems, thereby improving the retrieval accuracy of data maps.

[0029] For example, when the user's search intent is unclear, the keyword entered may be a description of the entity name. Therefore, for scenarios where the search intent is unclear, searching for the entity name can also accurately locate the entity name that the user wants to retrieve. When the user's search intent is clear, the keyword entered may be more of a description of the entity description or entity field name. Therefore, for scenarios where the search intent is clear, searching for multiple entity attributes (such as entity name, entity description, entity field name, entity field description, etc.) can accurately locate the detailed content of the data that the user wants to retrieve. It is suitable for searches in massive, complex and diverse data storage systems.

[0030] Figure 1 It is a schematic diagram of an application scenario applicable to the present disclosure. The application scenario includes a data processing system, and the data processing system includes multiple subsystems, such as a data management system, a data storage system, a data display system, etc. Among them, the data management system is used to collect metadata from multi-heterogeneous data, and to manage the collected metadata, for example, including metadata storage, metadata version management, metadata update, etc.; the data storage system is used to store the collected data; the data display system is used to display data in the form of a map and to display retrieved data. Optionally, the data management also provides a data retrieval function for retrieving required data from a data storage system (such as a data map).

[0031] The following specific embodiments are used to describe in detail the technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below in conjunction with the accompanying drawings.

[0032] Figure 2 Schematic diagram of the first embodiment of the present disclosure. Figure 2 As shown, the data processing method of the data map engine provided by the first embodiment of the present disclosure includes:

[0033] S201, determining the search intent of the search request.

[0034] Among them, the execution device of the embodiment of the present disclosure can be a terminal, a server, a cloud, etc., and the following description is taken as an example where the execution device is a terminal.

[0035] In this embodiment, the terminal provides a data map retrieval function, specifically providing a retrieval interface, and users and owners of data tables such as data analysis, data development, data mining, and data operations can retrieve the required data through the retrieval interface.

[0036] When a user triggers a search in the search interface or enters a keyword in the search window of the search interface, a search request is generated. The terminal analyzes the user's search intention according to the search request. Exemplarily, the search intention includes a clear search intention and an unclear search intention.

[0037] In an optional implementation, the search intent is determined in the following manner:

[0038] A1: Determine the character length of the search keyword in the search request;

[0039] The search request includes a search keyword.

[0040] A2: If the character length is greater than or equal to the first threshold, the search intent is clear;

[0041] Among them, the first threshold can be set according to the general word grouping structure, for example, general nouns (including Chinese and English) are 2 to 4 characters, and adjectives or adverbs (including Chinese and English) are 3 characters. Exemplarily, the first threshold can be set to 3. When the character length is greater than or equal to 3, the content described by the search keyword may include detailed content such as entity name, entity description, entity field, entity field description, etc., so the search intention is clear.

[0042] For example, taking a data table entity stored in a data storage system as an example, the entity name refers to the data table name, the entity field name refers to the name of a field in the data table, the entity description refers to the description of the data table, and the entity field description refers to the description of a field in the entity.

[0043] A3: If the character length is less than or equal to the second threshold, the search intent is unclear.

[0044] For example, when the search character length is less than 3, the content described by the search keyword may be an entity name, an entity field name, or an entity description, etc. Since there are too few search keywords, the search scope is large, so the search intention is unclear.

[0045] In the embodiment of the present disclosure, the identification of the search scenario (search intent) is achieved through the above method. It should be noted that the embodiment of the present disclosure is not limited to the above method for identifying the search intent.

[0046] S202: If the search intention is not clear, the entity name is searched in the data storage system according to the search keywords in the search request to obtain the query results.

[0047] The data storage system refers to a system for storing data, and the retrieval request in this embodiment refers to a request to retrieve the data required by the user from the data storage system. With the rapid increase in data volume and the complexity and diversity of data, when the user's retrieval intention is unclear, the query results returned by the data storage system are often noisy and the retrieval accuracy is low.

[0048] In the embodiment of the present disclosure, when the search intent is unclear, considering the scenario where the search character length is short, the search keyword is more of a description of the entity name, so in the data storage system, the entity name is searched to obtain query results whose entity names contain matches with the search keyword, and the query results are more accurate.

[0049] In an optional implementation, in order to further improve the accuracy of retrieval in scenarios where the retrieval intent is unclear, parsing is performed through a custom word segmenter to retrieve entity names to implement entity name retrieval completion capabilities.

[0050] Exemplarily, for scenarios where the search intent is unclear, a third word segmenter is deployed to segment the entity name SubField to achieve entity name search completion capabilities. Optionally, the third word segmenter includes a pattern word segmentation module, which uses regular expressions as word segmentation rules and segments search keywords based on regular expressions. Among them, the regular expression is set according to the specific application scenario. Exemplarily, in the field of data maps, most of the naming rules based on entity names are English and underscores, and English uses uppercase to distinguish two words, so the regular expression can be set to "(_|(?=[AZ]))", that is, word segmentation is performed according to underscores or before uppercase letters (AZ).

[0051] Optionally, the third word segmenter also includes a word segmentation filter module Edge-NGram Token Filter, which is used to generate an n-gram (n-tuple) sequence starting from the starting position. That is, the word segmentation filter is used to further segment the word segmentation output by the pattern word segmentation module, so that the word segmentation is divided into finer characters. By setting min_gram=N1, max_gram=N2, the word segmentation filter module is defined to switch lengths with word segmentation, that is, the maximum tuple is N2, and the minimum tuple is N1, indicating that each word segmentation will be further segmented into all possible prefixes from N1 characters to N2 characters. For example, in the disclosed embodiment, N1 is 1, N2 is 2, that is, the previous character or the first two characters are the word segmentation prefix.

[0052] Based on the above settings, during the search process, the entity name is searched and completed, and the specific implementation method for obtaining the query results is as follows:

[0053] B1: Based on the search keywords, the search keywords are segmented in the third segmenter to obtain the initial segmentation results;

[0054] In step B1, the pattern segmentation module of the third segmenter performs preliminary segmentation on the search keyword to obtain an initial segmentation result. For example, if the search keyword is "My_Name", the third segmenter performs segmentation on the search keyword, and the initial segmentation result obtained is "My", "_Name" (or "Name").

[0055] B2: Determine the word segmentation prefix in the initial word segmentation result according to the Edge-NGram Token Filter;

[0056] In step B2, based on the use of the word segmentation filter, the initial word segmentation result is further segmented to obtain the word segmentation prefix in the initial word segmentation result. Exemplarily, taking the above example, if the word segmentation filter is set to min_gram=1, max_gram=2, it means that each word segmentation will be further segmented into all possible prefixes from 1 character to 2 characters; therefore, "My", "_Name" (or "Name") is segmented into "M", "My", "_N", "_Na" (or "N", "Na"), and the switched words are the word segmentation prefixes in the initial word segmentation result.

[0057] B3: In the search server ES, after completing the entity name according to the segmentation prefix, the query result is output, where the query result includes an entity list.

[0058] Optionally, the specific implementation process of step B3 includes:

[0059] B31, query the search server ES for a list of entity names that match the word segmentation prefix, and output the list of entity names;

[0060] In the disclosed embodiment, data is stored in a search server ES (ES: Elasticsearch, a full-text search engine based on a search engine that provides distributed multi-user capabilities), and data is managed and searched based on the search server ES.

[0061] Therefore, in this step, after determining the word segmentation prefix corresponding to the search keyword, the entity name whose prefix contains the word segmentation prefix is ​​searched in ES to complete the entity name. Then all entity names are output to the interface in the form of a list, so that the user can view the completed entity names corresponding to the search keyword retrieved, and the user can select one of the completed entity names based on the entity name list.

[0062] B32: Obtain query results based on the target entity name selected in the entity name list.

[0063] In this step, after the user selects the target entity name (the completed entity name) based on the entity name list, the corresponding query result is obtained directly based on the inverted index of the target entity name.

[0064] In the above implementation, when the search intent is unclear, on the one hand, searching for entity names can simplify the search process. On the other hand, in the process of searching for entity names, the entity names are searched and completed. When the user cannot determine more search keywords, more information related to the search keywords is provided to the user, or the complete entity name is provided to the user, so that the user can perform accurate search. In the entity name search completion function, in the third word segmenter, the search keywords are preliminarily segmented, and then the segmented words are segmented. The segmentation prefix is ​​used to search the entity name and complete the entity name. In the scenario where the user's search intent is unclear, if the search keyword is partially wrong (such as spelling errors, combination errors, etc.), the problem of inaccurate search caused by the error can be corrected based on the segmentation prefix. The method adopted by the embodiment of the present disclosure provides search flexibility and user experience, enhances the search recall rate, and optimizes the search scenarios of spelling errors and abbreviations.

[0065] S203, if the search intention is clear, multiple entity attributes are searched in the data storage system according to the search keywords in the search request to obtain the query results;

[0066] The entity attributes refer to various attributes of an entity, such as entity name, entity description, entity field name, and entity field description.

[0067] Among them, the entity name is a simple description of the entity in the database; the entity description is a detailed description and definition of the entity in the database; each entity has multiple fields, the entity field name refers to a simple description of each field of the entity in the database, and the entity field description is a description of the entity's attributes or characteristics.

[0068] In this step, if the user's search intention is clear, for example, the user enters a search keyword of a certain character length, multiple entity attributes are searched in the data storage system to obtain query results. Among them, multiple entity attributes refer to entity attributes including 2 or more. Exemplarily, the entity name and entity description are searched, or the entity field name and entity field description are searched, or the entity name, entity description and entity field name are searched, or all entity attributes are searched, which are not listed here one by one.

[0069] In this step, the query result is a comprehensive result of the search results based on at least one entity attribute after searching multiple entity attributes, that is, the query result includes the search result of at least one entity attribute. Exemplarily, in the data storage system, the entity name 1 matching the search keyword is searched to obtain the search result 1, the entity description 2 matching the search keyword is searched to obtain the search result 2, and the search result 1 and the search result 2 are combined to obtain the final query result; or, the entity name 1 matching the search keyword is searched to obtain the search result 1, the entity description 2 matching the search keyword is searched, and there is no relevant description in the system, and no search result is obtained, then the search result 1 is the final query result.

[0070] If the search keywords include nouns that express entity names and adjectives that describe entity names, then by combining entity name search with entity description search, you can quickly and accurately locate entity names that meet the description, thereby improving search accuracy.

[0071] In some possible implementations, multiple entity attributes are retrieved in a data storage system, and the query results obtained are not simply obtained by stacking the retrieval results of each entity attribute. For specific implementation methods, please refer to the detailed description of the following embodiments.

[0072] In this embodiment, a search strategy corresponding to a matching search scenario is constructed according to different search scenarios. Specifically, it is identified whether the search intent is clear. When the search intent is unclear, a keyword search is performed on the entity name to clarify the search location and improve the search accuracy. When the search intent is clear, a keyword search is performed on multiple entity attributes, and the search results of multiple entity attributes are combined to obtain the final query result. The search intent is clarified through multiple entity attributes to improve the search accuracy. This is suitable for massive, complex and diverse data map engines.

[0073] Figure 3 is a schematic diagram according to the second embodiment of the present disclosure. Figure 3 As shown, the data processing method of the data map engine provided by the second embodiment of the present disclosure includes:

[0074] S301, determining the search intent of the search request;

[0075] S302: If the search intention is not clear, the entity name is searched in the data storage system according to the search keywords in the search request to obtain the query results.

[0076] Among them, the execution device of the embodiment of the present disclosure can be a terminal, a server, a cloud, etc., and the following description is taken as an example where the execution device is a terminal.

[0077] In the embodiments of the present disclosure, the implementation principles and technical effects of S301 to S302 may refer to the aforementioned embodiments and will not be described in detail.

[0078] S303: If the search intention is clear, the search keywords are segmented in the word segmenter according to multiple entity attributes to obtain multiple segmentation results.

[0079] In this implementation, the search keywords are segmented in the word segmenter according to multiple entity attributes, which means that the search keywords are segmented in the word segmenter according to the entity attributes to be searched. Exemplarily, if it is necessary to search whether the entity name contains the keyword search, the search keywords are segmented based on the entity name, and if it is necessary to search whether the entity description contains the keyword search, the search keywords are segmented based on the entity description. When searching for different entity attributes, the search keywords are segmented differently.

[0080] In an optional implementation, the word segmentation strategy laid out in the data map engine is: for different entity data, the word segmentation strategy of the word segmenter is different. For example, when searching for entity names or entity field names, the word segmentation rule of the word segmenter for the search keywords is: word segmentation in the form of phrases. When searching for entity descriptions or entity field descriptions, based on the description including names, adjectives, adverbs, etc., and these words have no fixed form, the word segmentation rule of the word segmenter for the search keywords is: word segmentation based on the finest granularity splitting strategy.

[0081] For example: Taking the search keyword "My_User data" as an example, the entity name or entity field name generally appears in the form of a phrase or noun, so the corresponding word segmentation result for the search of the entity name or entity field name is "My_User data". The presentation form of the entity description and entity field description is not fixed, so the corresponding word segmentation results for the search of the entity description and entity field description are "My", "User", "data", "User data" and "My_User data".

[0082] In an optional implementation, a word segmentation strategy is arranged in the data map engine: multiple word segmenters are arranged according to entity attributes, and each entity attribute corresponds to a word segmenter for word segmentation. Taking the entity attributes including entity name, entity description, entity field name and entity field description as an example, four word segmenters are deployed correspondingly, corresponding to the four entity attributes respectively, and the corresponding entity attributes are retrieved based on the word segmentation results of each word segmenter. Exemplarily, the search keyword is segmented in word segmenter 1 to obtain word segmentation result 1, which is used to retrieve the entity name. The search keyword is segmented in word segmenter 2 to obtain word segmentation result 2, which is used to retrieve the entity description. The search keyword is segmented in word segmenter 3 to obtain word segmentation result 3, which is used to retrieve the entity field. The search keyword is segmented in word segmenter 4 to obtain word segmentation result 4, which is used to retrieve the entity field description. Among them, the four word segmenters can be word segmenters of the same type or different types.

[0083] In this example, different word segmenters are used to segment the search keywords to improve the word segmentation speed and thus improve the search speed.

[0084] In an optional embodiment, if Figure 3a As shown, the word segmentation strategy is laid out in the data map engine: two types of word segmenters are deployed, namely the first word segmenter and the second word segmenter, wherein the first word segmenter is an N-gram word segmenter, that is, the first word segmenter uses N-grams Tokenizer for word segmentation. The second word segmenter is a Chinese word segmenter, such as ik_max_word word segmenter.

[0085] According to the characteristics of entity attributes, entity names and entity field names, the first word segmenter is used for word segmentation. The entity description and entity field description contain a large number of Chinese characters. Using the Chinese word segmenter for word segmentation can improve the accuracy of retrieval.

[0086] Based on the above-mentioned word segmentation strategy deployment, the specific implementation of step S303 can be: according to the entity name and the entity field name, the search keyword is segmented in the first word segmenter to obtain the corresponding first word segmentation result and the second word segmentation result; according to the entity description and the entity field description, the search keyword is segmented in the second word segmenter to obtain the corresponding third word segmentation result and the fourth word segmentation result. That is, the multiple word segmentation results include the first word segmentation result, the second word segmentation result, the third word segmentation result and the fourth word segmentation result.

[0087] Optionally, in some embodiments, the entity name and the entity field name may use the same first tokenizer, while the entity description and the entity field description may use the same second tokenizer, thereby reducing the deployment of tokenizers. Optionally, in other embodiments, two first tokenizers and two second tokenizers are deployed, wherein the entity name is tokenized by one first tokenizer, the entity field name is tokenized by another first tokenizer, the entity description is tokenized by one second tokenizer, and the entity field description is tokenized by another tokenizer, so as to improve the tokenization speed and thus improve the retrieval speed.

[0088] S304, obtaining a compound query condition according to the multiple word segmentation results.

[0089] The compound query condition refers to a query condition constructed based on multiple word segmentation results.

[0090] In an optional implementation, a compound query condition may be formed by constructing a logical relationship between multiple word segmentation results. For example, the compound query condition is a Boolean query.

[0091] In an optional implementation, the weight of the entity attribute can be set to reflect the importance of the word segmentation results corresponding to the entity attribute, and the query condition is constructed according to the importance of the word segmentation results. For example, important word segmentation results are sorted first, or the query results must contain the words in the important word segmentation results, etc. Exemplarily, after obtaining multiple word segmentation results, the importance of the word segmentation results corresponding to the entity attributes is obtained according to the weight of the entity attribute, and a composite query condition is obtained according to the importance of multiple word segmentation results. Exemplarily, the composite query condition is to determine the importance of the word segmentation results according to the weight of the entity attribute, and output the entity list (i.e., the query result) in order from high to low importance.

[0092] For example, the weights of various entity attributes are pre-set, for example, the entity name is P1, the entity description is P1, the entity field name is P3, and the entity field description is P4, where P1+P2+P3+P4=1. The first word segmenter performs word segmentation on the search keyword to obtain the first word segmentation result and the second word segmentation result, respectively, wherein the first word segmentation result is used to search for the entity name, and the second word segmentation result is used to search for the entity field name. The second word segmenter performs word segmentation on the search keyword to obtain the third word segmentation result and the fourth word segmentation result, respectively, wherein the third word segmentation result is used to search for the entity description, and the fourth word segmentation result is used to search for the entity field description. When a composite query condition is constructed based on multiple word segmentation results, the composite query condition is constructed based on the first word segmentation result and P1, the second word segmentation result and P2, the third word segmentation result and P3, and the fourth word segmentation result and P4. If P1 is the largest, the importance of the first word segmentation result is the largest, and if P4 is the smallest, the importance of the fourth word segmentation result is the smallest.

[0093] In an optional embodiment, a scoring rule is also set for the search results of entity attributes, and the importance of the entity attributes is determined according to the search scores of the entity attributes and the corresponding weights of the entity attributes, and the query results are output. Exemplarily, based on the word segmentation results of the entity attributes, the entity attributes are searched to obtain the search scores, and the total scores of the entity attributes are determined according to the search scores and the weights of the entity attributes. The entity list is output according to the order of the total scores from high to low. Exemplarily, assuming that the weight of the entity name is P1, the weight of the entity description is P2, the weight of the entity field name is P3, and the weight of the entity field description is P4, P1>P2>P3>P4, the entity name score*P1 is the total score of the entity name, the entity description score*P2 is the total score of the entity description, the entity field name score*P3 is the total score of the entity field description, and the entity field description score*P4 is the total score of the entity field description. Output the entity list retrieved based on each entity attribute in the order of the total scores from high to low. That is, in this embodiment, the compound query condition determines the importance of the entity attribute based on the total score of the entity attribute retrieval score and the entity attribute weight, and outputs the entity list in descending order according to the importance of the entity attribute.

[0094] It should be noted that when searching for entity attributes, some entity attributes may not be retrieved, in which case the search score of the entity attribute is 0. In this embodiment, if the search score of at least one entity attribute is greater than 0, the query result can be obtained.

[0095] When constructing a compound query condition, the first participle result may be used as a must-satisfy item, the fourth participle result as an item that may be satisfied but not necessarily satisfied, and so on. The third participle result and the fourth participle result, according to their importance, are used together with the first participle result as a must-satisfy item or together with the fourth participle result as an item that may be satisfied but not necessarily satisfied, to obtain a compound query condition; or, the first participle result may be used as a priority-satisfy item, the fourth participle result as a final-satisfy item, and so on. The third participle result and the fourth participle result, according to their importance, are used together with the first participle result as a priority-satisfy item or together with the fourth participle result as a second priority-satisfy item, to obtain a compound query condition.

[0096] Alternatively, compound query conditions may be constructed in other ways. The technical solution for constructing compound query conditions from multiple search results is claimed in the disclosed embodiments. The specific method for constructing compound query conditions may be set according to application scenarios and usage requirements. For example, compound query conditions of multiple word segmentation results may be constructed according to the relationship of logical AND or logical OR based on weight selection, etc., which are not listed here one by one.

[0097] S305: According to the compound query condition, query the search server ES for query results matching the compound query condition.

[0098] In an optional implementation, the data storage system includes a search server ES. That is, the data map engine in this embodiment uses the search server ES to store data.

[0099] After determining the compound query conditions, the search server ES matches the data in the database based on the word segmentation in the compound query conditions, obtains the query results, and highlights the query results to facilitate users to view the query results. The query results include metadata, metadata version, metadata details, etc.

[0100] In an optional implementation, in a search application scenario of a massive and complex database, in order to improve the search speed, an inverted index of the data is constructed during the data storage process. Therefore, during the search process, the searched data can be directly located based on the inverted index without traversing the data, which speeds up the processing speed.

[0101] Therefore, in this implementation, the specific implementation of step S305 is:

[0102] C1: Get the inverted index corresponding to the target word in the word segmentation result;

[0103] In step C1, the compound query condition includes multiple word segmentation results, each word segmentation result corresponds to at least one word segmentation. During the data storage process, an inverted index is constructed based on the word segmentation of the data, so in this step, the inverted index corresponding to the target word segmentation can be obtained based on the database.

[0104] C2: Based on the inverted index, the preliminary query results matching the target word are searched on the search server ES;

[0105] In step C2, the inverted index can quickly locate the entity attribute containing the segmentation. Therefore, after determining the inverted index of the target segmentation, the entity attribute containing the target segmentation can be directly located.

[0106] For example, when the target participle is used to search for entity names, the entity names containing the target participle can be quickly located based on the inverted index. Or when the target participle is used to search for entity descriptions, the entity descriptions containing the target participle can be quickly located based on the inverted index.

[0107] Among them, the data retrieved according to the inverted index of the target word is the preliminary query result.

[0108] C3: Based on the compound query conditions, the preliminary query results are integrated to obtain query results matching the compound query conditions.

[0109] In this implementation, all the word segments in the compound query condition are retrieved through step C1 and step C2 respectively, and all the query results obtained are preliminary query results. In order to improve the accuracy of the query, it is also necessary to synthesize the preliminary query results according to the relationship between the word segmentation results in the compound query condition to obtain the final query result matching the compound query condition. Among them, the relationship between the word segmentation results in the compound query condition can refer to the above description of the construction of the compound condition, which will not be repeated here.

[0110] In the disclosed embodiment, multiple word segmenters are used to segment keywords respectively, which on the one hand improves the efficiency of word segmentation, and on the other hand, word segmentation is performed according to the needs of entity attributes to improve the accuracy of word segmentation. The search results of multiple entity attributes are integrated to form query results. Multiple entity attributes predict user search intentions from multiple angles to improve the accuracy of retrieval.

[0111] Figure 4 is a schematic diagram according to the third embodiment of the present disclosure. Figure 4 As shown, the data processing method of the data map engine provided by the third embodiment of the present disclosure includes:

[0112] S401, upon receiving a metadata update request from a tenant, obtaining metadata to be updated;

[0113] The update request includes at least one of adding, deleting, replacing and changing a version, for example, adding metadata, deleting metadata, replacing metadata or changing a metadata version.

[0114] In an optional implementation, the metadata update request is initiated by the tenant, for example, the tenant manually adds metadata, changes metadata, deletes metadata, etc. in the system.

[0115] Alternatively, in an optional implementation, the metadata update request may also be triggered by the data storage system. For example, when the external interface of the data storage system collects data changes of the tenant, etc., it triggers a data update request. In this example, the data storage system determines which tenant's data to update based on the tenant to which the external interface that triggers the update request is connected.

[0116] Optionally, the update request includes metadata to be updated.

[0117] S402, obtaining the tenant index corresponding to the update request. In the search server ES, each tenant is created with at least one tenant index.

[0118] In this embodiment, the terminal deploys a public cloud storage solution and a private storage solution to store the metadata (data source + EDAPDatalake) basic information required for the data map retrieval page.

[0119] Public cloud storage solution: An ES cluster is set up in each region (such as Beijing, Yizhuang, Suzhou, etc.). Each tenant creates an independent index (namespace_edap_map_metadata) in the ES cluster, creates an index according to the tenant granularity, and stores the basic metadata information in ES. The deployment of the public cloud storage solution, on the one hand, when querying data, when querying a tenant's data, the data of other tenants will not be scanned, which realizes the isolation of tenant data. The cluster only loads the data of hot indexes into memory, and the resource consumption will be more refined. On the other hand, when writing data, the migration cost is small after the single tenant data volume reaches the upper limit (the single tenant data volume can be migrated separately when it reaches the upper limit). On the other hand, in the scenario of explosive growth of data volume, an ElasticSearch index cannot store massive data. Each tenant (customer) has an ElasticSearch index to store separately to achieve the storage of massive data. Therefore, when receiving a tenant's metadata update request, the tenant can be determined according to the update request, and then the tenant index can be determined according to the tenant.

[0120] Exemplarily, the configuration of the deployed public cloud storage solution is as follows: each single primary shard has 2 replica shards, where the size of a shard is in the range of [30GB-a; 30GB+b], where b is less than 20GB, that is, the maximum size of a shard (primary shard or replica shard) does not exceed 50GB; the number of doc entries is 100 million.

[0121] Based on the above public cloud storage solution deployment, when the data to be updated is public cloud data, the tenant index corresponding to the update request is obtained to facilitate storing the metadata to be updated in the corresponding ES.

[0122] Private storage solution: Store metadata basic information in ES. The private storage solution creates a single index, so metadata storage can be directly updated based on all.

[0123] Exemplarily, the deployed private storage solution is configured as follows: there are 5 primary shards and 2 replica shards.

[0124] Based on the above deployment, if the metadata to be updated is private data, the metadata to be updated can be directly stored in the tenant's ES.

[0125] Therefore, in this embodiment, the storage process of public cloud data is mainly described.

[0126] S403: In the search server ES, the metadata to be updated is updated under the tenant index.

[0127] In an optional implementation, if the update request is for newly added data, the metadata to be updated is written into the search server ES.

[0128] In an optional implementation, if the update request is to update metadata, the location of the historical metadata of the metadata to be updated is found in the search server ES, and the historical metadata is updated based on the metadata to be updated.

[0129] In an optional implementation, if the update request is to delete metadata, the location of the metadata to be deleted is found in the search server ES, and the metadata to be deleted is deleted.

[0130] Optionally, in some embodiments, the terminal deploys TafDB in public cloud storage to store historical version data of metadata, so when the update request is to delete metadata, it is also necessary to delete the metadata to be updated in ES and then delete the historical version data of the metadata to be updated in TafDB.

[0131] In an optional implementation, if the update request is to update the metadata version (such as version attribute), the historical version of the metadata to be updated is found in the search server ES, and the new version of the metadata is updated to the corresponding position.

[0132] Optionally, during the data update process, for newly added metadata or updated metadata version data, the metadata to be updated is updated by batch writing. The following example illustrates the specific update process:

[0133] like Figure 5As shown, the trigger adds metadata / metadata version data (basic information + schema), then starts the TafDB transaction, the system loops a single insert TafDB metadata version, after the loop ends, commits the transaction to confirm that the TafDB transaction is successfully executed; then determines whether the TafDB data batch metadata is written successfully, if so, executes the stored procedure; otherwise, an error is reported.

[0134] Optionally, after the batch metadata of TafDB data is written successfully, the execution of the storage process can be specifically as follows: determine whether the tenant index exists. If the tenant index exists, write the data batch to the ES tenant index. Determine whether the batch write is successful. If the full amount is successful, complete the creation of metadata and update the metadata version attribute; if not, extract the failed entry for writing to ES, open the log and roll back to the failed entry in TafDB. If the tenant index does not exist, create the tenant index according to the index template. If the creation is successful, write the data batch to the ES tenant index. If the creation fails, an error is reported.

[0135] In an optional implementation, when the update request is to add new metadata, the specific implementation method of step S404 is: segment the metadata to be updated in the word segmenter to obtain the segmentation result; construct an inverted index corresponding to each word in the segmentation result, and store the inverted index corresponding to the word segmentation in the search server ES.

[0136] In this implementation, when metadata is added to ES, an inverted index needs to be created for the newly added metadata so that when tenants retrieve data, they can directly locate the required data based on the inverted index created when the data is stored.

[0137] In this embodiment, the process of constructing an inverted index is also a process of entity indexing.

[0138] In this embodiment, Figure 3 The embodiments are the same, and the same tokenizer is used to tokenize the updated metadata, and different inverted indexes are constructed under different entity attributes.

[0139] For example, Figure 5a As shown, for entity names and entity field names, the first tokenizer is used for tokenization, and then an inverted index of the tokens output by the first tokenizer is constructed. For entity descriptions and entity field descriptions, the second tokenizer is used for tokenization, and then an inverted index of the tokens output by the second tokenizer is constructed. For entity name SubField, the third tokenizer is used for tokenization, and then an inverted index of the tokens output by the third tokenizer is constructed, and the constructed inverted index is stored in ES.

[0140] In this embodiment, during the data storage process, indexes are built for tenants, and the updated data is segmented according to the attributes of each entity, and then an inverted index of the segmentation is built. Based on the management of tenant indexes and the management of the segmentation inverted index under entity attributes, even if the amount of data increases dramatically, the above-mentioned management method can provide users with fast retrieval capabilities.

[0141] Figure 6 is a schematic diagram according to the fourth embodiment of the present disclosure. Figure 6 As shown, the data processing method of the data map engine provided by the fourth embodiment of the present disclosure includes:

[0142] S601, receiving a version change request from a tenant, determining the data to be updated, and collecting historical version data of the metadata to be updated;

[0143] In this embodiment, a storage solution is deployed for metadata version updates, which is different from the storage solution mentioned in the previous embodiment. Therefore, when the terminal receives the version change request from the tenant, it stores the metadata to be updated through the storage method of the previous embodiment, and also stores the historical version data of the metadata to be updated through the storage method provided by this embodiment. The following is a detailed description of the method for historical version data:

[0144] It should be noted that, in this step, the version change request is used to indicate the replacement of the version of the metadata that has been stored. Optionally, the method of initiating the version change request is the same as that listed in the third embodiment above, and will not be repeated here.

[0145] Optionally, the version change request includes the data to be updated to be replaced, and then the historical version data of the data to be updated is collected according to the data to be updated.

[0146] S602: If the historical version data is public cloud data, the historical version data is stored in a TafDB database. The data storage system also includes a TafDB database.

[0147] In this embodiment, for the storage of historical version data, the terminal deploys a public cloud storage solution: using the TafDB database as the storage engine. The collected historical version data will cause the amount of public cloud data to explode, so the TafDB database is used as the storage engine. Based on the linear expansion of the TafDB database, it can provide a metadata storage capacity of trillion-level metadata and tens of millions of QPS, which can meet the scalability and performance requirements of massive data storage and is suitable for the management and storage of massive data.

[0148] Based on the above deployment, if the historical version data collected is public cloud data, the historical version data is stored in the TafDB database. The metadata to be updated is stored in the corresponding ES through the tenant index based on the above embodiment.

[0149] S603: If the historical version data is private data, the historical version data is stored in a MySQL database. The data storage system also includes a MySQL database.

[0150] In this embodiment, for the storage of historical version data, the terminal deploys a private storage solution: using MYSQL database as the underlying storage engine.

[0151] Based on the above deployment, if the collected historical version data is privatized data, the historical version data is stored in the MYSQL database, and the metadata to be updated is updated to the ES of the private tenant based on the above embodiment.

[0152] MYSQL database technology has good compatibility, low cost and is suitable for private storage.

[0153] In this embodiment, when the metadata version is replaced, the historical version data of the public cloud metadata is stored in the TafDB database, and the historical version data of the privatized metadata is stored in the MYSQL database. The above deployment meets the scalability and performance requirements of massive data storage.

[0154] Figure 7 Schematic diagram of the fifth embodiment of the present disclosure. Figure 7 As shown, the data processing device 700 provided in the fifth embodiment of the present disclosure includes:

[0155] An analysis unit 701, used to analyze the search intent of the search request;

[0156] A retrieval unit 702, configured to retrieve the entity name in the data storage system according to the search keyword in the search request to obtain the query result when the search intention is unclear;

[0157] The retrieval unit 702 is also used to retrieve multiple entity attributes in the data storage system according to the search keywords in the search request when the search intention is clear, and obtain query results. The entity attributes include at least two of the entity name, entity description, entity field name and entity field description.

[0158] In some embodiments, the retrieval unit 702 includes: a first word segmentation module, which is used to segment the search keywords in the word segmenter according to multiple entity attributes to obtain multiple word segmentation results; a first acquisition module, which is used to obtain a compound query item according to the multiple word segmentation results; a retrieval module, which is used to query the query results matching the compound query conditions in the query search server ES according to the compound query conditions, and the data storage system includes the search server ES.

[0159] In some embodiments, the first word segmentation module includes: a first word segmenter, which is used to segment the search keywords according to the entity name and the entity field name, respectively, to obtain corresponding first word segmentation results and second word segmentation results; a second word segmenter, which is used to segment the search keywords according to the entity description and the entity field description, respectively, to obtain corresponding third word segmentation results and fourth word segmentation results.

[0160] In some embodiments, the first tokenizer is an N-gram tokenizer, and the second tokenizer is a Chinese tokenizer.

[0161] In some embodiments, the first acquisition module includes: a first acquisition sub-module, which is used to obtain the importance of the word segmentation results corresponding to the entity attributes according to the weight of the entity attributes, and the word segmentation results corresponding to the entity attributes are the word segmentation results obtained when the search keywords are segmented according to the entity attributes; a construction sub-module, which is used to construct a composite query condition based on the importance of multiple word segmentation results.

[0162] In some embodiments, the retrieval module includes: a second acquisition submodule, used to obtain the inverted index corresponding to the target word in the word segmentation results, and the compound query conditions include multiple word segmentation results; a retrieval submodule, used to query the search server ES for preliminary query results that match the target word based on the inverted index; a result synthesis submodule, used to synthesize the preliminary query results based on the compound query conditions to obtain query results that match the compound query conditions.

[0163] In some embodiments, the retrieval unit 702 is also used to complete the entity name in the data storage system according to the search keywords in the search request to obtain the query result when the search intention is unclear.

[0164] In some embodiments, the retrieval unit 702 also includes: a second word segmentation module, which is used to segment the search keywords in a third word segmenter according to the search keywords to obtain an initial word segmentation result, and the third word segmenter is a word segmenter that segments the entity name based on a preset regular expression; a determination module, which is used to determine the word segmentation prefix in the initial word segmentation result according to the Edge-NGram Token Filter; a retrieval module, which is used to complete the entity name according to the word segmentation prefix in the search server ES, and output the query result, and the query result includes an entity list.

[0165] In some embodiments, the analysis unit 701 includes: an identification module for determining the character length of the search keyword in the search request; a judgment module for judging whether the character length is greater than or equal to a first threshold, or a user judging whether the character length is less than or equal to a second threshold; a decision module, if the character length is greater than or equal to the first threshold, the search intention is clear, and if the character length is less than or equal to the second threshold, the search intention is unclear.

[0166] In some embodiments, the data processing device also includes a data update unit 703; the data update unit 703 includes: a first data acquisition module, which is used to obtain the metadata to be updated when receiving a metadata update request from the tenant, and the update request includes at least one of addition, deletion, replacement and version replacement; a second acquisition module, which is used to obtain the tenant index corresponding to the update request, and in the search server ES, each tenant creates at least one tenant index; the data update module is used to update the metadata to be updated under the tenant index in the search server ES.

[0167] In some embodiments, when the update request is for a new addition, the data update module includes: a third word segmentation module, which is used to segment the metadata to be updated in the word segmenter to obtain the word segmentation result; a construction module, which is used to construct an inverted index corresponding to each word in the word segmentation result, and store the inverted index corresponding to the word segmentation in the search server ES.

[0168] In some embodiments, when the update request is a version change, the data update unit 703 also includes: a second data acquisition module, used to acquire historical version data of the metadata to be updated; a first storage module, used to store the historical version data in the TafDB database when the historical version data is public cloud data, and the data storage system also includes the TafDB database; a second storage module, used to store the historical version data in the Mysql database when the historical version data is privatized data, and the data storage system also includes the Mysql database.

[0169] Figure 7The data processing device provided can execute the steps involved in the terminal in the above-mentioned corresponding method embodiments. Its implementation principles and technical effects are similar and will not be repeated here.

[0170] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the solution provided by any of the above embodiments.

[0171] According to an embodiment of the present disclosure, the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute a solution provided by any of the above embodiments.

[0172] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, which includes: a computer program, the computer program is stored in a readable storage medium, at least one processor of an electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program so that the electronic device executes the solution provided by any of the above embodiments.

[0173] According to an embodiment of the present disclosure, the present disclosure also provides a terminal, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the solution provided by any of the above embodiments.

[0174] Figure 8 8 is a schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit implementations of the present disclosure described and / or claimed herein.

[0175] like Figure 8 As shown, the electronic device 800 includes a computing unit 801, which can store data stored in a read-only memory (ROM) ( Figure 8 A computer program in ROM 802 is loaded from storage unit 708 to random access memory (RAM) ( Figure 8The computer programs in the RAM 803 are used to perform various appropriate actions and processes. The RAM 703 may also store various programs and data required for the operation of the electronic device 800. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. The input / output (I / O) interface ( Figure 8 The I / O interface 805 is also connected to the bus 804 .

[0176] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0177] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSP), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as the data processing method of the data map engine. For example, in some embodiments, the data processing method of the data map engine may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the data processing method of the data map engine described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the data processing method of the data map engine in any other appropriate manner (eg, by means of firmware).

[0178] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0179] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0180] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0181] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0182] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: Local Area Networks (LANs), Wide Area Networks (WANs), and the Internet.

[0183] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.

[0184] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0185] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A data processing method for a data map engine, comprising: Determine the search intent of the search request; If the search intention is unclear, the entity name is searched in the data storage system according to the search keywords in the search request to obtain the query results; If the search intention is clear, multiple entity attributes are searched in the data storage system according to the search keywords in the search request to obtain query results, and the entity attributes include at least two of the entity name, entity description, entity field name and entity field description.

2. The data processing method according to claim 1, wherein: The step of searching multiple entity attributes in the data storage system according to the search keyword in the search request to obtain the query result includes: According to the multiple entity attributes, the search keywords are segmented in a word segmenter to obtain multiple segmentation results; Obtaining a compound query condition according to the multiple word segmentation results; According to the compound query condition, a query result matching the compound query condition is searched in the search server ES, and the data storage system includes the search server ES.

3. The data processing method according to claim 2, wherein: According to the multiple entity attributes, the search keywords are segmented in a word segmenter to obtain multiple segmentation results, including: According to the entity name and the entity field name, the search keyword is segmented in a first segmenter to obtain a corresponding first segmentation result and a second segmentation result; According to the entity description and the entity field description, the search keyword is segmented in the second word segmenter to obtain corresponding third word segmentation results and fourth word segmentation results.

4. The data processing method according to claim 3, wherein: The first word segmenter is an N-gram word segmenter, and the second word segmenter is a Chinese word segmenter.

5. The data processing method according to any one of claims 2 to 4, wherein: The obtaining of a compound query condition according to the multiple word segmentation results includes: According to the weight of the entity attribute, the importance of the word segmentation result corresponding to the entity attribute is obtained, where the word segmentation result corresponding to the entity attribute is the word segmentation result obtained when the search keyword is segmented according to the entity attribute; The compound query condition is obtained according to the importance of multiple word segmentation results.

6. The data processing method according to any one of claims 2 to 5, wherein: The step of searching the search server ES for a query result matching the compound query condition according to the compound query condition includes: Obtaining an inverted index corresponding to a target word segmentation in the word segmentation result, wherein the compound query condition includes multiple word segmentation results; According to the inverted index, query the search server ES for preliminary query results that match the target word segmentation; The preliminary query results are integrated based on the compound query conditions to obtain query results matching the compound query conditions.

7. The data processing method according to any one of claims 1 to 6, wherein: If the search intention is unclear, the entity name is searched in the data storage system according to the search keyword in the search request to obtain the query result, including: If the search intention is not clear, the entity name is searched and completed in the data storage system according to the search keywords in the search request to obtain the query results.

8. The data processing method according to claim 7, wherein: The searching and completing the entity name in the data storage system according to the search keyword in the search request to obtain the query result includes: According to the search keyword, the search keyword is segmented in a third segmenter to obtain an initial segmentation result, wherein the third segmenter is a segmenter that segments the entity name based on a preset regular expression; Determining a segmentation prefix in the initial segmentation result according to an Edge-NGram Token Filter; In the search server ES, after completing the entity name according to the segmentation prefix, a query result is output, and the query result includes an entity list.

9. The data processing method according to any one of claims 1 to 8, wherein: Determining the search intent of the search request includes: Determining the character length of the search keyword in the search request; If the character length is greater than or equal to the first threshold, the search intention is clear; If the character length is less than or equal to the second threshold, the search intention is unclear.

10. The data processing method according to any one of claims 2 to 9, wherein: Also includes: When receiving a metadata update request from a tenant, obtaining metadata to be updated, wherein the update request includes at least one of adding, deleting, replacing, and changing a version; Obtain the tenant index corresponding to the update request. In the search server ES, each tenant has at least one tenant index created; In the search server ES, the metadata to be updated is updated under the tenant index.

11. The data processing method according to claim 10, wherein: When the update request is a new addition, updating the metadata to be updated in the search server ES includes: Segmenting the metadata to be updated in a word segmenter to obtain a word segmentation result; An inverted index corresponding to each word segmentation in the word segmentation result is constructed, and the inverted index corresponding to the word segmentation is stored in the search server ES.

12. The data processing method according to claim 10, wherein: When the update request is a version change, the method further includes: Collect historical version data of metadata to be updated; If the historical version data is public cloud data, the historical version data is stored in a TafDB database, and the data storage system also includes the TafDB database; If the historical version data is privatized data, the historical version data is stored in a MySQL database, and the data storage system also includes the MySQL database.

13. A data processing device, comprising: An analysis unit, used for analyzing the retrieval intention of the retrieval request; A retrieval unit, configured to retrieve the entity name in the data storage system according to the retrieval keyword in the retrieval request to obtain the query result when the retrieval intention is unclear; The retrieval unit is also used to retrieve multiple entity attributes in the data storage system according to the retrieval keywords in the retrieval request to obtain query results when the retrieval intention is clear, and the entity attributes include at least two of the entity name, entity description, entity field name and entity field description.

14. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the data processing method according to any one of claims 1 to 12.

15. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the data processing method according to any one of claims 1 to 12.

16. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the steps of the data processing method according to any one of claims 1 to 12 are implemented.