Methods, apparatus, computer devices, and storage media for a data system

By introducing a field extraction interface and multiple extraction rules into the data system, the field extraction process for unstructured data has been optimized, solving the problem of low efficiency in existing technologies and realizing efficient and intuitive field extraction operations.

CN116089596BActive Publication Date: 2025-12-30上海炎凰数据科技有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310020066.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-06
Publication Date
2025-12-30
Estimated Expiration
2043-01-06

AI Technical Summary

Technical Problem

Unstructured data has diverse formats and is difficult to standardize and understand. Existing technologies are cumbersome, time-consuming, and lack intuitive feedback in the field extraction process, resulting in a waste of human and material resources.

Method used

A data system is provided that receives user input through a field extraction interface, first locates and previews candidate data items, then selects sample data items, and finally performs field extraction operations based on conditions. The system optimizes the field extraction process by using rules such as key-value extraction, JSON extraction, IP address extraction, and regular expression extraction.

Benefits of technology

It improves the efficiency of field extraction from unstructured data, reduces resource waste and time costs caused by rule adjustments, provides intuitive user feedback, and reduces operational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116089596B_ABST
    Figure CN116089596B_ABST
Patent Text Reader

Abstract

A method for a data system for storing at least unstructured data and displaying a field extraction interface for performing a field extraction operation is provided, the method comprising: receiving, via the interface, a first input comprising a first condition for locating a candidate data item in the unstructured data for which the field extraction operation is to be performed; in response to receiving the first input, querying the unstructured data for the candidate data item that satisfies the first condition and providing, via the interface, a preview of the candidate data item; receiving, via the interface, a second input for locating a sample data item in the candidate data item; in response to receiving the second input, displaying the sample data item in the interface; receiving, via the interface, a third input comprising a second condition comprising a field extraction rule for performing the field extraction operation on the sample data item; and in response to receiving the third input, performing the field extraction operation on the sample data item based on the second condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for a data system. Background Technology

[0002] Data in computer information systems can be divided into structured data and unstructured data. Unstructured data is characterized by irregular or incomplete data structures, a lack of associated predefined data models, and inconvenience in representing it using two-dimensional logical tables in a database. Unstructured data can include documents, text, reports, images, logs, etc. Therefore, unstructured data has a wide variety of formats, and it is technically more difficult to standardize and understand than structured data.

[0003] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention

[0004] This disclosure provides a method, apparatus, computer device, computer-readable storage medium, and computer program product for a data system.

[0005] According to one aspect of this disclosure, a method is provided for a data system for storing at least unstructured data and displaying a field extraction interface for a user to perform field extraction operations. The method includes: receiving a first input via the field extraction interface, the first input including a first condition for locating candidate data entries in the unstructured data for which field extraction operations are to be performed; in response to receiving the first input, querying the unstructured data for candidate data entries that satisfy the first condition, and providing a preview of the candidate data entries via the field extraction interface; receiving a second input via the field extraction interface for locating sample data entries among the candidate data entries; in response to receiving the second input, displaying the sample data entries in the field extraction interface; receiving a third input via the field extraction interface, the third input including a second condition for performing field extraction operations on the sample data entries; and in response to receiving the third input, performing field extraction operations on the sample data entries based on the second condition.

[0006] According to another aspect of this disclosure, an apparatus for a data system is provided, the data system being used to at least store unstructured data and display a field extraction interface for a user to perform field extraction operations, the apparatus comprising: a first module for receiving a first input via the field extraction interface, the first input including a first condition, the first condition being used to locate candidate data entries in the unstructured data for which field extraction operations are to be performed; a second module for, in response to receiving the first input, querying the unstructured data for candidate data entries that satisfy the first condition, and providing a preview of the candidate data entries via the field extraction interface; a third module for receiving a second input via the field extraction interface, the second input being used to locate sample data entries among the candidate data entries; a fourth module for, in response to receiving the second input, displaying the sample data entries in the field extraction interface; a fifth module for receiving a third input via the field extraction interface, the third input including a second condition, the second condition including field extraction rules for performing field extraction operations on the sample data entries; and a sixth module for, in response to receiving the third input, performing field extraction operations on the sample data entries based on the second condition.

[0007] According to another aspect of this disclosure, a computer device is provided, comprising: at least one processor; and at least one memory having a computer program stored thereon, wherein when executed by the at least one processor, the computer program causes the at least one processor to perform the above-described method for a data system.

[0008] According to another aspect of this disclosure, a computer-readable storage medium is provided that stores a computer program thereon, which, when executed by a processor, causes the processor to perform the above-described method for a data system.

[0009] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, causes the processor to perform the above-described method for a data system.

[0010] These and other aspects of this disclosure will be apparent from the embodiments described below, and will be elucidated with reference to the embodiments described below. Attached Figure Description

[0011] Further details, features, and advantages of this disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0012] Figure 1 This is a schematic diagram illustrating an example system in which various methods described herein may be implemented according to exemplary embodiments;

[0013] Figure 2This is a flowchart illustrating a method for a data system according to an exemplary embodiment;

[0014] Figure 3 This is a schematic diagram illustrating an example field extraction interface according to an exemplary embodiment;

[0015] Figure 4 This is a flowchart illustrating the process of performing field extraction operations on sample data entries according to an exemplary embodiment;

[0016] Figure 5 This is a schematic diagram illustrating an example field extraction interface according to an exemplary embodiment;

[0017] Figure 6 This is a schematic diagram illustrating an example field extraction interface according to an exemplary embodiment;

[0018] Figure 7 This is a flowchart illustrating the process of adjusting the field extraction operation according to an exemplary embodiment;

[0019] Figure 8 This is a schematic diagram illustrating an example field extraction interface according to an exemplary embodiment;

[0020] Figure 9 This is a schematic block diagram illustrating an apparatus for a data system according to an exemplary embodiment; and

[0021] Figure 10 This is a block diagram illustrating an exemplary computer device that can be applied to an exemplary embodiment. Detailed Implementation

[0022] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0023] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. As used herein, the term "multiple" means two or more, and the term "based on" should be interpreted as "at least partially based on". Furthermore, the terms "and / or" and "at least one of..." cover any one of the listed items and all possible combinations thereof.

[0024] As shown above, compared to structured data, unstructured data has a much wider variety of formats, and it is also more difficult to standardize and understand at the technical level. Faced with such unstructured data, in most cases, it is necessary to extract key information from the data using certain fixed rules; this process is also known as field extraction for unstructured data. However, this process is often cumbersome, time-consuming, and lacks intuitive feedback, making it difficult to efficiently extract fields from unstructured data even with significant human and material resources.

[0025] In view of this, the present disclosure provides a method, apparatus, computer device, computer-readable storage medium, and computer program product for a data system.

[0026] Exemplary embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0027] Figure 1 This is a schematic diagram illustrating an example system 100 in which various methods described herein may be implemented according to exemplary embodiments.

[0028] refer to Figure 1 The system 100 includes a client device 110, a server 120, and a network 130 that communicatively couples the client device 110 and the server 120.

[0029] Client device 110 includes a display 114 and a client application (APP) 112 that can be displayed on the display 114. Client application 112 can be an application that needs to be downloaded and installed before running, or a lightweight application (liteapp). If client application 112 is an application that needs to be downloaded and installed before running, client application 112 can be pre-installed on client device 110 and activated. If client application 112 is a mini-app, user 102 can directly run client application 112 on client device 110 without installing it, by searching for client application 112 in the host application (e.g., by the name of client application 112) or scanning the graphic code of client application 112 (e.g., barcode, QR code, etc.). In some embodiments, client device 110 can be any type of mobile computing device, including mobile computers, mobile phones, wearable computing devices (e.g., smartwatches, head-mounted devices including smart glasses, etc.), or other types of mobile devices. In some embodiments, the client device 110 may alternatively be a fixed computer device, such as a desktop computer, server computer, or other type of fixed computer device.

[0030] Server 120 is typically a server deployed by an Internet Service Provider (ISP) or Internet Content Provider (ICP). Server 120 can represent a single server, a cluster of multiple servers, a distributed system, or a cloud server providing basic cloud services (such as cloud databases, cloud computing, cloud storage, and cloud communications). It will be understood that, although... Figure 1 The diagram shows that server 120 communicates with only one client device 110, but server 120 can provide background services to multiple client devices simultaneously.

[0031] Examples of network 130 include combinations of local area networks (LANs), wide area networks (WANs), personal area networks (PANs), and / or communication networks such as the Internet. Network 130 can be wired or wireless. In some embodiments, technologies and / or formats including Hypertext Markup Language (HTML), Extensible Markup Language (XML), etc., are used to process data exchanged through network 130. Furthermore, encryption technologies such as Secure Sockets Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can be used to encrypt all or some of the links. In some embodiments, custom and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.

[0032] For the purposes of this disclosure's embodiments, Figure 1 In the example, client application 112 may be a field extraction application that provides various functions related to field extraction from structured and unstructured data, such as displaying a field extraction interface for users to perform field extraction operations. Correspondingly, server 120 may be a server used in conjunction with the field extraction application. Server 120 may provide various field extraction services to client application 112 running on client device 110, such as receiving input from user 102 transmitted to client application 112 via network 130, processing the user input, and returning the processing results to client device 110 via network 130, etc.

[0033] Figure 2 This is a flowchart illustrating a method 200 for a data system according to an exemplary embodiment. Method 200 can be implemented on a client device (e.g., ...). Figure 1 The execution is performed at the client device 110 shown, that is, the execution entity of each step of method 200 can be... Figure 1 The client device 110 shown. In some embodiments, method 200 can be performed on a server (e.g., Figure 1The method 200 is executed at server 120 (as shown in the figure). In some embodiments, the method 200 may be executed in combination by a client device (e.g., client device 110) and a server (e.g., server 120).

[0034] According to some embodiments, the data system may be deployed at client device 110. According to some embodiments, the data system may be deployed at server 120.

[0035] According to some embodiments, the data system can be used to store at least unstructured data. In the example, the data system can store both structured and unstructured data.

[0036] According to some embodiments, the data system may be configured to provide a field extraction interface through which a user (e.g., user 102) can perform field extraction operations.

[0037] As an example and not a limitation, the data system is deployed on client devices (e.g., Figure 1 In the case of the client device 110 shown, method 200 can be executed at the client device. Specifically, the client application in the client device (e.g., Figure 1 The client application 112 shown (which may be a field extraction application) can retrieve data entries to be extracted from the data system in response to input from a user (e.g., user 102) to the client application, and perform field extraction operations on some or all of these data entries in response to subsequent input from the user. The data system may provide a field extraction interface via the client application to facilitate the user to perform field extraction operations using the client application.

[0038] As an example and not a limitation, in the context of data systems deployed on servers (e.g., Figure 1 In the case of server 120 shown, method 200 can be performed by a client device (e.g., Figure 1 The client device 110 shown herein) and the server are combined for execution. Specifically, the data system can be implemented via a network (e.g., Figure 1 The network 130 shown here sends data to the client application in the client device (e.g., Figure 1The client application 112 shown (which may be a field extraction application) provides a field extraction interface to facilitate user actions (e.g., providing user input). The client application may transmit multiple inputs from a user (e.g., user 102) to a server via a network. The server may process these user inputs (e.g., in response to user input, retrieve data entries from a data system for field extraction, perform field extraction on some or all of these data entries, etc.) and return the processing results to the client device via the network for display on the field extraction interface of the client application.

[0039] As an example and not a limitation, in the context of data systems deployed on servers (e.g., Figure 1 In the case of server 120 shown, method 200 can also be performed at the server. For example, the server may be equipped with input / output (I / O) devices, allowing users to provide input via input devices such as touch input devices, gesture input devices, mice, and keyboards. The server processes the user input and returns the processing results to the user via an output device such as a display. During field extraction operations, the data system may provide a field extraction interface (e.g., via a display) to facilitate users performing field extraction operations at the server.

[0040] The following text uses client device 110 as an example to describe the steps of method 200 in detail. Figure 2 As shown, method 200 includes:

[0041] Step S210: Receive a first input via the field extraction interface. The first input includes a first condition, which is used to locate candidate data entries in the unstructured data for which field extraction operation is to be performed.

[0042] Step S220: In response to receiving the first input, query the unstructured data for candidate data entries that meet the first condition, and provide a preview of the candidate data entries via the field extraction interface;

[0043] Step S230: Receive a second input via the field extraction interface. The second input is used to locate the sample data entry in the candidate data entry.

[0044] Step S240: In response to receiving the second input, display sample data entries in the field extraction interface;

[0045] Step S250: Receive a third input via the field extraction interface. The third input includes a second condition, which includes field extraction rules for performing field extraction operations on sample data entries; and

[0046] Step S260: In response to receiving the third input, perform field extraction operation on the sample data entries based on the second condition.

[0047] In step S210, a first input may be received via a field extraction interface. According to some embodiments, the field extraction interface may include various graphical interface elements that facilitate a user's field extraction operation; some example graphical interface elements will be described below with reference to the accompanying drawings. In the example, the first input may be provided by the user. In the example, the first input may include a first condition, which can be used to locate candidate data entries in the unstructured data for field extraction.

[0048] As used herein, the term candidate data entry refers to a data entry in the unstructured data stored by a data system from which field extraction operations are to be performed in order to extract useful information (e.g., desired fields). That is, a candidate data entry may be a subset of the stored unstructured data.

[0049] In the example, the first condition may include information for identifying candidate data entries. Thus, a user can conveniently locate candidate data entries in unstructured data for which a field extraction operation (e.g., extracting a desired field) is to be performed by providing a first input that includes the first condition.

[0050] In step S220, all data entries satisfying the first condition included in the received first input can be queried in the unstructured data, and these data entries can be used as candidate data entries. Further, a preview of the candidate data entries can be provided in the field extraction interface, for example, displaying the candidate data entries in a corresponding preview area of ​​the field extraction interface. Users can view these candidate data entries using operations such as mouse scrolling or page turning (e.g., to delete unwanted data entries).

[0051] In step S230, a second input can be received via a field extraction interface. In this example, the first input can be provided by the user. In this example, the second input can be used to locate a sample data entry among the candidate data entries.

[0052] As used herein, the term "sample data entry" refers to a data entry selected as a sample from candidate data entries. After the field extraction operation on the sample data entry is completed, the same or similar field extraction operation can be performed on the remaining candidate data entries other than the sample data entry. In related technologies, when it is necessary to adjust the field extraction rules (e.g., due to unsatisfactory field extraction results), it is necessary to roll back one or more previous field extraction operations performed on all candidate data entries, which wastes a lot of processing resources, manpower, and time. Compared with related technologies, performing field extraction on a single sample data entry before performing batch field extraction on a large number of candidate data entries can effectively save the waste of processing resources and manpower costs required for user inspection caused by adjusting field extraction rules during the field extraction operation, and also save the time required to perform the operation, especially when the number of candidate data entries is large and / or the composition of candidate data entries is complex.

[0053] In step S240, a sample data entry can be displayed in the field extraction interface in response to receiving the second input. In this example, the sample data entry can be renamed and saved for subsequent field extraction operations. For instance, the sample data entry can be renamed to sample_entry and saved.

[0054] In step S250, a third input can be received via a field extraction interface. In this example, the third input can be provided by the user. The third input may include a second condition, which may include field extraction rules for performing field extraction operations on sample data entries. Thus, the user can conveniently perform field extraction operations on sample data entries by providing a third input that includes the second condition.

[0055] In step S260, before performing field extraction operations on the remaining candidate data entries other than the sample data entries, the sample data entries may be extracted based on the second condition included in the received third input.

[0056] According to embodiments of this disclosure, method 200 first determines a subset of unstructured data as candidate data entries for field extraction in response to receiving a first input. Then, it determines sample data entries among the candidate data entries in response to receiving a second input. Finally, it performs field extraction on the sample data entries based on the received third input. Method 200 avoids the waste of processing resources, manpower, and time costs caused by rolling back one or more field extraction operations performed on all candidate data entries due to adjustments in field extraction rules. Furthermore, it provides intuitive feedback to the user during the field extraction process, thereby reducing the complexity of field extraction operations on unstructured data and improving the efficiency of field extraction from unstructured data.

[0057] According to some embodiments, the second condition may include at least one of the following field extraction rules: KV extraction rule, JSON extraction rule, IP address extraction rule, or regular expression extraction rule.

[0058] As used in this article, the term KV extraction can refer to extracting the field KEY from content of the form KEY=VALUE, with a value of VALUE.

[0059] As used in this article, JSON (JavaScript Object Notation) is a commonly used lightweight data exchange format, and the term JSON extraction can refer to the extraction of field values ​​from content that conforms to the JSON format.

[0060] As used in this article, the term IP address extraction can refer to extracting information such as country, province, city, and region from content containing IP addresses.

[0061] As used in this article, regular expression extraction refers to field extraction based on regular expressions. A regular expression, also known as a rule expression, is a text pattern. Regular expressions include ordinary characters (e.g., letters from a to z) and special characters (called "metacharacters"), and are a concept in computer science. Regular expressions use a single string to describe and match a series of strings that match a certain syntax rule, and are typically used to retrieve and replace text that conforms to a specific pattern (rule).

[0062] Figure 3 This is a schematic diagram illustrating an example field extraction interface 300 according to an exemplary embodiment. For example... Figure 3 As shown, the example field extraction interface 300 may include an extraction process area 310 and a field extraction preview area 320.

[0063] As shown in the figure, in the extraction process area 310, the current extraction process 312 includes field extraction rule 1 and field extraction rule 2. Field extraction rule 1 is collapsed. Users can expand field extraction rule 1 by clicking the "+" on the left side of field extraction rule 1 to browse, edit, save, or delete it. Unlike field extraction rule 1, field extraction rule 2 is expanded. Below field extraction rule 2, a source field drop-down box 314 and a select extraction rule drop-down box 316 are displayed. Delete and save icons are also displayed to the right of the select extraction rule drop-down box 316. As shown in the figure, the current source field drop-down box 314 displays "remote," indicating that the field for which field extraction rule 2 is applied is "remote." Furthermore, the current select extraction rule drop-down box 316 displays "IP address extraction," indicating that field extraction rule 2 is IP address extraction. It is understood that the layout of the extraction process area 310 can be adjusted according to actual needs, and this disclosure does not impose any restrictions on it.

[0064] As shown in the figure, the field extraction preview area 320 currently displays the extraction results tab. As an example, the field extraction preview area 320 may also include a data entry tab to display, for example, all data stored in the data system (e.g., unstructured data, structured data, or combinations thereof), candidate data entries filtered based on the received first input, sample data entries selected based on the received second input, etc. As mentioned earlier, in the extraction process area 310, field extraction rule 2 is expanded, and therefore the extraction results displayed in the extraction results tab of the field extraction preview area 320 are associated with field extraction rule 2. That is, the extraction results tab displays the results obtained by applying IP address extraction to the field remote. In the extraction results tab, as shown in the figure, a header 322 and extraction results 324 are displayed. As an example, the header 322 includes the name of the source field (e.g., remote in this case), country, province, city, and ISP (Internet Service Provider). It's important to note that the ellipsis in header 322 can indicate other information besides country, province, city, and ISP that can be obtained from the `remote` field (in this case, the IP address). Of course, the ellipsis in header 322 can also represent other fields in a data entry besides the `remote` field. For example, a data entry can be divided into various fields (different fields may overlap), and each field can have its own name (e.g., `remote`) to distinguish it from other fields in the data entry. In the example, the names of the fields extracted using JSON extraction rules and KV extraction rules can be automatically generated based on the content of the extracted fields. In the example, the names of the fields extracted using IP address extraction rules can be preset. Extraction result 324 shows only one extraction result, with its `remote` field being the IP address 116.179.33.81. The information obtained from extracting this IP address includes {Country: China, Province: Shanxi, City: Yangquan, ISP: China Unicom}. The ellipsis below extraction result 324 can indicate other extraction results. In the scenario shown, the extraction results include 91 entries, displayed across 5 pages, with a maximum of 20 entries per page. It is understood that the layout of the field extraction preview area 320 can be adjusted according to actual needs, and this disclosure does not impose any restrictions on it.

[0065] It should be understood that Figure 3 The example field extraction interface 300 shown is merely illustrative, and those skilled in the art, upon learning of the inventive concept of this disclosure, can make any suitable modifications and variations thereof, all of which are within the scope of this disclosure. The example field extraction interface described in this disclosure is not intended to be limiting.

[0066] As an example, not a limitation, the sample data entry is “116.179.33.81[01 / Aug / 2014:09:37:01 -0700]"GET / favicon.ico HTTP / 1.1"404 170"-""Mozilla / 5.0(Macintosh;U;PPC Mac OS X Mach-O;en-US;sl:1.7.7)Gecko / 00010100Firefox / 1.0.3"{"from":"apache","to":"access"}”. In the example, this sample data entry can be renamed to sample_entry and saved. Similar to the expanded field extraction rule 2 shown, it can be assumed that… Figure 3 The collapsed field extraction rule 1 in the sample_entry includes applying a regular expression extraction rule to extract the IP address portion "116.179.33.81". In the example, for the field that conforms to the above regular expression extraction rules (which is the IP address part "116.179.33.81"), it can be renamed to remote so that the IP address extraction rules can be applied to the remote field to extract information such as location and ISP, as described above; while for the field that does not conform to the above regular expression extraction rules (which is "[01 / Aug / 2014:09:37:01 -0700]"GET / favicon.ico HTTP / 1.1"404 170"-""Mozilla / 5.0(Macintosh;U;PPC Mac OSX Mach-O;en-US;sl:1.7.7)Gecko / 00010100Firefox / 1.0.3"{"from":"apache","to":"access"}"), it can be renamed to message_1.

[0067] According to some embodiments, the second condition may include a first field extraction rule, which can be selected from a group consisting of key-value extraction rules, JSON extraction rules, IP address extraction rules, and regular expression extraction rules. Further, performing field extraction on sample data entries based on the second condition may include: applying the first field extraction rule to the sample data entries to obtain one or more level-1 branch extracted fields, wherein the one or more level-1 branch extracted fields may include fields in the sample data entries that conform to the first field extraction rule.

[0068] Based on the example above, a first field extraction rule (i.e., a regular expression extraction rule) can be applied to the sample data entry `sample_entry` to obtain two level 1 branch extraction fields (i.e., the `remote` field and the `message_1` field). For these two level 1 branch extraction fields, the `remote` field is the field in the sample data entry `sample_entry` that conforms to the regular expression extraction rule, while the `message_1` field is the field in the sample data entry `sample_entry` that does not conform to the regular expression extraction rule. Therefore, field extraction rules can be conveniently applied to sample data entries to obtain one or more level 1 branch extraction fields, thereby improving the efficiency of field extraction from unstructured data.

[0069] Figure 4 This is a flowchart illustrating a process 400 of performing a field extraction operation on sample data entries according to an exemplary embodiment. Process 400 can be performed on a client device (e.g., Figure 1 The process is executed at the client device 110 shown, meaning that the execution entity of each step of process 400 can be... Figure 1 The client device 110 shown. In some embodiments, the process 400 can be performed on a server (e.g., Figure 1 The process 400 is executed at server 120 (as shown in the figure). In some embodiments, the process 400 may be executed in combination by a client device (e.g., client device 110) and a server (e.g., server 120).

[0070] In some embodiments, the second condition may include a first field extraction rule to an nth field extraction rule, each of which may be selected from a group consisting of a KV extraction rule, a JSON extraction rule, an IP address extraction rule, and a regular expression extraction rule, where n is an integer and n≥2.

[0071] In the following text, taking client device 110 as the executing entity as an example, the various steps of process 400 are described in detail. According to some embodiments, process 400 may be a further description of step S260 of method 200 described above. Figure 4 As shown, process 400 includes:

[0072] Step S410: Apply the first field extraction rule to the sample data entries to obtain one or more level 1 branch extraction fields, wherein the one or more level 1 branch extraction fields include fields in the sample data entries that conform to the first field extraction rule and fields in the sample data entries that do not conform to the first field extraction rule.

[0073] Step S420: Apply the nth field extraction rule to at least a portion of the extracted fields of one or more n-1 level branches to obtain one or more n-level branch extracted fields, wherein the one or more n-level branch extracted fields include fields in the n-1 level branch extracted fields that conform to the nth field extraction rule and fields in the n-1 level branch extracted fields that do not conform to the nth field extraction rule.

[0074] Continuing with the example above, the sample data entry is "116.179.33.81[01 / Aug / 2014:09:37:01 -0700]"GET / favicon.icoHTTP / 1.1"404170"-""Mozilla / 5.0(Macintosh;U;PPC Mac OS X Mach-O;en-US;sl:1.7.7)Gecko / 00010100Firefox / 1.0.3"{"from":"apache","to":"access"}". After applying the first field extraction rule (i.e., the regular expression extraction rule) to the sample data entry sample_entry to obtain two level 1 branch extraction fields (i.e., the remote field and the message_1 field), the second field extraction rule (here, the IP address extraction rule) is applied to the two level 1 branch extraction fields remote and message_1 respectively to obtain multiple level 2 branch extraction fields. These level 2 branch extraction fields include: the country field, province field, city field, and ISP field obtained by applying the IP address extraction rule to the level 1 branch extraction field remote, as well as the message_1 field itself (because the message_1 field does not contain the IP address part, it does not conform to the IP address extraction rule).

[0075] Furthermore, a third field extraction rule (e.g., a regular expression extraction rule) can be applied to the obtained second-level branch extracted field message_1 to extract the date and time portion "01 / Aug / 2014:09:37:01". In the example, for the field that conforms to the above regular expression extraction rules (which is the date and time part "01 / Aug / 2014:09:37:01"), it can be renamed to time so that the appropriate extraction rules can be applied to the time field to extract the specific date and time information. For the field that does not conform to the third field extraction rules (which is "[-0700]"GET / favicon.ico HTTP / 1.1"404 170"-""Mozilla / 5.0(Macintosh;U;PPC Mac OS X Mach-O;en-US;sl:1.7.7)Gecko / 00010100Firefox / 1.0.3"{"from":"apache","to":"access"}"), it can be renamed to message_2. Therefore, based on applying the first field extraction rule to the sample data entries to obtain one or more level 1 branch extraction fields, subsequent field extraction rules can be applied to obtain hierarchical branch extraction fields, thereby achieving hierarchical field extraction and improving the efficiency of field extraction for unstructured data.

[0076] In some embodiments, regular expression extraction rules may include extraction based on regular expressions. Further, the regular expression may be input by the user through the field extraction interface, or it may be automatically generated based on the content selected by the user through the field extraction interface. In the example, the regular expression may also be obtained by the user after partially modifying the automatically generated expression. In this case, the regular expression automatically generated by word selection may not be universal enough. The user further modifies the automatically generated regular expression to make it more universal, for example, applicable to more data entries to be extracted. Thus, unlike the complex operation of manually inputting regular expressions for field extraction in related technologies, this allows for word selection on the field extraction interface to automatically generate regular expressions. Users can extract partial content from sample data entries or branches at various levels. The system automatically generates regular expressions based on the selected content (including but not limited to the arrangement rules of numbers, letters, symbols, etc.) and the complete sample data entry (such as the specific order of the selected content in the complete sample data entry, whether the selected content contains certain special characters before and after it, etc.), thereby eliminating the complex operation of manually inputting regular expressions by the user.

[0077] Additionally, the content extracted via regular expressions can be highlighted in the field extraction interface so that users can quickly locate and modify it (e.g., re-extract).

[0078] Figure 5 This is a schematic diagram illustrating an example field extraction interface 500 according to an exemplary embodiment. For example... Figure 5 As shown, the example field extraction interface 500 may include an extraction process area 510. However, the example field extraction interface 500 may additionally include a field extraction preview area, similar to... Figure 3 The field extraction preview area 320 depicted in the document is not subject to any limitation in this disclosure.

[0079] As shown in the figure, in the extraction process area 510, the current extraction process 512 includes field extraction rules 1-3. Field extraction rules 1 and 2 are collapsed. Users can expand field extraction rules 1 and / or 2 by clicking the "+" to the left of field extraction rules 1 and / or 2 to browse, edit, save, or delete them. Unlike field extraction rules 1-2, field extraction rule 3 is expanded. Below field extraction rule 3, a source field dropdown box 514 and a select extraction rule dropdown box 516 are displayed. Delete and save icons are also displayed to the right of the select extraction rule dropdown box 516. As shown in the figure, the current source field dropdown 514 displays message_1, which means that the field to be applied to field extraction rule 3 is message_1. In addition, the current selection extraction rule dropdown 516 displays regular expression extraction, which means that field extraction rule 3 is regular expression extraction.

[0080] Below dropdown menus 514 and 516, a prompt message 1 is displayed: "You can select fields or edit regular expressions in the boxes below." Below this prompt message is an editable area 517, displaying the content of the message_1 field. Below this editable area 517, a prompt message 2 is displayed: "Hide Regular Expression," along with a regular expression editing box 518. As an example, and not a limitation, the "Edit Regular Expression" option in prompt message 1 can be clickable, so that when the user clicks it, the regular expression editing box 518 changes from a non-editable state (e.g., shown as gray) to an editable state. As an example, and not a limitation, prompt message 2 can also be clickable, so that when the user clicks it, the regular expression editing box 518 can be hidden to save display area on the field extraction interface.

[0081] In the editable area 517, the user can select "01 / Aug / 2014:09:37:01" by highlighting the text. At this point, a corresponding regular expression (displayed as "example regular expression A" in the current view) is automatically generated in the regular expression editing box 518 (regardless of whether it is in a non-editable or editable state). As mentioned above, the content extracted by the regular expression is highlighted in the field extraction interface so that the user can quickly locate and modify it. Suppose that the user unintentionally selects, for example, "01 / Aug / 2014:09:37:01 -07", the user can realize the error by checking the highlighted portion (e.g., in the editable area 517). The user can then re-select the correct field in the editable area 517, i.e., "01 / Aug / 2014:09:37:01". During this process, the regular expression automatically generated in the regular expression editing box 518 can automatically change according to the adjusted content. Similarly, users can also edit the regular expression in the regular expression editing box 518, causing the highlighted portion in the editable area 517 to change according to the user's adjustments to the regular expression. After completing the field extraction operation of field extraction rule 3, the selected fields (i.e., fields that satisfy field extraction rule 3) and / or the unselected fields (i.e., fields that do not satisfy field extraction rule 3) can be renamed and saved by clicking the save button. For example, “01 / Aug / 2014:09:37:01” can be renamed to “time” and saved, and “[-0700]"GET / favicon.ico HTTP / 1.1"404 170"-""Mozilla / 5.0(Macintosh;U;PPC Mac OS X Mach-O;en-US;sl:1.7.7)Gecko / 00010100Firefox / 1.0.3"{"from":"apache","to":"access"}” can be renamed to “message_2” and saved, as described above. It is understood that the layout of the extraction process area 510 can be adjusted according to actual needs, and this disclosure does not impose any restrictions on it.

[0082] It should be understood that Figure 5 The example field extraction interface 500 shown is merely illustrative, and those skilled in the art, upon learning of the inventive concept of this disclosure, can make any suitable modifications and variations thereof, all of which are within the scope of this disclosure. The example field extraction interface described in this disclosure is not intended to be limiting.

[0083] Figure 6This is a schematic diagram illustrating an example field extraction interface 600 according to an exemplary embodiment. For example... Figure 6 As shown, the example field extraction interface 600 may include an extraction process area 610. However, the example field extraction interface 600 may additionally include a field extraction preview area, similar to... Figure 3 The field extraction preview area 320 depicted in the document is not subject to any limitation in this disclosure.

[0084] As shown in the figure, in the extraction process area 610, the current extraction process 612 includes field extraction rules 1 to n. Currently, field extraction rule n is expanded, and below field extraction rule n are a source field dropdown box 614 and a select extraction rule dropdown box 616. Delete and save icons are also displayed to the right of the select extraction rule dropdown box 616. As shown in the figure, the current source field dropdown box 614 displays message_2, indicating that the field to be applied to field extraction rule n is message_2. Furthermore, the current select extraction rule dropdown box 616 displays regular expression extraction, indicating that field extraction rule n is regular expression extraction. In the example, the appropriate source field can be selected from several candidate fields that have been expanded downwards by clicking the triangle icon on the right side of dropdown box 614. As shown in the figure, the candidate fields include "remote", "message_1", "time", and "message_2".

[0085] Below dropdown lists 514 and 516, a tooltip 1 is displayed: "You can select fields or edit regular expressions in the boxes below." Below this tooltip is an editable area 617, which displays the content of the message_2 field. In the current view, tooltip 1 and part of the editable area 617 are covered by the expanded dropdown list 614. Below the editable area 617, a tooltip 2, "Hide Regular Expression," and a regular expression editing box 618 are displayed. For a description of the editable area 617 and the regular expression editing box 618, please refer to [link to documentation / section / etc.]. Figure 5 The descriptions of the editable area 517 and the regular expression edit box 518 are not repeated here.

[0086] In the editable area 617, users can select "{"from":"apache","to":"access"}" by highlighting words. At this point, a corresponding regular expression (displayed as "example regular expressionB" in the current view) is automatically generated in the regular expression editing box 618 (regardless of whether it is in a non-editable or editable state). After completing the field extraction operation according to field extraction rule n, the selected fields (i.e., fields satisfying field extraction rule n) and / or unselected fields can be renamed and saved by clicking the save button. For example, "{"from":"apache","to":"access"}" can be renamed to json and saved. It is understood that the layout of the extraction process area 610 can be adjusted according to actual needs, and this disclosure does not impose any restrictions on it.

[0087] It should be understood that Figure 6 The example field extraction interface 600 shown is merely illustrative, and those skilled in the art, upon learning of the inventive concept of this disclosure, can make any suitable modifications and variations thereof, all of which are within the scope of this disclosure. The example field extraction interface described in this disclosure is not intended to be limiting.

[0088] In some embodiments, the selected content can be highlighted in the field extraction interface. This makes it easier for users to locate and modify the content.

[0089] As a supplement to the above embodiments, in response to the completion of the field extraction operation on the sample data entry, the field extraction operation can be applied to the remaining candidate data entries other than the sample data entry. This avoids the waste of processing resources, manpower, and time costs caused by rolling back one or more field extraction operations performed on all candidate data entries due to adjustments in the field extraction rules, and improves the efficiency of the field extraction operation.

[0090] In the example, additionally, the field extraction interface can also preview the following events: after the field extraction operation on the sample data entry is applied to the remaining candidate data entries (or to all original data entries), the percentage or number of data entries from which the corresponding new field can be extracted (e.g., corresponding to the extraction result of the sample data entry).

[0091] Figure 7 This is a flowchart illustrating a process 700 for adjusting a field extraction operation according to an exemplary embodiment. Process 700 can be performed on a client device (e.g., ...). Figure 1 The process is executed at the client device 110 shown, meaning that the execution entity of each step of process 700 can be... Figure 1 The client device 110 shown. In some embodiments, process 700 can be performed on a server (e.g., Figure 1 The process 700 is executed at server 120 (as shown in the diagram). In some embodiments, the process 700 may be executed in combination by a client device (e.g., client device 110) and a server (e.g., server 120).

[0092] In the following text, taking client device 110 as the executing entity as an example, the various steps of process 700 are described in detail. According to some embodiments, process 700 may be a further supplement to the steps of method 200 described above. Figure 7 As shown, process 700 includes:

[0093] Step S710: After applying the field extraction operation to the remaining candidate data entries in the candidate data entries other than the sample data entries, a preview of the field extraction result obtained by applying the field extraction operation to the candidate data entries is provided via the field extraction interface.

[0094] Step S720: Receive a fourth input via the field extraction interface, the fourth input being used to adjust one or more field extraction rules for the corresponding data included in the second condition; and

[0095] Step S730: In response to receiving the fourth input, the field extraction operation is re-executed on the corresponding data according to the adjusted one or more field extraction rules.

[0096] The following is combined Figure 8 The process 700 is described in detail. Figure 8 This is a schematic diagram illustrating an example field extraction interface 800 according to an exemplary embodiment. For example... Figure 8 As shown, the example field extraction interface 800 may include an extraction process area 810 and a field extraction preview area 820.

[0097] As shown in the figure, in the extraction process area 810, the current extraction process 812 includes field extraction rules 1 to n+1. Field extraction rules 1 to n are collapsed; users can expand these rules by clicking the "+" icon to their left to view, edit, save, or delete some or all of them. Unlike field extraction rules 1 to n, field extraction rule n+1 is expanded. Below n+1 are a source field dropdown 814 and a select extraction rule dropdown 816, with delete and save icons displayed to the right of the select extraction rule dropdown 816. As shown in the figure, the current source field dropdown 814 displays JSON (as referenced above). Figure 6 (As described), this indicates that the field to be applied in field extraction rule n+1 is JSON. In addition, the current selection extraction rule dropdown 816 shows JSON extraction, which means that field extraction rule n+1 is JSON extraction.

[0098] In the example, users can click the right-hand triangle icon in dropdown 816 to select a suitable extraction rule from several candidate rules that expand downwards, thus adjusting the previously selected extraction rule. Assume that the previous field extraction rule n+1 for the corresponding data (here, a JSON field) was key-value extraction. Considering that JSON extraction rules might be more suitable for extracting the data structure contained in a JSON field than key-value extraction rules, users can click the right-hand triangle icon in dropdown 816 to reselect a suitable extraction rule from several candidate rules (i.e., JSON extraction).

[0099] As shown in the figure, the field extraction preview area 820 currently displays the extraction results tab. As an example, the field extraction preview area 820 may also include a data entry tab to display, for example, all data stored in the data system (e.g., unstructured data, structured data, or combinations thereof), candidate data entries filtered based on the received first input, sample data entries selected based on the received second input, etc. As mentioned earlier, in the extraction process area 810, field extraction rule n+1 is expanded, and therefore the extraction results displayed in the extraction results tab of the field extraction preview area 820 are associated with field extraction rule n+1. That is, the extraction results tab displays the results obtained by applying the JSON extraction rule to the field 'json'. In the extraction results tab, as shown in the figure, a header 822 and extraction results 824 are displayed. As an example, the header 822 includes the name of the source field (e.g., 'json' in this case), the source, and the target. It is important to note that the ellipsis in header 322 can indicate other information besides the source and target that can be obtained from the JSON extraction operation on the JSON field. In extraction result 824, only one extraction result is shown, with the JSON field "{"from":"apache","to":"access"}". The information obtained from the JSON extraction of this field includes {source:apache, target:access}. The ellipsis below extraction result 824 can indicate other extraction results. In the scenario shown, the extraction results include 91 entries, displayed across 5 pages, with a maximum of 20 entries per page. It is understood that the layout of the field extraction preview area 820 can be adjusted according to actual needs, and this disclosure does not impose any restrictions on it.

[0100] It should be understood that Figure 8 The example field extraction interface 800 shown is merely illustrative, and those skilled in the art, upon learning of the inventive concept of this disclosure, can make any suitable modifications and variations thereof, all of which are within the scope of this disclosure. The example field extraction interface described in this disclosure is not intended to be limiting.

[0101] Return to reference Figure 7In step S720, a fourth input can be received via a field extraction interface. In this example, the fourth input can be provided by the user. In this example, the fourth input can be used to adjust one or more field extraction rules for the corresponding data included in the second condition. In some embodiments, the corresponding data can be a sample data entry or a branch-extracted field obtained by performing a hierarchical field extraction operation on the sample data entry. It is understood that the corresponding data can be associated with one or more field extractions. That is, more than one field extraction rule can be applied to a sample data entry or a branch-extracted field derived from it.

[0102] In step S730, the field extraction operation can be re-executed on the corresponding data according to one or more adjusted field extraction rules. This allows users to specifically adjust unsuitable field extraction rules based on the inspection results, thereby improving the efficiency of the field extraction operation.

[0103] The following describes the complete process of extracting fields from the sample data entry “116.179.33.81[01 / Aug / 2014:09:37:01 -0700]"GET / favicon.ico HTTP / 1.1"404 170"-""Mozilla / 5.0(Macintosh;U;PPC Mac OS XMach-O;en-US;sl:1.7.7)Gecko / 00010100Firefox / 1.0.3"{"from":"apache","to":"access"}”.

[0104] Step a: Select the sample data item from the candidate data items through the field extraction interface;

[0105] Step b: Extract the content 116.179.33.81 using regular expressions through word selection. The field extraction interface generates the corresponding regular expressions. At the same time, the selected content is saved as the remote field, and the content not selected by word selection is saved as the message_1 field.

[0106] Step c: Further extract the remote field by extracting the IP address to obtain the country, province, city, region, and ISP fields corresponding to the IP address.

[0107] Step d: Extract the content 01 / Aug / 2014:09:37:01 from message_1 using regular expressions through word selection, generate the corresponding regular expression, save the selected content as the time field, and save the unselected content as the message_2 field.

[0108] Step e: Further extract the time field generated in step d, extracting 01, Aug, 2014, 09, 37, and 01 from the time field value to generate year, month, day, hour, minute, and second fields respectively.

[0109] Step f: Extract the content {"from":"apache","to":"access"} from message_2 using regular expressions, and save the selected content as a JSON field.

[0110] Step g: Further extract the JSON fields generated in step f, using the JSON extraction rule, to obtain two new fields: "Source" and "Target".

[0111] The field extraction results for the above sample data entries are shown in Table 1 below, based on the above steps:

[0112]

[0113]

[0114] Table 1

[0115] In some embodiments, the first condition may include at least one of the following: an SQL statement or a keyword. This allows for the rapid location of candidate data entries.

[0116] In some embodiments, the second input is the selection of a sample data entry from the candidate data entries. In the example, the user can select a sample data entry in a preview view of the candidate data entries. This simplifies the user operation.

[0117] As a supplement to the above embodiments, all field extraction rules included in the second condition can also be saved. Therefore, when field extraction operations are needed for new candidate data entries (e.g., from the same data source), the saved field extraction rules can be directly applied to the new candidate data entries, thereby effectively improving efficiency.

[0118] Understandably, in the case where no fourth input is received from the field extraction interface (i.e., the user does not expect to adjust any field extraction rules), the saved field extraction rules are the unadjusted field extraction rules included in the second condition. In the case where a fourth input is received via the field extraction interface (i.e., the user expects to adjust some or even all of the field extraction rules), the saved field extraction rules are the adjusted field extraction rules.

[0119] Although the operations are depicted in the accompanying drawings in a specific order, this should not be construed as requiring that the operations be performed in the specific order shown or in chronological order, nor should it be construed as requiring that all the operations shown be performed to obtain the desired result.

[0120] According to another aspect of this disclosure, an apparatus for a data system is provided.

[0121] Figure 9 This is a schematic block diagram illustrating an apparatus 900 for a data system according to an exemplary embodiment. Figure 9 As shown, device 900 may include:

[0122] The first module 910 is used to receive a first input via a field extraction interface. The first input includes a first condition, which is used to locate candidate data entries in unstructured data for field extraction.

[0123] The second module 920 is used to respond to receiving the first input, query candidate data entries that meet the first condition in the unstructured data, and provide a preview of the candidate data entries via the field extraction interface;

[0124] The third module 920 is used to receive a second input via the field extraction interface. The second input is used to locate the sample data entry in the candidate data entry.

[0125] The fourth module 940 is used to display sample data entries in the field extraction interface in response to receiving the second input;

[0126] The fifth module 950 is used to receive a third input via a field extraction interface. The third input includes a second condition, which includes field extraction rules for performing field extraction operations on sample data entries; and

[0127] The sixth module 960 is used to perform field extraction operations on sample data entries based on the second condition in response to receiving the third input.

[0128] It should be understood that Figure 9 The various modules of the device 900 shown can be connected to the reference. Figure 2 The steps in method 200 described correspond to each other. Therefore, the operations, features, and advantages described above for method 200 also apply to apparatus 900 and its included modules. For the sake of brevity, some operations, features, and advantages will not be repeated here.

[0129] While specific functions have been discussed above with reference to specific modules, it should be noted that the functions of the modules discussed herein can be divided into multiple modules, and / or at least some functions of multiple modules can be combined into a single module. The specific actions performed by the modules discussed herein include the specific module itself performing the action, or alternatively, the specific module calling or otherwise accessing another component or module that performs the action (or performs the action in conjunction with the specific module). Therefore, a specific module performing an action can include the specific module performing the action itself and / or another module that performs the action, called or otherwise accessed by the specific module.

[0130] It should also be understood that various techniques can be described in the general context of software hardware elements or program modules. The various modules described above with respect to 9 can be implemented in hardware or in hardware in combination with software and / or firmware. For example, these modules can be implemented as computer program code / instructions configured to execute in one or more processors and stored in a computer-readable storage medium. Alternatively, these modules can be implemented as hardware logic / circuit. For example, in some embodiments, one or more of the first modules 910 to the sixth modules 960 can be implemented together in a System on Chip (SoC). The SoC may include an integrated circuit chip (which includes a processor (e.g., a Central Processing Unit (CPU), microcontroller, microprocessor, digital signal processor (DSP), etc.), memory, one or more communication interfaces, and / or one or more components of other circuitry) and may optionally execute received program code and / or include embedded firmware to perform functions.

[0131] According to one aspect of this disclosure, a computer device is provided, including a memory, a processor, and a computer program stored in the memory. The processor is configured to execute the computer program to implement the steps of any of the method embodiments described above.

[0132] According to one aspect of this disclosure, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of any of the method embodiments described above.

[0133] According to one aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps of any of the method embodiments described above.

[0134] In the following text, combined with Figure 10Illustrative examples describing such computer devices, non-transitory computer-readable storage media, and computer program products.

[0135] Figure 10 An example configuration of a computer device 1000 that can be used to implement the methods described herein is shown. For example, Figure 1 The server 120 and / or client device 110 shown may include an architecture similar to computer device 1000. The aforementioned apparatus 900 may also be implemented wholly or at least partially by computer device 1000 or similar devices or systems.

[0136] Computer device 1000 can be a variety of different types of devices. Examples of computer device 1000 include, but are not limited to: desktop computers, server computers, laptop or netbook computers, mobile devices (e.g., tablets, cellular or other wireless phones (e.g., smartphones), notebook computers, mobile stations), wearable devices (e.g., glasses, watches), entertainment devices (e.g., entertainment appliances, set-top boxes communicatively coupled to a display device, game consoles), televisions or other display devices, automotive computers, and so on.

[0137] Computer device 1000 may include at least one processor 1002, memory 1004, multiple communication interfaces 1006, display device 1008, other input / output (I / O) devices 1010, and one or more mass storage devices 1012 capable of communicating with each other, such as via system bus 1014 or other suitable connections.

[0138] Processor 1002 may be a single processing unit or multiple processing units, and all processing units may include single or multiple computing units or multiple cores. Processor 1002 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any device that manipulates signals based on operating instructions. Among other capabilities, processor 1002 may be configured to acquire and execute computer-readable instructions stored in memory 1004, mass storage device 1012, or other computer-readable media, such as program code of operating system 1016, program code of application program 1018, program code of other program 1020, etc.

[0139] Memory 1004 and mass storage device 1012 are examples of computer-readable storage media for storing instructions executed by processor 1002 to perform the various functions described above. For example, memory 1004 may generally include both volatile and non-volatile memory (e.g., RAM, ROM, etc.). Furthermore, mass storage device 1012 may generally include hard disk drives, solid-state drives, removable media, including external and removable drives, memory cards, flash memory, floppy disks, optical disks (e.g., CDs, DVDs), storage arrays, network-attached storage, storage area networks, etc. Both memory 1004 and mass storage device 1012 may be collectively referred to herein as memory or computer-readable storage media, and may be non-transitory media capable of storing computer-readable, processor-executable program instructions as computer program code, which may be executed by processor 1002 as a specific machine configured to perform the operations and functions described in the examples herein.

[0140] Multiple programs may be stored on mass storage device 1012. These programs include operating system 1016, one or more application programs 1018, other programs 1020, and program data 1022, and they may be loaded into memory 1004 for execution. Examples of such application programs or program modules may include, for example, computer program logic (e.g., computer program code or instructions) for implementing the following components / functions: method 200, process 400, process 700 (including any suitable steps of method 200, process 400, process 700), and / or other embodiments described herein.

[0141] Although Figure 8 The modules 1016, 1018, 1020, and 1022, or portions thereof, are illustrated as being stored in memory 1004 of computer device 1000. However, modules 1016, 1018, 1020, and 1022 may be implemented using any form of computer-readable medium accessible by computer device 1000. As used herein, “computer-readable medium” includes at least two types of computer-readable media: computer-readable storage media and communication media.

[0142] Computer-readable storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, DVD, or other optical storage devices, magnetic cassettes, magnetic tapes, disk storage devices or other magnetic storage devices, or any other non-transmission medium that can be used to store information for access by a computer device. In contrast, communication media can embody computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms. Computer-readable storage media as defined herein do not include communication media.

[0143] One or more communication interfaces 1006 are used for exchanging data with other devices, such as via a network, direct connection, etc. Such communication interfaces can be one or more of the following: any type of network interface (e.g., a network interface card (NIC)), wired or wireless (such as IEEE 802.11 Wireless LAN (WLAN)) wireless interface, Wi-MAX interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth. TM Interfaces, near field communication (NFC) interfaces, etc. Communication interface 1006 can facilitate communication across various network and protocol types, including wired networks (e.g., LAN, cable, etc.) and wireless networks (e.g., WLAN, cellular, satellite, etc.), the Internet, etc. Communication interface 1006 can also provide communication with external storage devices (not shown) such as storage arrays, network-attached storage, storage area networks, etc.

[0144] In some examples, a display device 1008, such as a monitor, may be included for displaying information and images to the user. Other I / O devices 1010 may be devices that receive various inputs from the user and provide various outputs to the user, and may include touch input devices, gesture input devices, cameras, keyboards, remote controls, mice, printers, audio input / output devices, and so on.

[0145] The technologies described herein can be supported by these various configurations of computer device 1000, and are not limited to specific examples of the technologies described herein. For example, the functionality can also be implemented wholly or partially on a “cloud” using a distributed system. A cloud includes and / or represents a platform for resources. The platform abstracts the underlying functionality of the cloud’s hardware (e.g., servers) and software resources. Resources may include applications and / or data that can be used when performing computational processing on a server remote from computer device 1000. Resources may also include services provided via the Internet and / or via subscriber networks such as cellular or Wi-Fi networks. The platform can abstract resources and functionality to connect computer device 1000 to other computer devices. Therefore, the implementation of the functionality described herein can be distributed throughout the cloud. For example, the functionality may be implemented partly on computer device 1000 and partly through a platform that abstracts the functionality of the cloud.

[0146] Although this disclosure has been described and illustrated in detail in the accompanying drawings and the foregoing description, such description and illustration should be considered illustrative and suggestive, not restrictive; this disclosure is not limited to the disclosed embodiments. By studying the drawings, the disclosure, and the appended claims, those skilled in the art will be able to understand and implement variations of the disclosed embodiments in practice with respect to the claimed subject matter. In the claims, the word "comprising" does not exclude other elements or steps not listed, the indefinite article "a" or "an" does not exclude a plurality, the term "a plurality" means two or more, and the term "based on" should be interpreted as "at least partially based on". The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be beneficial.

Claims

1. A method for a data system for storing at least unstructured data and displaying a field extraction interface for a user to perform a field extraction operation, the method comprising: receiving, via the field extraction interface, a first input comprising a first condition for locating a candidate data entry in the unstructured data to be subjected to a field extraction operation; in response to receiving the first input, querying the unstructured data for the candidate data entry satisfying the first condition and providing a preview of the candidate data entry via the field extraction interface; receiving, via the field extraction interface, a second input for locating a sample data entry among the candidate data entries; in response to receiving the second input, displaying the sample data entry in the field extraction interface; receiving, via the field extraction interface, a third input comprising a second condition comprising a field extraction rule for performing a field extraction operation on the sample data entry; and in response to receiving the third input, performing the field extraction operation on the sample data entry based on the second condition, wherein the second condition comprises at least one of a KV extraction rule, a JSON extraction rule, an IP address extraction rule, or a regular extraction rule, wherein the second condition comprises a first field extraction rule to an n-th field extraction rule, each of the first field extraction rule to the n-th field extraction rule is selected from a group consisting of the KV extraction rule, the JSON extraction rule, the IP address extraction rule, and the regular extraction rule, n is an integer, n > 2, and wherein performing the field extraction operation on the sample data entry based on the second condition comprises: applying the first field extraction rule to the sample data entry to obtain one or more 1st-level branch extracted fields, wherein the one or more 1st-level branch extracted fields comprise fields in the sample data entry that satisfy the first field extraction rule and fields in the sample data entry that do not satisfy the first field extraction rule; applying the n-th field extraction rule to at least a portion of one or more (n-1)th-level branch extracted fields to obtain one or more n-th-level branch extracted fields, wherein the one or more n-th-level branch extracted fields comprise fields in the (n-1)th-level branch extracted fields that satisfy the n-th field extraction rule and fields in the (n-1)th-level branch extracted fields that do not satisfy the n-th field extraction rule. the second condition comprises a first field extraction rule selected from a group consisting of the KV extraction rule, the JSON extraction rule, the IP address extraction rule, and the regular extraction rule, and wherein performing the field extraction operation on the sample data entry based on the second condition comprises:

2. The method of claim 1, wherein, applying the first field extraction rule to the sample data entry to obtain one or more 1st-level branch extracted fields, ​ The one or more first-level branch extraction fields include fields in the sample data entry that meet the first field extraction rule.

3. The method of claim 1, wherein, The regular extraction rule includes extraction according to a regular expression, and the regular expression is input by the user via the field extraction interface or is automatically generated according to content drawn by the user via the field extraction interface.

4. The method of claim 3, wherein, The drawn content is highlighted in the field extraction interface.

5. The method of any of claims 1-4, further comprising: in response to the field extraction operation on the sample data entry being completed, applying the field extraction operation to the remaining candidate data entries of the candidate data entries other than the sample data entry.

6. The method of claim 5, further comprising: after applying the field extraction operation to the remaining candidate data entries of the candidate data entries other than the sample data entry, providing, via the field extraction interface, a preview of field extraction results obtained by applying the field extraction operation to the candidate data entries; receiving, via the field extraction interface, a fourth input for adjusting one or more field extraction rules included in the second condition for corresponding data; and in response to receiving the fourth input, re-executing the field extraction operation on the corresponding data according to the adjusted one or more field extraction rules.

7. The method of any one of claims 1-4, wherein, The first condition includes at least one of an SQL statement or a keyword.

8. The method of any one of claims 1-4, wherein, The second input is a selection of a sample data entry of the candidate data entries.

9. The method of any of claims 1-4, further comprising: saving all field extraction rules included in the second condition.

10. An apparatus for a data system for storing at least unstructured data and displaying a field extraction interface for a user to perform a field extraction operation, the apparatus comprising: a first module configured to receive, via the field extraction interface, a first input including a first condition for locating candidate data entries of the unstructured data to be subjected to a field extraction operation; a second module configured to, in response to receiving the first input, query the unstructured data for the candidate data entries satisfying the first condition and provide, via the field extraction interface, a preview of the candidate data entries; a third module configured to receive, via the field extraction interface, a second input for locating a sample data entry of the candidate data entries; a fourth module configured to, in response to receiving the second input, display the sample data entry in the field extraction interface; a fifth module configured to receive, via the field extraction interface, a third input including a second condition including field extraction rules for performing a field extraction operation on the sample data entry; and a sixth module configured to, in response to receiving the third input, perform the field extraction operation on the sample data entry based on the second condition, The second condition includes at least one of the following field extraction rules: a KV extraction rule, a JSON extraction rule, an IP address extraction rule, or a regular extraction rule. The second condition includes first field extraction rule to nth field extraction rule, each of the first field extraction rule to the nth field extraction rule is selected from the group consisting of the KV extraction rule, the JSON extraction rule, the IP address extraction rule, and the regular extraction rule, n is an integer, n≥2, and wherein the field extraction operation on the sample data entry based on the second condition includes: applying the first field extraction rule to the sample data entry to obtain one or more first-level branch extraction fields, wherein the one or more first-level branch extraction fields include fields in the sample data entry that meet the first field extraction rule and fields in the sample data entry that do not meet the first field extraction rule; applying the nth field extraction rule to at least part of one or more n-1 level branch extraction fields to obtain one or more n level branch extraction fields, wherein the one or more n level branch extraction fields include fields in the n-1 level branch extraction fields that meet the nth field extraction rule and fields in the n-1 level branch extraction fields that do not meet the nth field extraction rule.

11. A computer device comprising: at least one processor; and at least one memory having stored thereon a computer program, wherein the computer program, when executed by the at least one processor, causes the at least one processor to perform the method of any one of claims 1-9.

12. A computer-readable storage medium having stored thereon a computer program, the computer program, when executed by a processor, causing the processor to perform the method of any one of claims 1-9.

13. A computer program product comprising a computer program, the computer program, when executed by a processor, causing the processor to perform the method of any one of claims 1-9.

Citation Information

Patent Citations

  • Data query method and device, computer equipment and computer readable storage medium

    CN116150181A