Text analysis method, electronic equipment and storage medium
By using common field matching and text block concurrent processing in text parsing, the problem of efficiency reduction caused by the increase in matching conditions in the prior art is solved, and efficient text parsing is achieved.
Patent Information
- Application Number
- CN202410066092.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, as the number of matching conditions increases, the text parsing efficiency decreases, the calculation amount increases, making it difficult to efficiently process a large amount of text information.
By obtaining the common fields in the matching conditions for preliminary matching, the number of matches of non-public fields is reduced, and the text is split into multiple text blocks for concurrent processing is improved to improve matching efficiency.
By reducing the number of matches of non-common fields and concurrent processing of text blocks, the efficiency of text parsing is improved, the amount of calculation is reduced, and the system performance is improved.
Smart Images

Figure CN120337900A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information technology, and in particular, to a text parsing method, an electronic device, and a storage medium. Background Art
[0002] A large amount of text information is generated as the user uses the mobile operating system, and this text information can be used by developers for research and analysis. Among them, the text types can include log types, trace data types, etc.
[0003] In the prior art, relevant matching conditions are usually configured through a configuration file to parse the text information to obtain the parsing result required by the user. Among them, the matching conditions can include forms such as regular expressions, and the regular expressions can include one or more matching fields.
[0004] It can be understood that the user can configure multiple matching conditions in the configuration file. In this case, since each condition needs to match the entire text, as the number of matching conditions increases, the calculation amount will increase greatly, and the parsing efficiency will decrease. Summary of the Invention
[0005] This application provides a text parsing method, an electronic device, and a storage medium, which helps to improve the efficiency of text parsing.
[0006] In a first aspect, this application provides a text parsing method, including: obtaining a text to be parsed and a configuration file, where the configuration file configures at least multiple matching conditions, and the multiple matching conditions are used to match data records in the text to be parsed, and each matching condition in the multiple matching conditions includes one or more fields; obtaining a target matching condition with the same field in the multiple matching conditions, and performing a match in the text to be parsed based on the common field corresponding to the target matching condition to obtain a first matching result; based on the non-common fields in the target matching condition, performing a match in the first matching result to obtain a second matching result; and structurally outputting the second matching result.
[0007] This application performs matching in terms of the common fields and non-common fields in the matching conditions. Compared with the dimension of the entire matching condition in the prior art, for data with no relevant matching results in part, the non-common fields are no longer used for matching based on the matching result of the common field, thereby reducing the number of matches and improving the matching efficiency.
[0008] In one possible implementation, the structured output of the second matching result includes: combining a first data record subgroup and a second data record subgroup into a data record combination; and structurally outputting the data record combination, where the first data record subgroup includes data records with the same features in the second matching result, and the second data record subgroup includes data records with not completely the same features in the second matching result.
[0009] By combining and then structurally outputting the matching results, this application can reduce the input / output volume of data and improve system performance.
[0010] In one possible implementation, the first data record subgroup and the second data record subgroup are obtained through concurrent processing by different threads.
[0011] By obtaining the first data record subgroup and the second data record subgroup through processing by different threads, this application can improve data processing efficiency.
[0012] In one possible implementation, the obtaining of the first matching result by matching based on the common fields corresponding to the target matching condition in the text to be parsed includes: splitting the text to be parsed into multiple text blocks; and concurrently matching the multiple text blocks based on the common fields corresponding to the target matching condition to obtain the first matching result.
[0013] By concurrently processing multiple text blocks of the text to be parsed, this application can improve processing efficiency compared with the prior art that processes the text to be parsed line by line.
[0014] In one possible implementation, a delimiter is also configured in the configuration file, and the splitting of the text to be parsed into multiple text blocks includes: splitting the text to be parsed into multiple text blocks through the delimiter.
[0015] In one possible implementation, the data type of the text to be parsed includes at least log and trace data.
[0016] In one possible implementation, the form of the structured output includes at least file, database, command line, and network.
[0017] In a second aspect, this application provides a text parsing device, including one or more functional modules, where the one or more functional modules are used to implement the text parsing method as described in the first aspect.
[0018] In a third aspect, this application provides an electronic device, including: a processor and a memory, where the memory is used to store a program; and the processor is used to run the program to implement the text parsing method as described in the first aspect.
[0019] Fourthly, the present application provides a readable storage medium storing a program, which, when running on an electronic device, enables the electronic device to implement the text parsing method described in the first aspect.
[0020] Fifthly, the present application provides a program, which, when running on a processor of an electronic device, enables the electronic device to execute the text parsing method described in the first aspect.
[0021] In a possible design, the program in the fifth aspect may be stored in whole or in part on a storage medium packaged together with the processor, or may be stored in whole or in part on a memory not packaged together with the processor. Description of the Drawings
[0022] Figure 1 is a schematic structural diagram of the electronic device provided in an embodiment of the present application;
[0023] Figure 2 is a schematic flowchart of a method for parsing text according to an embodiment of the present application;
[0024] Figure 3 is a schematic structural diagram of the text parsing device provided in an embodiment of the present application. Detailed Embodiments
[0025] In the embodiments of the present application, unless otherwise specified, the character " / " indicates that the related objects before and after are in an "or" relationship. For example, A / B may represent A or B. "And / or" describes the relationship between related objects, indicating that there can be three relationships. For example, A and / or B may represent: A alone, A and B existing simultaneously, and B alone.
[0026] It should be noted that the terms "first", "second", etc. involved in the embodiments of the present application are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features, nor can they be understood as indicating or implying an order.
[0027] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. In addition, "at least one (item)" or its similar expression means any combination of these items, which may include any combination of single item (item) or plural items (items). For example, at least one (item) of A, B, or C may represent: A, B, C, A and B, A and C, B and C, or A, B, and C. Each of A, B, and C itself may be an element or a set containing one or more elements.
[0028] In the embodiments of the present application, terms such as "exemplary", "in some embodiments", and "in another embodiment" are used to present examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the term "exemplary" is intended to present concepts in a specific manner.
[0029] In the embodiments of the present application, the words "of", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings to be expressed are the same. In the embodiments of the present application, communication and transmission can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. For example, transmission can include sending and / or receiving, and can be a noun or a verb.
[0030] In the embodiments of the present application, the equality involved can be used in combination with greater than, applicable to the technical solutions adopted when it is greater than, or can also be used in combination with less than, applicable to the technical solutions adopted when it is less than. It should be noted that when equality is used in combination with greater than, it cannot be used in combination with less than; when equality is used in combination with less than, it is not used in combination with greater than.
[0031] A large amount of text information will be generated as the user uses the mobile phone operating system, and this text information can be used by developers for research and analysis. Among them, the text types can include log types, trace data types, etc.
[0032] In the prior art, relevant matching conditions are usually configured through a configuration file to parse the text information to obtain the parsing result required by the user. Among them, the matching conditions can include forms such as regular expressions, and the regular expressions can include one or more matching fields.
[0033] It can be understood that the user can configure multiple matching conditions in the configuration file. In this case, since each condition needs to match the entire text, as the number of matching conditions increases, the amount of calculation will increase significantly and the parsing efficiency will decrease.
[0034] Based on the above problems, the embodiment of the present application proposes a text parsing method, which is applied to an electronic device. The electronic device can be a fixed device, such as a large screen, a laptop computer, a desktop computer, etc. The electronic device can also be a mobile device, which can also be called a user equipment (UE), an access terminal, a user unit, a user station, a mobile station, a mobile station, a remote station, a remote terminal, a mobile terminal, a user terminal, a terminal, a wireless communication device, a user agent or a user device. The mobile device can be a station (STAION, ST) in a WLAN, a cellular phone, a cordless phone, a Session Initiation Protocol (SIP) phone, a Wireless Local Loop (WLL) station, a Personal Digital Assistant (PDA) device, a handheld device with wireless communication function, a computing device or other processing device connected to a wireless modem, a vehicle-mounted device, a vehicle networking terminal, a computer, a laptop computer, a handheld communication device, a handheld computing device, a satellite wireless device, a wireless modem card, a TV set top box (STB), a customer premise equipment (CPE) and / or other devices used to communicate on a wireless system and a next-generation communication system, for example, a mobile device in a 5G network or a mobile device in a future Public Land Mobile Network (PLMN) network. The mobile device can also be a wearable device. Wearable devices can also be called wearable smart devices, which are a general term for wearable devices that are intelligently designed and developed using wearable technology for daily wear, such as glasses, gloves, watches, clothing and shoes. Wearable devices are portable devices that are worn directly on the body or integrated into the user's clothes or accessories. Wearable devices are not just hardware devices, but also realize powerful functions through software support, data interaction, and cloud interaction.
[0035] Figure 1 First, a schematic structural diagram of the electronic device 100 is shown as an example.
[0036] The electronic device 100 may include at least one processor and at least one memory connected to the processor, wherein the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the method provided in the embodiment shown in this document.
[0037] Figure 1 A block diagram of an exemplary electronic device 100 suitable for implementing embodiments herein is shown.Figure 1 The illustrated electronic device 100 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments herein.
[0038] As Figure 1 shown, the components of the electronic device 100 may include, but are not limited to: one or more processors 110, a memory 120, a communication bus 140 connecting different system components (including the memory 120 and the processor 110), and a communication interface 130.
[0039] The communication bus 140 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus structures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnection (PCI) bus.
[0040] The electronic device 100 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the device, including volatile and non-volatile media, removable and non-removable media.
[0041] The memory 120 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. The device may further include other removable / non-removable, volatile / non-volatile computer system storage media. Although Figure 1Not shown in the figure, a disk drive for reading and writing to a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (such as: Compact Disc Read Only Memory (hereinafter referred to as: CD-ROM), Digital Video Disc Read Only Memory (hereinafter referred to as: DVD-ROM) or other optical media) can be provided. In these cases, each drive can be connected to the communication bus 140 through one or more data medium interfaces. The memory 120 can include at least one program product, which has a set (such as at least one) of program modules, and these program modules are configured to execute the functions of the embodiments herein.
[0042] A program / utility with a set (at least one) of program modules can be stored in the memory 120. Such program modules include—but are not limited to—an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules generally execute the functions and / or methods in the embodiments described herein.
[0043] The electronic device 100 can also communicate with one or more external devices (such as a keyboard, a pointing device, a display, etc.), and can also communicate with one or more devices that enable a user to interact with the device, and / or communicate with any device that enables the device to communicate with one or more other devices (such as a network card, a modem, etc.). Such communication can be carried out through the communication interface 130. And, the electronic device 100 can also communicate with one or more networks (such as a Local Area Network (hereinafter referred to as: LAN), a Wide Area Network (hereinafter referred to as: WAN) and / or a public network, such as the Internet) through a network adapter ( Figure 1 not shown in the figure), and the above network adapter can communicate with other modules of the device through the communication bus 140. It should be understood that although Figure 1 not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 100, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, Redundant Arrays of Independent Drives (hereinafter referred to as: RAID) systems, tape drives, and data backup storage systems, etc.
[0044] The processor 110 executes various functional applications and data processing by running the programs stored in the memory 120, such as implementing the methods provided by the embodiments herein.
[0045] It can be understood that the interface connection relationships among the modules illustrated in the embodiments of this article are only illustrative descriptions and do not constitute a structural limitation on the electronic device 100. In other embodiments of this article, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0046] Next, in combination with Figure 2 the text parsing method provided in the embodiments of this application will be described.
[0047] As Figure 2 shown in the flowchart of an embodiment of the text parsing method provided in this application, it specifically includes the following steps:
[0048] Step 201, obtain the text to be parsed and the configuration file.
[0049] Specifically, the text to be parsed may be text collected through a script, or the text to be parsed may be text collected through other means. The embodiments of this application do not make special limitations on the collection method of the text to be parsed.
[0050] Among them, the data type of the text to be parsed may include but is not limited to types such as logs and trace data, and the format of the text to be parsed includes but is not limited to formats such as json, xml, and txt.
[0051] It can be understood that the reading method of the text to be parsed may be in the form of a file, or the reading method of the text to be parsed may be in the form of a file stream. The embodiments of this application do not make special limitations on the reading method of the text to be parsed.
[0052] Taking the parsed text as a log as an example, Table 1 exemplarily shows the text information included when the text to be parsed is a log.
[0053] Table 1
[0054]
[0055] Referring to Table 1, the text to be parsed may include N log data, and one piece of data may include multiple lines, where N is a positive integer.
[0056] The format of the configuration file may be Tom's Obvious Minimal Language (TOML), json, xml, etc., or the format of the configuration file may be other formats. The embodiments of this application do not make special limitations on this.
[0057] The configuration file may include multiple matching conditions, which can be used to match data records in the text to be parsed. Among them, one matching condition may include one or more fields, and the one or more fields are used to match the text to be parsed to obtain the target data record in the text to be parsed.
[0058] Among them, the form of the matching condition can be in the form of a regular expression, or the form of the matching condition can be other forms, and the embodiments of the present application do not make special limitations on this.
[0059] It can be understood that the matching condition can also be called a matching pattern.
[0060] Exemplarily, the code example of the matching condition is as follows:
[0061] [[patterns]]
[0062] [patterns.delimiter_line_rest]
[0063] samples=
[0064] ”'1970-01-01 08:00:00,com.demo.app / app.launcher,JANK_FRAME,176”',]
[0065] inheritance=["clock_type","timestamp"]
[0066] pseudo_regex=”'${timestamp},${package} / ${activity},${type},${val}”'
[0067] [[patterns.delimiter_line_rest.casts]]
[0068] package="${package}"
[0069] activity="${activity}"
[0070] domain="performance"
[0071] source="event"
[0072] unit="ms"
[0073] measurement="measurement"
[0074] value = "${val}"
[0075] [[patterns.delimiter_line_rest.vars]]
[0076] name = "timestamp"
[0077] regex = " '\d{4}-\d{2}-\d{2}\s+\d{2}:\d{2}:\d{2}"'
[0078] [[patterns.delimiter_line_rest.vars]]
[0079] name = "package"
[0080] regex = " '[^ / ]+'"
[0081] [[patterns.delimiter_line_rest.vars]]
[0082] name = "activity"
[0083] regex = " '[^,]+'"
[0084] [[patterns.delimiter_line_rest.vars]]
[0085] name = "type"
[0086] regex = " 'JANK_FRAME'"
[0087] [[patterns.delimiter_line_rest.vars]]
[0088] name = "val"
[0089] regex = " '\d+'"
[0090] Referring to the above code, a configuration file may include multiple matching conditions. Each matching condition can be expressed by a regular expression. Exemplarily, the regular expression "regex = '\d{4}-\d{2}-\d{2}\s+\d{2}:\d{2}:\d{2}'" is used to match the data in the time field, the regular expression "regex = '[^ / ]+'" is used to match the data in the package field, the regular expression "regex = '[^,]+'" is used to match the data in the activity field, the regular expression "regex = 'JANK_FRAME'" is used to match the data in the type field, and the regular expression "regex = '\d+'" is used to match the data in the numerical field.
[0091] Step 202: Obtain the target matching condition with the same fields, and perform matching in the text to be parsed based on the common fields corresponding to the target matching condition to obtain the first matching result.
[0092] Specifically, as described above, the configuration file contains multiple matching conditions, and one matching condition contains one or more fields. In this case, there may be the same fields among the multiple matching conditions. Therefore, the target matching condition with the same fields can be found among the multiple matching conditions in the matching file, and the first matching result can be obtained by performing matching in the text to be parsed based on the common fields corresponding to the target matching condition.
[0093] Now, an exemplary description of the obtaining method of the first matching result will be given in combination with Table 2. Table 2 exemplarily shows 4 matching conditions in the configuration file. It can be understood that the number of matching conditions in Table 2 is only for exemplary illustration and does not constitute a limitation on the embodiments of the present application. In some embodiments, the number of matching conditions in the configuration file may be greater than 4 or less than 4.
[0094] Table 2
[0095] Matching condition Field Matching condition 1 Field 1, Field 2 Matching condition 2 Field 1, Field 3 Matching condition 3 Field 4, Field 5 Matching condition 4 Field 4, Field 6
[0096] Referring to Table 2, the configuration file configures 4 matching conditions, which can be respectively matching condition 1, matching condition 2, matching condition 3, and matching condition 4. Among them, matching condition 1 contains field 1 and field 2, matching condition 2 contains field 1 and field 3, matching condition 3 contains field 4 and field 5, and matching condition 4 contains field 4 and field 6.
[0097] Since matching condition 1 and matching condition 2 contain the same field (e.g., field 1), and matching condition 3 and matching condition 4 contain the same field (e.g., field 4), the target matching conditions with the same fields are matching condition group 1 composed of matching condition 1 and matching condition 2, and matching condition group 2 composed of matching condition 3 and matching condition 4. Among them, the common field in matching condition group 1 is field 1, and the common field in matching condition group 2 is field 4.
[0098] Next, it is possible to perform matching in the text to be parsed based on the common field (i.e., field 1) in matching condition group 1, and perform matching in the text to be parsed based on the common field (i.e., field 4) in matching condition group 2. Thus, a first matching result can be obtained, and the first matching result includes the result after matching based on field 1 and the result after matching based on field 4.
[0099] In some alternative embodiments, since the matching method in the prior art is to perform matching on the data in the text to be parsed line by line based on each matching condition, the efficiency is low. In the embodiments of the present application, the text to be parsed can also be split, and thus multiple text blocks can be obtained. Next, these multiple text blocks can be respectively handed over to corresponding threads for matching, and thus concurrent processing of multiple text blocks can be achieved, improving the matching efficiency.
[0100] Exemplarily, the above-mentioned splitting method of the text to be parsed can be based on a delimiter, and the delimiter can be set in a configuration file.
[0101] Among them, the delimiter can include multiple types. Exemplarily, the delimiter can include a time-type delimiter, or the delimiter can include a data-type delimiter. The embodiments of the present application do not make special limitations on this.
[0102] It can be understood that the delimiter can be used to separate two text blocks. That is to say, the delimiter can be used to distinguish the boundary between two text blocks, the delimiter can be used to represent the starting position of a text block, or the delimiter can be used to represent the ending position of a text block.
[0103] It can be understood that a text block can contain one or more lines of data in the text to be parsed. That is, through the delimiter, one or more lines of data in the text to be parsed can be formed into a text block, and all the data in the text to be parsed can be split into multiple text blocks.
[0104] Exemplarily, taking the delimiter as a time-type delimiter as an example, the code example of the delimiter is as follows:
[0105] [delimiter]
[0106] samples=["'2023-04-12 14:41:22.934H'"]
[0107] inheritance=[]
[0108] pseudo_regex=" '${timestamp}\s+H\s+'"
[0109] [[delimiter.casts]]
[0110] clock_type="perf_event"
[0111] timestamp="${timestamp}"
[0112] [[delimiter.vars]]
[0113] name="timestamp"
[0114] regex=" '\d{4}-\d{2}-\d{2}\s+\d{2}:\d{2}:\d{2}\.\d{3}'"
[0115] Among them, the delimiter can be recognized through the pseudo-regular expression and the regular expression.
[0116] Taking the code with the above delimiter as the delimiter of the time type as an example, the delimiter can be recognized through the pseudo-regular expression of "pseudo_regex = " '${timestamp}\s+H\s+'"", that is, in any line of data, the string that satisfies the combination of the timestamp + the end identifier H can be considered as a delimiter of the time type, and the timestamp in this delimiter satisfies the rule of the regular expression of "regex = " '\d{4}-\d{2}-\d{2}\s+\d{2}:\d{2}:\d{2}\.\d{3}'"".
[0117] Step 203: Based on the non-common fields in the target matching condition, perform matching in the first matching result to obtain the second matching result.
[0118] Specifically, after performing matching based on the common fields corresponding to the target matching condition, matching can be performed in the first matching result based on the non-common fields in the target matching condition to obtain the second matching result.
[0119] It can be understood that after matching based on the common fields corresponding to the target matching conditions, for some common fields, there may be no corresponding data records in the text to be parsed. For these common fields without data records, the corresponding non-common fields do not need to be matched either, thus reducing the number of matches and improving the matching efficiency. For some common fields, there are corresponding data records in the text to be parsed. For these data records, further matching can be performed based on the non-common fields in the target matching conditions corresponding to the data records to obtain a second matching result.
[0120] Taking the matching conditions in Table 2 as an example, the target matching conditions include Matching Condition Group 1 and Matching Condition Group 2. The common fields in Matching Condition Group 1 include Field 1, and the non-common fields in Matching Condition Group 1 include Field 2 and Field 3; the common fields in Matching Condition Group 2 include Field 4, and the non-common fields in Matching Condition Group 2 include Field 5 and Field 6.
[0121] Suppose that by matching through the common fields in the above target matching conditions (for example, Field 1 and Field 4), only the data records matching Field 1 are matched, but the data records matching Field 4 are not matched. Since Field 2 and Field 3 belong to Matching Condition Group 1 together with Field 1, therefore, based on Field 2 and Field 3 in Matching Condition Group 1, matching can be respectively performed on the data records matching Field 1, and thus a first matching result can be obtained. Since Field 5 and Field 6 belong to Matching Condition Group 2 together with Field 4, and there are no data records matching Field 4 in the text to be parsed, therefore, it is not necessary to match Field 5 and Field 6 in Matching Condition Group 2, thus reducing the number of matches and improving the matching efficiency.
[0122] In some alternative embodiments, target matching conditions with the same fields can be further obtained in the first matching result, and matching can be performed in the first matching result based on the common fields corresponding to these target matching conditions to obtain a third matching result. Among them, the specific way to obtain the third matching result can refer to the way to obtain the first matching result, which will not be elaborated here.
[0123] Then, based on the non-common fields in the target matching conditions corresponding to the third matching result, matching is respectively performed in the third matching result to obtain a second matching result.
[0124] That is to say, the common fields in the target matching conditions can be obtained in multiple rounds until there are no common fields in the target matching conditions.
[0125] Now, in combination with Table 3, an exemplary description is given of the method for obtaining the common fields in the matching conditions in multiple rounds.
[0126] Table 3
[0127] Matching condition Field Matching condition 1 Field 1, Field 2, Field 3 Matching condition 2 Field 1, Field 2, Field 4 Matching condition 3 Field 1, Field 5 Matching condition 4 Field 1, Field 6
[0128] Referring to Table 3, the configuration file contains 4 matching conditions, namely Matching Condition 1, Matching Condition 2, Matching Condition 3, and Matching Condition 4. Among them, Matching Condition 1 includes Field 1, Field 2, and Field 3; Matching Condition 2 includes Field 1, Field 2, and Field 4; Matching Condition 3 includes Field 1 and Field 5; Matching Condition 4 includes Field 1 and Field 6.
[0129] First, perform the first round of obtaining the common fields in the target matching conditions. Since Matching Condition 1, Matching Condition 2, Matching Condition 3, and Matching Condition 4 all have the same field, for example, Field 1, that is, the set of matching conditions 1 composed of Matching Condition 1, Matching Condition 2, Matching Condition 3, and Matching Condition 4 is the target matching condition, and Field 1 is the common field in this set of matching conditions 1. By matching through this common field of Field 1, the first matching result can be obtained.
[0130] Next, perform the second round of obtaining the common fields in the target matching conditions. Based on the first matching result, since Matching Condition 1 and Matching Condition 2 in the set of matching conditions 1 also contain the same field, for example, Field 2, that is, the set of matching conditions 2 composed of Matching Condition 1 and Matching Condition 2 is the target matching condition, and Field 2 is the common field in this set of matching conditions 2. By matching through this common field of Field 2 and the 2 non-common fields of Field 5 and Field 6 in the first matching result, the third matching result can be obtained.
[0131] Assume that there is no data record in the first matching result that matches Field 2, then there is no need to perform matching based on Field 3, Field 4, and Field respectively, which can reduce the number of matches and improve the matching efficiency.
[0132] It can be understood that the above example only provides an exemplary illustration of the method of obtaining the common fields in the matching conditions in 2 rounds, and does not constitute a limitation on the embodiments of the present application. In some embodiments, the common fields in the matching conditions can also be obtained through P rounds, where P is a positive integer greater than 2.
[0133] It should be noted that, in order to improve the matching efficiency, the matching method in this step 203 can also be executed through multiple threads. The specific method can refer to the relevant description in step 202 and will not be elaborated here.
[0134] Step 204, structurally output the second matching result.
[0135] Specifically, when the second matching result is obtained, the second matching result can be structurally output.
[0136] Among them, the forms of structured output may include, but are not limited to, files, databases, command lines, networks, etc.
[0137] In some alternative embodiments, in order to reduce the amount of data input and output, the data records in the second matching result may also be combined to obtain a data record combination, and the data record combination may be output in a structured manner.
[0138] Among them, the way of combining data records may be feature-based. The feature may be a field, or the feature may be other information in the data record. The embodiments of the present application do not make special limitations on this.
[0139] Exemplarily, data records with the same feature may be combined into a data record subgroup. For example, the combination of data records containing the same feature may be denoted as the first data record subgroup; then, data records with not completely the same features may be combined into another data record subgroup. For example, the combination of data records containing not completely the same features may be denoted as the second data record subgroup. Finally, the first data record subgroup and the second data record subgroup may be combined to obtain a data record combination, that is, the first data record subgroup and the second data record subgroup together form the data record combination.
[0140] It should be noted that the way of obtaining the first data record subgroup and the way of obtaining the second data record subgroup may be the MM combination method, that is, the method of combining every M objects, where M is a positive integer greater than 1, and the object may be a data record in the second matching result.
[0141] Now, taking M = 2 and combining Table 4, an exemplary description of the MM combination method is given.
[0142] Table 4
[0143]
[0144]
[0145] Since M = 2, that is, the MM combination method is 22 combination, that is, every two objects are combined. Referring to Table 4, the second matching record contains 4 data records, and these 4 data records may be respectively denoted as data record 1, data record 2, data record 3, and data record 4.
[0146] Among them, data record 1 contains field 1, field 2, and field 3. The data corresponding to field 1 is data information 11, the data corresponding to field 2 is data information 21, and the data corresponding to field 3 is data information 31.
[0147] Data record 2 contains fields 1, 2, and 3. The data corresponding to field 1 is data information 21, the data corresponding to field 2 is data information 22, and the data corresponding to field 3 is data information 32.
[0148] Data record 3 contains fields 4, 5, and 6. The data corresponding to field 4 is data information 41, the data corresponding to field 5 is data information 51, and the data corresponding to field 6 is data information 61.
[0149] Data record 4 contains fields 7 and 8. The data corresponding to field 7 is data information 71, and the data corresponding to field 8 is data information 81.
[0150] Since data records 1 and 2 correspond to the same fields, for example, the same fields can include fields 1, 2, and 3, data records 1 and 2 can be merged to obtain a new data record. For example, this new data record can be denoted as data record 5, and this data record 5 can also be considered a new object. In addition, since the fields corresponding to data records 3 and 4 are all different, data records 3 and 4 can be merged to obtain a new data record. For example, this new data record can be denoted as data record 6, and this data record 6 can also be considered a new object.
[0151] Next, since there are no other data records except data records 5 and 6, data record 5 can be considered as the first data record subgroup, and data record 6 can be considered as the second data record subgroup. Finally, data records 5 and 6 are merged to obtain the final data record combination, and this data record combination can be output in a structured manner.
[0152] It can be understood that since data records 1 and 2 have the same fields but different data corresponding to the fields, the data record 5 obtained after merging can be an object of an array with the same fields and containing multiple data. In addition, since data records 3 and 4 have different fields, the data record 5 obtained after merging can be an object containing multiple sub-objects, where data records 3 and 4 can be considered as one sub-object respectively.
[0153] Table 5 exemplarily shows the data structures of data record 5 and data record 6.
[0154] Table 5
[0155]
[0156] Referring to Table 5, the data record 5 contains the common fields of data record 1 and data record 2, namely, Field 1, Field 2, and Field 3. Among them, the data corresponding to Field 1 is an array containing {data information 11, data information 12}, the data corresponding to Field 2 is an array containing {data information 21, data information 22}, and the data corresponding to Field 3 is an array containing {data information 31, data information 32}. The data record 6 contains 2 sub-data records, namely data record 3 and data record 4, and these sub-data records can also be regarded as sub-objects in the data record 6.
[0157] It can be understood that when M is greater than 2, the specific MM merging method can refer to the 22 merging method, which will not be elaborated here.
[0158] In some alternative embodiments, during the merging process, multiple threads can be used for concurrent processing to improve the merging efficiency. Exemplarily, taking the data records in Table 4 as an example, the merging of data record 1 and data record 2 can be processed by thread 1, and the merging of data record 3 and data record 4 can be processed by thread 2. Through the concurrent processing of threads, the processing efficiency can be improved.
[0159] It should be noted that the text parsing method provided in the embodiments of the present application can be applied to the operating system of an electronic device or a distributed system, and the embodiments of the present application do not make special limitations in this regard.
[0160] Figure 3 This is a schematic structural diagram of an embodiment of the text parsing device of the present application. As Figure 3 shown, the above-mentioned text parsing device 30 may include: an acquisition module 31, a matching module 32, and an output module 33; among them,
[0161] The acquisition module 31 is used to acquire the text to be parsed and a configuration file. Among them, the configuration file configures at least multiple matching conditions, and the multiple matching conditions are used to match the data records in the text to be parsed. Each matching condition in the multiple matching conditions includes one or more fields;
[0162] The matching module 32 is used to obtain a target matching condition with the same fields among the multiple matching conditions, perform matching in the text to be parsed based on the common fields corresponding to the target matching condition to obtain a first matching result; perform matching in the first matching result based on the non-common fields in the target matching condition to obtain a second matching result;
[0163] The output module 33 is used to output the second matching result in a structured manner.
[0164] In one possible implementation manner, the output module 33 is specifically used to merge the first data record subgroup and the second data record subgroup into a data record combination;
[0165] Combining and structured output of the data records;
[0166] The first data record subgroup includes data records with the same characteristics in the second matching result, and the second data record subgroup includes data records with different characteristics in the second matching result.
[0167] In one possible implementation manner, the first data record subgroup and the second data record subgroup are obtained through concurrent processing by different threads.
[0168] In one possible implementation, the matching module 32 is specifically used to split the text to be parsed into multiple text blocks;
[0169] The multiple text blocks are matched concurrently based on the common fields corresponding to the target matching conditions to obtain a first matching result.
[0170] In one possible implementation, a separator is further configured in the configuration file, and the matching module 32 is further configured to split the text to be parsed into multiple text blocks according to the separator.
[0171] In one possible implementation, the data type of the text to be parsed includes at least log and tracking data.
[0172] In one possible implementation manner, the structured output is in the form of at least a file, a database, a command line, and a network.
[0173] Figure 3 The text parsing device 30 provided in the illustrated embodiment can be used to execute the technical solution of the method embodiment shown in the present application. Its implementation principle and technical effects can be further referred to the relevant description in the method embodiment.
[0174] It should be understood that the above Figure 3 The division of the various modules of the text parsing device 30 shown is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. And these modules can all be implemented in the form of software calling through processing elements; they can also be all implemented in the form of hardware; some modules can also be implemented in the form of software calling through processing elements, and some modules can be implemented in the form of hardware. For example, the detection module can be a separately established processing element, or it can be integrated in a chip of an electronic device. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together, or they can be implemented independently. In the implementation process, each step of the above method or each of the above modules can be completed by an integrated logic circuit of hardware in the processor element or instructions in software form.
[0175] For example, these modules above can be one or more integrated circuits configured to implement the above methods, such as: one or more Application Specific Integrated Circuits (ASICs), or, one or more Digital Signal Processors (DSPs), or, one or more Field Programmable Gate Arrays (FPGAs), etc. Additionally, these modules can be integrated together and implemented in the form of a System-On-a-Chip (SOC).
[0176] In the above embodiments, the processors involved can include, for example, CPUs, DSPs, microcontrollers or digital signal processors, and can also include GPUs, Neural-network Process Units (NPUs), and Image Signal Processings (ISPs). The processor can also include necessary hardware accelerators or logic processing hardware circuits, such as ASICs, or one or more integrated circuits for controlling the execution of the technical solution programs of this application. Furthermore, the processor can have the function of operating one or more software programs, and the software programs can be stored in a storage medium.
[0177] The embodiments of this application also provide a readable storage medium, in which a program is stored. When it runs on an electronic device, it causes the electronic device to execute the method provided by the embodiments shown in this application.
[0178] The embodiments of this application also provide a program product, which includes a program. When it runs on an electronic device, it causes the electronic device to execute the method provided by the embodiments shown in this application.
[0179] In the embodiments of this application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent the cases of A existing alone, A and B existing simultaneously, and B existing alone. Here, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.
[0180] Those of ordinary skill in the art can realize that the various units and algorithm steps described in the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0181] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0182] In several embodiments provided in this application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (hereinafter referred to as ROM), random access memories (hereinafter referred to as RAM), magnetic disks, or optical discs.
[0183] As described above, the foregoing are only specific embodiments of this application. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all such changes or substitutions should be covered by the protection scope of this application. The protection scope of this application shall be subject to the protection scope of the claims.
Claims
1. A text parsing method, characterized in that, The method includes: Obtaining the text to be parsed and a configuration file, where the configuration file is configured with at least multiple matching conditions for matching data records in the text to be parsed, and each of the multiple matching conditions includes one or more fields; Obtaining a target matching condition with the same fields among the multiple matching conditions, and performing matching in the text to be parsed based on the common fields corresponding to the target matching condition to obtain a first matching result; Performing matching in the first matching result based on the non-common fields in the target matching condition to obtain a second matching result; Structurally outputting the second matching result.
2. The method according to claim 1, wherein The structurally outputting the second matching result includes: Merging a first subgroup of data records and a second subgroup of data records into a data record combination; Structurally outputting the data record combination; Wherein, the first subgroup of data records includes data records with the same characteristics in the second matching result, and the second subgroup of data records includes data records with not completely the same characteristics in the second matching result.
3. The method according to claim 2, wherein The first subgroup of data records and the second subgroup of data records are obtained through concurrent processing by different threads.
4. The method according to any one of claims 1 to 3, characterized in that, The performing matching in the text to be parsed based on the common fields corresponding to the target matching condition to obtain a first matching result includes: Splitting the text to be parsed into multiple text blocks; Performing concurrent matching on the multiple text blocks based on the common fields corresponding to the target matching condition to obtain a first matching result.
5. The method according to claim 4, wherein A delimiter is also configured in the configuration file, and the splitting the text to be parsed into multiple text blocks includes: Splitting the text to be parsed into multiple text blocks through the delimiter.
6. The method according to any one of claims 1-5, characterized in that, The data type of the text to be parsed at least includes log and trace data.
7. The method according to any one of claims 1-6, characterized in that, The form of the structural output at least includes file, database, command line, and network.
8. An electronic device, characterized in that, Includes: A processor and a memory, where the memory is used to store a program; the processor is used to run the program to implement the text parsing method according to any one of claims 1-7.
9. A readable storage medium, characterized in that, The readable storage medium stores a program, and when the program runs on an electronic device, it implements the text parsing method according to any one of claims 1-7.