Analysis device, analysis method, and analysis program
The analysis device using V-PEG grammar with variable bindings efficiently performs syntax analysis by adding and extracting elements with attributes, reducing processing time for context-dependent patterns from exponential to polynomial levels.
Patent Information
- Application Number
- JP2022501600
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2020-02-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-02-21
AI Technical Summary
Conventional parsers for context-dependent patterns require an enormous amount of time for parsing, often leading to exponential-time operations with respect to data size.
An analysis device utilizing V-PEG (Parsing Expression Grammar with Variable Bindings) that includes an analysis unit, an addition unit, an extraction unit, and a determination unit to perform syntax analysis, adding elements with attributes, extracting the latest elements, and determining context matching, thereby reducing processing time to polynomial levels.
The proposed solution significantly shortens the time required for syntax analysis of context-dependent patterns, suppressing exponential processing times to polynomial times.
Smart Images

Figure 0007702933000001 
Figure 0007702933000002 
Figure 0007702933000003
Abstract
Description
Technical Field
[0001] The present invention relates to an analysis device, an analysis method, and an analysis program.
Background Art
[0002] A syntax analyzer that converts data into a form that can be processed by a computer is known. The syntax analyzer analyzes and converts data according to a language (hereinafter simply referred to as a language) that describes the pattern of the source of conversion.
[0003] For example, there is a syntax analyzer that describes a pattern in a language obtained by extending a regular expression and analyzes data using a syntax analysis algorithm based on backtracking. Further, for example, there is known a syntax analyzer that analyzes a context-dependent pattern using a syntax analysis algorithm called Stateful Packrat Parsing that describes a pattern in a language obtained by extending PEG (Parsing Expression Grammar) (see, for example, Non-Patent Documents 1 and 2).
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, there is a problem that a parser corresponding to a conventional context-dependent pattern may take an enormous amount of time for parsing. For example, in the parsers described in Non-Patent Documents 1 and 2, an exponential-time operation may be required with respect to the data size.
Means for Solving the Problems
[0006] In order to solve the above-described problems and achieve the object, an analysis device includes: an analysis unit that performs syntax analysis of a first character string based on a grammar described in PEG in which a variable is associated with a predetermined terminal symbol; an addition unit that adds an element obtained by assigning a predetermined attribute to a second character string that is a part of the first character string and is analyzed by the analysis unit to correspond to the terminal symbol, to the variable; an extraction unit that extracts the latest element of each attribute from the variable; and a determination unit that determines whether or not the element extracted by the extraction unit matches a predetermined condition related to context.
Advantages of the Invention
[0007] According to the present invention, it is possible to shorten the time required for syntax analysis corresponding to a context-dependent pattern.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
[0009] Hereinafter, embodiments of the analysis apparatus, analysis method, and analysis program according to the present application will be described in detail with reference to the drawings. Note that the present invention is not limited to the embodiments described below.
[0010] [Configuration of the First Embodiment] FIG. 1 is a diagram showing a configuration example of a display system according to the first embodiment. As shown in FIG. 1, the display system includes an analysis apparatus 10 and a display apparatus 20. The analysis apparatus 10 is, for example, a server. The display apparatus 20 is, for example, a personal computer.
[0011] The analysis device 10 accepts input of information on a language describing a pattern and data of a string in a predetermined format (hereinafter sometimes simply referred to as a string). The analysis unit 131 of the analysis device 10 performs syntax analysis of the string. Then, the display control unit 135 of the analysis device 10 generates web page information based on the result of the syntax analysis and transmits it to the display device 20.
[0012] The display device 20 displays a web page using functions such as a browser based on the web page information received from the analysis device 10. Note that the analysis device 10 may start syntax analysis processing in response to an access request to the web page by the display device 20.
[0013] In the example of FIG. 1, the analysis device 10 accepts input of information on a language describing a JSON pattern and JSON-formatted data. The analysis unit 131 extracts a person's name from the JSON-formatted data. The display control unit 135 performs rendering of a web page that displays the extracted person's name.
[0014] FIG. 2 is a diagram showing a configuration example of the analysis device according to the first embodiment. As shown in FIG. 2, the analysis device 10 includes an interface unit 11, a storage unit 12, and a control unit 13.
[0015] The interface unit 11 is an interface for input / output of data. The interface unit 11 accepts input of data via an input device such as a mouse or a keyboard, for example. Further, the interface unit 11 outputs data to an output device such as a display, for example. Also, the interface unit 11 may be a communication interface such as a NIC (Network Interface Card) for performing data communication with other devices.
[0016] The storage unit 12 is a storage device such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), or an optical disk. Note that the storage unit 12 may be a semiconductor memory capable of rewriting data, such as a RAM (Random Access Memory), a flash memory, or an NVSRAM (Non Volatile Static Random Access Memory). The storage unit 12 stores an OS (Operating System) and various programs executed by the analysis device 10. The storage unit 12 stores, for example, grammar information 121 and analysis result information 122.
[0017] The grammar information 121 is information on a language that describes a predetermined pattern. For example, the grammar information 121 is information described in V-PEG, which will be described later. Note that the grammar information 121 may be stored in the storage unit 12 in advance, or may be input to the analysis device 10 together with the character string to be analyzed.
[0018] The analysis result information 122 is information indicating the intermediate process and the final result of syntactic analysis. For example, the analysis result information 122 may include a memo table used in Packrat Parsing, which will be described later.
[0019] The control unit 13 controls the entire analysis device 10. The control unit 13 is, for example, an electronic circuit such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit), or an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). The control unit 13 also has an internal memory for storing programs and control data that define various processing procedures, and executes each process using the internal memory. Further, the control unit 13 functions as various processing units when various programs operate. For example, the control unit 13 includes an analysis unit 131, an addition unit 132, an extraction unit 133, a determination unit 134, and a display control unit 135.
[0020] The parsing unit 131 performs syntax analysis of the first string based on a grammar described in PEG (Parsing Expression Grammar) that associates variables with predetermined terminal symbols. The parsing unit 131 receives the grammar information 121 and the input to be analyzed, and outputs the analysis result information 122.
[0021] Here, in the present embodiment, a PEG (Parsing Expression Grammar) that associates variables with predetermined terminal symbols is called V-PEG (Parsing Expression Grammar with Variable Bindings). The grammar G in V-PEG is represented as G = (N, Σ, R, V, e S ). N is a finite set of non-terminal symbols. Σ is a finite set of terminal symbols. R is a finite set of rules. V is a finite set of variables. e S is the start expression. A rule is described as A = e. However, A ∈ N. Also, e is as shown in FIG. 3. FIG. 3 is a diagram showing the syntax of V-PEG.
[0022] FIG. 4 is a diagram for explaining the input and output of syntax analysis. As shown in FIG. 4, the parsing unit 131 receives the input of the information 201 described in V-PEG and the string 202. Then, the parsing unit 131 outputs the analysis result. Here, the parsing unit 131 receives the input of information in which the pattern of CSV is described in V-PEG, and outputs data in which item names such as date, name, and age are associated with each item value as the final analysis result.
[0023] Here, as a syntax analysis method for analyzing a context-independent pattern, Packrat Parsing is known. In Packrat Parsing, recursive descent syntax analysis, backtracking, and memoization are performed. A Packrat Parser, which is a syntax analyzer implementing Packrat Parsing, has a parse function for analysis corresponding to the grammar. The parse function is expressed as follows, where N is a set of non-terminal symbols and I is a set of positions on the input. parse : N × I → I
[0024] In Packrat Parsing, variables associated with terminal symbols are unnecessary. For example, in the grammar of HTML, it is described as follows. HTML::=‘<’Name‘>’HTML*‘< / ’Name‘>’|‘<’Name‘>’ Name::=[a-zA-Z]+
[0025] On the other hand, Stateful Packrat Parsing, a syntax analysis method that extends Packrat Parsing, analyzes context-dependent patterns. In Stateful Packrat Parsing, the grammar is described in Extended Backus-Naur Form (EBNF) and three functions (scope, bind, match) as follows. <ebnf> HTML = scope('<' bind(v, Name) '>' HTML* '< / ' match(v, Name) '>') | '<' Name '>' Name = [a-zA-Z]+
[0026] scope, bind, and match are functions used to examine the correspondence relationship of tags in an HTML file. At this time, for an HTML file with consecutive opening tags (for example <c>When the input is given as (...), it is known that the processing time becomes exponentially slower in proportion to the number of opening tags when using a conventional Stateful Packrat Parsing syntax parser.
[0027] Here, it is known that patterns dependent on the context required in practice can be expressed with just a few functions (see, for example, Reference 1). Reference 1: KURAMITSU, K. A symbol-based extension of parsing expression grammars and context-sensitive packrat parsing. In Proceedings of the 10th ACM SIGPLAN International Conference on Software Language Engineering (New York, NY, USA, 2017), SLE 2017, ACM, pp. 26-37.
[0028] Therefore, in this embodiment, instead of allowing arbitrary functions to be incorporated as in Stateful Packrat Parsing, a grammar is described using V-PEG that allows only the minimum necessary functions to be incorporated. For example, a grammar is described in V-PEG as follows. <v-peg> HTML = scope('<' bind(v, Name) '>' HTML* '< / ' match(v, Name) '>') / '<' Name '>' Name = [a-zA-Z]+
[0029] In this example, the grammar G is represented as follows. G = (N, Σ, R, V, e S ) N = {HTML, Name} Σ = {'<', '>', ' / ', 'a', …, 'z', 'A', …, 'Z'} R: The set of the above rules for HTML and Name V = {v} e S = HTML
[0030] In EBNF, multiple expressions are separated by the symbol '|', while in V-PEG, multiple expressions are separated by the symbol ' / '. For example, in the above V-PEG, the two expressions for HTML, "scope('<' bind(v, Name) '>' HTML* '< / ' match(v, Name) '>')" and "'<' Name '>'", are separated by the symbol ' / '.
[0031] For example, when there are two expressions, α and β, in EBNF, it is described as "α | β". On the other hand, in V-PEG, it is described as "α / β". Here, α | β tries whether the string matches β even if it matches α, while α / β does not try whether the string matches β when it matches α.
[0032] For example, in this embodiment, when the string does not match the expression "scope('<' bind(v, Name) '>' HTML* '< / ' match(v, Name) '>')", the analysis unit 131 analyzes whether it matches "'<' Name '>'".
[0033] The adding unit 132 adds, to a variable, an element that is a part of the first character string and to which a predetermined attribute is assigned to the second character string analyzed by the analyzing unit 131 as corresponding to an end symbol. For example, the adding unit 132 adds, to the right end of a variable that is an array, an element in a key-value format having the attribute as a key and the second character string as a value. The adding unit 132 receives an input of a character string and an attribute, and outputs an element or a variable to which the element is added.
[0034] Here, taking the case of HTML as an example, the attribute may be an "opening tag" and a "closing tag". In this case, the opening tag of HTML is "<" and ">" enclosing only [a-zA-Z]+, that is, one or more uppercase and lowercase alphabets. On the other hand, the closing tag of HTML is "< / " and ">" enclosing only [a-zA-Z]+, that is, one or more uppercase and lowercase alphabets.
[0035] For example, " < / c> <c>< / c> When analyzing the string "」" based on the grammar of HTML, the addition part 132 is a variable E that is an array. m Add elements such as (v1, a), (v2, b), and (v1, c) to m . v1 is the key corresponding to the attribute "opening tag". v2 is the key corresponding to the attribute "closing tag". Also, the addition part 132 adds new elements to the right side of the array.
[0036] The extraction part 133 extracts the latest element of each attribute from the variable. For example, the extraction part 133 extracts the rightmost element among the elements with the same array key. For example, consider the case where the variable E m = [(v1, a), (v2, b), (v1, c)]. In this case, there are two elements with v1 as the key, and the extraction part 133 extracts the newer (v1, c). The extraction part 133 accepts the input of the variable before extraction and outputs the variable after extraction. For example, the extraction part 133 accepts the input of the variable E m = [(v1, a), (v2, b), (v1, c)] and outputs the variable E m = [(v1, c), (v2, b)].
[0037] The determination part 134 determines whether the element extracted by the extraction part 133 meets a predetermined condition regarding the context. In the example of HTML, the determination part 134 determines whether the character string of the element extracted by the extraction part 133 matches the character string in the closing tag in the first character string. In the foregoing example, the extraction unit 133 extracts the element (v1, c). That is, the character string of the element extracted by the extraction unit 133 is c. c is the character string in the opening tag. On the other hand, the character string " <c>" is an example of the first character string. In this case, b is the character string in the closing tag.< / c>
[0038] The variable after extraction by the extraction unit 133 is E m =[(v1,c),(v2,b)], the determination unit 134 determines whether the character strings inside the opening tag and the closing tag are the same. In this case, the character string of the opening tag is c, and the character string of the closing tag is b. Therefore, the determination unit 134 determines that the character strings inside the opening tag and the closing tag are not the same. In the HTML syntax, the character strings inside the opening tag and the corresponding closing tag are the same. Therefore, this determination is based on the dependency relationship in the HTML context.
[0039] FIG. 5 is a diagram showing the algorithm of the parse function. The parse function of the present embodiment, as shown in the first line of FIG. 5, has the variable E m applied with the filter function. The parse function in the conventional Stateful Packrat Parsing (see Reference 2) does not include the filter function. The filter function represents the process executed by the extraction unit 133. The parse function of the present embodiment does not record another global variable E e . Reference 2: FORD, B. Packrat parsing:: Simple, powerful, lazy, linear time, functional pearl. In Proceedings of the Seventh ACM SIGPLAN International Conference on Functional Programming (New York, NY, USA, 2002), ICFP’02, ACM, pp. 36-47.
[0040] In the parse function in the conventional Stateful Packrat Parsing, the position i on the input, the non-terminal symbol A, and the whole of the global variables are recorded. On the other hand, in the parse function of the present embodiment, the position i on the input, the non-terminal symbol A, and a part of the global variables are recorded. M S is a memo table, a function that takes a quadruple as an argument and returns a triple (i’, E m ’, E e ’). dom is a function that returns the domain of M S . The bar-arrow between the key on the 5th line and (j, E m ’, E e ’) is a symbol indicating that the element of M S is to be replaced.
[0041] [Processing of the First Embodiment] FIG. 6 is a flowchart showing the processing flow of the analysis apparatus according to the first embodiment. As shown in FIG. 6, first, the analysis apparatus 10 receives an input of V-PEG and a character string (step S11). Next, the analysis apparatus 10 executes the parse function to perform analysis (step S12). Then, the analysis apparatus 10 outputs the analysis result (step S13).
[0042] FIG. 7 is a flowchart showing the processing flow of the parse function. FIG. 7 is a flowchart showing the details of the processing in step S12 of FIG. 6. Here, it is assumed that the analysis unit 131 executes parse(A, i, E m , E e ). It is assumed that the variable E m is an array. Also, the initial value of i is 0. Also, A is, for example, HTML.
[0043] First, the extraction unit 133 prepares a triple (A, i, filter(Em)) consisting of applying the filter function to A, i, and E m (step S101). Next, the determination unit 134 determines whether there is an element corresponding to (A, i, filter(Em)) in the memo table M s (step S102).
[0044] When there is an element corresponding to (A, i, filter(Em)) in the memo table M s (step S102, Yes), the analysis unit 131, in M s Returns the element corresponding to (A, i, filter(Em)) in it (step S103). On the other hand, the analysis unit 131 refers to the memo table M s If there is no element corresponding to (A, i, filter(E m )) in it (step S102, No), set A = e and execute parse(e, i, E m , E e ) (step S104).
[0045] Then, the analysis unit 131 records the return value (j, E' m , E' e ) of parse(e, i, E m , E' e ) in the memo table M s (step S105). Further, the analysis unit 131 returns the return value (j, E' m , E' e ) of parse(e, i, E m , E' e ) (step S106).
[0046] [Effect of the First Embodiment] As described above, the analysis unit 131 performs syntax analysis of the first character string based on the grammar described in PEG in which variables are associated with predetermined terminal symbols. Further, the addition unit 132 adds an element obtained by assigning a predetermined attribute to a second character string that is part of the first character string and is analyzed by the analysis unit 131 to correspond to a terminal symbol, to the variable. Further, the extraction unit 133 extracts the latest element of each attribute from the variable. Further, the determination unit 134 determines whether the element extracted by the extraction unit 133 matches a predetermined condition regarding the context. In this way, the analysis device 10 does not extract elements that are not the latest among the elements stored in the variable. Therefore, according to the present embodiment, it is possible to shorten the time required for syntax analysis corresponding to a context-dependent pattern. According to the present embodiment, it is possible to suppress the processing time, which has been increasing exponentially in the conventional technology, to polynomial time.
[0047] The adding part 132 adds an element in key-value format with an attribute as the key and the second string as the value to the right end of a variable that is an array. Also, the extraction part 133 extracts the rightmost element among the elements with the same key in the array. In this way, by using the array as a variable, the latest element can be easily extracted.
[0048] The parsing part 131 performs syntax analysis of the first string based on the HTML grammar. Also, the adding part 132 adds an element with an attribute indicating an opening tag to the variable for the string in the opening tag in the first string. Also, the extraction part 133 extracts the latest element among the elements with an attribute indicating an opening tag. Also, the determination part 134 determines whether the string of the element extracted by the extraction part 133 matches the string in the closing tag in the first string. Thereby, syntax interpretation along with the HTML syntax can be performed.
[0049] [Packrat Parsing] Here, for comparison with this embodiment, the conventional Packrat Parsing is described. Here, it is assumed that the HTML grammar is described as follows. HTML::=‘<’Name‘>’HTML*‘< / ’Name‘>’|‘<’Name‘>’ Name::=[a-zA-Z]+
[0050] The string to be parsed is “ Let it be so. Also, parse(HTML, i) is a function that parses HTML starting from position i in the input. Also, parse(Name, i) is a function that parses Name starting from position i in the input. When the arguments of the parse function are determined, the return value is uniquely determined. Also, in Packrat Parsing, all the parsing results of the parse function are recorded in a memo table, and when the parse function is called with arguments that have already been parsed, the recorded parsing result is returned without performing the parsing.
[0051] Figures 8 to 16 are diagrams showing an example of a memo table. In the upper table of Figure 8, each character of the string to be parsed and its position are shown. The lower table is a memo table that records the parsing result (H) of HTML and the parsing result (N) of Name. "?" is the initial value (for example, Null). Here, it is assumed that the parsing device 10a performs the parsing.
[0052] First, the parsing device 10a executes parse(HTML, 0). Here, since the "a" at position i = 1 corresponds to the terminal symbol Name, the parsing device 10a executes parse(Name, 1). Then, since the ">" at position i = 2 does not match Name, the parsing device 10a records 2 at position i = 1 in the Name table as the parsing result of parse(Name, 1), as shown in Figure 9.
[0053] Furthermore, since the "<" at position i = 3 corresponds to the terminal symbol HTML, the parsing device 10a executes parse(HTML, 3). Then, at position i = 4, the parsing device 10a further executes parse(Name, 4). Since the ">" at position i = 5 does not match Name, the parsing device 10a records 5 at position i = 4 in the Name table as the parsing result of parse(Name, 4), as shown in Figure 10.
[0054] Furthermore, since the "<" at the position where i = 6 corresponds to the Name of the end symbol, the analysis device 10a executes parse(HTML, 6). Then, at the position where i = 7, the analysis device 10a further executes parse(Name, 7). Since the " / " at the position where i = 7 does not match Name, as shown in FIG. 11, the analysis device 10a records fail at the position of i = 7 in the Name table as the analysis result of parse(Name, 7).
[0055] Although parse(HTML, 6) results in a match failure, due to the following conditions, the analysis device 10a tries the following conditions. The analysis device 10a backtracks to i = 6 and executes parse(Name, 7 when it advances to i = 7. As shown in FIG. 11, since the analysis result at the position of i = 7 already exists, the analysis device 10a does not perform the analysis again.
[0056] As a result, as shown in FIG. 12, the analysis device 10a records fail, which is the analysis result of parse(HTML, 6), at the position of i = 6 in the HTML table. Furthermore, when the analysis device 10a advances to i = 8, it executes parse(Name, 8). As a result of parse(Name, 8), since the ">" at the position of i = 9 does not match Name, as shown in FIG. 1 3 the analysis device 10a records 9 at the position of i = 8 in the Name table as the analysis result of parse(Name, 8).
[0057] Furthermore, since parse(HTML, 3) completely matches HTML at the position where i = 9, as shown in FIG. 1 4 the analysis device 10a records 10 at the position of i = 3 in the HTML table as the analysis result of parse(HTML, 3). Here, the analysis device 10a assumes the position of i = 10 and regards it as not matching HTML at the position of i = 10. Also, as shown in FIG. 16, the analysis result of parse(HTML, 10) is also recorded as fail.
[0058] The parsing device 10a backtracks to i = 0. As shown in FIG. 16, since the parsing result at the position of i = 4 already exists, the parsing device 10a does not perform parsing again. Then, the parsing device 10a records 3 at the position of i = 0 in the HTML table as the parsing result of parse(HTML, 0).
[0059] According to this parsing result, " 」 and 「 This means that each "」" corresponds to a separate HTML. However, in HTML, the string inside the opening and closing brackets must be the same. Therefore, in order to avoid such parsing results, the determination unit 134 of this embodiment determines whether or not it meets a predetermined condition related to the context. Furthermore, the extraction process by the extraction unit 133 can shorten the processing time for the determination.
[0060] [Other Embodiments] In the first embodiment, the case where the global variable is an array was described as an example. On the other hand, the global variable may be data other than an array. For example, the global variable may be a stack.
[0061] In this case, the addition unit 132 pushes the second string onto the stack corresponding to the attribute. Then, the extraction unit 133 extracts by popping the top of the stack. For example, assume that the stack corresponds to the opening bracket of HTML. At this time, the addition unit 132 adds by pushing the string inside the opening bracket as the second string onto the stack. Therefore, the latest element will exist at the top of the stack.
[0062] The determination unit 134 determines whether or not the strings inside the opening tag and the closing tag are the same by executing the following check function. Note that the extraction unit 133 pushes the top of the stack S when the string matches the closing tag. check(opening_tag): if S.empty(): return false closing_tag = S.top() return opening_tag == closing_tag
[0063] [System Configuration, etc.] Moreover, each component of each illustrated device is functionally conceptual and does not necessarily have to be physically configured as shown in the figures. That is, the specific forms of distribution and integration of each device are not limited to those shown in the figures, and all or part of them can be functionally or physically distributed or integrated in any unit according to various loads, usage situations, etc. Furthermore, each processing function performed by each device can be realized in whole or in any part by a CPU and a program analyzed and executed by the CPU, or can be realized as hardware by wired logic.
[0064] Also, among the various processes described in this embodiment, all or part of the processes described as being automatically performed can be manually performed, or all or part of the processes described as being manually performed can be automatically performed by a known method. In addition, regarding the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above documents and drawings, they can be arbitrarily changed unless otherwise specified.
[0065] [Program] As one embodiment, the analysis device 10 can be implemented by installing an analysis program that executes the above-described analysis process as package software or online software on a desired computer. For example, by causing the information processing device to execute the above analysis program, the information processing device can function as the analysis device 10. The information processing device mentioned here includes desktop or notebook personal computers. In addition, other information processing devices include mobile communication terminals such as smartphones, mobile phones, and PHS (Personal Handyphone System), and further slate terminals such as PDAs (Personal Digital Assistant) are included in this category.
[0066] Further, the analysis device 10 can also be implemented as an analysis server device that uses the terminal device used by the user as a client and provides services related to the above analysis processing to the client. For example, the analysis server device is implemented as a server device that takes graph data as input and outputs the results of graph signal processing or analysis of graph data. In this case, the analysis server device may be implemented as a web server, or may be implemented as a cloud that provides services related to the above analysis processing through outsourcing.
[0067] FIG. 17 is a diagram showing an example of a computer that executes an analysis program. The computer 1000 has, for example, a memory 1010 and a CPU 1020. The computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0068] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM 1012. The ROM 1011 stores, for example, a boot program such as a BIOS (BASIC Input Output System). The hard disk drive interface 1030 is connected to the hard disk drive 1090. The disk drive interface 1040 is connected to the disk drive 1100. A removable storage medium such as a magnetic disk or an optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to, for example, a mouse 1110 and a keyboard 1120. The video adapter 1060 is connected to, for example, a display 1130.
[0069] The hard disk drive 1090 stores, for example, the OS 1091, application programs 1092, program modules 1093, and program data 1094. That is, the programs that define each process of the analysis device 10 are implemented as program modules 1093 in which computer-executable code is described. The program modules 1093 are stored, for example, in the hard disk drive 1090. For example, program modules 1093 for executing processes similar to the functional configuration in the analysis device 10 are stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced by an SSD.
[0070] Also, the setting data used in the processes of the above-described embodiments is stored as program data 1094, for example, in the memory 1010 or the hard disk drive 1090. Then, the CPU 1020 reads out the program modules 1093 and program data 1094 stored in the memory 1010 or the hard disk drive 1090 into the RAM 1012 as necessary, and executes the processes of the above-described embodiments.
[0071] Note that the program modules 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090, and may be stored, for example, in a removable storage medium and read by the CPU 1020 via a disk drive 1100 or the like. Alternatively, the program modules 1093 and program data 1094 may be stored in another computer connected via a network (such as a LAN (Local Area Network) or WAN (Wide Area Network)). Then, the program modules 1093 and program data 1094 may be read by the CPU 1020 from another computer via the network interface 1070.
Explanation of Signs
[0072] 10 Analysis device 20 Display device 11 Interface section 12 Memory section 13 Control section 121 Grammar information 122 Parsing result information 131 Parsing section 132 Addition section 133 Extraction section 134 Judgment section 135 Display control section < / ebnf>
Claims
1. An expression for matching a character string, comprising an analysis unit that analyzes whether a first character string matches an expression in which a part corresponding to an attribute is indicated by a terminal symbol and a non-terminal symbol, an addition unit that, when the first character string matches the expression, adds an element obtained by assigning the attribute to a second character string that matches a non-terminal symbol among parts corresponding to the attribute of the first character string to a variable, an extraction unit that extracts the latest element of each attribute from the variable, a determination unit that determines whether an element extracted by the extraction unit meets a predetermined condition related to context, and an analysis device characterized by having the above.
2. The addition unit adds an element in key-value format having the attribute as a key and the second character string as a value to the right end of the variable which is an array, The extraction unit extracts the rightmost element among elements having the same key in the array. The analysis device according to claim 1, characterized by this.
3. The addition unit pushes the second character string onto a stack corresponding to the attribute, The extraction unit extracts by popping the top of the stack. The analysis device according to claim 1, characterized by this.
4. The analysis unit performs syntax analysis of the first character string based on the grammar of HTML, The addition unit adds an element obtained by assigning an attribute indicating an opening tag to a character string in the opening tag in the first character string to the variable, The extraction unit extracts the latest element among elements of the attribute indicating an opening tag, The determination unit determines whether a character string of an element extracted by the extraction unit matches a character string in a closing tag in the first character string. The analysis device according to any one of claims 1 to 3, characterized by this.
5. An analysis method executed by an analysis device, an analysis step of analyzing whether a first character string matches an expression in which a part corresponding to an attribute is indicated by a terminal symbol and a non-terminal symbol, which is an expression for matching a character string, an addition step of, when the first character string matches the expression, adding an element obtained by assigning the attribute to a second character string that matches a non-terminal symbol among parts corresponding to the attribute of the first character string to a variable, an extraction step of extracting the latest element of each attribute from the variable, a determination step of determining whether an element extracted by the extraction step meets a predetermined condition related to context, and including the above. The analysis method is characterized by this.
6. An analysis program for causing a computer to function as the analysis device according to any one of claims 1 to 4.