Xml file reading method, device, equipment and storage medium

CN116644213BActive Publication Date: 2026-09-25SHENZHEN FULIN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310672194.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-07
Publication Date
2026-09-25
Estimated Expiration
2043-06-07

AI Technical Summary

Technical Problem

但是,这种传统的文件读取方法不仅需要占用较多服务器内存,在加载XML文件的过程中还需要耗费大量的时间,导致XML文件的读取效率较低

Benefits of technology

[0046]本发明实施例中,首先通过利用预定义的类函数定义XML文件的Root节点,可以确定占用XML文件最大内存的Root节点之下的单个完整节点的大小,减少了读取XML文件的内存占用时间,便于提高后续的XML文件读取效率;其次,通过获取Root节点之后的XML文件中的内容元素,能够逐个读取文档中的内容元素,实现流式处理XML文件的元素,提高后续XML文件读取的效率,利用预定义的栈变量定义内容元素中的父子元素关系,能够通过栈变量捕捉不同内容元素的父子关系,正确的表示XML文件的层次结构,通过利用预设的解码器对内容元素进行遍历,得到内容标签,能够便于后续依照XML的实际结构捕捉每个元素的完整内容;最后,通过识别内容标签对应的标签类型,并判断标签类型是否为结束标签,当标签类型不为结束标签,能够随时读取标签类型对应的标签内容,得到XML文件的读取结果,以实现在XML文件读取的过程中无需加载完整的XML文件,减少XML文件读取所需的时间,提高XML文件的读取效率。因此本发明提出的XML文件读取方法、装置、设备及存储介质可以提高XML文件的读取效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116644213B_ABST
    Figure CN116644213B_ABST
Patent Text Reader

Abstract

The application belongs to the field of data analysis, and relates to an XML file reading method, which comprises the following steps: acquiring an XML file to be read, defining a Root node of the XML file by using a class function; acquiring content elements after the Root node by using an element function, and defining parent-child element relationships in the content elements by using a stack variable; traversing the content elements by using a preset decoder to obtain content tags; identifying a tag type corresponding to the content tags by using a type statement, and judging whether the tag type is an end tag according to the parent-child element relationships; when the tag type is not the end tag, acquiring a tag order of the tag type in the XML file, reading contents in the tag type in sequence according to the tag order, and obtaining a reading result of the XML file; and when the tag type is the end tag, stopping the reading of the XML file. The application further provides an XML device, a computer device and a storage medium. The application can improve the reading efficiency of the XML file.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data analysis technology, and in particular to an XML file reading method, apparatus, device, and storage medium. Background Technology

[0002] XML is a structured data exchange format that can define and encode data, enabling data to be easily transmitted and stored between different systems. The traditional method of reading XML files is to use an XML parser to read the entire XML file into memory, convert the XML file into an XML DOM object, and provide a programming interface for programmers to read the XML file through the interface.

[0003] With the development of big data, XML files are becoming increasingly large, and reading large XML files has become the mainstream approach. However, this traditional file reading method not only consumes a lot of server memory but also takes a significant amount of time to load the XML file, resulting in low efficiency in reading XML files. Summary of the Invention

[0004] This invention provides an XML file reading method, apparatus, device, and storage medium, the main purpose of which is to improve the reading efficiency of XML files.

[0005] To address the aforementioned technical problems, this application provides an XML file reading method, employing the following technical solution:

[0006] Obtain the XML file to be read, and define the root node of the XML file using a predefined class function;

[0007] The content elements after the Root node are obtained using a predefined element function, and the parent-child relationship of the content elements is defined using a predefined stack variable;

[0008] The content elements are traversed using a preset decoder to obtain content tags;

[0009] The preset type statement is used to identify the tag type corresponding to the content tag, and the parent-child element relationship is used to determine whether the tag type is an end tag;

[0010] When the tag type is not an ending tag, obtain the tag order of the tag type in the XML file, and read the content of the tag type in the tag order to obtain the reading result of the XML file;

[0011] If the tag type is an end tag, then reading the XML file will stop.

[0012] Furthermore, defining the parent-child element relationship in the content element using predefined stack variables includes:

[0013] The content element is pushed onto the stack using the stack variable to store it in the stack variable, and the current content element that needs to be read is identified from the content elements.

[0014] Determine whether the current content element has a parent element. If the current content element has a parent element, add the current content element to the parent element's child element array to obtain the parent-child element relationship of the current content element.

[0015] Furthermore, the step of using a predefined element function to obtain the content elements after the Root node includes:

[0016] The element function call cursor is used to read any element after the Root node, and the element object corresponding to the cursor is output.

[0017] Determine whether the element object is an end identifier;

[0018] When the element object is an end identifier, all element objects in the XML file are output, and the element object is used as the content element;

[0019] If the element object is not an end identifier, the cursor continues to read the next element object in the XML file until the element object is an end identifier. Then, all element objects in the XML file are output, and the element object is used as the content element.

[0020] Furthermore, determining whether the tag type is an end tag based on the parent-child element relationship includes:

[0021] Obtain the content ending tag in the tag type, and obtain the tag parent element of the content ending tag according to the parent-child element relationship;

[0022] Determine whether the output content of the parent element of the tag is an end identifier;

[0023] If the output content of the parent element of the tag is not an end identifier, then the tag type is determined to be not an end tag;

[0024] If the output content of the tag's parent element is an end identifier, then the tag type is determined to be an end tag.

[0025] Furthermore, the step of sequentially reading the content of the tag types according to the tag order to obtain the reading result of the XML file includes:

[0026] The preset second token function is used to parse the content attribute values ​​of the content start tag, the text content of the text tag, and the comment content of the comment tag in the XML file.

[0027] Based on the parent-child element relationship and the tag order, the content attribute values, text content, and annotation content will be output as the reading result of the XML file.

[0028] Furthermore, after parsing the content attribute values ​​of the content start tag, the text content of the text tag, and the comment content of the comment tag in the XML file using the preset second token function, the method further includes:

[0029] Determine whether the name attribute value in the content attribute value belongs to the root node;

[0030] If the name attribute value in the content attribute value does not belong to the Root node, then the content attribute value is determined to be the reading result of the content start tag;

[0031] If the name attribute value in the content attribute value belongs to the Root node, then the content corresponding to the content start tag does not need to be read.

[0032] Furthermore, the step of traversing the content elements using a preset decoder to obtain content tags includes:

[0033] The content element's marker is identified one by one using a preset first token function. When the marker is an end identifier, all content tags corresponding to the content element are obtained. The decoder includes the first token function.

[0034] To address the aforementioned technical problems, this application also provides an XML file reading device, which employs the following technical solution:

[0035] The acquisition module is used to acquire the XML file to be read and to define the root node of the XML file using predefined class functions;

[0036] The definition module is used to obtain the content elements after the Root node using predefined element functions, and to define the parent-child element relationship in the content elements using predefined stack variables;

[0037] The traversal module is used to traverse the content elements using a preset decoder to obtain content tags;

[0038] The identification module is used to identify the tag type corresponding to the content tag using a preset type statement, and to determine whether the tag type is an end tag based on the parent-child element relationship; and

[0039] The reading module is used to obtain the tag order of the tag type in the XML file when the tag type is not an end tag, and read the content of the tag type in the tag type sequentially according to the tag order to obtain the reading result of the XML file; when the tag type is an end tag, the reading of the XML file is stopped.

[0040] To address the aforementioned technical problems, this application also provides a computer device that employs the following technical solution:

[0041] Memory, storing at least one computer program; and

[0042] The processor executes the computer program stored in the memory to perform the XML file reading described above.

[0043] To address the aforementioned technical problems, this application also provides a computer-readable storage medium, employing the technical solution described below:

[0044] The computer-readable storage medium stores at least one computer program, which is executed by a processor in an electronic device to perform the XML file reading described above.

[0045] Compared with the prior art, this application has the following main advantages:

[0046] In this embodiment of the invention, firstly, by defining the root node of the XML file using predefined class functions, the size of a single complete node under the root node, which occupies the largest amount of memory in the XML file, can be determined, reducing the memory usage time of reading the XML file and improving the efficiency of subsequent XML file reading. Secondly, by obtaining the content elements in the XML file after the root node, the content elements in the document can be read one by one, realizing streaming processing of the XML file elements and improving the efficiency of subsequent XML file reading. By defining the parent-child relationship of the content elements using predefined stack variables, the parent-child relationship of different content elements can be captured through the stack variables, correctly representing the hierarchical structure of the XML file. By traversing the content elements using a preset decoder, the content tags are obtained, which facilitates the subsequent capture of the complete content of each element according to the actual structure of the XML. Finally, by identifying the tag type corresponding to the content tag and determining whether the tag type is an end tag, the tag content corresponding to the tag type can be read at any time when the tag type is not an end tag, so as to obtain the reading result of the XML file. This achieves the goal of not loading the complete XML file during the XML file reading process, reducing the time required for XML file reading and improving the efficiency of XML file reading. Therefore, the XML file reading method, apparatus, device, and storage medium proposed in this invention can improve the reading efficiency of XML files. Attached Figure Description

[0047] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 This is an exemplary system architecture diagram to which this application can be applied;

[0049] Figure 2 This is a flowchart illustrating one implementation of the XML file reading method according to this application;

[0050] Figure 3 This is a structural diagram of an embodiment of the control terminal described in the XML file reading system according to this application;

[0051] Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation

[0052] The data format determination method provided in this invention is applied to a data processing system. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing specific embodiments only and is not intended to limit the application. The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or accompanying drawings of this application are used to distinguish different objects, not to describe a specific order.

[0053] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0054] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0055] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0056] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social online platform software.

[0057] Terminal devices 101, 102, and 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 players (Moving Picture Experts Group Audio Layer IV), laptops, and desktop computers, etc.

[0058] Server 105 can be a server that provides various services, such as a backend server that supports the pages displayed on terminal devices 101, 102, and 103.

[0059] It should be noted that the XML file reading method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the XML file reading system is generally set in the server / terminal device.

[0060] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0061] Continue to refer to Figure 2 A flowchart illustrating an embodiment of the XML file reading method according to this application is shown. The XML file reading method includes the following steps:

[0062] S210 obtains the XML file to be read and defines the root node of the XML file using a predefined class function.

[0063] In this embodiment of the invention, the XML file to be read refers to an Extensible Markup Language (XML) file, which is typically used to mark up data and define data types. It is a source language file that allows users to define their own markup language, and the XML file to be read can be a streaming file.

[0064] In this embodiment of the invention, the Root node refers to the outermost element of the XML file, which includes the entire content of the XML file and is the starting point of the XML file.

[0065] In this embodiment of the invention, by obtaining the XML file to be read and defining the root node of the XML file using a predefined class function, the size of a single complete node under the root node that occupies the largest memory of the XML file can be determined, which reduces the memory usage time of reading the XML file and facilitates the improvement of subsequent XML file reading efficiency.

[0066] In one embodiment of the present invention, the predefined class function can be the NewXmlScanner function. This function defines the XML file and returns an XmlScanner object, which is the rootTag. The root element character can be defined through the returned rootTag, and the XML file stream can be passed into the XML parser through this function, which facilitates the subsequent reading of the XML file.

[0067] S220 uses a predefined element function to obtain the content elements after the Root node, and uses a predefined stack variable to define the parent-child element relationship in the content elements.

[0068] In this embodiment of the invention, the predefined element function can be the NextElement function, which reads the elements after the Root node one by one by calling a cursor or pointer, and returns the Element object corresponding to each element. The content element refers to a single complete XML element that needs to be read under the Root element, including but not limited to the tag name, XML attributes and attribute values, XML text content, XML child element array and parent element array, etc.

[0069] In this embodiment of the invention, the stack variable is a data structure with the "last in first out" characteristic. Through the "last in first out" characteristic of the stack variable, the child elements of the content element can be included in the parent element, thereby ensuring the correct hierarchical structure of the content elements and accurately representing the structure of the XML file.

[0070] In this embodiment of the invention, the parent-child element relationship is the tag relationship in the XML file. For example, when the tag "project" is nested within the tag "description" in the XML file, then the tag "description" is the child element and the tag "project" is the parent element.

[0071] In this embodiment of the invention, by using predefined element functions to obtain content elements after the Root node, the content elements in the document can be read one by one, realizing streaming processing of XML file elements, improving the efficiency of subsequent XML file reading, and using predefined stack variables to define the parent-child element relationship in the content elements, the parent-child relationship of the content elements can be captured, ensuring the correct hierarchical structure of the content elements, so as to accurately represent the structure of the XML file, which facilitates the improvement of the accuracy of subsequent XML file reading.

[0072] In this embodiment of the invention, obtaining the content elements after the Root node using a predefined element function includes:

[0073] The element function call cursor is used to read any element after the Root node, and the element object corresponding to the cursor is output.

[0074] Determine whether the element object is an end identifier;

[0075] When the element object is an end identifier, all element objects in the XML file are output, and the element object is used as the content element;

[0076] If the element object is not an end identifier, the cursor continues to read the next element object in the XML file until the element object is an end identifier. Then, all element objects in the XML file are output, and the element object is used as the content element.

[0077] Specifically, the element function call cursor is used to read any element after the Root node and returns the element object (i.e., Element object) of that element object. By reading the element objects one by one through the cursor, it is not necessary to load the entire XML file, thus improving the efficiency of subsequent XML file reading.

[0078] In one embodiment of the present invention, the end identifier can be the nil symbol. By calling the cursor to move down and reading elements, if the returned element object is nil, it means that the element object is the end identifier, and all element objects read before nil are used as content elements; if the returned element object is not nil, it means that the reading of the XML file content is not complete, and the cursor needs to be called again to read until nil is returned, so as to obtain the complete content elements and ensure the accuracy of subsequent XML file reading.

[0079] Furthermore, in this embodiment of the invention, defining the parent-child element relationship in the content element using a predefined stack variable includes:

[0080] The content element is pushed onto the stack using the stack variable to store it in the stack variable, and the current content element that needs to be read is identified from the content elements.

[0081] Determine whether the current content element has a parent element. If the current content element has a parent element, add the current content element to the parent element's child element array to obtain the parent-child element relationship of the current content element.

[0082] The push operation can be performed by calling stack.push(content element) to store the content element in a stack variable and define the parent-child relationship of the content elements through the stack variable.

[0083] In one embodiment of the present invention, when the current content element A is parsed, it is first determined whether the current content element A has a parent element B. If the current content element A has a parent element B, then the current content element A has a parent element B, and the current content element A is added to the Children array of content element B. If the current content element A does not have a parent element B, it means that the current content element A is the last parent element and there is no parent-child element relationship. By subsequently determining whether the parent element is the last parent element, the end element of the XML file being read can be determined.

[0084] S230 uses a preset decoder to traverse the content elements and obtain content tags.

[0085] In this embodiment of the invention, the preset decoder can be `encoding / xml.NewDecoder` from the Golang system library. This decoder can decode all content elements in an XML file, thereby determining the content tags corresponding to each content element, thus enabling the reading of the XML file. The content tags refer to the markers applied to the content elements by the decoder, such as `StartElement`, `EndElement`, `CharData`, and `Comment`.

[0086] In this embodiment of the invention, by using a preset decoder to traverse the content elements and obtain content tags, it is easier to capture the complete content of each element according to the actual structure of XML.

[0087] In this embodiment of the invention, the step of traversing the content elements using a preset decoder to obtain content tags includes:

[0088] The content element's marker is identified one by one using a preset first token function. When the marker is an end identifier, all content tags corresponding to the content element are obtained. The decoder includes the first token function.

[0089] The first token function can be the RawToken function; the token symbol is the content tag. The RawToken function marks the content elements of the input decoder and returns the corresponding content tag of the content element. This continues until the NextToken function marks the content element and returns an end identifier (such as nil), which means that the end of the XML file has been read.

[0090] S240 uses a preset type statement to identify the tag type corresponding to the content tag, and determines whether the tag type is an end tag based on the parent-child element relationship.

[0091] In this embodiment of the invention, the preset type statement can be a switch statement, which is mainly used to determine the tag type of the content tag. The tag type can include content start tag, content end tag, text tag, and comment tag, etc. The end tag refers to the end tag corresponding to the content element, which can be an EndElement token.

[0092] In one embodiment of the present invention, if the StartElement in the content tag is captured using a switch statement, it means that the start symbol of the XML file has been captured, and the content start tag (such as StartElement token) is output; if the EndElement in the content tag is captured, it means that the end symbol of the XML file has been captured, and the content end tag (such as EndElement token) is output; if the CDATA text string in the content tag is captured, the text tag (such as CharData token) is output; if the Comment in the content tag is captured, it means that the comment block symbol of the XML file has been captured, and the comment tag (such as CharData token) is output.

[0093] In this embodiment of the invention, by using a preset type statement to identify the tag type corresponding to the content tag, and by determining whether the tag type is an end tag based on the parent-child element relationship, the tag content corresponding to the tag type can be read at any time to obtain the reading result of the XML file. This achieves the goal of not needing to load the complete XML file during the XML file reading process, reducing the time required for XML file reading, and improving the reading efficiency of XML file.

[0094] As an embodiment of the present invention, determining whether the tag type is an end tag based on the parent-child element relationship includes:

[0095] Obtain the content ending tag in the tag type, and obtain the tag parent element of the content ending tag according to the parent-child element relationship;

[0096] Determine whether the output content of the parent element of the tag is an end identifier;

[0097] If the output content of the parent element of the tag is not an end identifier, then the tag type is determined to be not an end tag;

[0098] If the output content of the tag's parent element is an end identifier, then the tag type is determined to be an end tag.

[0099] The content end tag can be an EndElement token. By obtaining the parent-child element relationship of the content element corresponding to the content end tag, the tag parent element of the content tag can be found.

[0100] In this embodiment of the invention, a content ending tag C is called stack.pop(C) to pop the top element C from the stack. If the parent element of C is not an ending identifier (nil), it means that the current tag type is not an ending tag; if the parent element of C is nil, it means that the tag type is an ending tag.

[0101] S250 When the tag type is not an end tag, obtain the tag order of the tag type in the XML file, and read the content of the tag type in sequence according to the tag order to obtain the reading result of the XML file.

[0102] In this embodiment of the invention, the tag order refers to the order of content elements in the XML file. For example, an XML file contains elements such as <tag>, <script>, and <script>. <cdata> 、 <description>and<Project id="2”> The output, based on the order of the XML files, should be Project id="2", description, and CDATA content.

[0103] As an embodiment of the present invention, the step of sequentially reading the content of the tag types according to the tag order to obtain the reading result of the XML file includes:

[0104] The preset second token function is used to parse the content attribute values ​​of the content start tag, the text content of the text tag, and the comment content of the comment tag in the XML file.

[0105] Based on the parent-child element relationship and the tag order, the content attribute values, text content, and annotation content will be output as the reading result of the XML file.

[0106] The preset second token function can be a token function, and the second token function can be the same as or different from the first token function, depending on the actual scenario.

[0107] In one embodiment of the present invention, the Token function can parse the content starting tag (such as StartElementtoken) to obtain the content attribute values ​​including Name and Attr, where the Name attribute represents the tag name and the Attr attribute represents all attributes and attribute values ​​of the current XML tag; the text content can be parsed using the Token function to obtain a CDATA text string, and the text content can be saved by calling Element.setText(text); the comment content can be parsed using the Token function to obtain a complete segment comment. When it is necessary to obtain information from the comments of the XML file, the comment content can be split by newline characters and the comment text array can be traversed.

[0108] In this embodiment of the invention, by parsing the content in the tag type using the Token function, it can be ensured that the returned StartElement token and EndElement token will be correctly nested and matched. If the Token encounters an unexpected end element, or encounters an EOF before all expected end elements, the Token will return an error, thereby ensuring the accuracy of reading the XML file.

[0109] Furthermore, in this embodiment of the invention, after parsing the content attribute values ​​of the tag type start tag, the text content of the tag type text tag, and the comment content of the tag type comment tag in the XML file using a preset second token function, the method further includes:

[0110] Determine whether the name attribute value in the content attribute value belongs to the root node;

[0111] If the name attribute value in the content attribute value does not belong to the Root node, then the content attribute value is determined to be the reading result of the content start tag;

[0112] If the name attribute value in the content attribute value belongs to the Root node, then the content corresponding to the content start tag does not need to be read.

[0113] In this embodiment of the invention, the name of the Root node can be any valid XML element name. By determining whether the name attribute value in the content attribute value belongs to the Root node, the Root can be filtered out in the Token function processing. Since the Root node is the starting point of the XML file and its main function is to describe the content or purpose of the XML file, the Root node does not contain XML content during the reading process of the XML file, but it still occupies a certain amount of memory. Therefore, filtering the Root during the reading process of the XML file can ensure the accuracy of the read XML file while improving the reading efficiency of the XML file.

[0114] S260 When the tag type is an end tag, the reading of the XML file is stopped.

[0115] In this embodiment of the invention, when the tag type is an end tag, it indicates that the last end symbol of the XML file has been read, and the complete XML line content corresponding to the end tag can be returned, indicating that the XML file reading is complete.

[0116] Compared with the prior art, the embodiments of this application have the following main advantages:

[0117] In this embodiment of the invention, firstly, by defining the root node of the XML file using predefined class functions, the size of a single complete node under the root node, which occupies the largest amount of memory in the XML file, can be determined, reducing the memory usage time of reading the XML file and improving the efficiency of subsequent XML file reading. Secondly, by obtaining the content elements in the XML file after the root node, the content elements in the document can be read one by one, realizing streaming processing of the XML file elements and improving the efficiency of subsequent XML file reading. By defining the parent-child relationship of the content elements using predefined stack variables, the parent-child relationship of different content elements can be captured through the stack variables, correctly representing the hierarchical structure of the XML file. By traversing the content elements using a preset decoder, the content tags are obtained, which facilitates the subsequent capture of the complete content of each element according to the actual structure of the XML. Finally, by identifying the tag type corresponding to the content tag and determining whether the tag type is an end tag, the tag content corresponding to the tag type can be read at any time when the tag type is not an end tag, thus obtaining the reading result of the XML file. This achieves the goal of not loading the complete XML file during the XML file reading process, reducing the time required for XML file reading and improving the efficiency of XML file reading. Therefore, the XML file reading method proposed in this embodiment of the invention can improve the reading efficiency of XML files.

[0118] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0119] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0120] Further reference Figure 3 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of an XML file reading device 310, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0121] An embodiment of the present invention provides an XML file reading system, the XML file reading system comprising:

[0122] The acquisition module 311 is used to acquire the XML file to be read and to define the root node of the XML file using a predefined class function;

[0123] Module 312 is defined to obtain the content elements after the Root node using a predefined element function, and to define the parent-child element relationship in the content elements using a predefined stack variable;

[0124] Traversal module 313 is used to traverse the content elements using a preset decoder to obtain content tags;

[0125] The identification module 314 is used to identify the tag type corresponding to the content tag using a preset type statement, and to determine whether the tag type is an end tag based on the parent-child element relationship; and

[0126] The reading module 315 is used to obtain the tag order of the tag type in the XML file when the tag type is not an end tag, and read the content of the tag type in the tag type in sequence according to the tag order to obtain the reading result of the XML file; when the tag type is an end tag, the reading of the XML file is stopped.

[0127] Compared with the prior art, the embodiments of this application have the following main advantages:

[0128] In this embodiment of the invention, firstly, by defining the root node of the XML file using predefined class functions, the size of a single complete node under the root node, which occupies the largest amount of memory in the XML file, can be determined, reducing the memory usage time of reading the XML file and improving the efficiency of subsequent XML file reading. Secondly, by obtaining the content elements in the XML file after the root node, the content elements in the document can be read one by one, realizing streaming processing of the XML file elements and improving the efficiency of subsequent XML file reading. By using predefined stack variables to define the parent-child relationship in the content elements, the parent-child relationship of different content elements can be captured through the stack variables, correctly representing the hierarchical structure of the XML file. By using a preset decoder to traverse the content elements, the content tags are obtained, which facilitates the subsequent capture of the complete content of each element according to the actual structure of the XML. Finally, by identifying the tag type corresponding to the content tag and determining whether the tag type is an end tag, if the tag type is not an end tag, the tag content corresponding to the tag type can be read at any time to obtain the reading result of the XML file. This achieves the goal of not loading the complete XML file during the XML file reading process, reducing the time required for XML file reading and improving the efficiency of XML file reading. Therefore, the XML file reading device proposed in this embodiment of the invention can improve the reading efficiency of XML files.

[0129] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.

[0130] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0131] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0132] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may also include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as program code for XML file reading methods. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.

[0133] In some embodiments, the processor 42 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run program code stored in the memory 41 or process data, for example, to run the program code for the XML file reading method.

[0134] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.

[0135] This application also provides another embodiment, namely, providing a computer-readable storage medium storing the XML file reading method program, which can be executed by at least one processor to cause the at least one processor to perform the steps of the XML file reading method as described above.

[0136] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware online platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0137] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0138] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.< / description> < / cdata>

Claims

1. A method for reading XML files, characterized in that, Includes the following steps: Obtain the XML file to be read, and define the root node of the XML file using a predefined class function; The content elements following the Root node are obtained using predefined element functions, and the parent-child element relationships in the content elements are defined using predefined stack variables; wherein, the content element refers to a single complete XML element that needs to be read under the Root element, and the stack variable is a data structure with last-in-first-out characteristics; The content elements are traversed using a preset decoder to obtain content tags; The preset type statement is used to identify the tag type corresponding to the content tag, and the parent-child element relationship is used to determine whether the tag type is an end tag; When the tag type is not an ending tag, obtain the tag order of the tag type in the XML file, and read the content of the tag type in the tag order to obtain the reading result of the XML file; If the tag type is an end tag, then reading the XML file will stop; The step of defining the parent-child element relationship in the content element using a predefined stack variable includes: The content element is pushed onto the stack using the stack variable to store it in the stack variable, and the current content element that needs to be read is identified from the content elements. Determine whether the current content element has a parent element. If the current content element has a parent element, add the current content element to the child element array of the parent element to obtain the parent-child element relationship of the current content element. The step of determining whether the current content element has a parent element includes: obtaining whether the current content element has an upper-level nested content element; if the current content element has an upper-level nested content element, then determining that the current content element has a parent element and adding the current content element to the subarray of the nested content element; if the current content element does not have an upper-level nested function, then determining that the current content element is the last parent element and that the current content element does not have a parent-child element relationship. The step of determining whether the tag type is an end tag based on the parent-child element relationship includes: Obtain the content ending tag in the tag type, and obtain the tag parent element of the content ending tag according to the parent-child element relationship; Determine whether the output content of the parent element of the tag is an end identifier; If the output content of the parent element of the tag is not an end identifier, then the tag type is determined to be not an end tag; If the output content of the tag's parent element is an end identifier, then the tag type is determined to be an end tag; The step of reading the content of the tag types sequentially according to the tag order to obtain the reading result of the XML file includes: The preset second token function is used to parse the content attribute values ​​of the content start tag, the text content of the text tag, and the comment content of the comment tag in the XML file. The second token function is a Token function. The content start tag parsed by the Token function contains the Name and Attr content attribute values, where the Name attribute represents the tag name and the Attr attribute represents all attributes and attribute values ​​of the current XML tag. The text content is parsed using the Token function to obtain a CharData token containing CDATA, and the text content is saved by calling Element.setText(text). The comment content is parsed using the Token function to obtain a complete segment comment. When information needs to be obtained from the comments in the XML file, the comment content is split by newline characters, and the comment text array is traversed. Based on the parent-child element relationship and the tag order, the content attribute values, text content, and annotation content will be output as the reading result of the XML file.

2. The XML file reading method according to claim 1, characterized in that, The step of retrieving the content elements after the Root node using a predefined element function includes: The element function call cursor is used to read any element after the Root node, and the element object corresponding to the cursor is output. Determine whether the element object is an end identifier; When the element object is an end identifier, all element objects in the XML file are output, and the element object is used as the content element; If the element object is not an end identifier, the cursor continues to read the next element object in the XML file until the element object is an end identifier. Then, all element objects in the XML file are output, and the element object is used as the content element.

3. The XML file reading method according to claim 1, characterized in that, After parsing the content attribute values ​​of the content start tag, the text content of the text tag, and the comment content of the comment tag in the XML file using the preset second token function, the method further includes: Determine whether the name attribute value in the content attribute value belongs to the root node; If the name attribute value in the content attribute value does not belong to the Root node, then the content attribute value is determined to be the reading result of the content start tag; If the name attribute value in the content attribute value belongs to the Root node, then the content corresponding to the content start tag does not need to be read.

4. The XML file reading method according to claim 1, characterized in that, The step of traversing the content elements using a preset decoder to obtain content tags includes: The content element's marker is identified one by one using a preset first token function. When the marker is an end identifier, all content tags corresponding to the content element are obtained. The decoder includes the first token function.

5. An XML file reading device, characterized in that, include: The acquisition module is used to acquire the XML file to be read and to define the root node of the XML file using predefined class functions; The definition module is used to obtain the content elements after the Root node using predefined element functions, and to define the parent-child element relationship in the content elements using predefined stack variables; wherein, the content element refers to a single complete XML element that needs to be read under the Root element, and the stack variable is a data structure with last-in-first-out characteristics; The traversal module is used to traverse the content elements using a preset decoder to obtain content tags; The identification module is used to identify the tag type corresponding to the content tag using a preset type statement, and to determine whether the tag type is an end tag based on the parent-child element relationship; and The reading module is used to obtain the tag order of the tag type in the XML file when the tag type is not an end tag, and read the content of the tag type in the XML file sequentially according to the tag order to obtain the reading result of the XML file; when the tag type is an end tag, the reading of the XML file is stopped. The definition module includes: The push operation submodule is used to perform a push operation on the content element using the stack variable to store the content element in the stack variable and identify the current content element that needs to be read from the content elements. The parent element determination submodule is used to determine whether the current content element has a parent element. If the current content element has a parent element, the current content element is added to the parent element's child element array to obtain the parent-child element relationship of the current content element. The parent element determination submodule includes: An element acquisition unit is used to determine whether the current content element has an upper-level nested content element; The first judgment unit is used to determine if the current content element has a parent element if the current content element has a parent nested content element, and to add the current content element to the subarray of the nested content element. The second judgment unit is used to determine that if the current content element does not have a nested function, then the current content element is the last parent element and the current content element does not have a parent-child element relationship. The identification module includes: The parent element retrieves the child module, which is used to retrieve the content ending tag in the tag type and retrieve the tag parent element of the content ending tag according to the parent-child element relationship; The identifier determination submodule is used to determine whether the output content of the tag's parent element is an end identifier; The first identifier determination submodule is used to determine that the tag type is not an end tag when the output content of the tag's parent element is not an end identifier; The second identifier determination submodule is used to determine that the tag type is an end tag when the output content of the tag's parent element is an end identifier; The reading module includes: The content parsing submodule is used to parse the content attribute values ​​of the content start tag, the text content of the text tag, and the comment content of the comment tag in the XML file using the preset second token function. The second token function is a Token function. The content start tag parsed by the Token function contains the Name and Attr content attribute values, where the Name attribute represents the tag name and the Attr attribute represents all attributes and attribute values ​​of the current XML tag. The text content is parsed using the Token function to obtain a CharData token containing CDATA, and the text content is saved by calling Element.setText(text). The comment content is parsed using the Token function to obtain a complete segment comment. When information needs to be obtained from the comments in the XML file, the comment content is split by newline characters, and the comment text array is traversed. The result output submodule is used to output the content attribute values, text content, and annotation content as the reading results of the XML file according to the parent-child element relationship and the tag order.

6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the XML file reading method as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the XML file reading method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • XML data analyzing method, generating method and processing system

    CN105868257A

  • Generalized data analysis method based on xml file

    CN110457526A