Data analysis method and device

By processing the index of the parsed data and establishing a parsing tree, the FastJson framework's low efficiency when parsing JSON strings is solved, and efficient data parsing is achieved.

CN120146027APending Publication Date: 2025-06-13BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311704924.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, the FastJson framework needs to parse JSON strings layer by layer when parsing JSON strings, resulting in cumbersome and inefficient parsing process.

Method used

By processing the index of the data to be parsed based on the attribution relationship between the data to be parsed, the values ​​corresponding to the new index and the new index are obtained, and a parsing tree is established to obtain the parsing results, so that data analysis does not need to be parsed layer by layer.

Benefits of technology

A simple, convenient and fast data analysis method is realized, effectively improving data analysis efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146027A_ABST
    Figure CN120146027A_ABST
Patent Text Reader

Abstract

The invention discloses a data analysis method and device, and relates to the technical field of computers. A specific embodiment of the method comprises the steps of obtaining to-be-analyzed data in response to a received data analysis request; according to the attribution relation between the to-be-analyzed data, processing the index of the to-be-analyzed data to obtain a corresponding new index and a value corresponding to the new index; and according to the new index and the value corresponding to the new index, establishing a parsing tree so as to obtain a parsing result corresponding to the data parsing request by reading the parsing tree. According to the implementation mode, a simple, convenient and rapid data analysis method is realized, data does not need to be analyzed layer by layer, and the data analysis efficiency is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method and device for data parsing. Background Art

[0002] In the scenario of network communication interaction, the received byte stream is usually a JSON string. The receiving party needs to parse the received JSON string into a Java object before subsequent processing and application. Currently, the open-source parsing library FastJson framework is mainly used to parse JSON strings.

[0003] In the process of implementing the present invention, the inventors found the following problems in the prior art:

[0004] The FastJson framework requires that the established Java object needs to be consistent with the hierarchical structure of the JSON data. When parsing JSON, it is necessary to parse layer by layer according to the hierarchy, extract the values therein, and then assign them to the corresponding Java object. This method has a cumbersome parsing process and low efficiency. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a method and device for data parsing, which implement a simple, convenient, and fast data parsing method, without the need to parse data layer by layer, effectively improving the data parsing efficiency.

[0006] To achieve the above object, according to one aspect of the embodiments of the present invention, a method for data parsing is provided, including:

[0007] Responding to a received data parsing request, and obtaining data to be parsed;

[0008] According to the attribution relationship between the data to be parsed, processing the index of the data to be parsed to obtain a corresponding new index and the value corresponding to the new index;

[0009] According to the new index and the value corresponding to the new index, establishing a parsing tree, so as to obtain a parsing result corresponding to the data parsing request by reading the parsing tree.

[0010] Optionally, processing the index of the data to be parsed according to the attribution relationship between the data to be parsed to obtain a corresponding new index includes: performing a structure parsing on the data to be parsed to determine the attribution relationship between each data in the data to be parsed and the connection symbol corresponding to the attribution relationship; and performing a flattening process on the index of the data to be parsed according to the attribution relationship and the connection symbol corresponding to the attribution relationship to generate a new index of the data to be parsed.

[0011] Optionally, when the data to be parsed is in the JSON string format, the structure parsing of the data to be parsed includes: performing structure parsing by identifying symbols in the data to be parsed according to the symbol features of the JSON string format.

[0012] Optionally, after establishing the parse tree, the method further includes: generating a corresponding class dictionary tree according to the parse tree, so as to obtain the parse result corresponding to the data parsing request by reading the class dictionary tree.

[0013] Optionally, when there are non-integer type values in the parse tree, after obtaining the corresponding new index, the method further includes: generating a corresponding integer type encoding for the new index according to the encoding rule; establishing a first association relationship between the new index and the corresponding encoding, and a second association relationship between the encoding and the value corresponding to the new index; establishing a parse tree according to the new index and the value corresponding to the new index, including: establishing a parse tree according to the new index and the encoding corresponding to the new index, so as to obtain the parse result corresponding to the data parsing request according to the encoding and the second association relationship therein on the premise of meeting the data type requirements of the class dictionary tree.

[0014] Optionally, the second association relationship is established and maintained in the form of a list, so as to query the corresponding value from the list according to the encoding.

[0015] Optionally, the encoding rule is: initializing the starting encoding, and sequentially determining the encoding corresponding to the new index through the self-increment operation of the starting encoding; the list is constructed in the following way: sequentially storing the value corresponding to the new index into a newly created list file, so as to determine the parse result according to the positions where each value is stored in the list, and the positions have the information of the encoding.

[0016] According to the second aspect of the embodiments of the present invention, there is provided a data parsing device, including:

[0017] A data acquisition module, configured to acquire data to be parsed in response to receiving a data parsing request;

[0018] An index processing module, configured to process the index of the data to be parsed according to the attribution relationship between the data to be parsed, to obtain a corresponding new index and the value corresponding to the new index;

[0019] A parse tree establishment module, configured to establish a parse tree according to the new index and the value corresponding to the new index, so as to obtain the parse result corresponding to the data parsing request by reading the parse tree.

[0020] According to a third aspect of an embodiment of the present invention, there is provided an electronic device for data parsing, including:

[0021] One or more processors;

[0022] A storage device for storing one or more programs,

[0023] When the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the first aspect of the embodiment of the present invention.

[0024] According to a fourth aspect of an embodiment of the present invention, there is provided a computer-readable medium having a computer program stored thereon, and when the program is executed by a processor, the method provided in the first aspect of the embodiment of the present invention is implemented.

[0025] One embodiment of the invention has the following advantages or beneficial effects: By responding to a received data parsing request, obtaining data to be parsed; processing the index of the data to be parsed according to the attribution relationship between the data to be parsed, obtaining a corresponding new index and the value corresponding to the new index; establishing a parsing tree according to the new index and the value corresponding to the new index, so as to obtain the parsing result corresponding to the data parsing request by reading the parsing tree. The technical solution realizes a simple, convenient and fast data parsing method, without the need to parse the data layer by layer, and effectively improves the data parsing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention. Among them:

[0027] Figure 1 is a schematic diagram of the main process of the data parsing method according to the embodiment of the present invention;

[0028] Figure 2 is a schematic diagram of the code execution process of JSON string data parsing according to the embodiment of the present invention;

[0029] Figure 3 is a schematic diagram of the principle process of the JSON string data parsing method according to the embodiment of the present invention;

[0030] Figure 4 is a schematic diagram of the data parsing principle of JSON string data according to the embodiment of the present invention;

[0031] Figure 5 is a schematic diagram of the main modules of the data parsing device according to the embodiment of the present invention;

[0032] Figure 6 is an exemplary system architecture diagram to which the embodiment of the present invention can be applied;

[0033] Figure 7 It is a schematic structural diagram of a computer system of a terminal device or a server suitable for implementing the embodiments of the present invention. Detailed implementation manners

[0034] It should be noted that in the technical solutions of the present disclosure, the acquisition, storage, application, etc. of the user's personal information all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0035] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0036] The existing data parsing method, the FastJson framework, requires that the Java objects to be established be consistent with the hierarchical structure of the JSON data. When parsing JSON, it parses layer by layer according to the hierarchy, extracts the values therein, and assigns them to the corresponding Java objects. The parsing process is cumbersome and inefficient, and cannot well meet the actual application.

[0037] To solve the above problems existing in the prior art, the present invention proposes a data parsing method. Based on the attribution relationship between the data to be parsed, the index of the data to be parsed is processed to obtain a new index that is convenient for constructing a parsing tree. Then, a parsing tree is established according to the new index and the value corresponding to the new index, so as to obtain the parsing result by reading the parsing tree, realizing simple, convenient, and fast data parsing without parsing the data layer by layer, and effectively improving the data parsing efficiency.

[0038] In the introduction of the embodiments of the present invention, the nouns involved and their meanings are as follows:

[0039] FastJson: It is an open-source JSON parsing library that can parse JSON-formatted strings;

[0040] TreeMap: It is a data structure that stores and organizes objects according to key-value pairs. It provides a method for quickly finding specific elements and an effective method for finding the keys associated with the elements;

[0041] FST: It is a data structure similar to a trie, but it is in a k-v structure and can quickly query according to the index.

[0042] Figure 1 It is a schematic diagram of the main process of the data parsing method according to the embodiments of the present invention, as Figure 1As shown in the figure, the method for data parsing according to the embodiment of the present invention includes the following steps S101 to S103.

[0043] Step S101: In response to receiving a data parsing request, obtain the data to be parsed.

[0044] Specifically, according to the received data parsing request, obtain the data to be parsed from the data parsing request. In the case where the data volume of the data to be parsed is large, the data to be parsed can also be temporarily stored in a specified storage space, and the requesting party that receives the parsing request obtains the data to be parsed according to the data storage address in the parsing request.

[0045] Step S102: According to the attribution relationship between the data to be parsed, process the index of the data to be parsed to obtain a corresponding new index and the value corresponding to the new index.

[0046] Specifically, considering that the data to be parsed is usually a data set, and there is a certain attribution relationship between the data included therein. For example, data object A includes data objects A1 and A2, then A1 and A2 both belong to A. According to the attribution relationship between the data to be parsed, the embodiment of the present invention processes the indexes corresponding to each data included in the data to be parsed to obtain a new index that can represent the attribution relationship and the value corresponding to the new index. It can be understood that the index here usually refers to the key in key-value. In addition, each attribute name in the data to be parsed can also be used as an index, or a unique identification code is assigned to each attribute, and this identification code is used as an index, and the value corresponding to the index of the data to be parsed remains unchanged, that is, the value corresponding to the new index, which is also the value corresponding to the index of the data to be parsed corresponding to it.

[0047] According to an embodiment of the present invention, processing the index of the data to be parsed according to the attribution relationship between the data to be parsed to obtain a corresponding new index includes: performing a structure analysis on the data to be parsed to determine the attribution relationship between each data in the data to be parsed and the connection symbol corresponding to the attribution relationship; according to the attribution relationship and the connection symbol corresponding to the attribution relationship, performing a flattening process on the index of the data to be parsed to generate a new index of the data to be parsed.

[0048] Specifically, perform a structure analysis on the obtained data to be parsed, such as the data objects, arrays included therein, and the composition of each data object and array, etc., to obtain the attribution relationship between each data in the data to be parsed and determine the connection symbol corresponding to the attribution relationship. According to the attribution relationship and the connection symbol corresponding to the attribution relationship, the embodiment of the present invention performs a flattening process on the indexes corresponding to the data included in the data to be parsed, flattening the originally structured hierarchical data into one level, and the characteristics of the original structured level are represented by the new index.

[0049] According to another embodiment of the present invention, when the data to be parsed is in the JSON string format, the structural parsing of the data to be parsed includes: performing structural parsing by identifying symbols in the data to be parsed according to the symbol features of the JSON string format.

[0050] Specifically, when the data to be parsed is in the JSON string format, according to the syntax rules of the JSON string, the symbol features of the JSON string are obtained. For example, what is stored inside square brackets is an array, that is, the square brackets are the flags of the array; what is stored inside curly brackets is an object, that is, the curly brackets are the flags of the object; what follows a colon is a value, that is, the colon is the flag of the value; the content inside double quotes is an index or the value corresponding to the index, that is, the double quotes are the flags of the index or the value, etc. Based on these symbol features, each character of the JSON string to be parsed is judged and identified, and the composition of each data in the data to be parsed is obtained through the identified symbols, and then the attribution relationship between the data and the corresponding connectors are determined.

[0051] Furthermore, the indexes identified by the above symbol recognition are written into a custom index stack. For the indexes corresponding to data with an attribution relationship, the connectors corresponding to the attribution relationship also need to be written into the index stack, so that when the data exits the index stack by using the stack structure, the indexes with an attribution relationship are connected by the connectors to obtain a new flattened index.

[0052] In addition, the present invention embodiment also defines a symbol stack for storing symbols, which is mainly used to determine the composition of the data to be parsed through symbol matching during structural parsing. For example, the left curly bracket "{" and the right curly bracket "}" are matched; and a position stack for recording the array position sequence number information included in the connectors is defined. For example, if array A includes object 1 and object 2, then the connector corresponding to object 1 is &0, and the connector corresponding to object 2 is &1. The position sequence numbers 0 and 1 are determined by this position stack.

[0053] Step S103, establish a parsing tree according to the new index and the value corresponding to the new index, so as to obtain the parsing result corresponding to the data parsing request by reading the parsing tree.

[0054] Specifically, the embodiment of the present invention is based on the above-mentioned new flattened index and the value corresponding to the new index, associates the value corresponding to the new index with the node of the new index, and establishes a parsing tree TreeMap. In this way, when parsing the data to be parsed, it only needs to clarify the index of the specific data and the attribution relationship of the index as needed, and query the corresponding index from the parsing tree to obtain the corresponding value.

[0055] According to an embodiment of the present invention, after the parsing tree is established, the method further includes: generating a corresponding class dictionary tree according to the parsing tree, so as to obtain the parsing result corresponding to the data parsing request by reading the class dictionary tree.

[0056] Specifically, according to the newly generated index representing the attribution relationship, it can be known that the newly generated indexes corresponding to multiple data with the same attribution relationship have the same prefix. Embodiments of the present invention utilize this feature. After the parsing tree is established, based on the parsing tree, a class dictionary tree FST similar to a dictionary tree is constructed in such a way that each node represents a character of the newly generated index. The newly generated indexes in the parsing tree are added to the respective nodes of the class dictionary tree character by character, and the value corresponding to the newly generated index is placed on the last ending character of the newly generated index, obtaining the class dictionary tree corresponding to the data to be parsed. This can not only share the nodes with the same characters to the greatest extent and optimize the occupied memory space. For example, for newly generated indexes with the same prefix, the nodes of the prefix characters can be shared, and for newly generated indexes with the same suffix, the nodes of the suffix characters can be shared; moreover, it also supports fuzzy query of newly generated indexes to meet the needs of more usage scenarios.

[0057] According to another embodiment of the present invention, in the case where there are non-integer type values in the parsing tree, after obtaining the corresponding newly generated index, the method further includes: generating a corresponding integer type encoding for the newly generated index according to the encoding rule; establishing a first association relationship between the newly generated index and the corresponding encoding, and a second association relationship between the encoding and the value corresponding to the newly generated index; establishing a parsing tree according to the newly generated index and the value corresponding to the newly generated index, including: establishing a parsing tree according to the newly generated index and the encoding corresponding to the newly generated index, so as to obtain the parsing result corresponding to the data parsing request according to the encoding and the second association relationship under the premise of meeting the data type requirements of the class dictionary tree.

[0058] Specifically, considering that the data type of the class dictionary tree FST currently supports integer type or long integer type, in the case where there are non-integer type values in the parsing tree, the current parsing tree is not suitable for direct conversion to the class dictionary tree FST. Therefore, after obtaining the newly generated index, according to the encoding rule, a corresponding integer type encoding is generated for the newly generated index. Specifically, an encoding toolkit can be used to randomly generate a corresponding encoding for the newly generated index. At the same time, a first association relationship between the newly generated index and the encoding, and a second association relationship between the encoding and the value corresponding to the newly generated index are established, that is, the generated encoding is used as an intermediate transfer number for associating the newly generated index and the value corresponding to the newly generated index.

[0059] Further, according to the new index and the encoding corresponding to the new index, construct a parse tree according to the above method, and then add the content in the parse tree to the class trie character by character to ensure that the data in the class trie is all integer or long integer. Further, when querying data, query the encoding corresponding to the new index from the class trie, and then according to the queried encoding, query the corresponding value from the second association relationship. This can break the data type limit of the class trie, expand the applicable range, ensure that all parse trees can be converted into class tries, and further improve the performance of the data parsing method of the embodiments of the present invention, save storage space, and also support fuzzy queries in the full range.

[0060] According to another embodiment of the present invention, the second association relationship is established and maintained in the form of a list, so as to query the corresponding value from the list according to the encoding.

[0061] Specifically, for the convenience of management and maintenance in the embodiments of the present invention, the second association relationship is saved to a list List file, that is, the association relationship between the encoding and the value corresponding to the new index is stored in the List. Then, the class trie FST and the list List form the final parsed new data object.

[0062] According to another embodiment of the present invention, the encoding rule is: initialize the starting encoding, and sequentially determine the encoding corresponding to the new index through the increment operation of the starting encoding; the list is constructed in the following way: sequentially store the values corresponding to the new index into a newly created list file, so as to determine the parsing result according to the positions where the values are stored in the list, and the positions have the information of the encoding.

[0063] Specifically, in order to further optimize the data parsing method, considering that the encoding only serves to associate the new index and the value corresponding to the new index, the embodiments of the present invention simplify the encoding rule to the increment rule of positive integers, with the initialized starting encoding as the increment starting point, and assign the corresponding encoding to the new index by incrementing the encoding by 1 each time. Preferably, taking 1 as the starting encoding, and then assigning the encodings 2, 3, 4, 5... to the corresponding new indexes respectively. Based on the above encoding rule, in the construction of the list List, the values corresponding to the new index can also be directly stored sequentially. In this way, when querying the encoding corresponding to the new index from the class trie, according to the value of the encoding, directly query the value stored at the position corresponding to the encoding from the list List. For example, if the encoding is 10, then just find the 10th value from the list List as the target value.

[0064] Exemplarily, there is JSON data:

[0065]

[0066]

[0067] The corresponding new indexes and corresponding codes after its flattening process are as follows:

[0068]

[0069] Among them, "code", "data.&0.skuId", "data.&0.skuName", "data.&0.skuTags.&0.tagId", "data.&0.skuTags.&1.tagId" are new indexes; 1, 2, 3, 4, 5 are the corresponding codes.

[0070] Figure 2 It is a schematic diagram of the code execution process for parsing JSON string data in an embodiment of the present invention. Among them, kStack is an index stack for storing keys, sStack is a symbol stack for storing symbols, and listStack is a position stack for storing the position numbers of connectors. For the JSON string to be parsed, it is traversed and parsed character by character. If it is a left curly brace {, it is marked as reading the index key. Then, the top character of the symbol stack is judged. If it is a left square bracket [, it means that the read index key is an object in an array. At this time, the connector &num is pushed into the kStack stack, where num is the position number used to mark which object in the array this key is. Subsequently, this left curly brace is pushed into the symbol stack sStack. If it is not a square bracket, then there is no array structure involved, and this left curly brace is directly pushed into the symbol stack sStack. If a left double quote is recognized, the content inside the double quotes is read. The content inside the double quotes is usually an index key or the value corresponding to the key. If the previous step is marked as reading the key, it means that the content of this double quote is the index key, and the read content is added to the index stack kStack. If it is not reading the key, it means that the content here is the value. At this time, it is necessary to traverse and read the keys and connectors in the index stack kStack to obtain the new index pathkey. At the same time, an incrementing code id is generated for the new index. The new index pathkey and the corresponding code id are saved to the parsing tree TreeMap, and the currently read content is saved to the list List. After processing, the top of the index stack kStack is popped to complete the reading of a new index and the corresponding value.

[0071] Similarly, if a colon is recognized, it indicates that the following content is the value, and it is marked as reading the value; if it is a left square bracket '[', the top of the internal counter of the square bracket, that is, the top of the position stack listStack, is set to -1 to facilitate the position serial number after the first accumulation to be 0, and the left square bracket is added to the top of the symbol stack sStack; if a comma is recognized and it is determined that the top of the current sStack stack is a left square bracket, it indicates that this is the separator for multiple objects in the array, and the next content is the index, so it is marked as reading the key; if a right square bracket is recognized, it indicates that the array reading is completed, and the position serial number marked by the position stack listStack is taken out and assigned to num; if a right curly bracket is recognized, it indicates that the reading is completed, and the concatenation symbol stored at the top of the index stack kStack is popped, and the left curly bracket at the top of the symbol stack sStack is popped; if 'default' is recognized, directly read the following content and determine whether it is the index key or the value.

[0072] If the recognition is not completed, continue to recognize each symbol of the unrecognized JSON string according to the above method. If the recognition is completed, convert the obtained parse tree TreeMap into a finite state transducer (FST) similar to a trie. The specific conversion process is to first construct an object of the FST, and add the content of the parse tree TreeMap to the FST tree of the trie-like structure character by character to generate the parsed trie-like FST. Finally, save the trie-like FST and the list List into a new JSON parsing object to complete the parsing.

[0073] Figure 3 It is a schematic diagram of the principle flow of the JSON string data parsing method in the embodiment of the present invention. Traverse and read each character of the JSON string, and process the index of the JSON string according to the attribution relationship between each data in the JSON string to obtain a new index and the corresponding value value. For the new index among them, according to the encoding rule of increasing the encoding by itself, assign a corresponding encoding to each new index; save the new index and the corresponding encoding into the parse tree TreeMap; generate a trie-like structure according to the parse tree. For the value value among them, save the value in the list List in the order of the corresponding encoding. Finally, save the List and the trie-like FST into a new JSON parsing object.

[0074] Figure 4It is a schematic diagram of the data parsing principle of the JSON string data in the embodiment of the present invention. The index of the JSON string is processed to obtain a new index and the value corresponding to the new index; a parsing tree is constructed according to the new index, and then the parsing tree TreeMap is converted into a class dictionary tree FST; the value value corresponding to the obtained new index is stored in the list List in sequence; the class dictionary tree FST containing the new index and the list List containing the value corresponding to the new index are saved into a new JSON parsing object.

[0075] Figure 5 It is a schematic diagram of the main modules of the data parsing device according to the embodiment of the present invention. As Figure 5 shown, the data parsing device 500 mainly includes a data acquisition module 501, an index processing module 502, and a parsing tree establishment module 503.

[0076] The data acquisition module 501 is used to obtain the data to be parsed in response to receiving a data parsing request;

[0077] The index processing module 502 is used to process the index of the data to be parsed according to the attribution relationship between the data to be parsed, to obtain the corresponding new index and the value corresponding to the new index;

[0078] The parsing tree establishment module 503 is used to establish a parsing tree according to the new index and the value corresponding to the new index, so as to obtain the parsing result corresponding to the data parsing request by reading the parsing tree.

[0079] According to an embodiment of the present invention, the index processing module 502 is further used to: perform a structure analysis on the data to be parsed, determine the attribution relationship between each data in the data to be parsed, and the connection symbol corresponding to the attribution relationship; according to the attribution relationship and the connection symbol corresponding to the attribution relationship, perform a flattening process on the index of the data to be parsed to generate a new index of the data to be parsed.

[0080] According to another embodiment of the present invention, when the data to be parsed is in the JSON string format, the index processing module 502 is further used to: perform a structure analysis by identifying the symbols in the data to be parsed according to the symbol characteristics of the JSON string format.

[0081] According to still another embodiment of the present invention, the data parsing device 500 further includes a class dictionary tree generation module (not shown in the figure), which is used to: after establishing a parsing tree, generate a corresponding class dictionary tree according to the parsing tree, so as to obtain the parsing result corresponding to the data parsing request by reading the class dictionary tree.

[0082] According to another embodiment of the present invention, in the case where there are non-integer type values in the parsing tree, the data parsing device 500 further includes an encoding generation module (not shown in the figure), which is configured to: after obtaining the corresponding new index, generate a corresponding integer type encoding for the new index according to the encoding rule; establish a first association relationship between the new index and the corresponding encoding, and a second association relationship between the encoding and the value corresponding to the new index; the parsing tree establishment module 503 is further configured to: establish a parsing tree according to the new index and the encoding corresponding to the new index, so as to obtain the parsing result corresponding to the data parsing request according to the encoding and the second association relationship under the premise of meeting the data type requirements of the class dictionary tree.

[0083] According to another embodiment of the present invention, the second association relationship is established and maintained in the form of a list, so as to query the corresponding value from the list according to the encoding.

[0084] According to still another embodiment of the present invention, the encoding rule is: initialize the starting encoding, and sequentially determine the encoding corresponding to the new index through the increment operation of the starting encoding; the list is constructed in the following manner: sequentially store the value corresponding to the new index into a newly created list file, so as to determine the parsing result according to the positions where each value is stored in the list, and the positions have the information of the encoding.

[0085] Figure 6 It is an exemplary system architecture diagram to which the embodiments of the present invention can be applied.

[0086] As Figure 6 shown, the system architecture 600 may include terminal devices 601, 602, 603, a network 604, and a server 605. The network 604 is used to provide a medium for communication links between the terminal devices 601, 602, 603 and the server 605. The network 604 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0087] Users can use the terminal devices 601, 602, 603 to interact with the server 605 through the network 604 to receive or send messages, etc. Various communication client applications, such as data parsing applications (only for example), may be installed on the terminal devices 601, 602, 603.

[0088] The terminal devices 601, 602, 603 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0089] The server 605 may be a server that provides various services. For example, it can be a background management server (only an example) that supports data parsing performed by users using the terminal devices 601, 602, and 603. The background management server can, in response to receiving a data parsing request, obtain the data to be parsed; process the index of the data to be parsed according to the attribution relationship between the data to be parsed, to obtain a corresponding new index and the value corresponding to the new index; establish a parsing tree according to the new index and the value corresponding to the new index, so as to obtain the parsing result corresponding to the data parsing request through reading the parsing tree and other processing, and feedback the processing result (such as the parsing result, etc.--only an example) to the terminal device.

[0090] It should be noted that the data parsing method provided by the embodiments of the present invention is generally executed by the server 605. Correspondingly, the data parsing device is generally set in the server 605.

[0091] It should be understood that Figure 6 the numbers of the terminal devices, networks, and servers in

[0092] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 7 is a schematic structural diagram of a computer system suitable for implementing the terminal device or server of the embodiments of the present invention. Figure 7 The terminal device or server shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0093] As Figure 7 shown, the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage section 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the system 700 are also stored. The CPU 701, ROM 702, and RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0094] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed on the drive 710 as needed so that a computer program read therefrom is installed into the storage section 708 as needed.

[0095] Specifically, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product including a computer program carried on a computer-readable medium, the computer program including program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by a central processing unit (CPU) 701, the above-described functions defined in the system of the present invention are performed.

[0096] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0098] The units involved in the embodiments of the present invention can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a data acquisition module, an index processing module, and a parsing tree building module.

[0099] Among them, the names of these modules do not constitute a limitation on the modules themselves in some cases. For example, the data acquisition module can also be described as "a module for acquiring data to be parsed in response to receiving a data parsing request".

[0100] On the other hand, the present invention also provides a computer-readable medium. The computer-readable medium can be included in the device described in the embodiments; or it can exist separately without being assembled into the device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the device includes: acquiring data to be parsed in response to receiving a data parsing request; processing the index of the data to be parsed according to the attribution relationship between the data to be parsed to obtain a corresponding new index and the value corresponding to the new index; and building a parsing tree according to the new index and the value corresponding to the new index, so as to obtain a parsing result corresponding to the data parsing request by reading the parsing tree.

[0101] According to the technical solution of the embodiments of the present invention, the following advantages or beneficial effects are achieved: by acquiring data to be parsed in response to receiving a data parsing request; processing the index of the data to be parsed according to the attribution relationship between the data to be parsed to obtain a corresponding new index and the value corresponding to the new index; and building a parsing tree according to the new index and the value corresponding to the new index, so as to obtain a parsing result corresponding to the data parsing request by reading the parsing tree, a simple, convenient, and fast data parsing method is realized, without the need to parse the data layer by layer, effectively improving the data parsing efficiency.

[0102] The specific implementation manners do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for data parsing, characterized in that, comprising: Upon receiving a data parsing request, obtaining the data to be parsed; According to the attribution relationship between the data to be parsed, processing the index of the data to be parsed to obtain a corresponding new index and the value corresponding to the new index; According to the new index and the value corresponding to the new index, building a parsing tree so as to obtain the parsing result corresponding to the data parsing request by reading the parsing tree.

2. The method according to claim 1, characterized in that, According to the attribution relationship between the data to be parsed, processing the index of the data to be parsed to obtain a corresponding new index, including: Performing a structure parsing on the data to be parsed to determine the attribution relationship between each data in the data to be parsed and the connector corresponding to the attribution relationship; According to the attribution relationship and the connector corresponding to the attribution relationship, performing a flattening process on the index of the data to be parsed to generate a new index of the data to be parsed.

3. The method according to claim 2, characterized in that, When the data to be parsed is in the JSON string format, performing a structure parsing on the data to be parsed, including: According to the symbol characteristics of the JSON string format, performing a structure parsing by identifying the symbols in the data to be parsed.

4. The method according to claim 1, characterized in that, After building the parsing tree, the method further includes: According to the parsing tree, generating a corresponding class dictionary tree so as to obtain the parsing result corresponding to the data parsing request by reading the class dictionary tree.

5. The method according to claim 4, characterized in that, When there are non-integer type values in the parsing tree, after obtaining the corresponding new index, the method further includes: according to the encoding rule, generating a corresponding integer type encoding for the new index; establishing a first association relationship between the new index and the corresponding encoding, and a second association relationship between the encoding and the value corresponding to the new index; According to the new index and the value corresponding to the new index, building a parsing tree, including: building a parsing tree according to the new index and the encoding corresponding to the new index, so as to obtain the parsing result corresponding to the data parsing request according to the encoding and the second association relationship under the premise of meeting the data type requirements of the class dictionary tree.

6. The method according to claim 5, characterized in that, The second association relationship is established and maintained in the form of a list so as to query the corresponding value from the list according to the encoding.

7. The method according to claim 6, characterized in that, The encoding rule is: initializing a starting encoding, and sequentially determining the encoding corresponding to the new index through an increment operation on the starting encoding; The list is constructed in the following way: sequentially storing the value corresponding to the new index into a newly created list file so as to determine the parsing result according to the positions where each value is stored in the list, and the positions have the information of the encoding.

8. A data parsing device, characterized in that, comprising: A data acquisition module, configured to acquire data to be parsed in response to receiving a data parsing request; An index processing module, configured to process the index of the data to be parsed according to the attribution relationship between the data to be parsed, so as to obtain a corresponding new index and the value corresponding to the new index; A parsing tree building module, configured to build a parsing tree according to the new index and the value corresponding to the new index, so as to obtain a parsing result corresponding to the data parsing request by reading the parsing tree.

9. A mobile electronic device terminal, characterized in that, it includes: one or more processors; a storage device, configured to store one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.

10. A computer-readable medium, on which a computer program is stored, characterized in that, when the program is executed by a processor, the method according to any one of claims 1-7 is implemented.