Binary XML Scalar Extraction via Node Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current streaming evaluation techniques for XPath expressions on binary-encoded XML data are inefficient, leading to increased demand on computer resources and poor performance during evaluation.
Innovation Solution
The technique involves identifying whether nodes in binary-encoded XML data are simple or complex, allowing for minimal processing to extract scalar values, and using node information stored in the binary-encoded XML data or XML schema to optimize XPath evaluation by avoiding unnecessary decoding and parsing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If streaming evaluation techniques are used for XPath expressions on binary-encoded XML data, then evaluation can be performed without complete decoding, but processing overhead increases and performance deteriorates
Solution Approach 1:
The patent segments the XML node processing into two distinct categories: simple nodes and complex nodes. This segmentation allows the system to apply different processing strategies to each type, optimizing performance by avoiding unnecessary processing steps for simple nodes while maintaining proper evaluation for complex nodes.
Solution Approach 2:
The patent applies partial action by performing only the necessary processing for each node type. For simple nodes, it extracts scalar values directly without complete decoding or parsing, performing less action than traditional methods. For complex nodes, it performs the necessary decoding and parsing actions. This selective approach reduces overall processing overhead while maintaining evaluation accuracy.
2Reliability
If complete decoding and parsing is performed on all XML nodes, then accurate XPath evaluation is achieved, but processing overhead increases and performance decreases
Solution Approach 1:
The patent applies local quality by making the processing intensity dependent on the local characteristics of each XML node. Simple nodes receive minimal processing (direct scalar value extraction), while complex nodes receive full processing (decoding and parsing). This localized adaptation of processing quality maintains evaluation accuracy for complex nodes while optimizing performance across the entire XML structure.
Solution Approach 2:
The patent changes the processing parameter based on node type classification. For simple nodes, the processing parameter is set to minimal extraction operations. For complex nodes, the processing parameter is set to full decoding and parsing operations. This dynamic parameter adjustment based on node characteristics resolves the contradiction between accuracy and performance.
Data Source
AI summary
Techniques are provided for efficiently extracting scalar values from binary-encoded XML data. Node information is stored in association with binary-encoded XML data to indicate whether one or more nodes of an XML document are simple or complex. A node is simple if the node has no child elements and no attributes. The node information of a particular node is used to determine whether a particular node, identified in a query, is simple or complex. If the particular node is simple, then the scalar value of the particular node is identified without performing any operations other than possibly converting the scalar value to a non-binary-encoded format or converting the scalar value to a value of a different data type.


