A distributed group circle selection SQL automatic generation and calculation method based on multi-layer nested JSON parsing

CN122364258BActive Publication Date: 2026-09-29LIAONING EXPRESSWAY SMART TRAVEL CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610833037.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-10
Publication Date
2026-09-29
Estimated Expiration
2046-06-10

AI Technical Summary

Technical Problem

[0005]本发明的目的在于提供一种基于多层嵌套JSON解析的分布式群体圈选SQL自动生成与计算方法,以解决背景技术中指出的现有处理方式在应对跨实体多层嵌套规则时因缺乏前置过滤拦截机制引发底层引擎拉取冗余数据,进而造成系统内存与算力开销增加以及任务处理耗时偏长的问题

Benefits of technology

本发明通过对群体圈选规则内部包含的跨实体域条件节点执行关联拆解处理,生成带有参数列表的降维结构子树,降低了底层执行引擎处理多表联合查询初期的逻辑复杂度。本发明依据计算获得的嵌套层级深度参量对各个逻辑分支进行排序,将浅层深度的逻辑分支提取为前置独立片段优先触发异步执行,利用返回的初始运算数据在内存中构建特征存在性映射表。本发明利用该特征存在性映射表对后置依赖片段进行运行前探针约束追加处理,强制分布式计算引擎在执行全量硬盘表扫描动作前优先校验数据行的身份主键命中状态。这种前置拦截机制在数据读取阶段提前过滤未命中特征状态的无效数据行,避免了冗余记录直接参与后续消耗算力的多表关联环节,缩减了集群运算过程中的数据吞吐基数,均衡了分布式节点的内存与处理器负荷,改善了系统在面临高并发复杂查询时的任务流转效率与计算耗时。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122364258B_ABST
    Figure CN122364258B_ABST
Patent Text Reader

Abstract

The application provides a distributed group circle selection SQL automatic generation and calculation method based on multi-layer nested JSON parsing, relates to the technical field of distributed data processing, and comprises the following steps: obtaining rules organized in a multi-layer nested JSON format and implementing parsing processing; converting the processed rules into structured query language text and issuing the text for execution to output results. The implementation of the parsing processing comprises the following steps: identifying cross-entity domain nodes and disassembling to generate a dimension reduction subtree; calculating branch nesting depth and sorting, extracting shallow depth branches as pre-independent fragments, and retaining deep branches as post-dependent fragments. The conversion of the text comprises the following steps: controlling asynchronous execution of the pre-independent fragments in priority, using returned data to build a feature existence mapping table in memory; and using the mapping table to add a probe constraint to the post-dependent fragments to obtain a final optimized text. The application can discard invalid data rows in advance and reduce the operation resource consumption of a bottom node when multiple tables are joined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of distributed data processing technology, and in particular to a method for automatically generating and calculating distributed group selection SQL based on multi-level nested JSON parsing. Background Technology

[0002] In current digital operations and data analytics, technical personnel often need to retrieve specific target audience sets from the underlying data warehouse based on various conditional rules. The audience selection rules configured on the business side typically have a multi-layered nested logical structure and cover multiple different business entity data tables, as well as various logical intersection and union constraints.

[0003] When handling such complex group selection tasks, the conventional approach is to parse the overall rules uniformly and directly convert them into structured query statements containing multi-table joins and multi-level subqueries, which are then executed uniformly by the backend distributed computing nodes. Upon receiving such query instructions containing complex relationships, the underlying engine typically initiates extensive disk reads and table scans on the relevant underlying tables, pulling the underlying data into its working memory to complete field comparisons and filtering verifications.

[0004] When faced with massive amounts of underlying data and high-concurrency computational requests, a huge number of data records that have not been filtered by basic conditions will directly participate in the subsequent resource-intensive correlation and statistical processes. This computational mechanism consumes a significant amount of memory and processor scheduling time in the distributed system, objectively prolonging the queuing and processing cycle of query tasks, and thus limiting the data processing efficiency of the cluster when dealing with complex business rules. Summary of the Invention

[0005] The purpose of this invention is to provide a distributed group selection SQL automatic generation and calculation method based on multi-level nested JSON parsing, in order to solve the problem pointed out in the background art that the existing processing methods, when dealing with multi-level nested rules across entities, lack a pre-filtering and interception mechanism, causing the underlying engine to pull redundant data, which in turn leads to increased system memory and computing power consumption and long task processing time.

[0006] This invention provides a method for automatically generating and calculating distributed group selection SQL based on multi-level nested JSON parsing, including: Obtain the group selection rules carrying the set conditions, wherein the group selection rules are organized in a multi-level nested JSON format; The group selection rules are parsed: cross-entity domain condition nodes exist within the group selection rules and cross-entity domain association decomposition is performed to generate a dimensionality-reduced subtree with a parameter list; the nesting depth parameter of each logical branch within the dimensionality-reduced subtree is calculated and sorted, and logical branches at shallow depths are extracted as front-end independent SQL fragments, while logical branches at deep depths are retained as back-end dependent SQL fragments. The parsed rules are converted into Structured Query Language (SQL) statement text: the preceding independent SQL fragments are executed asynchronously first, and a feature existence mapping table is constructed in memory using the returned initial operation data; the following SQL fragments are appended with pre-run probe constraints using the feature existence mapping table, and the final optimized SQL statement text is assembled. The final optimized SQL statement text is used as the structured query language SQL statement text and sent to the distributed computing engine for execution to obtain and output the individual result set of the members.

[0007] Preferably, the group selection rules cover a top-level logical structure, a secondary group structure, a third-level conditional structure, and a fourth-level parameter structure that are nested and related to each other; The top-level logical structure records the logical AND or logical OR operators that represent the global joint decision attributes. The secondary group structure includes a group type identifier and a subordinate condition array; The third-level condition structure records the logical AND or logical OR operators that control the combined state of the subordinate condition array; The fourth-level parameter structure is configured as follows: comparison dimension, data table name, field name, field type, comparison operator, and comparison value.

[0008] Preferably, before obtaining the group selection rules carrying the set conditions, the method further includes: The distributed columnar data warehouse was used to read the individual vehicle profile tables, the electronic non-stop toll collection individual profile tables, and the highway travel application individual profile tables, respectively. The statistical script calculates the number of people covered by multiple profile tables read by tag dimension, and updates the total number of people covered by each tag dimension to the tag statistics result table. The rule parameter enumeration values ​​recorded in the business configuration library are synchronously pushed with the tag statistics result table, and cross-node joint storage is performed in the federated ad-hoc query node to respond to front-end display requests.

[0009] Preferably, the step of sending the final optimized SQL statement text as the Structured Query Language (SQL) statement text to the distributed computing engine for execution, obtaining and outputting the individual member result set, covers the following underlying computational processing: The group clustering requests generated based on the group selection rules and the pending execution status are ordered sequentially according to multiple grouping dependencies; The distributed computing engine is controlled to read multiple subquery instructions included in the final optimized SQL statement text at the computation layer; The multiple subquery instructions are concatenated into a combined view by using a join query statement that retains duplicate records; Nested outside the collection view are grouping aggregation instructions targeting individual primary identifiers; The number comparison interception command is appended to the end of the group aggregation command to perform intersection and union verification calculation, and finally the set of individual member results is output and stored in the offline relation table.

[0010] Preferably, after the step of storing the set of individual member results in an offline relation table, the following steps are included: The parameters of the number of people covered by each audience group and the latest calculation status update time are statistically analyzed in the offline relationship table. The individual member result set, the number of people covered, and the latest calculation status update time parameter are synchronized to the ad-hoc federated query database. The latest snapshot data of the business primary key mapping is retained in the ad hoc federated query database, which enables the application server to initiate fast cross-database retrieval of multiple tables.

[0011] Preferably, the step of identifying cross-entity domain condition nodes within the group selection rules and performing cross-entity domain association decomposition processing to generate a dimensionality-reduced subtree with a parameter list includes: The group selection rules are analyzed to identify multiple independent group type identifiers. When a judgment situation where different independent group type identifiers are intertwined within the nested hierarchy is detected, the logical node that triggers the intertwining operation is marked as the cross-entity domain condition node. In an in-memory computing environment, the range of the source entity data table to which the cross-entity domain condition node is attached is extracted, and a pilot probe instruction pointing to the source entity data table is generated. The pilot detection command is triggered to collect the source primary key, and the collected source verification is encapsulated into a specific group identifier column vector through the primary key.

[0012] Preferably, generating the dimensionality-reduced subtree with a parameter list includes: Obtain the pairwise mapping network table of group entities generated daily by offline computation; Using the specific group identifier column vector as the input retrieval source, the primary key set of the corresponding target entity data table after mapping transformation is obtained by searching the pairwise mapping relationship network table of the group entities; The primary key set after the mapping transformation is converted into a hard-coded constant constraint condition that includes the matching syntax of the set. In the in-memory computing environment, the complex multi-table association branches caused by the cross-entity domain condition nodes are removed from the original nested tree structure, and the hard-coded constant constraints are directly concatenated to the removal position to complete the reshaping of the multidimensional topology into a flattened structure, and the dimensionality-reduced subtree is output.

[0013] Preferably, calculating the nesting depth parameter of each logical branch within the dimensionality-reduced subtree includes: A depth-first traversal scan is initiated starting from the root node of the reduced-dimensional subtree; A depth counter is assigned to each tree node. When crossing a level boundary, the depth counter is incremented by a fixed step value to obtain the nesting level depth parameter corresponding to each logical branch. Sort the logical branches, extract the shallowest logical branches as independent SQL fragments, and retain the deepest logical branches as dependent SQL fragments, including: The logical branches are reorganized using an ascending order sorting mechanism; Set a hierarchical splitting threshold parameter, and convert logical branches whose nesting depth parameter is lower than the hierarchical splitting threshold parameter into the preceding independent SQL fragments so as to trigger asynchronous operations first; The logical branches whose nesting depth parameter is greater than or equal to the level splitting critical parameter are converted into the post-dependent SQL fragments.

[0014] Preferably, the step of constructing a feature existence mapping table in memory using the returned initial computation data includes: The distributed computing engine receives the initial computation data returned after processing the preceding independent SQL fragment, the initial computation data including each verified primary key; Allocate a contiguous bit array space in the application server's memory; By calling various different hash mapping calculation functions, each verification is hash-encoded using the primary key to obtain multiple address coordinates; The storage bits corresponding to the multiple addressing coordinates in the continuous bit array space are adjusted to the active recognition state, and the bit array space adjusted to the active recognition state together constitute the feature existence mapping table.

[0015] Preferably, the step of using the feature existence mapping table to perform pre-run probe constraint appending processing on the post-dependency SQL fragment, and assembling it to obtain the final optimized SQL statement text, includes: At the beginning of the conditional filtering code segment of the post-dependent SQL fragment, insert a pre-read probe constraint function that calls the feature existence mapping table; The pre-read probe constraint function is used to force the distributed computing engine to first verify whether the identity primary key of the data row being read hits the storage bit in the active identification state before performing a full disk table scan. Discard invalid data rows that fail to meet the activation recognition status, and perform subsequent correlation and statistical actions on the data rows that pass the verification to combine them into the final optimized SQL statement text.

[0016] The present invention has achieved the following beneficial effects: This invention reduces the logical complexity of the underlying execution engine in handling multi-table join queries by performing association decomposition processing on cross-entity domain condition nodes contained within the group selection rules, generating a dimensionality-reduced subtree with a parameter list. Based on the calculated nesting depth parameter, this invention sorts each logical branch, extracting shallow-depth logical branches as pre-emptive independent fragments for asynchronous execution, and using the returned initial computation data to construct a feature existence mapping table in memory. This invention uses this feature existence mapping table to add pre-run probe constraints to subsequent dependent fragments, forcing the distributed computing engine to prioritize verifying the identity primary key hit status of data rows before performing a full disk table scan. This pre-interception mechanism filters invalid data rows that do not hit the feature status during the data reading stage, preventing redundant records from directly participating in subsequent computationally intensive multi-table join processes, reducing the data throughput cardinality during cluster computation, balancing the memory and processor load of distributed nodes, and improving the task flow efficiency and computation time when facing high-concurrency complex queries.

[0017] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.

[0018] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0019] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1This is a schematic diagram of the system module composition structure of the application cluster environment in an embodiment of the present invention; Figure 2 This is a flowchart of the preprocessing data loading and configuration flow steps of the base table in an embodiment of the present invention; Figure 3 This is a flowchart illustrating the main logic of the distributed group selection SQL automatic generation and calculation method based on multi-layer nested JSON parsing in this embodiment of the invention. Figure 4 This is a flowchart illustrating the specific steps involved in parsing the group selection rules in this embodiment of the invention. Figure 5 This is a flowchart illustrating the specific steps involved in converting rules into Structured Query Language (SQL) statement text in an embodiment of the present invention. Figure 6 This is a flowchart illustrating the specific steps involved in sending the Structured Query Language (SQL) statement text to the distributed engine for execution in this embodiment of the invention. Figure 7 This is a schematic diagram illustrating the configuration of multi-layered nested group selection rules for the front-end interactive interface in an embodiment of the present invention; Figure 8 This is a schematic diagram of the interface for group segmentation task estimation and execution request issuance in an embodiment of the present invention; Figure 9 This is a schematic diagram showing the group calculation details and status display after obtaining and outputting the individual member result set in an embodiment of the present invention. Detailed Implementation

[0020] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0021] This application discloses a method for automatically generating and calculating distributed group selection SQL based on multi-level nested JSON parsing. For example... Figure 1 As shown, the method is applied to a cluster environment comprising front-end interaction nodes, back-end business service nodes, distributed columnar data warehouse nodes, federated ad-hoc query nodes, and distributed computing engine nodes. The aforementioned nodes communicate via an internal local area network, using a transmission control protocol to establish connections and ensure reliable message delivery.

[0022] Before executing the core parsing and calculation process, the system first performs preprocessing data loading and configuration flow steps for the base table. Specifically, such as... Figure 2 As shown, the preprocessing process includes the following steps: Step S010: Use the distributed columnar data warehouse to read the vehicle individual profile base table, the electronic non-stop toll collection individual profile base table, and the highway travel application individual profile base table respectively.

[0023] The underlying distributed computing cluster of the system pre-initializes the data storage state and starts the background process for reading the distributed columnar storage data warehouse. On each physical storage node, the vehicle individual profile table is configured with data columns such as vehicle license plate string, vehicle color identification code, and vehicle axle load value; the electronic non-stop toll collection individual profile table is configured with data columns such as device network physical address, historical binding status bit, and toll settlement history timestamp; the highway travel application individual profile table is configured with data columns such as user registration unique identity code, front-end login frequency parameter, and behavioral preference feature vector. Each of the above data columns is stored independently in contiguous disk sectors. The columnar storage engine maintains independent file blocks for each field at the physical level and writes a metadata area containing minimum value, maximum value, and dictionary encoding mapping table in the header area of ​​the file block.

[0024] Step S020: Calculate the number of people covered by the multiple profile tables read by the statistical script according to the tag dimension, and update the total number of people covered by each tag dimension to the tag statistics result table.

[0025] The distributed task scheduler periodically sends preprocessing instructions to the computing node cluster according to a preset time-round scheduling algorithm. Upon receiving the preprocessing instructions, the computing nodes use a streaming method to load the base table data into pre-allocated data cache areas in random access memory. Within the cache area, an aggregation statistical algorithm based on hash mapping is executed for each tag dimension field. Specifically, the computing nodes establish a hash slot array in the process heap memory allocated by the operating system, read the data content byte stream within the fields, calculate the memory offset constant using an unencrypted hash mapping function, use the corresponding tag enumeration attribute value as the retrieval key, and call the hardware atomic instructions supported by the processor to perform concurrent accumulation operations on the count value variable. After each physical computing node completes the hash statistics of its local partition, it triggers a synchronous exchange operation based on network data packets, distributing the local statistical payload data to pre-designated global reduction computing nodes. The global reduction computing nodes perform merge and summation logical operations on the received data groups to obtain the total number of audience members covered by each tag dimension in the full dataset. Subsequently, the global reduction computing nodes update the total number of audience members parameter and write it into the tag statistics result table. The label statistics result table is stored in a key-value pair memory model within a distributed file system.

[0026] Step S030: Synchronously push the rule parameter enumeration values ​​recorded in the business configuration library with the tag statistics result table, and perform cross-node joint storage within the federated ad hoc query node to respond to front-end display requests.

[0027] The internal data synchronization service program connects to the relational configuration library storing business metadata via a connection socket. It extracts the Chinese name mapping dictionary of each tag rule required for the front-end page configuration items, as well as the array of optional parameter ranges. This data constitutes the rule parameter enumeration values. Next, the rule parameter enumeration values ​​are concatenated with the tag statistics result table in memory, encapsulated into a structured data message with a hypertext header, and sent as a data push stream to the federated ad-hoc query node cluster. Upon receiving the structured data message sequence, the federated ad-hoc query node cluster constructs an inverted index data structure for the key-value pairs and a bitmap filter in the dynamic random access memory of each worker node server. When the front-end business terminal interface sends a query request for the optional range and coverage of a specific business tag, the network interface of the federated ad-hoc query node intercepts the query request. By reading the inverted index parameters residing in the memory block, it returns the matching enumeration option array and the corresponding total number of audience members. This processing operation traps the computational load of the metadata statistics request in the memory cache layer, blocking the operation process of issuing scan commands to the underlying hard disk.

[0028] After completing the aforementioned preprocessing data loading and configuration flow steps, the application system enters the core stage of group rule parsing and structured query language text generation and execution. Specifically, for example... Figure 3 As shown, the main logic of the distributed group selection SQL automatic generation and calculation method based on multi-level nested JSON parsing includes the following steps: Step S100: Obtain the group selection rules carrying the set conditions. The group selection rules are organized in JSON format, which is a multi-level nested JavaScript object representation.

[0029] After the front-end computing system assembles the filtering criteria, it sends a hypertext request data packet carrying business logic parameters to the back-end service application layer node via the network transmission control protocol layer. The reverse proxy component of the back-end service application layer node extracts the main payload and reads the group selection rules carrying the set conditions. The group selection rules internally encompass a nested and interconnected top-level logical structure, a secondary group structure, a third-level condition structure, and a fourth-level parameter structure.

[0030] Specifically, the top-level logical structure records the logical AND or logical OR operators representing the globally joint judgment attributes. This top-level logical structure is located at the root node of the multi-layered nested JavaScript object representation JSON format parsing abstract tree. The logical AND or logical OR operators are configured as string variables identifying logical intersection or logical union operations. The secondary group structure includes a group type identifier and a subordinate condition array. The group type identifier is configured using a fixed-length numerical encoding format to uniquely point to a specific physical business table in the distributed storage layer. The subordinate condition array is represented in memory as a sequential set of object elements, used to store data filtering condition object entities directly applied to the currently involved group type. The third-level condition structure records the logical AND or logical OR operators controlling the combination state of the subordinate condition array, indicating the sub-intersection or sub-union Boolean operation relationship of the condition array elements within a locally defined logical scope. The fourth-level parameter structure is configured with specific comparison dimensions, table names, field names, field types, comparison operators, and comparison values. The comparison dimension corresponds to a specific label attribute parameter; the data table name indicates the physical identifier name of the target data table; the field type declares the memory data storage alignment structure format of the current operation column; the comparison operator includes constant symbols for equality matching and range matching of set or prefix strings; the comparison value is stored in the working memory area as a comparison scalar during conditional judgment execution.

[0031] Step S200: The group selection rules are parsed.

[0032] The backend service application layer components initialize a lexical scanner instance and an abstract syntax analysis engine instance in the process heap memory area allocated by the operating system. The lexical scanner instance, following the JavaScript object representation JSON syntax specification standard, identifies curly braces, square brackets, double quotes, colons, and commas as delimiters in the input sequence, segmenting the continuous plain text sequence into multiple streams of basic lexical token fragments containing independent data type definitions. The abstract syntax analysis engine instance receives these basic lexical token fragment streams, dynamically allocates structure data blocks in the memory address space based on hierarchical indentation and wrapping relationships, attaches physical address reference pointers, and instantiates an abstract syntax tree model reflecting conditional dependency topology.

[0033] like Figure 4 As shown, step S200 is specifically divided into the following execution logic steps to implement deep structure transformation: Step S210: Identify cross-entity domain condition nodes existing within the group selection rules, perform cross-entity domain association decomposition processing on the cross-entity domain condition nodes, and generate a dimensionality-reduced subtree with a parameter list.

[0034] After the abstract syntax tree model is constructed, the system compilation scheduling module allocates execution worker threads and starts the depth-first traversal scanning operation. It parses the multiple independent group type identifiers covered within the group selection rules. When it detects a situation where different independent group type identifiers are intertwined within a nested level, it marks the logical node that triggered the intertwining operation as the cross-entity domain condition node. Specifically, the depth-first traversal scanning operation is configured with a state stack array, and reads the set of all group type identifiers recorded within the subtree of each traversal node by following the node reference pointer. When the detection logic detects that under the control domain of a logical AND or OR operator node defined in the same third-level condition structure, the subordinate branch child node structure object is configured with mutually inconsistent group type identifiers, the program determines that different independent group type identifiers are intertwined. The system opens an abnormal node registration table in the memory control block, records the logical AND or logical OR node that triggers the cross-domain query operation to the memory abnormal handling table, and adds a specific cross-entity domain condition node identity tag to the physical node involved, while terminating the downlink structure merging action of the associated branch.

[0035] In an in-memory computing environment, the range of the source entity data table to which the cross-entity domain condition node is attached is extracted, and a preliminary probe instruction pointing to the source entity data table is generated. The compilation scheduling module initiates separate instruction stripping execution logic for the cross-entity domain condition node in the exception handling table. The program reads the judgment attribute data of all leaf nodes under the corresponding node, selects branches with shallow nesting levels and equality filtering restrictions as source query branches, and parses the data table name, field name, comparison operator, and comparison value configured within the source query branch. Using a standard relational query builder, the extracted parameter elements are assembled into a character control query string targeting a single physical data table, thereby generating the preliminary probe instruction. The preliminary probe instruction embeds a selection declaration parameter during construction, limiting the return of only specific primary key field content, excluding the retrieval of irrelevant attribute columns.

[0036] The system triggers the pilot probe command to collect source primary keys, and encapsulates the collected source verification into a specific group identifier column vector using the primary key. The system control module calls the underlying database network connection socket component to send the constructed pilot probe command to the underlying columnar data warehouse. The columnar data warehouse management process initiates the underlying physical partition index lookup and comparison operation, obtains the binary set of primary key row identifier records that meet the retrieval constraints from the storage medium, and transmits it back to the backend service application layer using the streaming control channel. The memory allocation module of the backend service application layer dynamically allocates a 32-bit numerical array contiguous memory structure. The receiving thread transfers the primary key data in the data stream to the numerical array contiguous memory structure one by one. After the receiving stream ends, the array sorting and deduplication function library function is triggered to perform redundant item removal on the memory block, forming the specific group identifier column vector stored in the linear address space.

[0037] The system retrieves a network table of pairwise mapping relationships between group entities generated daily through offline computation. A distributed data extraction scheduled job framework is deployed on the cluster system base. During the set offline computation window, the job framework reads historically generated vehicle registration details, road network equipment registration details, and user information relationship details. In the memory processing pool of the computing nodes, the framework executes an equi-join computation model, comparing the association fields between different base tables one by one, and extracting a mapping graph topology structure that stores primary and foreign key mapping relationships. This mapping graph topology structure stores a pairing sequence of specific primary keys from the vehicle base table and their corresponding specific primary keys from the user registration base table. The job framework stores the pairing sequences in a physical partition file with high-density Bloom filter index markers, thus forming the network table of pairwise mapping relationships between the group entities.

[0038] Using the specific group identifier column vector as the input retrieval source, the corresponding target entity data table's mapped and transformed primary key set is retrieved from the pairwise mapping relationship network table of the group entities. The application service module loads the mapping read routine, using the specific group identifier column vector resident in memory as the probe input parameter. The mapping read routine employs a batch fragmentation traversal transmission strategy, delivering the input parameter sequence to the dedicated retrieval engine module of the pairwise mapping relationship network table of group entities. The retrieval engine module performs memory-level hash matching operations to find the target transformation record paired with the input source and outputs the target transformation primary key data stream. After obtaining the returned target transformation primary key data stream, the application service module performs a second memory set operation within a local worker thread, filtering out identical return values ​​and retaining the remaining non-duplicate target element set state to form the mapped and transformed primary key set.

[0039] The mapped primary key set is converted into hard-coded constant constraints containing matching syntax. The system calls a memory-extended concatenation method library to convert each independent primary key in the mapped primary key set into a continuous text block separated by commas. Then, left and right parentheses are written to the beginning and end memory locations of the generated continuous text blocks, and constant character variables representing inclusion operations in Structured Query Language are added at the beginning. These are combined to generate static constraint fragments that apply to individual table queries, thus forming the hard-coded constant constraints.

[0040] In the in-memory computing environment, the complex multi-table association branches triggered by the cross-entity domain condition node are removed from the original nested tree structure. The hard-coded constant constraints are directly appended to the removal location, completing the reshaping of the multi-dimensional topology into a flattened structure and outputting the reduced-dimensional subtree. The compilation and parsing component accesses the previous original abstract nested tree structure object and locates the memory pointer marked with the identity of the cross-entity domain condition node based on the physical address. The program scheduling instruction forcibly unmaps and binds the downlink physical address pointer of the involved node pointing to the complex multi-table association query subtree, and overwrites it with a null constant state, implementing the branch removal instruction. Simultaneously, the compilation and parsing component allocates physical space for the newly created leaf node object in the process heap area and assigns a copy of the generated hard-coded constant constraint content text to the attribute slot of the newly created leaf node. Subsequently, the downlink physical address pointer of the cross-entity domain condition node is remapped and bound to the newly created leaf node unit. By rewriting pointers on the memory tree structure, the original tree-like branches containing dependencies are reshaped and transformed into a flattened decision model that only involves matching a single-level set of constants. The final output memory tree diagram object is confirmed as the dimensionality-reduced subtree.

[0041] Step S220: Calculate the nesting level depth parameter of each logical branch inside the dimensionality-reduced subtree.

[0042] A depth-first traversal scan is initiated starting from the root node of the reduced-dimensional subtree. A depth counter is assigned to each tree node, and the depth counter is incremented by a fixed step size when crossing level boundaries to obtain the nesting level depth parameter corresponding to each logical branch. Within the virtual memory area allocated by the application process, the backend computing control module configures an independent data operation stack space queue and a backtracking stack pointer. The control module retrieves the physical address from the root node object of the reduced-dimensional subtree and pushes it to the bottom of the stack. In the attribute slots within the class template defining the data structure of each tree node, a predefined unsigned short integer data area is used as the depth counter variable. The system initializes the depth counter of the root node structure instance to a constant value of zero, representing the starting level baseline. The backend computing control module initiates a continuously looping stack popping control system for obtaining node data. During the loop execution of the probe, if the popped current operation node contains a set of downward child node pointer references, the branch child node entity objects are accessed based on the child node pointers. During the process flow transition from a parent node to an associated child node, the backend computation control module extracts the depth counter value constant stored in the structure of the currently operating parent node and transmits it to the central arithmetic logic unit (ALU). The ALU performs addition logic on this value constant, with the accumulated operand being a pre-set fixed step size constant value factor.

[0043] Specifically, the fixed step size constant value factor is explicitly configured as a positive integer constant 1 in this embodiment. Its underlying physical and logical setting is based on the following: each crossing of the level boundary from the parent node to the associated child node in the abstract dimensionality reduction subtree represents an increase in the nested filtering depth of logical intersection or union in the distributed multi-table joint query plan by the underlying computing engine; through this atomically incrementing constant 1, the system accurately transforms the abstract logical nested topology, which is difficult to compare directly, into a numerical physical depth scalar (i.e., the nested level depth parameter), thus serving as an unambiguous scaling parameter for subsequently evaluating the computational cost of the underlying scanning of logical branches and performing asynchronous SQL fragment decomposition.

[0044] The arithmetic logic unit overwrites and stores the summation result into the depth counter data area of ​​the corresponding branch child node instance. Then, it pushes the branch child node instance to the top of the stack, awaiting the next iteration. This continues until all tree-like node elements have been exhausted. At this point, the leaf condition nodes at the bottom of the reduced-dimensional subtree, as well as the depth counter data areas of the subquery node units containing independent aggregation operation boundaries, all hold quantized calculation values ​​accumulated according to the hierarchy. The numerical variables stored within specific operation node instances are extracted and combined to form the nested hierarchy depth parameter index set corresponding to each logical branch.

[0045] Step S230: Sort the logical branches according to the nesting depth parameter, extract the logical branches in the shallow depth as the preceding independent SQL fragments, and retain the logical branches in the deep depth as the following dependent SQL fragments.

[0046] An ascending order mechanism is used to reorganize the various logical branches. The system creates a one-dimensional paired sequence matrix structure in memory. Each matrix element has a reference field storing the memory address of a specific logical branch object instance, and a numerical field storing the corresponding determined nesting depth parameter. The system's main thread loads general comparison and sorting library functions, sets the numerical field as the comparison scale, and executes a quicksort algorithm based on partition swapping to control the flow. The sorting algorithm establishes an ascending order by iteratively comparing numerical values ​​and swapping array elements, thus reorganizing the data hierarchy referenced by each logical branch object.

[0047] A hierarchical segmentation threshold parameter is set. Logical branches with a nesting depth parameter lower than the threshold parameter are converted into the preceding independent SQL fragments to prioritize asynchronous computation. Logical branches with a nesting depth parameter greater than or equal to the threshold parameter are converted into the following dependent SQL fragments. The backend resource coordination service collects the available random access memory (RAM) parameters of the currently deployed server's operating system and the task queue length parameters of the underlying execution engine. These parameters are then substituted into the internally set resource calculation formula, and the constant result (integer bits) is used to confirm the hierarchical segmentation threshold parameter.

[0048] In detail, to avoid a surge in scheduling overhead due to the introduction of complex nonlinear calculation models, the internal resource calculation formula is set as a linear mathematical derivation formula that takes into account both memory sufficiency and queuing congestion status: ; in, The calculated output is the critical parameter for hierarchical segmentation; This is the minimum disassembly level limit for mandatory protection of the system, and its value is a constant 1. The baseline segmentation level constant is preset for the system, with an initial value of 3; The collected available random access memory (RAM) margin parameters (unit: GB). The total physical memory capacity of the execution node is a constant; The task queue length parameter for the underlying execution engine (i.e., the total number of concurrent queries currently pending in the system). The maximum queuing capacity threshold allowed by the system (pre-configured as a constant 500); This is the positive gain coefficient for memory margin, and its value range is configured as follows: Floating-point numbers within; The reverse penalty coefficient for queuing congestion is configured with a value range of [value range missing]. Floating-point numbers within; This represents the floor function. This indicates taking the maximum of the two options.

[0049] Based on this specific judgment logic: when the system hardware has sufficient available memory and the concurrent queued tasks are sparse, the formula will derive a larger critical parameter, thereby allowing more deep logic branches to be assigned to the front independent SQL fragments for asynchronous acceleration priority operation; conversely, when resources are congested and scarce, the formula can quickly reduce the critical parameter to shrink the asynchronous extraction range, effectively preventing memory overflow crashes and computing node service suspensions caused by excessive partitioning from the operating system scheduling level.

[0050] The system reads the memory pointer sequentially along the data sequence after it has been rearranged in ascending order, and checks the numerical field parameters. It extracts logical branch node combination entity objects whose nesting depth parameter value is lower than the level segmentation threshold parameter. The system calls the statement dialect mapping conversion package, following the syntax construction rules of the target relational database, to sequentially concatenate the parameters of the extracted entity objects and encapsulate them into the aforementioned independent SQL fragment involving independent form relationship judgments. The system starts an isolated sending thread within the multi-threaded concurrency management component, allocates an underlying network sending handle, and sends the packaged independent SQL fragment to the distributed cluster's subordinate computing gateway, entering the priority trigger asynchronous operation phase. For the remaining logical branch entity set read at the end of the sequence and determined to have a nesting depth parameter greater than or equal to the level segmentation threshold parameter, the system calls the mapping conversion package to generate a composite string text block code with table join and subquery attributes, converting it into the subsequent dependent SQL fragment. The system transfers the character buffer block containing the subsequent dependent SQL fragment into a delayed commit stack container created by the runtime context to perform a suspended state setting.

[0051] After completing the deep dimensionality reduction and control flow stripping of the above rules, the system executes step S300, converting the parsed rules into Structured Query Language (SQL) statement text. For example... Figure 5 As shown, step S300 is divided into the following underlying execution stages: Step S310: Control the distributed computing engine to execute the preceding independent SQL fragment asynchronously first, and use the returned initial computation data to build a feature existence mapping table in memory.

[0052] The distributed computing engine receives the initial computation data returned after processing the preceding independent SQL fragment. This initial computation data includes each validated primary key. Since the preceding independent SQL fragment is issued independently on shallow nodes, the underlying execution engine schedules the disk controller to initiate an index page filtering scan and extracts the primary key content fields of data rows that meet the basic constraints. The extracted byte payload stream is pushed to the application server's waiting pool via the internal exchange interface and the Transmission Control Protocol channel. The application server's receiving socket process captures the arriving network data packets, performs deserialization processing, and parses and extracts a set of string arrays containing the content of each validated primary key.

[0053] A contiguous bit array space is allocated in the application server's memory. The backend storage allocation scheduler calculates the critical memory capacity required to construct the Bloom filter state model based on the total number of entries in the returned array sequence and the collision tolerance formula.

[0054] Specifically, the step of deriving the critical value of the memory capacity required to construct the Bloom filter state model by combining the collision tolerance calculation formula covers the following precise calculation process: extracting the total number of entries contained in the returned array sequence as the sample base. The system's false positive target collision tolerance rate is set to 10%. To balance high-precision data interception accuracy with minimal memory usage of computing nodes, this embodiment will... The value range is limited to the interval between 0.001 and 0.01; the base number is... Collision tolerance with target Substitute into the storage space calculation formula based on error rate constraints This allows for the precise determination of the total number of bits required to allocate the contiguous bit array space. ; (in the formula) (This represents a rounding up operation). The system then further calculates the total number of bits obtained. Divide by the 8-bit boundary and perform rounding up to convert to the total byte length, and finally confirm it as the critical value of memory capacity for requesting the allocation of physical page frame boundaries from the operating system addressing layer.

[0055] The system kernel sends a low-level request to the host system kernel to allocate a contiguous range of physical page frame addresses of the corresponding length. Upon receiving the response, it locks the address region and forcibly overwrites and initializes all binary storage bits within that range to a constant value of 0, representing an inactive or inactive state, thus completing the process of allocating a contiguous bit array space.

[0056] The application server invokes various hash mapping calculation functions to perform hash encoding on each verification using its primary key, yielding multiple addressing coordinates. The server then calls concurrent processing threads to process the primary key array sequence. These concurrent processing threads run various heterogeneous hash conversion rules from the component library to process the primary key content of each verification, outputting the corresponding number of 32-bit unsigned integer hash intermediate values.

[0057] Specifically, the aforementioned process involves the corresponding number (set as a positive integer parameter). The calculation method is based on the conversion formula. The derivation is performed automatically and synchronously. Based on the obtained parameters... The system instantiation calls are equal in number (i.e.) The system uses different unencrypted fast hash algorithms (preferably extracted from internal variant libraries of Murmur Hash 3, City Hash, or FNV-1a algorithms) as the hash conversion rules for the various heterogeneous hash algorithms. Each algorithm injects distinct large prime numbers generated by a hardware clock as seeds for random discrete computation masks. By combining these various heterogeneous hash operators, the hash avalanche effect of the underlying bit mapping can be fully stimulated, effectively dispersing primary keys containing regular text clusters (such as license plate number segments with fixed location prefixes, or MAC address manufacturer prefix segments of devices of the same brand) that are prone to appear in the underlying business profile data. This avoids the accumulation and collision penetration of continuous bit array spatial hash points caused by similar primary key segments at the physical addressing level.

[0058] Furthermore, the arithmetic logic unit performs modulo operations on the acquired intermediate hash values ​​and the total addressing span of the allocated bit array space. The output remainders determine the memory offset information of the corresponding primary key's addressing coordinates within the memory space.

[0059] The storage bits corresponding to the multiple addressing coordinates in the continuous bit array space are adjusted to the active identification state, and the bit array spaces adjusted to the active identification state together constitute the feature existence mapping table. The system addressing instruction modifies the memory access pointer to jump to a specific physical address endpoint based on the calculated displacement vector of the multiple addressing coordinates. The binary bit-level mask control operation logic is initiated to rewrite and invert the level data stored in the specified physical storage bit. The flag variable is changed from the inactive state of digit 0 to the active identification state of digit 1. After the overwrite and reversal operation of all data in the primary key array is completed, the previously reset and cleared continuous page frame physical space is transformed into a control area containing the active bit data, together constituting the entity architecture of the feature existence mapping table.

[0060] Step S320: Use the feature existence mapping table to perform pre-run probe constraint appending processing on the post-dependent SQL fragment, and assemble it to obtain the final optimized SQL statement text.

[0061] At the beginning of the conditional filtering code segment of the post-dependency SQL fragment, a pre-read probe constraint function calling the feature existence mapping table is inserted. The application server's packaging component initiates serialization encoding, packaging the constructed feature existence mapping table and its associated addressing boundary metadata into a binary read-only shared variable payload. This payload is deployed and sent down through the internal management gateway node to each worker node of the distributed computing engine base participating in the computation, and instructs it to reside in the resident memory isolation area of ​​the underlying process. The application server's text compiler extracts the preserved post-dependency SQL fragment from the deferred commit stack container. The compiler parses the code logic structure within the fragment body, locating and tracing to the boundary of the header code declaration. The system pre-sets a fixed identifier control string pointing to the underlying user-defined function, which is registered and configured as a pre-read probe constraint function. Its underlying encapsulation logic includes accepting primary key string input, hash conversion, and verifying the state of specific memory bits. The text compiler writes the identifier control string and the passed primary key reference field name to the header boundary of the post-dependency SQL fragment through a string insertion operation. After completing the pre-run probe constraint appending process, the assembled long string data block becomes the final optimized SQL statement text.

[0062] The prefetching probe constraint function forces the distributed computing engine to prioritize verifying whether the identity primary key of the data row being read matches a storage bit in an active identification state before performing a full disk table scan. Finally, the optimized SQL statement text is sent to the storage scan layer after the underlying distributed computing engine generates an execution tree. The storage scan layer reads data row tuples from the underlying disk sectors into the system cache. Following the execution tree control flow, the processor prioritizes filling the identity primary key content of the currently buffered data row into the execution environment of the called prefetching probe constraint function. The prefetching probe constraint function internally calculates the storage bit and checks the identification status of the memory-resident mapping table.

[0063] Invalid data rows that fail to meet the activation identification state are discarded. Subsequent correlation and statistical actions are performed on the data rows that pass the verification, thereby assembling the final optimized SQL statement text. If a storage location is found to be inactive during the detection and verification process, the logic judgment engine outputs a logical negation instruction scalar to the caller. Upon receiving the negation instruction, the distributed computing engine terminates the input / output loading requests for reading other columns of data in that row record, erases the remaining data in the buffer, and discards the invalid data rows that fail to meet the activation identification state. Conversely, when all array storage locations meet the activation identification state, the data row object verification is allowed. The allowed data row body is transferred to the subsequent correlation and statistical action pipeline, which contains high physical resource consumption such as join and aggregation processing. After header interception and filtering, the underlying data scale is effectively reduced, completing the design of this query instruction control flow. In system execution step S330, the final optimized SQL statement text is used as the structured query language SQL statement text. At this point, only the compilation and assembly of the SQL instruction is completed, and it resides on the application as the final executable text, awaiting unified scheduling and execution in subsequent stages.

[0064] After the structured query text is constructed, the system proceeds to step S400, whereby the structured query language (SQL) statement text is sent to the distributed computing engine for execution, obtaining and outputting the individual member result set. For example... Figure 6 As shown, this step includes the following underlying scheduling and computation operations: Step S410: Sort the group grouping requests generated based on the group selection rules in the execution state according to multiple grouping dependencies.

[0065] The application server process constructs a directed acyclic graph (DAG) data topology object in the heap memory area. For each group request in the received execution state, corresponding entity vertex data units are initialized and created in the topology. The scheduling engine scans the internal rules of each group request to see if there are relational instructions that reference the output results of preceding group requests. If a dependency instruction is detected, a directed edge network physical pointer is established. The cursor at the start of the directed edge is attached to the preceding group request entity vertex that provides the data computation source, and the cursor at the end of the directed edge points to the derived group request entity vertex that depends on it for computation. After completing all directed edge connections, the computation scheduler uses graph theory algorithms to calculate the in-degree parameter value of the entity vertex data units. The in-degree parameter value is the cumulative number of upstream directed edges pointing to a specific vertex. The scheduler establishes a ready state buffer queue in shared memory and pushes entity vertices with an in-degree parameter value that is always equal to the constant 0 into the ready state buffer queue. Then, a polling daemon process is started, sequentially extracting executable tasks from the queue and sending them to the network communication gateway. After each task vertex is stripped and dispatched, the daemon searches for downstream derived group entity vertices along the directed edges. Atomic subtraction arithmetic instructions are performed on the in-degree parameter of the derived group entity vertices. When an in-degree parameter decrements to zero, the affected derived group entity vertices are synchronously pushed into the ready state buffer queue. The aforementioned method based on in-degree graph theory establishes a sequentially ordered output queue for task objects, avoiding deadlock in computation execution.

[0066] Step S420: Control the distributed computing engine to read multiple subquery instructions contained in the final optimized SQL statement text at the operation layer.

[0067] The physical execution compiler deployed at distributed nodes receives network request packets, performs parsing and splitting operations according to lexical nodes, and maps the identified and decoupled independent logical extraction statement fragments into scanning and reading task operators that are executed in parallel in different storage communication channels.

[0068] Step S430: Use a join query statement that retains duplicate records to concatenate the multiple subquery instructions into the set view.

[0069] To reorganize data sets extracted from different sources, the instruction orchestration mechanism specifically configures the invocation of a union statement. The producer worker threads allocated by the computation layer initiate direct access to the corresponding physical sectors based on task operators. Once a unidirectional match retrieves an array of individual primary identifiers that meet the criteria, the producer thread forcibly skips the deduplication cycle performed locally within the node, and pushes it in batches to a pre-allocated virtual channel buffer slot in memory in its original structural state. By omitting the deduplication operation, the compliant identifier records output by each sub-query source are linearly concatenated and superimposed. The data combination set stored in the memory slots constitutes the logically defined union view.

[0070] Step S440: Nest a grouping aggregation instruction for the individual master identifier outside the collection view.

[0071] As the concatenated and aggregated identifier records flood into the memory region, the consumer worker processing thread invokes the distributed grouping and aggregation memory operation control flow. To avoid physical segment overflow caused by a simple hash table scheme when facing a large cardinality, the engine kernel enables a two-phase routing cardinality evaluation algorithm module. In the initial stage of concurrent aggregation, the working system allocates a small number of probe memory panes to calculate the distribution characteristic density of the merged primary key data stream, identify data records whose frequency exceeds the conventional preset standard line, and define them as hot hash keys.

[0072] Specifically, regarding the hotspot data identification process, the steps of calculating and merging the primary key data stream distribution characteristic density and determining the conventional preset standard line in this technical solution are explicitly constrained to a lightweight detection mechanism based on a sliding ratio threshold: during the initialization phase, the system is configured to extract a dataset containing a constant number of detection record rows (e.g., setting a data sampling window limit). A streaming micro-sliding detection window (for line records) is used; with the help of a memory pipeline counter residing within this window, the local occurrence count parameter of each independent primary key flowing in is captured in real time. Calculate the frequency proportion feature density corresponding to each primary key. The system hardcodes the conventional preset standard line as a precise proportional judgment constant threshold (this constant is set within the range of 2% to 5%); when the engine's comparator verifies the frequency proportion feature density of a specific hash key value... When the value is greater than or equal to the default standard line ratio constant threshold, an anomaly label is directly thrown, the data record is intercepted, and it is labeled as the hotspot hash key. This window ratio extraction judgment rule effectively eliminates the variance interference caused by the fluctuation of the population base in the selection calculation, ensuring that high-frequency data records are robustly stripped to a dedicated lock-free counting register before causing congestion on the memory routing bus.

[0073] For defined hotspot hash keys, the system allocates local dedicated memory addresses as register units for the counter accumulator. Lock-free incrementing calculations are performed using hardware mutex instructions. For scattered, non-hotspot hash keys that do not meet the criteria, a global addressing allocation strategy is invoked. The hash calculation mapping routes the keys to the parallel hash table slots for key-value comparison and summation calculations. This hash data isolation mechanism eliminates the blocking caused by concurrent worker threads competing for the memory bus.

[0074] Step S450: Append a quantity comparison interception instruction to the end of the group aggregation instruction to perform intersection and union check calculation, and finally output the individual member result set and store the individual member result set in the offline relationship table. At the same time, trigger the group subsequent statistics and accelerated synchronization process for the calculation result.

[0075] The underlying environment binds the quantity comparison interception trigger code to the end of the hash addressing slot refresh write-back link. The system monitors the aggregate accumulator register status in real time. When a refresh carry occurs in the addressing slot or local accumulator, the trigger code immediately captures the latest accumulator status parameter. The central processing unit's arithmetic logic core compares the captured latest accumulator status parameter with the baseline limit threshold set and hard-coded during the parsing and compilation phase in a comparison instruction cycle. For the intersection branch established by logical AND in the group rules, it checks whether the accumulator status parameter is equal to the total number of sub-branches participating in the comparison and merging, configured constant. When the result is equal, it means that the primary identifier meets the intersection filtering criteria, the system marks the identifier as verified and pushes the memory address cursor into the write-back data collection queue. For the union branch established by logical OR, it determines whether the accumulator status parameter is greater than the baseline constant of 0, representing no match. Once a result greater than 0 is detected, it indicates that the target constraint rule is met, a verification pass mark is made, and the data is pushed into the output queue. Data entries that fail to meet the conditions are centrally cleared and released by the memory reclamation process. The remaining intercepted subject identifier data sequence is transformed into the final output set of individual member results.

[0076] Regarding the operational details of storing the set of individual member results into an offline relational table: the persistent write component adheres to the distributed block-based columnar storage layout specification, establishing file operation control handles for local and remote associated storage devices. The output set of member sequence records is segmented by a preset buffer control component according to a constant byte physical block size limit. During the serialization encoding of the generated disk file, the built-in engine detects the cardinality parameter of the data fields.

[0077] Specifically, the underlying rules and triggering logic configuration for the cardinality parameter of the probe data field are as follows: Within the currently processed fixed physical buffer block (e.g., a storage boundary set to 64 megabytes), the persistent write component calls the hash calculation sub-compute to calculate the total number of non-repeating independent element types contained in the target data field. This number is then divided by the total number of record rows covered by the buffer block to obtain a floating-point cardinality ratio parameter between 0 and 1. The system sets a constant of 0.05 (i.e., a 5% ratio) as the critical threshold for triggering the dictionary compression strategy in its underlying configuration. When the cardinality ratio parameter is detected to be strictly less than 0.05, the physical engine determines that the data field exhibits a highly regular and redundant distribution characteristic, and directly classifies it as the single-enumerated feature attribute data column, safely triggering the hash mapping dictionary table conversion. This clearly quantified cardinality threshold evaluation mechanism prevents the mis-sentenced high-dispersion business primary keys to the dictionary compressor, thereby avoiding the system crash risk of Java Virtual Machine (JVM) heap memory overflow (OOM) due to a surge in the size of the memory dictionary table.

[0078] For single-enumeration feature attribute data columns, the engine generates a localized hash mapping dictionary table, converting string text into unsigned short integer index identifiers. For timestamp long integer numerical variables with incremental characteristics, the engine loads a bit differential encoding mechanism, with the first record serving as the baseline, and subsequent records only writing the difference parameter obtained by subtracting the baseline. After dual encoding and compression, the physical payload block is appended with an assembled parity check bit array. Subsequently, the file physical unit is concurrently copied and written to multiple non-volatile server nodes deployed in different rack segments via network protocol. The file physical unit resides in the outer track structure area of ​​the disk, with a nested metadata header offset index block attached to the end of the segment, thereby completing the storage of the aforementioned individual member result set in the offline relation table.

[0079] After ensuring the secure delivery of the physical calculation output to disk, the system executes subsequent status synchronization and system query integration workflows. This includes the following sub-steps: Step S510 involves performing specific statistics on the group dimension based on the calculation results of step S450 (i.e., the individual set objects calculated by each rule). This yields the coverage parameters for each audience group and the latest calculation status update time parameter, replacing real-time querying to accelerate front-end query response. A status monitoring daemon is configured in the backend application system. The daemon monitors directory changes on the distributed file system name nodes mounted by the system. When it captures the confirmation flag text file written by an offline storage transaction, it activates the distributed statistics routine script. This script bypasses the conventional read mechanism, locates the end of a specific data physical file, and extracts the nested mounted offset index metadata block structure data. It reads the row count accounting scalar variable from the metadata block, which is dedicated to counting the number of valid records in that block. The statistics routine schedules concurrent coroutines to collect the accounting scalars from each data shard and performs lock-free integer addition aggregation operations. The cumulative sum is then confirmed as the coverage parameter corresponding to the business rule. Within the same period, the statistical routine reads the most recently modified clock attribute parameters from the operating system kernel file descriptor via system call instructions, and uses a formatting library to convert and finalize them into the latest computational state update time parameter that can be understood and read externally. This design avoids the time-consuming long-range re-retrieval and reading process at the table level.

[0080] Step S520: Synchronize the set of individual member results generated in S450, the number of people covered, and the latest calculation status update time parameter to the federated ad-hoc query node.

[0081] The cluster console starts an asynchronous communication pipeline microservice. The microservice instance allocates a large-capacity relay exchange buffer in the host memory. It extracts and re-decodes the individual member result sets and accompanying statistical parameters from the file system and loads them into the buffer. The microservice calls a two-dimensional array matrix row-column converter component to transform continuous long column segments suitable for offline batch processing into a data row format sequence suitable for fast online query retrieval. After format conversion, the microservice sends a network handshake write request to the ad-hoc federated query database. Upon successful verification, a direct transmission pipeline payload stream is initiated. The operating system kernel takes over the physical network interface card (NIC) device control mode, instructing the controller to skip application-layer replication and spontaneously read the specified blocks from the memory buffer. The byte sequence is packaged into the physical NIC device's transmit queue register and continuously delivered to the ad-hoc federated query database's backend storage node.

[0082] Step S530: The latest snapshot data of the business primary key mapping is retained in the federated ad hoc query node, which enables the application server to directly initiate multi-table joint cross-database fast retrieval, thereby achieving extremely accelerated querying of massive cluster results through the federated ad hoc node.

[0083] Within the ad-hoc federated query database that receives and loads synchronized data, an append-only log append-write mechanism is used for multi-version concurrent transaction control algorithm architecture isolation. Upon receiving formatted data streams, the underlying storage manager initiates segmented append-write operations to the physical extension file from an empty data address pool. After the file entity is written, the transaction management module issues a scalar value containing a globally unidirectional unsigned incrementing pattern as a version number to the newly added data segment. The system core redirects and overwrites the memory pointers mapped to the business key-value queries, pointing the address target to the starting address of the physical block bound to the highest fixed-quota version number scalar. Newly received records are naturally stored as the latest snapshot data, which is read with the highest priority. For historical data entities with expired or incomplete version number tags, the merge and reclaim thread initiates a forced erase and clear command during off-peak periods to archive them. Based on this physical storage generational replacement mechanism, when an application requests a cross-source query, the query read operation and the background disk write operation can be carried out in parallel, avoiding access lock queuing issues.

[0084] When the application server issues a cross-table join analysis retrieval request involving heterogeneous physical distribution model form data, the cross-source operation pushdown optimizer within the ad hoc federated query engine takes over the instruction payload. The cross-source operation pushdown optimizer analyzes the physical location of the join comparison forms. For the latest snapshot data residing within the memory processing module, the system constructs a workflow directly reaching the memory binary search tree; for the instruction portion that backtracks to query low-speed archived forms on the hard drive, the system reconstructs the lower bound boundary of the request and pushes the instruction dispatch to the underlying physical hard drive scan control layer operator processing. The two retrieval task data streams are received and processed in the hash table calculation and exchange area reserved in the engine coordination center. After performing union and intersection matching verification, standard messages are generated and combined into an entity return sequence, which is finally presented to the front-end interface.

[0085] To ensure data consistency under network fluctuations and node failures in the computing system, a checkpoint-based resume recovery mechanism is deployed in the architecture. During the distributed engine's computation execution cycle, the task supervisor module divides parallel operation tasks into different levels according to the execution tree slices. After intermediate results are output from the execution segment, the main control module collects the temporary physical configuration path of the output results and writes it along with the task status constant values ​​into a configuration service cluster with high availability and consistency deployment specifications, establishing a task execution checkpoint snapshot. If a momentary interruption occurs in the hardware network communication components and the main control module does not receive an acknowledgment reply packet within the response time limit, it determines that the computing execution node connection has weakened and been interrupted.

[0086] In this distributed physical connection status determination logic, the response timeout period is not a broad, manually estimated waiting time, but a calculated boundary threshold strictly derived from the underlying network probe mechanism: the underlying protocol stack of the system's main control module is pre-configured to cyclically send probe heartbeat packets to the worker nodes at a fixed constant time interval (e.g., preferably 3000 milliseconds), and a network hardware jitter packet loss tolerance threshold constant is set (configured for 3 consecutive packet losses); the boundary limit of the response timeout period is explicitly derived as: the fixed constant time interval multiplied by the packet loss tolerance threshold constant, and finally superimposed with the estimated maximum one-way propagation delay constant of the hardware physical link (set to 200 milliseconds). Thus, a disconnection timeout boundary of 9200 milliseconds is statically calculated in the compilation configuration. Only when the CPU hardware tick clock on the main control node's motherboard detects that no acknowledgment data has been received after exceeding this exact boundary threshold is a connection interruption judgment issued and a task reorganization instruction initiated. This fundamentally shields the process from forced termination and invalid restarts caused by microsecond-level network physical jitter.

[0087] The main control module intervenes to configure the service cluster to load the most recent snapshot volume of the failed task record. The system initiates a search to retrieve the data referenced by the snapshot volume address pointer, inputs the intermediate buffer results to the new working node, and issues an instruction to load and restart the remaining computational logic flow from the physical breakpoint, thus blocking the re-tracing overhead caused by the transmission of a single task failure.

[0088] The system employs a majority consensus voting procedure to prevent data inconsistency anomalies caused by network splits. The core scheduler and its subordinate worker nodes undergo election arbitration. If a network switch outage causes the cluster to split into disconnected subgroups, the subgroup holding more than half the votes elects a control node to take over the computation and distribution queue. Isolated nodes that become the minority, upon losing heartbeat responses, proactively suspend all write request interfaces to the underlying disk. Consensus voting prevents multiple worker nodes from concurrently overwriting the same physical partition in the relational table.

[0089] The underlying system is configured with a garbage collection intervention control program. The virtual machine's memory garbage collector tracks and records the physical parameters of the old generation memory space saturation in the background. When the saturation physical parameters cross a set threshold constant and the system is in an asynchronous transmission waiting idle period, the resource monitoring program actively sends a garbage compaction trigger request beacon to the kernel.

[0090] To further clarify, the triggering conditions here all have strictly quantified physical monitoring indicators: the saturation physical parameter is specifically defined as the percentage floating-point value between the number of physical bytes currently occupied by long-lifetime large objects (such as resident memory feature existence mapping tables) in the virtual machine's old generation memory region and the upper limit of the total available physical byte capacity pre-allocated by the operating system; the set threshold constant is hard-coded and fixed as an 80% capacity warning high-water level constant in this embodiment. Meanwhile, to prevent global application pauses (Stop-The-World) caused by garbage collection actions from cutting off ongoing high-concurrency distributed disk communication, the asynchronous transmission waiting idle period is explicitly quantified as follows: the resource monitoring program, through kernel calls, continuously probes the number of backlogged network packets in the downlink receive hardware queue (Rx queue) and uplink transmit hardware queue (Tx queue) of the physical network card of the computing node. Within a continuous 50-millisecond sliding operation monitoring window, the number of packets in the above queues remains at a system physical idle state of zero. Only when both the memory warning and network card physical idle preconditions are met is the trigger request beacon approved for issuance.

[0091] The processor scheduler uses a mark-and-sweep algorithm to reclaim the byte space units occupied by lost or detached objects in parallel, and performs memory page frame physical address regularization and migration operations on scattered surviving objects, supporting and ensuring the execution of large-capacity hash allocation requests.

[0092] The business application gateway deploys a high-concurrency access rate limiting and interception component. A token bucket algorithm is configured at the port where the front-end business system extracts, reads, and updates statistical values. Tokens are generated and accumulated in the cache stack at a constant system clock frequency. When a data packet containing query details arrives, it claims a token and is allowed to cross the security layer to the back-end query engine for execution. If a short-term surge in traffic depletes the token pool, the rate limiting component immediately outputs an interception code response and returns a snapshot estimate of the degradation parameters in the system's degradation cache.

[0093] The snapshot estimation degradation value parameter returned in the system degradation cache is not a predicted value derived from a dynamic model, but has a definite physical storage extraction mapping path: after the control and interception component intercepts the blocking instruction caused by concurrent overload, it directly cuts off the complex downlink execution plan involving cross-table scans, extracts the business query hash primary key contained in the trigger rate limiting message header, and reverses the address in the shared memory inverted index structure of the operating system; it directly extracts the true value of the historical coverage number parameter of the same group that was statistically completed and solidified on disk by the offline calculation batch of the previous natural day (i.e., T-1 calculation cycle) in the previous step S510. The control and interception component directly copies the historical static true value and outputs it as the snapshot estimation degradation value parameter of the current blocking request, and adds the Boolean flag bit of the historical non-real-time degradation protection data to the response header. This explicit engineering design only calls the memory address pointer addressing mapping action with a time complexity of O(1), and uses the tiny second-level data timeliness error as the exchange cost to build a low-level physical anti-overload security protection layer that eliminates disk seek load.

[0094] The control and interception mechanism ensures that the underlying storage array queues are protected from system overload.

[0095] For the low-level verification calculation steps that utilize probe constraint appending processing, the system invokes hardware-level vectorized filtering operators to assist in instruction execution. With the support of the low-level processor register stream hardware single-instruction multiple-data parallel extended instruction set, the engine skips the line-by-line scalar decomposition single-step operation mode during operator reading. The vectorized data reading component loads column tuple contents in batches from the system memory columnar format storage area into the CPU's multi-level cache static storage area. Relying on the high-bandwidth hardware register array structure, the computation core can concurrently compare and test the hit status of multiple master identifier hash calculation features against the physical bits of the feature existence mapping table within a single computation clock cycle. Through this compact memory alignment model and coordinated scheduling, the instruction-data processing pipeline reduces the main memory access latency bottleneck.

[0096] During the input parsing engine's handling of structured sequence payloads, the system security module initiates syntax format verification and loop blocking detection to prevent abnormal damage to the input structure. When the parser encounters non-standard character variable permutations that do not conform to the defined set, the control endpoint forcibly throws an interrupt exception structure object. The program control flow immediately halts, and the analysis logic calculation operation stack extends downwards. The system tracing and capture module takes over the intervention process, collecting current micro-environment operating parameters, physical stack call flow, and the original byte stream of the triggering blocking code. It formats, encapsulates, and backs up this data in a background disk audit log text file, and sends out the de-identified abstract formatted diagnostic code via network response feedback. This isolation action prevents unauthorized probes, rewrites, and unauthorized tampering of the underlying files in the storage area by illegal assembly statements.

[0097] Combined with actual application front-end interaction scenarios, such as Figure 7 The diagram illustrates the configuration of multi-layered nested group selection rules in the front-end interactive interface of this invention. In the business operation environment, the application system's front-end view component provides users with a visual condition orchestration panel. As shown, operators can dynamically add multiple cross-entity domain condition nodes such as "vehicle information satisfied," "electronic toll collection (ETC) user information satisfied," and "highway travel application (highway pass) user information satisfied." The front-end page component listens to the user's dropdown selection and click switching actions of logical associations (such as "AND" and "OR"), and constructs a multi-dimensional tree-like object model in real time in the browser's memory. After configuration, the front-end component automatically serializes this complex object model, containing cross-entity node associations and nested levels, into the multi-layered nested JavaScript object representation JSON format group selection rules and submits it to the upstream gateway, thereby providing a precise data input source for the underlying cross-entity domain association decomposition processing and dimensionality reduction subtree generation.

[0098] Furthermore, such as Figure 8 The diagram shows the interface for group segmentation task estimation and execution request issuance in an embodiment of the present invention. After assembling the group selection rules, the system provides a pre-emptive probe estimation mechanism. When the "estimate" probe command in the diagram is triggered, the application server first intercepts the shallower logical branch nodes in the rules and issues a pilot probe command. By initiating a local scan of the columnar data warehouse, it quickly estimates the scale of hit records under the current composite conditions. This pre-emptive estimation action not only provides business personnel with data feedback on the scale of the selected group, but also the initial computation data generated by its asynchronous request can be used to construct the aforementioned feature existence mapping table (Bloom filter state model) in memory, laying the foundation for subsequent full-scale probe constraint interception. After confirming that the estimation result is correct, the grouping request is encapsulated and submitted to a distributed ready state buffer queue, waiting for the core computing engine to schedule and execute it.

[0099] like Figure 9The diagram illustrates the group computation details and status display after obtaining and outputting the individual member result set in this embodiment of the invention. After the distributed computing engine completes the full multi-table join, intersection and union verification calculations, and successfully stores the individual member result set in the offline relation table, the system background application extracts the underlying metadata through a monitoring daemon. As shown in the details panel, the front-end system intuitively presents the "group size" (i.e., the number of people covered), "computation baseline time" (i.e., the latest computation status update time parameter), and current running status of the involved group by calling the latest snapshot data in the ad-hoc federated query database. This interface not only provides visual feedback on the final result of the entire closed-loop distributed computing process but also demonstrates that the target audience group, calculated by reducing the dimensionality of multi-layered nested logic, filtering, and then calculating, has been successfully solidified into a static group asset that can be reused for secondary business applications.

[0100] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for automatically generating and calculating distributed group selection SQL based on multi-level nested JSON parsing, characterized in that, include: Obtain the group selection rules carrying the set conditions, wherein the group selection rules are organized in a multi-level nested JSON format; The group selection rules are parsed: cross-entity domain condition nodes exist within the group selection rules and cross-entity domain association decomposition is performed to generate a dimensionality-reduced subtree with a parameter list; The calculation of the nesting level depth parameter of each logical branch within the dimensionality-reduced subtree includes: initiating a depth-first traversal scan starting from the root node of the dimensionality-reduced subtree; assigning a depth counter to each tree node, and incrementing the depth counter by a fixed step size when crossing level boundaries to obtain the nesting level depth parameter corresponding to each logical branch; reorganizing the logical branches using an ascending order sorting mechanism; setting a level splitting critical parameter, converting logical branches with nesting level depth parameters lower than the level splitting critical parameter into front-end independent SQL fragments to prioritize asynchronous operations; converting logical branches with nesting level depth parameters greater than or equal to the level splitting critical parameter into back-end dependent SQL fragments; the backend resource coordination service collects the available random access memory (RAM) balance parameters of the currently deployed server operating system and the task queuing length parameters of the underlying execution engine, substitutes them into the internally set resource calculation formula, and outputs a constant result of integer data bits to confirm the level splitting critical parameter; the internally set resource calculation formula is set as a linear mathematical derivation formula that takes into account both memory sufficiency and queuing congestion status. ; in, The calculated output is the critical parameter for hierarchical segmentation; This is the minimum disassembly level limit for mandatory protection of the system, and its value is a constant 1. The baseline segmentation level constant is preset for the system, with an initial value of 3; The collected parameters refer to the available random access memory (RAM) margin. The total physical memory capacity of the execution node is a constant; The task queue length parameter for the underlying execution engine. This is the maximum allowed queuing capacity threshold of the system; This is the positive gain coefficient for memory margin, and its value range is configured as follows: Floating-point numbers within; The reverse penalty coefficient for queuing congestion is configured with a value range of [value range missing]. Floating-point numbers within; This represents the floor function. This indicates taking the maximum of the two options; The parsed rules are converted into Structured Query Language (SQL) statement text: the preceding independent SQL fragments are executed asynchronously first, and a feature existence mapping table is constructed in memory using the returned initial computation data; a pre-read probe constraint function calling the feature existence mapping table is inserted at the beginning of the conditional filtering code segment of the following dependent SQL fragments; the application server's text compiler extracts the retained following dependent SQL fragments from the deferred commit stack container, and the text compiler parses the code logic structure within the fragment body, locating and tracing to the boundary of the initial code declaration; the system pre-sets a fixed identifier control string pointing to the underlying user-defined function, which is registered and configured as a pre-read probe constraint function, and its underlying encapsulated logic package... The process includes accepting primary key string input, hashing, and verifying the state of specific memory bits. The text compiler writes the identifier control string and the passed primary key reference field name to the header boundary of the subsequent dependent SQL fragment through string insertion. After completing the pre-run probe constraint appending process, the assembled long string data block becomes the final optimized SQL statement text. The pre-read probe constraint function forces the distributed computing engine to first verify whether the identity primary key of the data row being read hits the storage bit in the active identification state before performing a full disk table scan. Invalid data rows that do not hit the active identification state are discarded, and subsequent association statistics are performed on the data rows that pass the verification, thereby combining the final optimized SQL statement text. The final optimized SQL statement text is used as the structured query language SQL statement text and sent to the distributed computing engine for execution to obtain and output the individual result set of the members.

2. The method according to claim 1, characterized in that, The group selection rules cover a top-level logical structure, a secondary group structure, a third-level conditional structure, and a fourth-level parameter structure that are nested and related to each other. The top-level logical structure records the logical AND or logical OR operators that represent the global joint decision attributes. The secondary group structure includes a group type identifier and a subordinate condition array; The third-level condition structure records the logical AND or logical OR operators that control the combined state of the subordinate condition array; The fourth-level parameter structure is configured as follows: comparison dimension, data table name, field name, field type, comparison operator, and comparison value.

3. The method according to claim 2, characterized in that, Before obtaining the group selection rules carrying the set conditions, the method further includes: The distributed columnar data warehouse was used to read the individual vehicle profile tables, the electronic non-stop toll collection individual profile tables, and the highway travel application individual profile tables, respectively. The statistical script calculates the number of people covered by multiple profile tables read by tag dimension, and updates the total number of people covered by each tag dimension to the tag statistics result table. The rule parameter enumeration values ​​recorded in the business configuration library are synchronously pushed with the tag statistics result table, and cross-node joint storage is performed in the federated ad-hoc query node to respond to front-end display requests.

4. The method according to claim 3, characterized in that, The process of sending the final optimized SQL statement text as the Structured Query Language (SQL) statement text to the distributed computing engine for execution, obtaining and outputting the individual member result set, encompasses the following underlying computational processing: The group clustering requests generated based on the group selection rules and the pending execution status are ordered sequentially according to multiple grouping dependencies; The distributed computing engine is controlled to read multiple subquery instructions included in the final optimized SQL statement text at the computation layer; The multiple subquery instructions are concatenated into a combined view by using a join query statement that retains duplicate records; Nested outside the collection view are grouping aggregation instructions targeting individual primary identifiers; The number comparison interception command is appended to the end of the group aggregation command to perform intersection and union verification calculation, and finally the set of individual member results is output and stored in the offline relation table.

5. The method according to claim 4, characterized in that, After the step of storing the set of individual member results into an offline relation table, the following steps are included: The parameters of the number of people covered by each audience group and the latest calculation status update time are statistically analyzed in the offline relationship table. The individual member result set, the number of people covered, and the latest calculation status update time parameter are synchronized to the ad-hoc federated query database. The latest snapshot data of the business primary key mapping is retained in the ad hoc federated query database, which enables the application server to initiate fast cross-database retrieval of multiple tables.

6. The method according to claim 1, characterized in that, The process of identifying cross-entity domain condition nodes within the group selection rules and performing cross-entity domain association decomposition to generate a dimensionality-reduced subtree with a parameter list includes: The group selection rules are analyzed to identify multiple independent group type identifiers. When a judgment situation where different independent group type identifiers are intertwined within the nested hierarchy is detected, the logical node that triggers the intertwining operation is marked as the cross-entity domain condition node. In an in-memory computing environment, the range of the source entity data table to which the cross-entity domain condition node is attached is extracted, and a pilot probe instruction pointing to the source entity data table is generated. The pilot detection command is triggered to collect the source primary key, and the collected source verification is encapsulated into a specific group identifier column vector through the primary key.

7. The method according to claim 6, characterized in that, The generation of the dimensionality-reduced subtree with a parameter list includes: Obtain the pairwise mapping network table of group entities generated daily by offline computation; Using the specific group identifier column vector as the input retrieval source, the primary key set of the corresponding target entity data table after mapping transformation is obtained by searching the pairwise mapping relationship network table of the group entities; The primary key set after the mapping transformation is converted into a hard-coded constant constraint condition that includes the matching syntax of the set. In the in-memory computing environment, the complex multi-table association branches caused by the cross-entity domain condition nodes are removed from the original nested tree structure, and the hard-coded constant constraints are directly concatenated to the removal position to complete the reshaping of the multidimensional topology into a flattened structure, and the dimensionality-reduced subtree is output.

8. The method according to claim 1, characterized in that, The step of constructing a feature existence mapping table in memory using the returned initial computation data includes: The distributed computing engine receives the initial computation data returned after processing the preceding independent SQL fragment, the initial computation data including each verified primary key; Allocate a contiguous bit array space in the application server's memory; By calling various different hash mapping calculation functions, each verification is hash-encoded using the primary key to obtain multiple address coordinates; The storage bits corresponding to the multiple addressing coordinates in the continuous bit array space are adjusted to the active recognition state, and the bit array space adjusted to the active recognition state together constitute the feature existence mapping table.

Citation Information

Patent Citations

  • Database operating system and method based on JSON and SQL

    CN117807106A

  • Data contract-oriented strategy real-time compiling and consistency hot deployment method and system

    CN121635903A