Method, device, storage medium and terminal for constructing a field bloodline tree
By constructing a field lineage tree, executable calculation functions are automatically generated and data dependencies are displayed, solving the problem of poor data quality in existing technologies, achieving accurate identification and reliability of data dependencies, and improving data management efficiency and intelligent service capabilities.
Patent Information
- Application Number
- CN202511280353.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-09-09
AI Technical Summary
Existing data production solutions suffer from narrow knowledge scope, simplistic instructions, and heavy reliance on manual annotation, resulting in poor data quality and impacting the training effectiveness and reliability of machine learning models.
Construct a field lineage tree by obtaining the target field, database query statement, and upstream lineage information, generating an executable calculation function, parsing the input parameters, recursively calling to generate a complete lineage tree, and displaying it through a visualization interface.
It enables automated and accurate identification of data dependencies, reduces the cost of manual sorting, ensures the reliability and consistency of data quality, improves the efficiency of data understanding and management, and lays the foundation for intelligent data services.
Smart Images

Figure CN120780710B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present specification relate to the technical field of data processing, and particularly relate to a method and apparatus for constructing a field blood relation tree, a storage medium, and a terminal. BACKGROUND
[0002] Currently, data has become indispensable information relied on by a large number of production and operation activities of enterprises, especially in the era of rapid development of various machine learning models, the quality of model training data directly affects the convergence efficiency and accuracy of the model. The traditional data production scheme still has problems such as narrow knowledge and single instruction, and because the traditional scheme is highly dependent on manual annotation, not only does it consume human cost, but the deviation between the data generated by manual work and the real data in the application scenario also accumulates and magnifies, thereby damaging the training effect of the model. Under this background, how to obtain reliable and high-quality data has become a key problem that needs to be solved in this field. SUMMARY
[0003] Embodiments of the present specification provide a method and apparatus for constructing a field blood relation tree, a storage medium, and a terminal, which can solve the technical problem of poor data quality used in the production process in related technologies.
[0004] In a first aspect, embodiments of the present specification provide a method for constructing a field blood relation tree, which comprises:
[0005] obtaining a target field in a target table, a database query statement corresponding to the target table, and upstream blood relation information on which the target field depends;
[0006] taking the target field as a root node of a blood relation tree, generating a first executable calculation function corresponding to the target field according to the target field, the database query statement, and the upstream blood relation information, the first executable calculation function being used to realize the calculation logic of the target field;
[0007] parsing the input parameters of the first executable calculation function, and determining the index field corresponding to the input parameters as a child node of the root node;
[0008] performing a recursive call operation on each child node, the recursive call operation being used to generate a second executable calculation function corresponding to each child node and parse the input parameters of the second executable calculation function as a lower-level child node, until all leaf nodes of the blood relation tree are determined, the leaf node being an index field from an original base table and having no child node;
[0009] outputting the blood relation tree containing all nodes and the executable calculation functions corresponding to each node as a complete blood relation tree corresponding to the target field, and displaying the complete blood relation tree through a visual interface.
[0010] In a possible implementation, the generating, according to the target field, the database query statement, and the upstream lineage information, of the first executable computing function corresponding to the target field includes: generating, according to the target field, the database query statement, and the upstream lineage information, a first computing logic package corresponding to the target field, the first computing logic package at least including a pseudo-database query statement corresponding to the target field, the first executable computing function, and a document string; and packing the first executable computing function as a get function into a get service, and storing the document string as a knowledge document in a knowledge base.
[0011] In a possible implementation, the generating, according to the target field, the database query statement, and the upstream lineage information, of the first computing logic package corresponding to the target field includes: filling the target field, the database query statement, and the upstream lineage information into a first prompt template to obtain a first prompt, and controlling a large language model to generate the first computing logic package corresponding to the target field based on the first prompt; and the first prompt is used to specify a recognition purpose and an output rule of the large language model.
[0012] In a possible implementation, the performing, on each sub-node, of the recursive call operation includes: filling the target field and the first computing logic package into a second prompt template to obtain a second prompt, and controlling the large language model to perform the recursive call operation on each sub-node based on the second prompt; and the second prompt is used to specify a recognition purpose and an output rule of the large language model.
[0013] In a possible implementation, the obtaining of the target field in the target table, the database query statement corresponding to the target table, and the upstream lineage information on which the target field depends includes: obtaining meta information of the target table, the meta information including a database query statement for generating the target table, a database definition language file, all field names of the target table, and descriptions of each field; and determining, according to the meta information, the target field in the target table that meets a preset condition and the upstream lineage information on which the target field depends.
[0014] In a possible implementation, the determining, according to the meta information, of the target field in the target table that meets the preset condition includes: obtaining, according to the meta information, upstream lineage information of each field in the target table; calculating, based on the upstream lineage information of each field, a lineage score corresponding to each field, and determining a field whose lineage score meets a preset screening condition as the target field; and the lineage score is a weighted sum of a number of field lineages and a field lineage operation score.
[0015] In a possible implementation, the calculating of the bloodline score corresponding to each field based on the upstream bloodline information of each field comprises: identifying, based on the upstream bloodline information of each field, a field operation adopted by an upstream field corresponding to each field in a calculation process, and performing weighted calculation on the field operation corresponding to each field according to an operation weight value allocated to each field operation in advance to obtain a field bloodline operation score corresponding to each field.
[0016] In a possible implementation, the method further comprises: allocating a first type of weight value to a field operation for performing a mathematical operation and a logical operation, and allocating a second type of weight value to a field operation for performing a string operation, the first type of weight value being greater than the second type of weight value.
[0017] In a possible implementation, after the bloodline tree containing all nodes and executable calculation functions corresponding to the nodes is output as a complete bloodline tree corresponding to the target field, and the complete bloodline tree is displayed through a visual interface, the method further comprises: automatically generating a question-answer pair for machine learning model training based on the calculation logic of each node in the complete bloodline tree and the bloodline dependency relationship between the nodes.
[0018] In a possible implementation, the automatically generating of the question-answer pair for machine learning model training comprises: organizing, in a natural language form, a data value of a target leaf node corresponding to a target node in the complete bloodline tree and an indicator field identifier into a first type of question corresponding to the target node, the first type of question being used to indicate calculation of the data value of the target node; calling an executable calculation function bound to the target node and inputting the data value, and taking an output result as a first type of answer corresponding to the first type of question.
[0019] In a possible implementation, the automatically generating of the question-answer pair for machine learning model training comprises: organizing, in a natural language form, at least one specified primary key and a data value of each specified primary key into a second type of question corresponding to a target node, the second type of question being used to indicate calculation of the data value of the target node and in which all leaf node data values are hidden; obtaining data values of corresponding leaf nodes from an original base table according to the at least one specified primary key and the data value of each specified primary key, calling a corresponding executable calculation function from the leaf nodes according to the complete bloodline tree, until the data value of the target node is calculated, and taking the data value as a second type of answer corresponding to the second type of question.
[0020] In a possible implementation, the first type of question and the first type of answer are used to train an explainability prediction model, and the second type of question and the second type of answer are used to train an intelligent agent model with data query and logical reasoning capabilities.
[0021] In a possible implementation, when the question-answer pairs for training the machine learning model are generated, the data values of the nodes used are simulation data values; or when the question-answer pairs for training the machine learning model are generated, the data values of the nodes used are real data values obtained from a real online environment.
[0022] In a second aspect, an apparatus for constructing a field lineage tree is provided, and the apparatus includes:
[0023] a root node obtaining module configured to obtain a target field in a target table, a database query statement corresponding to the target table, and upstream lineage information on which the target field depends;
[0024] a computing logic analysis module configured to take the target field as a root node of a lineage tree, and generate a first executable computing function corresponding to the target field according to the target field, the database query statement, and the upstream lineage information, the first executable computing function being used to implement a computing logic of the target field;
[0025] a parameter analysis module configured to analyze input parameters of the first executable computing function, and determine an index field corresponding to the input parameters as a child node of the root node;
[0026] a child node completion module configured to perform a recursive calling operation on each child node, the recursive calling operation being used to generate a second executable computing function corresponding to each child node, and analyze input parameters of the second executable computing function as a lower-level child node, until all leaf nodes of the lineage tree are determined, the leaf nodes being index fields from original base tables and having no child nodes;
[0027] a lineage tree construction module configured to output a lineage tree containing all nodes and executable computing functions corresponding to the nodes as a complete lineage tree corresponding to the target field, and display the complete lineage tree through a visual interface.
[0028] In a third aspect, a computer program product containing instructions is provided, and when the computer program product is run on a computer or a processor, the computer or the processor performs the steps of the method described above.
[0029] In a fourth aspect, a computer storage medium is provided, and the computer storage medium stores a plurality of instructions, the instructions being adapted to be loaded by a processor and perform the steps of the method described above.
[0030] In a fifth aspect, a terminal is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being adapted to be loaded by the processor and perform the steps of the method described above.
[0031] The technical solutions provided by some embodiments of the present specification have at least the following beneficial effects:
[0032] The embodiment of the present specification provides a method for constructing a field blood relationship tree, obtaining a target field in a target table, a database query statement corresponding to the target table and upstream blood relationship information dependent on the target field; taking the target field as a root node of the blood relationship tree, generating a first executable calculation function corresponding to the target field according to the target field, the database query statement and the upstream blood relationship information, the first executable calculation function being used for realizing the calculation logic of the target field; parsing the input parameters of the first executable calculation function, determining the index field corresponding to the input parameters as the child node of the root node; performing a recursive call operation on each child node, the recursive call operation being used for generating a second executable calculation function corresponding to each child node and parsing the input parameters of the second executable calculation function as a lower-level child node, until all leaf nodes of the blood relationship tree are determined, the leaf node being an index field from an original base table and having no child node; outputting the blood relationship tree containing all nodes and the executable calculation function corresponding to each node as a complete blood relationship tree corresponding to the target field, and displaying the complete blood relationship tree through a visual interface. In the embodiment of the present specification, by obtaining the target field, the associated database query statement and the upstream blood relationship information, an accurate starting point is established for the entire process, ensuring that all subsequent analyses are based on accurate metadata, guaranteeing the pertinence and reliability of the blood relationship analysis from the source, and avoiding the tracing deviation caused by incomplete or incorrect information. On this basis, taking the target field as the root node, the first executable calculation function is generated according to the calculation logic thereof, the function being a standardized and machine-executable encapsulation of the field calculation rule, which makes the calculation logic implicit in the complex query statement or service code explicit and modular, which not only makes the calculation process of any field transparent and understandable, but also can be used as an independent and callable service unit to provide technical support for subsequent automated data services. Further, by parsing the input parameters of the calculation function, the upstream index fields directly dependent on the target field can be intelligently identified and determined as the child nodes in the blood relationship tree, this step uses the function parameter characteristics of the programming language to automatically discover the dependency relationship, realizes the automated and accurate identification of the dependency relationship, and significantly reduces the cost and error rate when manually combing the large and complex data dependencies. Then, the same operation is recursively performed on each child node to generate the executable calculation function corresponding thereto and parse the input parameters thereof, and the calculation is traced layer by layer downward until all leaf nodes are reached, that is, the basic index fields directly derived from the original base table and no longer dependent on other field calculations, a complete calculation chain from the target field to the most original data source is automatically constructed through the recursive mechanism, ensuring the comprehensiveness and depth of the blood relationship, which reveals the complete conversion path of the data from generation to final application, so that any calculation deviation or data quality problem in the intermediate link can be quickly located and traced.In addition, the constructed complete lineage tree with each node calculation function is outputted and displayed through a visualization interface to convert the results of the foregoing technical process into knowledge assets that can be directly understood and interacted by users. On the one hand, the visual tree structure provides intuitive insight into complex data dependency relationships, greatly improving the efficiency of data understanding. On the other hand, the outputted lineage tree structure containing executable calculation functions is a knowledge framework that can be directly called by downstream systems, thereby laying a solid foundation for building an intelligent data management and service system. BRIEF DESCRIPTION OF DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present specification or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present specification, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0034] Figure 1 An exemplary system architecture diagram of a method for constructing a field lineage tree provided by an embodiment of the present specification;
[0035] Figure 2 A flowchart of a method for constructing a field lineage tree provided by an embodiment of the present specification;
[0036] Figure 3 An exemplary diagram of a visual field lineage tree provided by an embodiment of the present specification;
[0037] Figure 4 A flowchart of another method for constructing a field lineage tree provided by an embodiment of the present specification;
[0038] Figure 5 A flowchart of still another method for constructing a field lineage tree provided by an embodiment of the present specification;
[0039] Figure 6 A structural block diagram of an apparatus for constructing a field lineage tree provided by an embodiment of the present specification;
[0040] Figure 7 A structural diagram of a terminal provided by an embodiment of the present specification. DETAILED DESCRIPTION
[0041] In order to enable the features and advantages of the embodiments of the present specification to be more obvious and easy to understand, the technical solutions in the embodiments of the present specification will be clearly and completely described below in conjunction with the drawings in the embodiments of the present specification. Obviously, the described embodiments are only a part of the embodiments of the present specification, rather than all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present specification.
[0042] In the following description with reference to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments are not meant to represent all implementations consistent with the present specification. Rather, they are merely examples in accordance with some aspects of the embodiments of the specification, as detailed in the appended claims. And in the description of the embodiments of the present specification, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B: "and / or" in the text only describes the relationship between the associated objects, which means that there can be three relationships, for example, A and / or B, which means that there are three cases of A alone, A and B together, and B alone. In addition, in the description of the embodiments of the present specification, "multiple" means two or more than two.
[0043] Hereinafter, the terms "first" and "second" are used only for descriptive purposes and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more of the features.
[0044] As various machine learning models and agent architectures are increasingly applied in various vertical fields, the requirements for data quality, diversity, and logical complexity of the model training process have also significantly increased. However, the current data supply system gradually reveals its limitations. Under the existing mode, the data produced for training models often fails to meet the requirements of machine learning models for semantic breadth, reasoning depth, and logical consistency, significantly restricting the model training effect and application landing.
[0045] At the vertical field data level, the knowledge covered by existing data is relatively narrow, and there is a large gap between the richness of instruction types and real scenarios. This leads to a single semantic context encountered by the model during the training phase, limiting its generalization ability and adaptability in actual complex situations. Secondly, the existing data production process relies heavily on artificial design rules and templates, making it difficult to automatically generate high-value data with multi-step reasoning chains and deep logical associations, which are the core elements to support agent (Agentic) models to achieve autonomous decision-making and task planning.
[0046] More importantly, even if a large amount of manpower is invested in data labeling and processing, there can still be significant differences between the results and the real records in the underlying database. These uncontrollable differences not only reduce the credibility of the training data, but also introduce difficult-to-detect biases and noise, making it difficult to ensure the consistency and depth of the reasoning logic in the model, and thus damaging the credibility and robustness of the model. At the same time, the artificial production mode is also accompanied by high proofreading and rework costs, causing a mismatch between resource input and output efficiency.
[0047] In summary, the current vertical data faces multiple dilemmas in terms of production scale, complexity, authenticity, and sustainability, so the embodiments of the present specification provide a method for constructing a field blood tree to solve the technical problem of poor data quality used in the above production process.
[0048] Referring to Figure 1 , Figure 1 An exemplary system architecture diagram of a method for constructing a field blood tree provided by the embodiments of the present specification is shown.
[0049] As Figure 1 shown, the system architecture can include a terminal 101, a network 102, and a server 103. The network 102 is used to provide a communication link medium between the terminal 101 and the server 103. The network 102 can include various types of wired communication links or wireless communication links, for example: wired communication links include optical fiber, twisted pair or coaxial cable, wireless communication links include Bluetooth communication links, Wireless-Fidelity (Wi-Fi) communication links or microwave communication links, etc.
[0050] The terminal 101 can interact with the server 103 through the network 102 to receive messages from the server 103 or send messages to the server 103, or the terminal 101 can interact with the server 103 through the network 102 to receive messages or data sent by other users to the server 103. The terminal 101 can be hardware or software. When the terminal 101 is hardware, it can be various electronic devices, including but not limited to smartwatches, smartphones, tablet computers, laptop computers, and desktop computers, etc. When the terminal 101 is software, it can be installed in the above-mentioned electronic devices, which can be implemented as multiple software or software modules (for example: used to provide distributed services), or as a single software or software module, which is not specifically limited here.
[0051] Optionally, in the embodiment of the present specification, the terminal 101 first acquires the target field in the target table, the database query statement corresponding to the target table, and the upstream blood relationship information relied on by the target field; taking the target field as the root node of the blood relationship tree, the terminal 101 further generates a first executable calculation function corresponding to the target field according to the target field, the database query statement, and the upstream blood relationship information, and the first executable calculation function is used to realize the calculation logic of the target field. Based on this, the terminal 101 parses the input parameters of the first executable calculation function, and determines that the index field corresponding to the input parameters is the child node of the root node. In order to complete the entire blood relationship tree, the terminal 101 performs a recursive call operation on each child node, and the recursive call operation is used to generate a second executable calculation function corresponding to each child node and parse the input parameters of the second executable calculation function as a lower-level child node, until all leaf nodes of the blood relationship tree are determined. The leaf node is an index field from the original base table and has no child node. Finally, the terminal 101 outputs the blood relationship tree containing all nodes and the executable calculation function corresponding to each node as a complete blood relationship tree corresponding to the target field, and displays the complete blood relationship tree through a visual interface.
[0052] The server 103 can be a server providing various services. It should be noted that the server 103 can be hardware or software. When the server 103 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When the server 103 is software, it can be implemented as multiple software or software modules (for example, used to provide distributed services), or as a single software or software module, which is not specifically limited here.
[0053] Alternatively, the system architecture can also not include the server 103, in other words, the server 103 can be an optional device in the embodiment of the present specification, that is, the method provided by the embodiment of the present specification can be applied to a system structure including only the terminal 101, and the embodiment of the present specification does not limit this.
[0054] It should be understood that Figure 1 The number of terminals, networks, and servers in the system architecture is only illustrative, and can be any number of terminals, networks, and servers according to the needs of implementation.
[0055] Please refer to Figure 2 , Figure 2 A flowchart of a method for constructing a field blood relationship tree provided by the embodiment of the present specification. The execution subject of the embodiment of the present specification can be a terminal executing the construction of the field blood relationship tree, a processor in the terminal executing the method of constructing the field blood relationship tree, or a service in the terminal executing the method of constructing the field blood relationship tree. For the convenience of description, the specific execution process of the method of constructing the field blood relationship tree is introduced below taking the execution subject as the processor in the terminal.
[0056] As Figure 2 shown, the method of constructing the field blood relationship tree can at least include:
[0057] S202, obtaining the target field in the target table, the database query statement corresponding to the target table and the upstream blood relationship information dependent on the target field.
[0058] Optionally, in the automatic construction method of the data blood relationship tree, first, the target field in the target table, the database query statement corresponding to the target table and the upstream blood relationship information dependent on the target field need to be obtained. The target field is the core data indicator concerned by the user and needs to be traced and analyzed, such as the composite indicators of "merchant responsibility cancellation rate compared with last month" and "shop rating" in the intelligent business scenario. The database query statement, also known as the SQL statement, is an instruction for retrieving, filtering and aggregating data from the database. The database query statement of the target table includes the SQL code for generating or calculating the target field, which encapsulates the extraction logic and complete calculation rules of the data and is an important basis for analyzing the calculation logic of the target field. The upstream blood relationship information is information describing the dependency relationship between the target field and other fields, which clearly shows that the target field is obtained from which upstream fields through what kind of operation or processing, and reflects the source and flow path of the data. These information is usually provided by the data governance platform or the metadata management system.
[0059] Optionally, in the actual implementation process, the target field and its associated SQL statement can be automatically identified by parsing the metadata of the target table or interfacing with the metadata service of the enterprise. Specifically, the system can automatically scan the structure information of the target table, identify each field therein, and determine the target field according to the service needs of the user or the preset rules. For the database query statement, it can be obtained by accessing the query log, stored procedure or view definition of the database, etc. These information contains the extraction and calculation logic of the target field. The upstream blood relationship information dependent on the target field can be analyzed by the pre-defined blood relationship extraction rules or with the help of syntax analysis tools (such as SQL parser) to perform lexical, syntax and semantic analysis on the SQL statement, such as using the blood relationship analysis tool provided by the database itself or by parsing the field reference relationship in the database query statement to obtain, for example, by analyzing the JOIN, WHERE, SELECT and other clauses in the SQL statement to determine the association between the target field and other fields, thereby extracting all the upstream fields and data tables dependent on the target field and forming the blood relationship information. The automatic processing mechanism has significant speed and accuracy advantages compared to the traditional manual data dependency relationship analysis method, effectively avoiding the problem of incomplete and inaccurate blood relationship information caused by manual misjudgment or omission, and ensuring the reliability of the subsequent construction process from the data source.
[0060] S204, taking the target field as a root node of a bloodline tree, generating a first executable calculation function corresponding to the target field according to the target field, the database query statement and the upstream bloodline information, the first executable calculation function being used for implementing the calculation logic of the target field.
[0061] Optionally, in the prior art, the dependency relationship of the data field is often unclear, lacking systematic analysis and presentation, resulting in great difficulty in data tracing and understanding. Therefore, in the embodiments of the present application, the dependency relationship between data is clearly and intuitively displayed by constructing a bloodline tree with the target field as the root node, so as to realize the visualization and systematic management of data bloodline.
[0062] Specifically, the target field is taken as the root node of the bloodline tree, and the first executable calculation function corresponding to the target field can be generated according to the identification of the target field, the database query statement and the upstream bloodline information obtained in the previous step. The bloodline tree is a model for intuitively presenting the dependency relationship between fields in a tree structure, wherein the root node is the target field, and other nodes are the upstream fields of each level on which the target field depends. The bloodline tree can clearly trace the source of data flow. The first executable calculation function is a function that can be directly run and implement the calculation logic of the target field. It is converted based on the database query statement and the upstream bloodline information, and can automatically complete the calculation of the target field.
[0063] Specifically, first, the target field is set as the root node of the bloodline tree, and the starting point of the construction of the entire bloodline tree is determined. Then, the database query statement is parsed and analyzed to extract the calculation logic of the target field, including the upstream fields, operator symbols, function calls and the like. In combination with the upstream bloodline information, the specific dependency relationship between each upstream field and the target field is further determined. Based on these information, a suitable programming language (such as Python) is selected to convert the calculation logic into an executable function, i.e. the first executable calculation function. During the function generation process, it is necessary to ensure that the function can accurately receive the upstream fields as input and perform operations according to the parsed calculation logic, and finally output the value of the target field. By automatically generating the first executable calculation function, the calculation logic of the target field is solidified into automatically executable code, which not only improves the calculation efficiency, but also guarantees the accuracy and consistency of the calculation results.
[0064] Further, in the process of generating the executable calculation function, metadata description information (such as a document string Docstring) can be further injected to explain the purpose, parameters and output results of the function, thereby enhancing its explainability. This not only facilitates the understanding and maintenance of developers, but also lays a solid foundation for building an enterprise-level data knowledge base and improves the manageability of data assets.
[0065] S206, parse the input parameters of the first executable computing function, and determine that the index fields corresponding to the input parameters are the child nodes of the root node.
[0066] Optionally, in the process of automatically constructing the data lineage tree, after generating the first executable computing function corresponding to the root node, the input parameters of the function need to be further parsed to determine the directly dependent upstream index fields of the root node, and these fields are established as the child nodes of the root node. The input parameters are external input variables necessary for the computing function when performing computing, which essentially represent the dependent items in the target field computing logic; the child nodes are the direct dependent items of the root node (target field) in the lineage tree, which reflect the direct association between the target field and the upstream data.
[0067] In a feasible implementation, after obtaining the first executable computing function, the parameter list of the function can be parsed by a code static analysis tool (such as an abstract syntax tree parser), the input parameter names and data types declared in the function definition are extracted, and then the input parameters are mapped to the corresponding index fields in combination with the meta information of the target table and the upstream lineage information, so as to determine that these index fields are the child nodes of the root node in the lineage tree. For example, if the first executable computing function is "target field = A + B" (where A and B are input parameters), then A and B are determined to be the child nodes of the root node after parsing. In specific implementation, the field mapping relationship in the upstream lineage information can also be combined to verify the correspondence accuracy of the parameters and the index fields, to ensure the accuracy and integrity of the child nodes. Through static code analysis or reflection mechanism, the parameter list in the function signature can be automatically extracted and then mapped to the specific service index fields. The conversion from computing logic to dependent relationship is realized, the data dependent link between adjacent nodes in the lineage tree is accurately and efficiently identified, and the high error rate and low efficiency problems caused by complex logic or large scale in the traditional manual carding method are solved.
[0068] S208, perform a recursive call operation on each child node, the recursive call operation is used to generate a second executable computing function corresponding to each child node and parse the input parameters of the second executable computing function as lower-level child nodes, until all leaf nodes of the lineage tree are determined, and the leaf node is an index field from the original bottom table and has no child node.
[0069] Optionally, on the basis of having extracted the direct dependent sub-nodes of the target field, these sub-nodes may also need to be dependent on other fields for calculation, and therefore, in order to complete the complete calculation link from the original field of the bottom table to the target field, further tracing of the sub-nodes is needed until all leaf nodes of the bloodline tree are determined, that is, the leaf nodes. The data in the original bottom table is the original data directly collected without the need for a calculation process; the leaf nodes are the basic data fields directly derived from the original bottom table without being calculated by any other field, and they serve as the source of the data bloodline, marking the end of the dependency chain. In actual implementation, whether a field is a leaf node can be determined by judging whether it is associated with calculation logic or whether it is directly mapped to a bottom table field, and the originality of the data table to which the field belongs and whether there is a lower-level dependency are checked to accurately define the leaf node and ensure the uniqueness of the data source.
[0070] Further, the process of completing the bloodline tree can be implemented by performing a recursive call operation for each sub-node. That is, in a top-down manner, the same processing procedure as the root node is repeatedly applied to each newly discovered sub-node: first, a second executable calculation function corresponding to the sub-node is generated, which encapsulates the calculation logic of the index field corresponding to the sub-node itself; then, the input parameters of the second executable calculation function are parsed to determine the index field of the next level of dependency, and it is taken as a sub-node (i.e., a lower node in the bloodline tree) of the current sub-node. The recursive process continues to iterate until it reaches the leaf node that no longer has any dependency relationship. The recursive traversal mechanism realizes the automation, full-link discovery and modeling of large-scale complex data dependency relationships, can deeply mine multi-level, nested calculation dependency relationships between fields, and finally constructs a complete and deep bloodline tree from the target field to all underlying source fields. In effect, it can change the traditional bloodline construction mode that highly depends on manual interviews, document review and experience inference, greatly improving the efficiency, coverage and accuracy of bloodline analysis; secondly, the complete dependency chain is conducive to the user's understanding of the complete conversion process of data from generation to consumption, providing support capabilities for data tracing, impact analysis, error root cause positioning and other scenarios.
[0071] S210, output the bloodline tree containing all nodes and the executable calculation functions corresponding to the nodes as a complete bloodline tree of the target field, and display the complete bloodline tree through a visual interface.
[0072] Optionally, after the construction of the bloodline tree is completed, the complete tree structure containing all nodes and their corresponding executable computing functions is further output. The complete bloodline tree corresponding to the target field combines the topology and executable code in its computing process, and as a data asset, each node of the complete bloodline tree not only records the meta information of the field, but also binds the specific implementation of the computing logic of the field. Finally, the structure is displayed to the user in a graphical manner through a visualization interface. The interface usually provides interactive operation capabilities, such as expanding / collapsing nodes, viewing function details, and tracing along links. Please refer to Figure 3 , Figure 3 for an example diagram of a visualized field bloodline tree provided by an embodiment of the present specification. As shown in Figure 3 , the “store rating” is taken as the target field, and the direct dependent child nodes of the target field include, for example, “store collection number”, “product rating”, “delivery rating”, “purchase quantity”, “evaluation number”, and “evaluation score”. The “product rating” is obtained by averaging the scores of all products (A1, A2, A3, A4) in the store, and the “evaluation score” is obtained by performing a preset logical operation on the comment data (D1, D2, D3, D4, D5) of the store in the past period of time. In the visualized complete bloodline tree, the identification of each field is the node of the bloodline tree, and the description of the node includes the computing logic of the node (if any). The relationship between the nodes is the dependency relationship between the fields. In this way, the bloodline relationship and the executable computing function are combined and output and visualized, which is equivalent to outputting the computable bloodline map corresponding to the target field. Not only is the path of data flow displayed, but each node along the path also has the ability of instant verification and simulation execution, greatly improving the usability and practical value of data bloodline. Users can directly explore the computing logic of a specific node on the visualization interface, and this also lays a foundation for building an intelligent data operation system. For example, based on this structure, data quality checking, precise impact analysis, or the ability to dynamically generate query statements for self-service data retrieval services can be automatically triggered.
[0073] Specifically, the above step flow provided in the embodiments of the present specification realizes a systematic, automated and deep data bloodline analysis method. It realizes the full-link automatic tracing from the target field to the original data source by recursively generating executable functions and analyzing the dependency relationship, and finally outputs a bloodline tree structure with visualization and computational ability. It improves the traceability of data by constructing a bloodline tree for the target field. The source and calculation process of the target field can be clearly traced through the bloodline tree, which facilitates data auditing and problem troubleshooting. The automated processing mechanism can improve data processing efficiency. The automatic generation of calculation functions reduces manual intervention, speeds up the calculation of target fields, and also enhances the consistency and accuracy of data, avoids errors that may occur in manual calculation, ensures the reliability of the calculation results of target fields, and provides a solid data foundation for subsequent data applications (such as model training, data analysis, etc.), which helps to improve the effect and quality of data application. Traditional data bloodline analysis can only provide a static dependency relationship view, while the executable calculation function generated by the present scheme makes the calculation process of each calculation link transparent, verifiable and directly callable service unit. For example, after generating the calculation function and bloodline tree for the target field "near 7-day merchant responsibility cancellation rate compared with the previous period", the function can be directly integrated and called by the subsequent data quality detection platform, self-service data retrieval service or machine learning feature engineering pipeline, thereby significantly improving the automation level and consistency of data application.
[0074] In the embodiments of the present specification, a method for constructing a field blood relation tree is provided, a target field in a target table, a database query statement corresponding to the target table and upstream blood relation information dependent on the target field are obtained; the target field is taken as a root node of the blood relation tree, a first executable computing function corresponding to the target field is generated according to the target field, the database query statement and the upstream blood relation information, the first executable computing function is used to realize the computing logic of the target field; input parameters of the first executable computing function are parsed, and an index field corresponding to the input parameters is determined as a child node of the root node; a recursive calling operation is performed on each child node, the recursive calling operation is used to generate a second executable computing function corresponding to each child node and parse input parameters of the second executable computing function as a lower-level child node, until all leaf nodes of the blood relation tree are determined, the leaf node is an index field from an original base table and has no child node; the blood relation tree containing all nodes and the executable computing function corresponding to each node is output as a complete blood relation tree corresponding to the target field, and the complete blood relation tree is displayed through a visual interface. In the embodiments of the present specification, by obtaining the target field, the associated database query statement and the upstream blood relation information, an accurate starting point is established for the entire process, ensuring that all subsequent analyses are based on accurate metadata, ensuring the pertinence and reliability of the blood relation analysis from the source, and avoiding the tracing deviation caused by incomplete or incorrect information. On this basis, the target field is taken as the root node, and the first executable computing function is generated according to the computing logic thereof, the function is a standardized and machine-executable encapsulation of the field computing rule, which makes the computing logic implicit in the complex query statement or service code explicit and modular, which not only makes the computing process of any field transparent and understandable, but also provides technical support for subsequent automated data services as an independent and callable service unit. Further, by parsing the input parameters of the computing function, the upstream index fields directly dependent on the target field can be intelligently identified and determined as child nodes in the blood relation tree, this step uses the function parameter characteristics of the programming language to automatically discover the dependency relationship, realizes the automatic and accurate identification of the dependency relationship, and significantly reduces the cost and error rate when manually combing large and complex data dependencies. Then, the same operation is recursively performed on each child node to generate the corresponding executable computing function and parse the input parameters thereof, and the blood relation tree is recursively constructed from the target field to the original data source, ensuring the comprehensiveness and depth of the blood relation, which reveals the complete conversion path of data from generation to final application, so that any computing deviation or data quality problem in the intermediate link can be quickly located and traced.In addition, the constructed complete bloodline tree with each node calculation function is outputted and displayed through a visualization interface to convert the results of the foregoing technical process into knowledge assets that can be directly understood and interacted by users. On the one hand, the visual tree structure provides intuitive insight into complex data dependency relationships, greatly improving the efficiency of data understanding. On the other hand, the output bloodline tree structure containing executable calculation functions is a knowledge framework that can be directly called by downstream systems, thereby laying a solid foundation for building an intelligent data management and service system.
[0075] Please refer to Figure 4 , Figure 4 Another flowchart of a method for constructing a field bloodline tree provided by an embodiment of the present specification.
[0076] As Figure 4 indicated, the method for constructing a field bloodline tree can at least include:
[0077] S402, obtaining meta information of the target table, the meta information including a database query statement for generating the target table, a database definition language file, all field names of the target table, and descriptions of each field.
[0078] Optionally, in the process of data governance and analysis, automatically identifying key target fields in the target table with multiple index fields is an important prerequisite for constructing an effective data bloodline model. In order to determine the target fields with important analysis value in the target table, the importance of each field can be analyzed according to the meta information of the target table.
[0079] Specifically, a connection can be established with the database, and the meta data query interface provided by the database can be used to obtain the meta information of the target table. Specifically, the meta information of the target table is metadata describing the data table structure and generation method, including a database query statement (SQL query statement) for generating the table, a database definition language (table DDL) file, the names (identifiers) of all fields in the table, and the description text of each field. These meta information constitutes the technical and service context of the data table and is the data basis for automated analysis. In actual implementation, automatic extraction can be realized by integrating a metadata management system, a database system directory or a data directory tool to ensure the completeness and accuracy of the obtained information.
[0080] S404, determining target fields in the target table that meet a preset condition and upstream bloodline information dependent on the target fields according to the meta information.
[0081] Optionally, after obtaining the meta information, upstream lineage information of each field in the table is further acquired. The information is used to describe the source and dependency relationship of the field data, including the underlying table fields directly or indirectly dependent on, the intermediate fields involved in the calculation process and the conversion logic. By parsing the SQL query statement and the DDL structure, and with the help of a syntax parser and a lineage inference algorithm, a dependency graph at the field level can be automatically constructed. After obtaining the lineage information of all fields, important fields can be determined as target fields for which it is necessary to construct a lineage tree according to the quality, depth, importance and other elements of the lineage information of each field.
[0082] In a feasible implementation, for each field, its lineage score can be quantitatively calculated according to the upstream lineage information, and the field lineage score can be defined as a weighted combination of two core components. Specifically, the field lineage score = field lineage number + field lineage operation score, where the field lineage number is the number of upstream fields directly or indirectly dependent on the field in the calculation process, reflecting the depth and breadth of the data source of the field; the field lineage operation score is used to quantify the importance and complexity of the operation experienced by the field in the calculation process, for example, operations involving aggregation operations (such as SUM, AVG), conditional judgment (such as CASE WHEN) or service rule conversion will be given a higher weight, while simple field mapping or type conversion operation will be given a lower weight. Further, the fields whose lineage scores meet the preset screening conditions are determined as target fields, and the scoring mechanism is made to have dual considerations of dependency scale and logic complexity through weighted summation, so as to more comprehensively and objectively evaluate the key degree of the field in the data system.
[0083] In actual operation, the field blood relationship operation score of each field is calculated. Based on the upstream blood relationship information of each field, the field operation adopted by the corresponding upstream field of each field in the calculation process is identified, and the operation weight value allocated to each field operation in advance is weighted to obtain the field blood relationship operation score corresponding to each field. First, in the field blood relationship meta information table, the operation of the field blood relationship is recorded. Then, according to the upstream blood relationship information of the target field, all field operation types involved in the calculation process of the target field can be identified. Field operation is the specific operation performed when generating or converting the field data, such as mathematical operation (such as addition, subtraction, multiplication, division, aggregation function), logical operation (such as conditional judgment CASE WHEN, IF-THEN-ELSE), string processing (such as CONCAT, SUBSTRING, REPLACE) or data type conversion, etc. For example: [COALESCE, trunc, CaseWhenResult, BinaryOP], wherein BinaryOP is a binary operation of addition, subtraction, multiplication and division, CaseWhenResult is casewhen, and IFResult is IF else. Each operation type reflects the functional characteristics and complexity of a certain aspect of field calculation logic.
[0084] On this basis, the identified field operation is weighted according to the operation weight value allocated in advance for different operation types, so as to obtain the field blood relationship operation score of the field. This weighting process is not simply counting, but by introducing a weight mechanism, the contribution difference of different operation types to the complexity and importance of field calculation is distinguished. Specifically, the first type of weight value is allocated to the field operation for mathematical operation and logical operation, and the second type of weight value is allocated to the field operation for string operation, wherein the first type of weight value is greater than the second type of weight value. The weight allocation idea of each type of operation is that mathematical operation (such as sum, average, ratio calculation, etc.) and logical operation (such as conditional branching, multiple judgment), such as BinaryOP, CaseWhenResult, IFResult, etc., usually reflect the core calculation logic and decision process in the service rule, and its correctness has a significant impact on the final value of the data, and the complexity is also relatively high, so a higher weight is given to highlight its importance. In contrast, general string operations (such as extraction, concatenation, replacement, case conversion), such as jsonreplace, replace, concat, etc. usually involve data formatting or cleaning, and do not directly affect the calculation result of the core service index, so a relatively low weight is given. Through such weight allocation strategy, the essential complexity and service value of the calculation logic behind the field can be more accurately quantified, so as to avoid making a one-sided judgment based on the number of operations.
[0085] In a feasible implementation, the field bloodline operation score of the field is determined. The determination can be implemented by using a set of standardized operation type classification dictionary and weight mapping table. By automatically parsing the upstream SQL or program logic, all operators and functions are identified, classified into corresponding operation types, and the weight values are accumulated, and finally the field bloodline operation score is generated, so as to realize the automation and objectivity of scoring, and improve the explainability and configurability of the evaluation system. In the embodiment of the present specification, the calculation logic complexity of each field is converted into a quantitative score output by using the weighting calculation mechanism, thereby providing a reliable basis for accurately evaluating the field importance and screening the key target field.
[0086] Further, based on the quantitative calculation of the bloodline score of the target field, at least one field with a bloodline score meeting the requirement can be selected from all fields as the target field according to a preset screening condition (such as a score threshold or a score ranking), for example, a field with the highest score, or the top K fields, or a field with a score higher than a preset threshold, for subsequent deep bloodline analysis or model training. By introducing the quantifiable bloodline evaluation system and the automatic screening mechanism, the core field with high value, high complexity and high dependence in the data set can be intelligently and efficiently identified, the problem that important fields are ignored or secondary fields are overemphasized is effectively avoided, and the accuracy and efficiency of data asset management are significantly improved. Not only the subjectivity and cost of manual intervention are reduced, but also the data management system is endowed with the ability to identify key data nodes and evaluate data asset value, which provides a solid foundation for building high-quality and interpretable data bloodline, and supports the implementation of various advanced application scenarios such as automatic data document generation, impact analysis and fault root cause positioning.
[0087] S406, generating a first calculation logic package corresponding to the target field according to the target field, the database query statement and the upstream bloodline information, the first calculation logic package at least including a pseudo-database query statement corresponding to the target field, a first executable calculation function and a document string.
[0088] Optionally, in the data lineage automatic processing flow, generating the first computing logic package corresponding to the target field is a key link to realize the field-level logic reuse. The first computing logic package is a structured and multi-modal logic description unit, which at least includes three core contents: pseudo-database query statement, first executable computing function and document string. Among them, the pseudo-database query statement is a query logic expression form convenient for human understanding, which is used to clearly show the core computing logic. In the embodiments of the present specification, the computing path and service rules of the target field can be intuitively displayed, but it does not depend on the syntax details of the specific database system and does not pay attention to the specific database execution details, which facilitates the understanding of the reader and improves the readability and communicability of the computing logic; the first executable computing function encapsulates the computing logic of the target field into an independent and callable program unit (such as Python or Java function), which realizes the coding and automatic execution ability of the computing process; the document string (Docstring) is a structured text form to explain the function, input, output and service meaning of the function, which is used to form a persistent and retrievable technical document, which is equivalent to an explanation document.
[0089] In the specific implementation process, the automatic generation of the first computing logic package can be realized by a large language model. The target field, related database query statement and its upstream lineage information are filled into a pre-designed first prompt template, thereby generating a standardized first prompt. The first prompt is a structured instruction for the large language model, and its content is pre-set and can be used to clearly specify the identification target of the model (such as “analyze the computing logic in the SQL statement”), the output format specification (such as “the output needs to include pseudo-SQL, Python function and Docstring”) and the domain knowledge constraint (such as “follow the e-commerce indicator calculation specification”). By calling the large language model service and inputting the prompt, a first computing logic package that meets the requirements and has high quality can be automatically generated. This method effectively overcomes the shortcomings of low efficiency and easy errors of traditional manual code writing and document writing methods, and avoids the problems of inconsistent and inaccurate computing logic description caused by personnel replacement or service complication.
[0090] In the instant e-commerce scenario, the first prompt word can be designed in two stages to generate the first computing logic package of the target field. For example, the first prompt word template can be as follows:
According to the given SQL code, target field, and upstream blood relationship information of the target field, summarize the calculation expression of the target field. Requirements: 1. The calculation expression output is in quotation marks; 2. When explaining the calculation logic, give it in the format of a mathematical expression. If it is a logic calculation, give it in a SQL-like logic code; 4. The calculation expression must include all the given upstream fields; 4. Chinese characters do not appear in the calculation expression, and the names of the fields are aligned with SQL. Special note: 1. Each indicator may have different weights, reference values, and calculation functions for these same-named but different-sense indicator fields. Given content: <SQL code> {sql}; <target field> {target_filed}; <target field dependent upstream field> {blood_fileds}; <calculation logic>
Convert the given SQL logic pseudo-code, target field, into a standard python function. Requirements: 1. The function name is the target field. 2. The python function is wrapped with '```python ```'. 4. Dates are uniformly represented as "ds". 4. The id of the store is uniformly represented as "shop_id". 5. If you involve import external libraries, such as import datetime, please put the import operation in the function. <SQL code> {sql}; <target field> {target_filed}; <python function>
[0091] It should be noted that the first prompt word template is exemplary, and in actual application process, adaptive adjustment can be made according to different needs and scenes, and the specific content of the first prompt word template is not limited in the embodiments of the present application.
[0092] S408, pack the first executable computing function as a data extraction function into the data extraction service, and store the document string as a knowledge document in the knowledge base.
[0093] Optionally, after generating the first computing logic package, the first executable computing function and the document string therein are respectively deployed and archived automatically. On the one hand, the first executable computing function is packaged and published as a data service interface that can be called in real time by means of an automated script or a continuous integration tool, so that the computing process of the target field is standardized and reusable, and the agility and response efficiency of the data service are significantly improved. On the other hand, the document string is stored as a structured knowledge document in the enterprise knowledge base, and is usually indexed in association with related fields, tables or service indicators. This not only forms a data asset document that can be accumulated continuously, but also provides key semantic resources for subsequent collaborative development, audit verification and model training. In summary, this step realizes the transformation from the original database query statement to the executable, interpretable and reusable data service component by introducing a large language model driven logic package automatic generation mechanism. The core invention point lies in deeply integrating natural language processing technology and traditional data management processes to extract, encapsulate and document the computing logic in an automated and intelligent manner, greatly improving the governance efficiency, consistency and usability of data assets.
[0094] S410, parse the input parameters of the first executable computing function, and determine that the indicator field corresponding to the input parameters is a child node of the root node.
[0095] For step S410, please refer to the detailed description in step S206, which will not be repeated here.
[0096] S412, perform a recursive call operation on each child node, and the recursive call operation is used to generate a second executable computing function corresponding to each child node and parse the input parameters of the second executable computing function as a lower-level child node, until all leaf nodes of the blood relationship tree are determined.
[0097] Optionally, after determining the child node directly dependent on the target field, the computing logic of all child nodes in the blood relationship tree can be further recursively generated until all leaf nodes are found. In the specific implementation process, the data dependency relationship can also be automatically processed through the code generation and logical reasoning ability of the large language model. Specifically, the determined target field and the corresponding first computing logic package are filled into a second prompt template designed in advance, and then a precise instruction, i.e., a second prompt, is generated.
[0098] Here, the second prompt word is used to strictly define the task execution framework of the large language model, clearly specifying the identification target of the model (such as "analyze the calculation logic in the SQL statement"), the output format specification (such as "the output needs to include pseudo-SQL, Python functions and Docstring") and the domain knowledge constraint (such as "follow the e-commerce index calculation specification"), so as to ensure that the output of the large language model is highly consistent with the overall system requirements in structure, semantics and function. In the embodiments of the present specification, the large language model service can be called through an application programming interface (API) and the second prompt word can be submitted, so that a required calculation logic package for the current sub-node can be obtained, which includes the pseudo-query statement, executable calculation function and supporting document corresponding to the sub-node. Similarly, in the aforementioned instant e-commerce scenario, the second prompt word can still be designed in two stages to generate the second calculation logic package of each sub-field, and the specific examples are similar to the first prompt word template, which will not be described here. This step realizes the automatic discovery and logic reconstruction of large-scale and deep-level dependency relationships, solving the complexity and scale problems that manual processing cannot handle; secondly, the large language model is precisely controlled through the prompt word engineering, ensuring the standardization, usability and reliability of the output results, and avoiding the uncertainty of generated content; in addition, this process accelerates the knowledge encapsulation process from raw data to reusable data assets, providing core power for building computable, interpretable and operable intelligent data systems.
[0099] S414, output the blood relationship tree containing all nodes and the executable calculation functions corresponding to each node as a complete blood relationship tree of the target field, and display the complete blood relationship tree through a visual interface.
[0100] For step S414, please refer to the detailed description in step S210, which will not be repeated here.
[0101] In the embodiments of the present specification, a method for constructing a field lineage tree is provided. Automatic extraction is realized through an integrated metadata management system, a database system directory or a data directory tool, ensuring the completeness and accuracy of the obtained information. By introducing a quantifiable lineage evaluation system and an automated screening mechanism, the core fields with high value, high complexity and high dependence in the data set can be intelligently and efficiently identified, effectively avoiding the problem of important fields being ignored or secondary fields being overemphasized, and significantly improving the accuracy and efficiency of data asset governance. The first computing logic package is automatically generated by a large language model, and a first computing logic package that meets the requirements and has high quality is automatically generated, realizing the transformation from the original database query statement to an executable, interpretable and reusable data service component. The first executable computing function is taken as a data extraction function, which is packaged and published to a data extraction service platform through an automated script or a continuous integration tool, so that it becomes a data service interface that can be called in real time, which makes the calculation process of the target field standardized and reusable, and significantly improves the agility and response efficiency of the data service. The document string is stored as a structured knowledge document in the enterprise knowledge base, and is usually associated and indexed with related fields, tables or service indicators, which not only forms a sustainable accumulation of data asset documents, but also provides key semantic resources for subsequent collaborative development, audit verification and model training. Through the code generation and logical reasoning ability of the large language model, the data dependency relationship is automatically processed layer by layer, ensuring the standardization, usability and reliability of the lineage tree.
[0102] The complete lineage tree obtained through the above steps integrates the topological structure and rich semantic data assets, and each node of the tree not only records the meta information of the field, but also binds the computing logic, i.e., an executable function or pseudo code that encapsulates the calculation rule of the field, realizing the complete description of the service rule and technical path of data generation. At the same time, the parent-child relationship and edge direction in the tree clearly express the lineage dependency relationship, i.e., the complete conversion and flow path of data from the original source to the final indicator. In a feasible implementation manner, for the target field lineage tree that has been analyzed, in the actual downstream application link, high-quality question-answer pairs for machine learning model training can be automatically generated based on the constructed complete lineage tree.
[0103] Therefore, please refer to Figure 5 , Figure 5 Another flowchart of a method for constructing a field lineage tree provided by the embodiments of the present specification is provided. As Figure 5As shown, in the method of constructing the field bloodline tree, after step S210, outputting the bloodline tree containing all nodes and the executable computing functions corresponding to each node as a complete bloodline tree corresponding to the target field, and displaying the complete bloodline tree through a visualization interface, at least the following can be further included: based on the computing logic of each node in the complete bloodline tree and the bloodline dependency relationship between each node, automatically generating question and answer pairs for machine learning model training.
[0104] In the specific implementation process, different forms of questions are generated by traversing the nodes in the bloodline tree and according to the types and positions thereof. The computing rules encapsulated in the computing logic of any non-leaf node (i.e., a field obtained through calculation) in the bloodline tree and the information of the dependent child nodes thereof provide rich materials for constructing questions. For example, the system can extract the service meaning of the node (such as "7-day merchant cancellation rate change"), the directly upstream field (such as "merchant cancellation order number" and "order number entering the merchant background") dependent thereon, and combine the computing steps described by the corresponding "pseudo-database query statement" to assemble a question requiring the calculation of the index value in a natural language form. The question can include specific numerical conditions (whether simulated data or real data obtained through a unique identifier from a bottom table in real time), or can only provide a service context and a unique identifier, thereby adapting to different types of machine learning tasks.
[0105] Furthermore, the answer generation process is fully automated and highly reliable, requiring no additional manual calculations or annotations. Instead, it directly calls the "executable computation function" corresponding to the node. By passing the required parameters to the function (whether simulated data or real data queried in real time from the underlying table using a unique identifier), the output of the function is the standard answer. This mechanism ensures the accuracy of the answer, fundamentally eliminating the errors and ambiguities that may be introduced by traditional manual annotation. This step automatically transforms data lineage into structured knowledge rich in logic and context that can be directly used to train machine learning models. It automates and intelligentizes the training data generation process, enabling large-scale, high-efficiency production of question-answer pairs covering complex service logic, effectively solving the bottleneck problems of scarce vertical domain data and insufficient instruction richness. Secondly, by utilizing the inherent hierarchical structure and computational logic of the lineage tree, the generated questions naturally possess different levels of reasoning depth and complexity, meeting the diverse training needs of agentic models, from simple queries to multi-step reasoning and even those requiring external query capabilities. Furthermore, the data generated using the above methods maintains a high degree of consistency with the computational logic of the real service system, greatly reducing the gap between training data and production environment data (GAP), thus laying a solid foundation for improving the accuracy and reliability of the model in practical applications. The core value of computable lineage trees is fully realized, providing high-quality and highly reliable data support for training the next generation of machine learning models with deep reasoning capabilities and service cognition levels.
[0106] In one feasible implementation, embodiments of this specification provide an application method for an interpretable prediction model, namely, generating a first type of question and a first type of answer for training the interpretable prediction model.
[0107] S502. Organize the data values of the target leaf nodes corresponding to the target nodes in the complete lineage tree and the indicator field identifiers into a first type of question corresponding to the target node in the form of natural language. The first type of question is used to indicate the calculation of the data values of the target nodes.
[0108] Optionally, when providing high-quality, structured training data for the interpretable prediction model, based on the characteristics of the interpretable prediction model, it is necessary to use the complete lineage tree constructed in the aforementioned steps to automatically generate question-answer pairs in a specific format, namely, type I questions and type I answers.
[0109] In the implementation process, when constructing the first type of problem, the system locates the target node (i.e. the core indicator field required to be learned by the prediction model to be trained) from the complete blood relationship tree, and traces back to all target leaf nodes it depends on. These leaf nodes are directly derived from the fields of the original base table, representing the most original data input required for calculation. By obtaining the data values of these leaf nodes, combined with the indicator field identifier (i.e. the service indicator name represented by the target node, such as "merchant cancellation rate compared with the previous period"), the problem statement is organized in a clear, standardized natural language form. For example: "Given that the number of orders rejected by merchants in the last 7 days is X, the number of orders entering the merchant's back office in the last 7 days is Y, …, please calculate the value of the merchant cancellation rate compared with the previous period in the last 7 days." The essence of the first type of problem is to provide a complete context instruction for the model that contains all the necessary input data and requires it to calculate the target indicator value. Its design aims to simulate the complete scenario of predicting indicators based on given data in a real environment.
[0110] S504, calling the executable calculation function bound to the target node and inputting the data value, and taking the output result as the first type of answer corresponding to the first type of problem.
[0111] For the first type of answer corresponding to the first type of problem, its construction process also does not require human intervention or calculation. The system directly calls the executable calculation function that has been firmly bound to the target node when constructing the blood relationship tree. This function is a pre-encapsulated, machine-executable code unit that internally completely encodes the calculation logic from input (leaf node value) to output (target node value). Then the leaf node data value extracted from the question is passed to this function as an input parameter, and the calculation result obtained after executing the function is directly taken as the first type of answer. The complex calculation knowledge implied in the data blood relationship is converted into standardized "question-answer" instances that can be used for supervised machine learning model training. In the process of constructing the question, the original data input and the target indicator are contextualized and packaged in a natural language form understandable by humans, providing specific context and all feature information required for model learning. In the process of automatically solving the answer construction, the existing and verified executable calculation logic is further utilized to efficiently generate accurate answers, thereby providing supervision signals for model training.
[0112] In a feasible implementation, in the process of automatically generating machine learning question and answer pairs based on the blood tree, the selection of the source of the data value affects the characteristics and application scenarios of the generated data. In a specific implementation, there can be two data assignment strategies to adapt to different training objectives and environmental requirements. The first strategy is to use mock data for the data values of each node. Mock data refers to fictitious data generated by artificial generation or algorithm synthesis, which conforms to the value range and distribution rules of specific fields, and is similar to real data in semantics and statistical characteristics, but does not come from actual service operations. In actual application, data values with reasonable can be generated by random generation or based on pre-defined rules according to the data types (such as integer, floating point) defined in the field meta information, value range constraints (such as percentage between 0 and 1), and historical statistical characteristics (such as mean, variance). The core advantage of this method is its high controllability and safety: on the one hand, it can flexibly construct data covering corner cases or extreme scenarios to enhance the robustness of the machine learning model; on the other hand, since it does not involve real service data at all, it avoids the risk of sensitive information leakage and meets the requirements of data privacy supervision. Therefore, this method is suitable for model initial development, algorithm verification, or isolated environments where production data cannot be directly accessed.
[0113] The second strategy is to use real data values (Real Production Data) obtained from real online environment for the data values of each node. This kind of data is directly obtained from online databases, data warehouses or real-time data streams, and is real value generated and recorded in actual service. Specifically, snapshot data at the current or historical time can be obtained on demand through a secure data access interface (such as a desensitized database connection or data service call). The value of real data lies in its reflection of the actual state and complexity of the service, which contains naturally occurring noise, bias and correlation, so that the model trained based on it can better fit the data distribution in the actual application scenario, and improve the generalization performance and decision accuracy of the model after going online. Therefore, this method can be used for the final training and fine-tuning stage of the model to ensure the high consistency of the model output with the production environment. The introduction of both mock data and real data modes enables the framework to adapt to the full life cycle requirements from research and development verification to production deployment. When using mock data, the safety and flexibility of the data generation process can be ensured, and when using real data, the high fidelity and practicality of the training data can be ensured.
[0114] In addition to the application mode for explainable prediction models in steps S502-S504, or the embodiments of the present specification also provide an application mode for agent models, that is, generating second type questions and second type answers for training agent models with data query and logical reasoning capabilities.
[0115] S506, organize the second type question corresponding to the target node according to at least one specified primary key and the data value of each specified primary key in the form of natural language, the second type question is used to indicate the calculation of the data value of the target node and the data value of all leaf nodes is hidden therein.
[0116] Optionally, in the agent-oriented model, the core training data for training the agent model with autonomous data query and complex logical reasoning capability is provided by generating the second type question and the second type answer. The focus of the question and answer pair construction in this scenario is to simulate the task scenario that the agent faces in the real environment, that is, only the target task is known, rather than directly obtaining all input data, so as to force the model to learn how to actively obtain data and perform multi-step reasoning to solve the problem.
[0117] In the specific implementation process, the system first generates a second type question. The construction of this question is different from other types of questions, and one or more specified primary keys need to be determined, which are the key fields that can uniquely identify service entities (such as specific dates, store IDs, etc.). Subsequently, the system takes these primary keys and their specific data values (such as "service date is 20250813, store ID is 999999") as the core elements, combines the service meaning of the target node (that is, the identifier of the index field to be solved), and organizes a complete task instruction in the form of natural language that omits all intermediate data and leaf node specific values. For example: "In the context of , the service date is 20250813, the store ID is 999999, and the value of the near 7-day merchant responsibility cancellation rate is required." Here, the purpose of hiding the data values of all leaf nodes is to simulate the challenge to the agent in the real world, and the agent only knows the task to be completed and the available positioning information (primary key), and must determine how to obtain the required specific data. This question form forces the model to have the ability to understand the task intent, identify data requirements, and plan the path to obtain the data.
[0118] S508, obtain the data value of the corresponding leaf node from the original base table according to at least one specified primary key and the data value of each specified primary key, call the corresponding executable calculation function from the leaf node according to the complete bloodline tree, until the data value of the target node is calculated, and take the data value as the second type answer corresponding to the second type question.
[0119] After generating the second type of question, a second type of answer matching the question is further automatically generated. The generation process of the answer itself is a complete simulation of the behavior that the expected agent should exhibit. First, according to the specified primary key and its data value provided in the question, a query operation on the original base table is automatically initiated. The original base table referred to at this time is the data table that stores the most basic and most fine-grained service records, which is the source of data bloodline. Through the execution of a precise database query, all necessary leaf node data values related to the service entity (uniquely determined by the primary key) are obtained. Then, according to the dependency relationship and calculation path defined by the complete bloodline tree, starting from the obtained leaf node data, the corresponding executable calculation function is called from bottom to top and layer by layer. Each step of calling uses the function of the current layer to calculate the results of the directly dependent upper layer, and gradually infers until the data value of the target node is finally calculated. This final calculation result is taken as the standard answer to the second type of question.
[0120] In summary, through the above steps, a closed-loop training instance capable of training an agent model to have the ability of "perception-planning-action-reasoning" is constructed. In the process of defining tasks and contexts, an explicit goal is set for the agent and a sparse environment containing only positioning information (primary key) but not all input data, which forces the model to learn to actively interact with the environment (database) to obtain information. In the process of simulating the execution trajectory of the agent, not only an answer is generated, but also each step of operation that an ideal agent should perform when solving the problem is more completely simulated, that is, from querying the original data according to the primary key to step by step deducing the final result according to the known calculation logic. This process defines a preferred behavior sequence for machine learning. Using this specific method, the core ability of the agent model can be directly cultivated, and by hiding the input data and providing only the primary key, the incomplete information environment of real-world tasks is successfully simulated, thereby training the model's key ability, i.e., after understanding the task intent, the model can autonomously issue a data query request to obtain necessary information, rather than passively accepting all inputs. In addition, the answer generation process strictly reproduces the full-link calculation logic from the original data to the final result, which provides a demonstration for the model to learn the correct and interpretable multi-step reasoning path, so that the model not only knows what the answer is, but also learns how to get the answer through a series of correct operations, which is crucial for cultivating the model's logical rigor and decision transparency.
[0121] In the embodiments of the present specification, a method for constructing a field blood relation tree is provided. This blood relation tree-based question and answer production mechanism can realize the automation of the training data generation process, greatly improving the efficiency and reducing the labor cost; by generating answers through the blood relation tree and the calculation function derived from the real service logic, the consistency of the calculation between the training data and the production environment is ensured, effectively eliminating the gap between the data and the real situation, thereby greatly improving the prediction accuracy and service reliability of the model after training; the question is presented in the form of natural language, so that the model can learn and understand the specific semantics and calculation context of the service indicators, which directly enhances the explainability of the model, even for complex prediction results, the decision basis can be traced back to these explicit input data.
[0122] Please refer to Figure 6 , Figure 6 A structural block diagram of a device for constructing a field blood relation tree is provided in the embodiments of the present specification. As shown in Figure 6 , the device 600 for constructing a field blood relation tree comprises:
[0123] The root node acquisition module 610 is configured to acquire a target field in a target table, a database query statement corresponding to the target table, and upstream blood relation information on which the target field depends;
[0124] The calculation logic analysis module 620 is configured to take the target field as a root node of a blood relation tree, generate a first executable calculation function corresponding to the target field according to the target field, the database query statement, and the upstream blood relation information, and the first executable calculation function is used to realize the calculation logic of the target field;
[0125] The input parameter analysis module 630 is configured to analyze the input parameters of the first executable calculation function, and determine that the index field corresponding to the input parameters is a child node of the root node;
[0126] The child node completion module 640 is configured to perform a recursive call operation on each child node, the recursive call operation is used to generate a second executable calculation function corresponding to each child node and analyze the input parameters of the second executable calculation function as a lower-level child node, until all leaf nodes of the blood relation tree are determined, and the leaf node is an index field from an original base table and has no child node;
[0127] The blood relation tree construction module 650 is configured to output the blood relation tree containing all nodes and the executable calculation functions corresponding to each node as a complete blood relation tree corresponding to the target field, and display the complete blood relation tree through a visual interface.
[0128] Optionally, the computing logic analysis module 620 is further configured to generate a first computing logic package corresponding to the target field according to the target field, the database query statement, and the upstream blood relationship information, the first computing logic package at least including a pseudo-database query statement corresponding to the target field, a first executable computing function, and a document string; package the first executable computing function as a data fetching function and write it into a data fetching service, and store the document string as a knowledge document in a knowledge base.
[0129] Optionally, the computing logic analysis module 620 is further configured to fill the target field, the database query statement, and the upstream blood relationship information into a first prompt word template to obtain a first prompt word, and control a large language model to generate a first computing logic package corresponding to the target field based on the first prompt word; the first prompt word is used to specify the identification purpose and output rule of the large language model.
[0130] Optionally, the sub-node completion module 640 is further configured to fill the target field and the first computing logic package into a second prompt word template to obtain a second prompt word, and control the large language model to perform a recursive call operation on each sub-node based on the second prompt word; the second prompt word is used to specify the identification purpose and output rule of the large language model.
[0131] Optionally, the root node acquisition module 610 is further configured to acquire meta information of the target table, the meta information including a database query statement for generating the target table, a database definition language file, all field names of the target table, and descriptions of each field; and determine a target field satisfying a preset condition in the target table and upstream blood relationship information relied on by the target field according to the meta information.
[0132] Optionally, the root node acquisition module 610 is further configured to acquire upstream blood relationship information of each field in the target table according to the meta information; calculate a blood relationship score corresponding to each field based on the upstream blood relationship information of each field, and determine a field satisfying a preset screening condition as the target field; wherein the blood relationship score is a weighted sum of a field blood relationship number and a field blood relationship operation score.
[0133] Optionally, the root node acquisition module 610 is further configured to identify a field operation adopted by an upstream field corresponding to each field in a computing process based on the upstream blood relationship information of each field, and perform weighted calculation according to an operation weight value pre-assigned to each field operation to obtain a field blood relationship operation score corresponding to each field.
[0134] Optionally, the device 600 for constructing a field blood relationship tree further includes a weight assignment module configured to assign a first type of weight value to a field operation for performing mathematical operation and logical operation, and assign a second type of weight value to a field operation for performing string operation, the first type of weight value being greater than the second type of weight value.
[0135] Optionally, the device 600 for constructing a field blood relation tree further comprises a question and answer pair generation module configured to automatically generate question and answer pairs for machine learning model training based on the calculation logic of each node in the complete blood relation tree and the blood relation dependency between the nodes.
[0136] Optionally, the question and answer pair generation module is further configured to organize the data value of the target leaf node corresponding to the target node in the complete blood relation tree and the indicator field identifier into a first type question corresponding to the target node in a natural language form, and the first type question is used to indicate calculation of the data value of the target node; and the question and answer pair generation module is further configured to call an executable calculation function bound to the target node and input the data value, and take the output result as a first type answer corresponding to the first type question.
[0137] Optionally, the question and answer pair generation module is further configured to organize at least one specified primary key and the data value of each specified primary key into a second type question corresponding to the target node in a natural language form, and the second type question is used to indicate calculation of the data value of the target node while hiding all leaf node data values therein; and the question and answer pair generation module is further configured to obtain the data value of the corresponding leaf node from the original base table according to the at least one specified primary key and the data value of each specified primary key, call the corresponding executable calculation function from the leaf node according to the complete blood relation tree, until the data value of the target node is calculated, and take the data value as a second type answer corresponding to the second type question.
[0138] Optionally, the first type question and the first type answer are used to train an explainability prediction model, and the second type question and the second type answer are used to train an agent model with data query and logical reasoning capabilities.
[0139] Optionally, when generating the question and answer pairs for machine learning model training, the data values of the nodes used are simulation data values; or, when generating the question and answer pairs for machine learning model training, the data values of the nodes used are real data values obtained from a real online environment.
[0140] Please refer to Figure 7 , Figure 7 A structural schematic diagram of a terminal is provided for an embodiment of the present specification. As shown in Figure 7 , the terminal 700 can include at least one terminal processor 701, at least one network interface 704, a user interface 703, a memory 705, and at least one communication bus 702.
[0141] The communication bus 702 is configured to realize the connection communication between the components.
[0142] The user interface 703 can include a display screen (Display) and a camera (Camera). Optionally, the user interface 703 can further include a standard wired interface and a wireless interface.
[0143] The network interface 704 can optionally include a standard wired interface, a wireless interface (e.g., a WI-FI interface).
[0144] The terminal processor 701 can include one or more processing cores. The terminal processor 701 connects various parts within the terminal 700 via various interfaces and lines, executes various functions of the terminal 700 and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 705, and calling data stored in the memory 705. The terminal processor 701 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The terminal processor 701 can be integrated with a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU is mainly responsible for processing an operating system, a user interface, and an application program, etc.; the GPU is responsible for rendering and drawing content to be displayed on a display screen; and the modem is responsible for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the terminal processor 701, but can be implemented by a separate chip.
[0145] The memory 705 can include a random access memory (RAM) and a read-only memory (ROM). Optionally, the memory 705 includes a non-transitory computer-readable storage medium. The memory 705 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 705 can include a program storage area and a data storage area. The program storage area can store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playing function, an image playing function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store data involved in the above-mentioned various method embodiments, etc. The memory 705 can optionally be at least one storage device located away from the terminal processor 701. As shown in the figure, the memory 705 as a computer storage medium can include an operating system, a network communication module, a user interface module, and a program for constructing a field bloodline tree. Figure 7 As shown in the figure, the memory 705 as a computer storage medium can include an operating system, a network communication module, a user interface module, and a program for constructing a field bloodline tree.
[0146] In Figure 7 In the terminal 700 shown, the user interface 703 is mainly used to provide an interface for user input, and obtain data input by the user. The terminal processor 701 can be used to call a program for constructing a field blood relation tree stored in the memory 705, and specifically perform the following operations:
[0147] Obtain a target field in a target table, a database query statement corresponding to the target table, and upstream blood relation information on which the target field depends;
[0148] Take the target field as a root node of a blood relation tree, and generate a first executable calculation function corresponding to the target field according to the target field, the database query statement, and the upstream blood relation information. The first executable calculation function is used to implement the calculation logic of the target field.
[0149] Parse the input parameters of the first executable calculation function, and determine that the index field corresponding to the input parameters is a child node of the root node;
[0150] Perform a recursive call operation on each child node. The recursive call operation is used to generate a second executable calculation function corresponding to each child node, and parse the input parameters of the second executable calculation function as a lower-level child node, until all leaf nodes of the blood relation tree are determined. The leaf node is an index field from an original base table and has no child node.
[0151] Output the blood relation tree containing all nodes and the executable calculation functions corresponding to the nodes as a complete blood relation tree corresponding to the target field, and display the complete blood relation tree through a visualization interface.
[0152] In some embodiments, when the terminal processor 701 performs the operation of generating a first executable calculation function corresponding to a target field according to the target field, a database query statement, and upstream blood relation information, the terminal processor 701 specifically performs the following steps: generates a first calculation logic package corresponding to the target field according to the target field, the database query statement, and the upstream blood relation information. The first calculation logic package at least includes a pseudo-database query statement corresponding to the target field, the first executable calculation function, and a document string; packs the first executable calculation function as a data retrieval function and writes it into a data retrieval service, and stores the document string as a knowledge document in a knowledge base.
[0153] In some embodiments, when the terminal processor 701 performs the operation of generating a first calculation logic package corresponding to a target field according to the target field, a database query statement, and upstream blood relation information, the terminal processor 701 specifically performs the following steps: fills the target field, the database query statement, and the upstream blood relation information into a first prompt word template to obtain a first prompt word, and controls a large language model to generate a first calculation logic package corresponding to the target field based on the first prompt word; the first prompt word is used to specify the identification purpose and output rule of the large language model.
[0154] In some embodiments, the terminal processor 701, when performing the recursive call operation on each sub-node, specifically performs the following steps: fills the target field and the first calculation logic package into the second prompt word template to obtain a second prompt word, and controls the large language model to perform the recursive call operation on each sub-node based on the second prompt word; the second prompt word is used to specify the identification purpose and output rule of the large language model.
[0155] In some embodiments, the terminal processor 701, when performing the operation of obtaining the target field in the target table, the database query statement corresponding to the target table, and the upstream blood relationship information dependent on the target field, specifically performs the following steps: obtaining the meta information of the target table, the meta information including the database query statement for generating the target table, the database definition language file, all field names of the target table and the description of each field; determining the target field in the target table that meets the preset condition and the upstream blood relationship information dependent on the target field according to the meta information.
[0156] In some embodiments, the terminal processor 701, when performing the operation of determining the target field in the target table that meets the preset condition according to the meta information, specifically performs the following steps: obtaining the upstream blood relationship information of each field in the target table according to the meta information; based on the upstream blood relationship information of each field, calculating the blood relationship score corresponding to each field, and determining the field whose blood relationship score meets the preset filtering condition as the target field; wherein the blood relationship score is the weighted sum of the field blood relationship number and the field blood relationship operation score.
[0157] In some embodiments, the terminal processor 701, when performing the operation of calculating the blood relationship score corresponding to each field based on the upstream blood relationship information of each field, specifically performs the following steps: based on the upstream blood relationship information of each field, identifying the field operation adopted by the upstream field corresponding to each field in the calculation process, and performing weighted calculation according to the operation weight value pre-allocated to each field operation to obtain the field blood relationship operation score corresponding to each field.
[0158] In some embodiments, the terminal processor 701 further specifically performs the following steps: allocating a first type of weight value to the field operation for performing mathematical operation and logical operation, and allocating a second type of weight value to the field operation for performing string operation, the first type of weight value being greater than the second type of weight value.
[0159] In some embodiments, after the terminal processor 701 performs the operation of outputting the blood relationship tree containing all nodes and the executable calculation function corresponding to each node as a complete blood relationship tree of the target field, and displaying the complete blood relationship tree through a visual interface, it further specifically performs the following steps: based on the calculation logic of each node in the complete blood relationship tree and the blood relationship dependent relationship between each node, automatically generating a question and answer pair for machine learning model training.
[0160] In some embodiments, the terminal processor 701, when automatically generating the question-answer pairs for machine learning model training, specifically performs the following steps: organizes the data values of the target leaf nodes corresponding to the target node in the complete bloodline tree and the indicator field identification into a first type of question corresponding to the target node in a natural language form, and the first type of question is used to indicate the calculation of the data values of the target node; calls the executable calculation function bound to the target node and inputs the data values, and outputs the results as the first type of answer corresponding to the first type of question.
[0161] In some embodiments, the terminal processor 701, when automatically generating the question-answer pairs for machine learning model training, specifically performs the following steps: organizes the data values of the target leaf nodes corresponding to the target node in the complete bloodline tree and the indicator field identification into a first type of question corresponding to the target node in a natural language form, and the first type of question is used to indicate the calculation of the data values of the target node; calls the executable calculation function bound to the target node and inputs the data values, and outputs the results as the first type of answer corresponding to the first type of question.
[0162] In some embodiments, the first type of question and the first type of answer are used to train an explainability prediction model, and the second type of question and the second type of answer are used to train an agent model with data query and logical reasoning capabilities.
[0163] In some embodiments, when generating the question-answer pairs for machine learning model training, the data values of each node used are simulation data values; or, when generating the question-answer pairs for machine learning model training, the data values of each node used are real data values obtained from a real online environment.
[0164] The embodiments of the present specification provide a computer program product containing instructions, which, when the computer program product is run on a computer or a processor, causes the computer or the processor to perform the steps of the method of any one of the above embodiments. The embodiments of the present specification also provide a computer storage medium, which can store a plurality of instructions, and the instructions are suitable for being loaded and executed by a processor to perform the steps of the method of any one of the above embodiments. Wherein, the apparatus, computer readable storage medium, computer program product or chip provided by the embodiments of the present specification are used to execute the corresponding method provided above, and thus the beneficial effects that can be achieved are referred to the beneficial effects in the corresponding method provided above, which will not be described here.
[0165] In several embodiments provided in the present specification, it should be understood that the disclosed apparatus and method can be implemented in other manners. For example, the division of the apparatus embodiments is merely a logical function division, and there can be another division manner for the actual implementation, for example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different modules can be indirect couplings or communication connections through some interfaces, and there can be electric, mechanical or other forms.
[0166] The modules illustrated as separated components can or can not be physically separated, and the components illustrated as modules can or can not be physical modules, i.e., can be located in one place, or can be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment.
[0167] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The above computer program product includes one or more computer instructions. When the above computer program instructions are loaded and executed on a computer, all or part of the above processes or functions described in the embodiments of the present specification are generated. The above computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices. The above computer instructions can be stored in a computer readable storage medium or transmitted by the above computer readable storage medium. The above computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through a wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The above computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The above available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital versatile disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)) and the like.
[0168] It should be noted that, for the foregoing method embodiments, for the convenience of description, they are all described as a combination of a series of actions, but those skilled in the art should know that the embodiments of the present specification are not limited by the order of the described actions, because according to the embodiments of the present specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of the present specification.
[0169] In addition, it should be noted that the technical solutions of the embodiments of the present specification can be applied to transactions, distribution services of instant e-commerce platforms, such as Taobao flash shopping, Tao Fresh, Eleme takeout, and retail, etc. The information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of the present specification are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0170] The above describes specific embodiments of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in an order different than the order in the embodiments and still achieve the desired result. In addition, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous or possible.
[0171] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0172] The above is a description of a method for constructing a field blood tree, an apparatus, a storage medium, and a terminal provided by the embodiments of the present specification. For those skilled in the art, according to the idea of the embodiments of the present specification, there will be changes in specific implementation and application range. In summary, the content of the present specification should not be understood as a limitation of the present application.
Claims
1. A method for constructing a field lineage tree, characterized in that, The method includes: Obtain the target field from the target table, the database query statement used to generate the target table, and the upstream lineage information on which the target field depends; Using the target field as the root node of the lineage tree, a first executable calculation function corresponding to the target field is generated based on the target field, the database query statement, and the upstream lineage information. The first executable calculation function is used to implement the calculation logic of the target field. Parse the input parameters of the first executable calculation function and determine that the index field corresponding to the input parameters is a child node of the root node; A recursive call operation is performed on each child node. The recursive call operation is used to generate a second executable calculation function corresponding to each child node and parse the input parameters of the second executable calculation function as the next-level child node, until all leaf nodes of the lineage tree are determined. The leaf nodes are indicator fields from the original base table that have no child nodes. The lineage tree containing all nodes and the executable computation functions corresponding to each node is output as the complete lineage tree corresponding to the target field, and the complete lineage tree is displayed through a visualization interface; The step of generating a first executable calculation function corresponding to the target field based on the target field, the database query statement, and the upstream lineage information includes: Based on the target field, the database query statement, and the upstream lineage information, a first computational logic package corresponding to the target field is generated. The first computational logic package includes at least a pseudo-database query statement corresponding to the target field, a first executable computation function, and a document string. The document string is a persistent and searchable technical document in structured text form that describes the function, input, output, and service meaning of the first executable computation function. The first executable calculation function is packaged as a data retrieval function and written into the data retrieval service, and the document string is stored as a knowledge document in the knowledge base.
2. The method according to claim 1, characterized in that, The step of generating a first calculation logic package corresponding to the target field based on the target field, the database query statement, and the upstream lineage information includes: The target field, the database query statement, and the upstream lineage information are filled into the first prompt word template to obtain the first prompt word. The large language model is then controlled to generate the first calculation logic package corresponding to the target field based on the first prompt word. The first prompt word is used to specify the recognition purpose and output rules of the large language model.
3. The method according to claim 2, characterized in that, The recursive call operation on each child node includes: The target field and the first calculation logic package are filled into the second prompt word template to obtain the second prompt word, and the large language model is controlled to perform recursive call operations on each child node based on the second prompt word; The second prompt word is used to specify the recognition purpose and output rules of the large language model.
4. The method according to claim 1, characterized in that, The step of obtaining the target field in the target table, the database query statement used to generate the target table, and the upstream lineage information on which the target field depends includes: Obtain the metadata of the target table, which includes the database query statement that generated the target table, the database definition language file, all field names of the target table, and the description of each field; Based on the metadata, the target fields in the target table that meet the preset conditions and the upstream lineage information on which the target fields depend are determined.
5. The method according to claim 4, characterized in that, The step of determining the target fields in the target table that meet preset conditions based on the metadata includes: Based on the metadata, obtain the upstream lineage information for each field in the target table; Based on the upstream lineage information of each field, calculate the lineage score corresponding to each field, and determine the field whose lineage score meets the preset screening conditions as the target field; Wherein, the lineage score is the weighted sum of the number of field lineages and the field lineage operation score; the field lineage operation score is used to quantify the importance and complexity of the field operations that the field undergoes during the calculation process, and the score of the field lineage operation score is positively correlated with the importance and complexity of the field operations.
6. The method according to claim 5, characterized in that, The calculation of the kinship score corresponding to each field based on the upstream kinship information of each field includes: Based on the upstream lineage information of each field, the field operation used by the upstream field corresponding to each field in the calculation process is identified, and the field lineage operation score corresponding to each field is obtained by weighting the operation operation according to the operation weight value pre-assigned to each field operation.
7. The method according to claim 6, characterized in that, The method further includes: A first type of weight value is assigned to field operations used for mathematical and logical operations, and a second type of weight value is assigned to field operations used for string operations, wherein the first type of weight value is greater than the second type of weight value.
8. The method according to claim 1, characterized in that, After outputting the lineage tree containing all nodes and the executable computation functions corresponding to each node as the complete lineage tree corresponding to the target field, and displaying the complete lineage tree through a visualization interface, the process further includes: Based on the computational logic of each node in the complete lineage tree and the lineage dependencies between nodes, question-answer pairs for training machine learning models are automatically generated.
9. The method according to claim 8, characterized in that, The automatically generated question-answer pairs for machine learning model training include: The data values and indicator field identifiers of the target leaf nodes corresponding to the target nodes in the complete lineage tree are organized into a first type of question corresponding to the target node in natural language form. The first type of question is used to indicate the calculation of the data values of the target nodes. Call the executable calculation function bound to the target node and input the data value, then use the output result as the first type of answer corresponding to the first type of question.
10. The method according to claim 9, characterized in that, The automatically generated question-answer pairs for machine learning model training include: The second type of question is organized in natural language form according to at least one specified primary key and the data values of each specified primary key into a second type of question corresponding to the target node. The second type of question is used to instruct the calculation of the data value of the target node and hides the data values of all leaf nodes therein. Based on the at least one specified primary key and the data values of each specified primary key, the data values of the corresponding leaf nodes are obtained from the original base table. According to the complete lineage tree, the corresponding executable calculation function is called starting from the leaf node until the data value of the target node is calculated, and the data value is used as the second type answer corresponding to the second type of question.
11. The method according to claim 10, characterized in that, The first type of question and the first type of answer are used to train an interpretable prediction model, and the second type of question and the second type of answer are used to train an agent model with data query and logical reasoning capabilities.
12. The method according to claim 10, characterized in that, When generating question-answer pairs for training machine learning models, the data values used for each node are simulated data values; Alternatively, when generating question-answer pairs for training machine learning models, the data values used for each node are real data values obtained from a real online environment.
13. An apparatus for constructing a field lineage tree, characterized in that, The device includes: The root node acquisition module is used to acquire the target field in the target table, the database query statement for generating the target table, and the upstream lineage information on which the target field depends. The computational logic analysis module is used to generate a first executable computation function corresponding to the target field, based on the target field, the database query statement, and the upstream lineage information, with the target field as the root node of the lineage tree. The first executable computation function is used to implement the computational logic of the target field. The input parameter parsing module is used to parse the input parameters of the first executable calculation function and determine that the index field corresponding to the input parameter is a child node of the root node; The child node completion module is used to perform recursive call operations on each child node. The recursive call operation is used to generate the second executable calculation function corresponding to each child node and parse the input parameters of the second executable calculation function as the next-level child node, until all leaf nodes of the lineage tree are determined. The leaf nodes are indicator fields from the original base table that have no child nodes. The lineage tree construction module is used to output the lineage tree containing all nodes and the executable calculation function corresponding to each node as the complete lineage tree corresponding to the target field, and to display the complete lineage tree through a visualization interface; The computational logic analysis module is further configured to generate a first computational logic package corresponding to the target field based on the target field, the database query statement, and the upstream lineage information. The first computational logic package includes at least a pseudo-database query statement corresponding to the target field, a first executable computation function, and a document string. The document string is a persistent and searchable technical document that describes the function, input, output, and service meaning of the first executable computation function in a structured text format. The first executable computation function is packaged as a data retrieval function and written into the data retrieval service, and the document string is stored as a knowledge document in the knowledge base.
14. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions adapted for loading by a processor and executing the steps of the method as described in any one of claims 1 to 12.
15. A terminal, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the steps of the method as described in any one of claims 1 to 12.
Citation Information
Patent Citations
Question and answer model training method and device, computer equipment and storage medium
CN119513253A
Method and device for data blood relationship analysis, storage medium and electronic equipment
CN119988460A