Evaluation data generation method and device and product

By generating knowledge graphs and utilizing large-scale artificial intelligence models, the problem of generating high-quality evaluation datasets has been solved, achieving efficient and accurate evaluation data generation and improving the quality and credibility of the evaluation data.

CN122020672APending Publication Date: 2026-05-12BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2026-02-12
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

The process of collecting and labeling high-quality evaluation datasets is time-consuming and labor-intensive, especially in long-tail scenarios where data collection is difficult, and existing technologies struggle to generate evaluation data efficiently.

Method used

By generating knowledge graphs from raw business data through structured processing, and then using large-scale artificial intelligence models to generate evaluation data based on these knowledge graphs, the problems of entity confusion, logical breaks, and ambiguous relationships have been solved, thus improving the quality and efficiency of generation.

Benefits of technology

It has enabled the generation of high-quality and reliable evaluation data, improved generation efficiency, solved the problems of entity confusion and logical breaks, and enhanced the accuracy and reliability of evaluation data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122020672A_ABST
    Figure CN122020672A_ABST
Patent Text Reader

Abstract

The invention provides an evaluation data generation method and device, electronic equipment, a storage medium and a computer program product, relates to the technical field of computers, in particular to the technical fields of large models, knowledge maps and the like, and can be applied to an evaluation data generation scene. According to the specific implementation scheme, service native data is processed in a structured mode, and a knowledge graph representing the incidence relation between entities in the service native data is generated; and through the artificial intelligence large model, according to the knowledge graph, generating evaluation data of the to-be-evaluated model. Based on the structural advantages of the knowledge graph, the problems of entity confusion, logic fracture, relation fuzziness and the like can be solved, the efficient generation capability of the artificial intelligence large model is superimposed, and the generation quality, the credibility and the generation efficiency of the evaluation data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, specifically to the fields of large models and knowledge graphs, and in particular to a method, apparatus, electronic device, storage medium, and computer program product for generating evaluation data, which can be applied to evaluation data generation scenarios. Background Technology

[0002] Evaluation datasets are the core support for evaluating the performance of large model products. However, the process of collecting and labeling high-quality datasets is time-consuming and labor-intensive, especially in the face of long-tail scenarios that are difficult to collect. It is urgent to use technical means to achieve the screening and generation of evaluation data. Summary of the Invention

[0003] This disclosure provides a method, apparatus, electronic device, storage medium, and computer program product for generating evaluation data.

[0004] According to the first aspect, a method for generating evaluation data is provided, including: structuring business native data to generate a knowledge graph representing the relationships between entities in the business native data; and generating evaluation data for the model to be evaluated based on the knowledge graph using a large artificial intelligence model.

[0005] According to the second aspect, an evaluation data generation device is provided, comprising: a graph generation unit configured to process business native data in a structured manner and generate a knowledge graph representing the relationships between entities in the business native data; and a data generation unit configured to generate evaluation data of the model to be evaluated based on the knowledge graph using an artificial intelligence big data model.

[0006] According to a third aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method as described in any implementation of the first aspect.

[0007] According to a fourth aspect, a non-transitory computer-readable storage medium is provided that stores computer instructions for causing a computer to perform the method described in any implementation of the first aspect.

[0008] According to a fifth aspect, a computer program product is provided, comprising: a computer program that, when executed by a processor, implements the method as described in any implementation of the first aspect.

[0009] According to the technology disclosed herein, a method and apparatus for generating evaluation data are provided. By structuring business native data, a knowledge graph representing the relationships between entities in the business native data is generated. Based on the knowledge graph, evaluation data for the model to be evaluated is generated using an artificial intelligence big data model. Thus, based on the structuring advantages of the knowledge graph, it helps to solve problems such as entity confusion, logical breaks, and ambiguous relationships. Furthermore, the efficient generation capability of the artificial intelligence big data model is superimposed, improving the generation quality, credibility, and efficiency of the evaluation data.

[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is an exemplary system architecture diagram that can be applied to an embodiment of this disclosure; Figure 2 This is a flowchart of an embodiment of the method for generating evaluation data according to this disclosure; Figure 3 This is an architecture diagram of the evaluation data generation system according to this embodiment; Figure 4 This is a flowchart of the evaluation data generation process in the data query scenario according to this embodiment; Figure 5 This is a timeline diagram of the evaluation data generation process in the data query scenario of this embodiment; Figure 6 This is a flowchart of the evaluation data generation process in the document question-and-answer scenario according to this embodiment; Figure 7 This is a timeline diagram of the evaluation data generation process in the document question-and-answer scenario according to this embodiment; Figure 8 This is a schematic diagram illustrating an application scenario of the evaluation data generation method according to this embodiment; Figure 9 This is a flowchart of yet another embodiment of the method for generating evaluation data according to this disclosure; Figure 10 This is a flowchart of yet another embodiment of the method for generating evaluation data according to this disclosure; Figure 11 This is a structural diagram of an embodiment of the evaluation data generation apparatus according to the present disclosure; Figure 12 This is a schematic diagram of the structure of a computer system suitable for implementing the embodiments of the present disclosure. Detailed Implementation

[0012] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0013] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0014] Figure 1 An exemplary architecture 100 is shown that can be used to generate evaluation data according to the present disclosure.

[0015] like Figure 1 As shown, the system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. The communication connections between terminal devices 101, 102, and 103 form a network topology. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.

[0016] Terminal devices 101, 102, and 103 can be hardware or software that supports network connectivity for data interaction and processing. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices that support network connectivity, information acquisition, interaction, display, and processing functions, including but not limited to smartphones, tablets, e-book readers, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices. They can be implemented as, for example, multiple software programs or software modules to provide distributed services, or as a single software program or software module. No specific limitations are imposed here.

[0017] Server 105 can be a server that provides various services, such as acquiring native business data uploaded by users through terminal devices 101, 102, and 103, and generating evaluation data based on a knowledge graph representing the native business data using a large-scale artificial intelligence model. As an example, server 105 could be a cloud server.

[0018] It should be noted that a server can be either hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software programs or software modules (such as software programs or software modules used to provide distributed services), or as a single software program or software module. No specific limitations are made here.

[0019] It should also be noted that the evaluation data generation method provided in the embodiments of this disclosure is generally executed by a server, but the possibility of it being executed by a terminal device, or by the server and the terminal device cooperating with each other, is not excluded. Accordingly, the various parts (e.g., various units) included in the evaluation data generation apparatus can be all located in the server, all located in the terminal device, or separately located in the server and the terminal device.

[0020] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included. When the electronic devices on which the evaluation data generation method runs do not require data transmission with other electronic devices, the system architecture may only include the electronic devices (e.g., servers or terminal devices) on which the evaluation data generation method runs.

[0021] Please refer to Figure 2 , Figure 2 A flowchart illustrating a method for generating evaluation data according to an embodiment of this disclosure. (Continue to refer to...) Figure 3 The diagram illustrates the architecture of the evaluation data generation system. Process 200 includes the following steps: Step 201: Structure the original business data and generate a knowledge graph that represents the relationships between entities in the original business data.

[0022] In this embodiment, the entity executing the evaluation data generation method (e.g., Figure 1 The server in the system can be accessed remotely via wired or wireless network connection. It will retrieve the original business data from the local machine, process the original business data in a structured manner, and generate a knowledge graph that represents the relationship between entities in the original business data.

[0023] Business-native data refers to the raw data directly generated, collected, or defined by an enterprise, organization, or information system in its core operations, production, service, or management activities (e.g., transportation and government-related businesses) to record business facts, states, or processes. Essentially, it is the initial information carrier of business activities in the digital domain, unprocessed for specific analytical or model training purposes. This data may take various technical forms, carrying the original descriptions of things, rules, behaviors, and relationships within the business domain.

[0024] An entity is a business object or concept described in the native business data, possessing independent identification and / or clearly defined attributes. Each entity represents a specific thing (such as a product, a customer, or an order) or an abstract concept (such as a department, a state, or a rule) in the business domain, and is the basic unit constituting business knowledge.

[0025] The relationships between entities refer to specific semantic bonds that are explicitly defined or implicitly present in the native business data, describing how two or more entities are related or interact with each other. These relationships define structural connections (such as membership, inclusion), procedural connections (such as creation, flow), logical connections (such as constraints, dependencies), or other business semantic connections between entities.

[0026] As an example, first, the explicit declarations or definitions of business object types, their attribute composition, and inter-object relationship constraints are read and parsed from the native business data. This information typically exists in the form of metadata or configuration. Then, from the above structural definitions, all defined entity types are identified (e.g., identifying the object types "order" and "product"), and the relationship definitions connecting these entity types are extracted (e.g., extracting the relationship definition "order 'contains' product"). Finally, based on the extracted entity types and relationship definitions, a knowledge graph describing "what relationships might exist between entity types" is constructed.

[0027] As another example, natural language processing is performed on the text content of the raw business data to identify and extract named entities representing specific business things or concepts (e.g., identifying "Customer A" and "Product XYZ" from the text), as well as key descriptive information representing the attributes or states of these entities. Then, based on the identified information units, the semantics of the text are further analyzed to extract the specific relationships between these entities expressed by the text (e.g., extracting the "purchase" relationship between "Customer A" and "Product XYZ" from "Customer A purchased Product XYZ"). The extracted entities and relationships are then combined to form discrete fact triples with "(entity, relationship, entity)" or "(entity, attribute, value)" as the basic unit. Finally, the resulting triples are combined to generate a knowledge graph.

[0028] In some optional implementations of this embodiment, the execution entity performs step 201 as follows: it uses the processing method corresponding to the generation scenario of the evaluation data to structure the original business data and generate a knowledge graph representing the relationship between entities in the original business data.

[0029] The generation scenario refers to the specific evaluation objective, context, and environment served by the evaluation data generation process. This scenario is jointly defined by the core functions, interaction patterns, and business requirements handled by the model or service to be evaluated (i.e., the "model to be evaluated"). Different generation scenarios determine what kind of knowledge should be extracted from the original business data, in what form the knowledge graph should be organized, and ultimately what evaluation objectives should be generated for the evaluation data.

[0030] As an example, firstly, the core features of the generation scenario corresponding to the current task are parsed and identified. These features may be obtained through configuration input, task description analysis, or model interface definition, with the aim of clarifying the intended use and format requirements of the evaluation data. Then, based on the identification of the generation scenario features, a target strategy matching the pre-defined set of multiple candidate data processing strategies mapped to different scenario features is selected. This target strategy fully defines the specific parsing rules for the business native data in the current scenario, the identification or definition criteria for entities and relationships, and the organizational paradigm of knowledge elements. Subsequently, the input business native data is parsed and transformed strictly according to the logic and steps specified by the selected target strategy. This process systematically transforms the knowledge components in the business native data that meet the needs of the current scenario into a set of structured knowledge components with clear semantics, using "entity-relationship-entity" or "entity-attribute-value" as basic units. These knowledge components and their relationships together constitute the basic knowledge graph structure serving this specific generation scenario.

[0031] In this implementation, by identifying and applying specific processing strategies that match the specific generation scenario, targeted structured transformation of the business's native data is achieved, which improves the matching degree between the knowledge graph and the generation scenario and helps to further improve the quality of evaluation data generated based on the knowledge graph.

[0032] In some optional implementations of this embodiment, the generation scenario includes a data query scenario, where the native business data is descriptive data of the database structure. (Continue to refer to...) Figure 4 This illustrates a flowchart of the evaluation data generation process in a data query scenario. (Continue to refer to...) Figure 5 This is a timeline diagram showing the evaluation data generation process in a data query scenario.

[0033] The data query scenario, also known as the data inquiry scenario, refers to a specific context in which the generation of evaluation data aims to assess the model's ability to perform structured data retrieval and querying. In this scenario, the evaluation data aims to simulate the interactive process of users exploring and obtaining information from structured databases through formal or semi-formal query languages ​​(such as SQL (Structured Query Language)).

[0034] Database structure description data is metadata that defines the logical and physical structure of the database. It does not contain specific business data instances, but rather describes, in a formal or declarative manner, the types of data objects contained in the database (such as tables), the attribute composition of each object (such as fields), the data types and constraints of the attributes, and the relationships between different objects (such as foreign keys). For example, database structure description data is schema data.

[0035] Users upload schema data via their terminal devices, such as a schema file named "Vehicle Management.txt". The aforementioned execution entity receives and verifies the schema file format.

[0036] In this implementation, the aforementioned execution entity performs the knowledge graph construction process in the data query scenario in the following manner: The first step is to determine the relationships between entities at multiple levels in the descriptive data, generating multiple triples. Each triple represents the relationship between two entities.

[0037] The second step is to combine multiple triples to generate a knowledge graph for data query scenarios.

[0038] As an example, firstly, the descriptive data is parsed (e.g., reading CREATE TABLE statements and extracting table structure definitions) to identify multiple hierarchical structural entity types belonging to different levels of abstraction or having containment relationships, as well as the structural relationships between these entity types. These structural relationships include not only associations between entity types at the same level but also membership or reference relationships between entity types at different levels. Then, based on the parsed hierarchical entity type system and its structural relationships, a series of structural triples are generated. Each triple represents a specific structural relationship fact, with its subject and object being structural entity types from the same or different levels, and its predicate being the specific structural relationship connecting them. Finally, by integrating all generated structural triples and associating and merging the parts pointing to the same structural entity type, a unified knowledge graph network that can fully reflect the multi-level structural entities and their complex relationships in the database schema is constructed.

[0039] After obtaining the knowledge graph for the data query scenario, it can be displayed through a front-end. Users can view the knowledge graph on the front-end, perform modification operations such as deleting irrelevant nodes or adding missing relationships, and conduct graph review. The reviewed knowledge graph can be stored in a database, such as Neo4j. Neo4j is a native graph database specifically designed for the efficient storage, management, and querying of highly interconnected graph-structured data.

[0040] In this implementation, a knowledge graph is generated by automatically parsing the description data of the database structure. This makes the implicit database structure relationships explicit into a computable knowledge network, providing an accurate semantic and relational foundation for the subsequent generation of query evaluation data that fits the real data structure and complexity.

[0041] In some optional implementations of this embodiment, the execution entity performs the first step described above to generate triples in the following manner: First, it extracts entities at multiple levels and the relationships between entities at multiple levels from the descriptive data. The multiple levels include data table entities, field entities in the data table, field value entities under the fields, and constraint rule entities under the fields. Then, based on the relationships, it generates multiple triples.

[0042] Taking the transportation sector as an example, data table entities include "Vehicle Daily Table" and "Traffic Record Table"; field entities include "License Plate Number", "Traffic Route", and "Date"; field value entities include enumerated values ​​for traffic routes such as "East Second Ring Road", "West Second Ring Road", and "South Second Ring Road"; and constraint rule entities include "License Plate Number Not Null Constraint" and "Date Value Range: 2024-01-01~2024-12-31".

[0043] As an example, firstly, lexical, syntactic, and semantic analysis is performed on the descriptive data to identify and extract all defined data table entities. Each data table entity corresponds to a logically independent data set unit. For each identified data table entity, its internal structure is further analyzed to extract all field entities belonging to that table. Each field entity corresponds to a named data column in the table, and its basic data type and other attributes are recorded. Next, for each field entity, its additional definitions are analyzed to extract two specific types of entities: firstly, field value entities, which represent discrete or enumerated values ​​with specific semantics that may appear in a business instance; secondly, constraint rule entities, which represent integrity constraints or business rules (such as NOT NULL, unique, foreign key references, and range checks) applied to the field. While completing the above multi-level entity extraction, the explicitly defined structural relationships between these entities are simultaneously parsed and recorded. Subsequently, all extracted structural relationships are traversed, and a corresponding structural triple is instantiated for each relationship. This triple, in the form of "(subject entity, structural relationship, object entity)," formally represents the aforementioned multi-level entities and their associations.

[0044] In this implementation, by automatically parsing the descriptive data of the database structure, the system systematically extracts multi-granular entities and their complex relationships, from tables and fields to value domains and rules, and encodes them into unified triples. This achieves explicit, standardized, and machine-readable representation of the implicit, multi-level structural knowledge of the database, providing accurate and reliable data basis for the evaluation set generation process in data query scenarios.

[0045] In some optional implementations of this embodiment, the association relationships include the association relationships between data table entities and field entities, the association relationships between field entities and constraint rule entities, the association relationships between field entities and field value entities, and the association relationships between data table entities.

[0046] Specifically, the relationships include the containment relationship between data table entities and field entities, the "optional value" relationship between field entities and field value entities, the "constrained by" relationship between field entities and constraint rule entities, and the foreign key relationship between data table entities.

[0047] For example, the containment relationship between data table entities and field entities is "Vehicle Daily Table - Contains - License Plate Number, Route, Date", the relationship between field entities and constraint entities is "License Plate Number - Has - NOT NULL Constraint, Unique Constraint", the relationship between field entities and field values ​​is "Route - Available - East Second Ring Road, West Second Ring Road, South Second Ring Road", and the foreign key relationship between data table entities is "Vehicle Daily Table.License Plate Number - Foreign Key Association - Traffic Record Table.License Plate Number".

[0048] In this implementation, by precisely defining multi-level structural relationships such as tables, fields, constraints, value ranges, and between tables, a fine-grained knowledge network that can fully reflect the semantics of database design and business logic is constructed, thereby improving the data comprehensiveness and accuracy of the knowledge graph.

[0049] In some optional implementations of this embodiment, the execution entity performs the second step described above to generate a knowledge graph in the following manner: merging multiple triples that include the same entity, and using the data table entity as the first-level node, the field entity as the second-level node, and the constraint rule entity and the field value entity as the third-level nodes to generate a knowledge graph for the data query scenario.

[0050] As an example, first, all triples are traversed, and those containing the same entity are identified and merged. Specifically, all relation edges in all triples sharing the same entity, where that entity is the subject or object, are aggregated to a single node representing that entity. For example, for a data table entity named "Order Table", all relations in triples where "Order Table" is the subject (e.g., "Contains - Field A") or "Order Table" is the object (e.g., "Associates - Customer Table") will be connected to the "Order Table" node. Second, based on the entity type, each node is assigned a hierarchical label: all "Data Table Entities" are labeled as first-level nodes, all "Field Entities" as second-level nodes, and all "Constraint Rule Entities" and "Field Value Entities" as third-level nodes. Finally, based on the merged relation connections and hierarchical labels, the system constructs a directed graph with a clear hierarchical structure. In this knowledge graph, first-level nodes (data table entities) serve as the root or center, second-level nodes (field entities) serve as their child nodes or subordinate nodes, and are connected to the first-level nodes through relational edges; third-level nodes (constraint / value entities) further serve as child nodes of second-level nodes, and are connected to the second-level nodes through relational edges.

[0051] In this implementation, by merging discrete triples into entities and assigning them a hierarchical structure, an organized and hierarchical graph is generated. This enables the subsequent query generation process to efficiently and accurately utilize data along the semantic path of "table -> field -> constraint / value", which helps to further improve the accuracy and generation efficiency of the evaluation data.

[0052] In some optional implementations of this embodiment, the generation scenario includes a document question-and-answer scenario, where the native business data is unstructured text data. (Continue to refer to...) Figure 6 This illustrates a flowchart of the evaluation data generation process in a document question-and-answer scenario. (Continue to refer to...) Figure 7 The diagram shows a timeline of the evaluation data generation process in a document question-and-answer scenario.

[0053] Document question-answering scenarios, also known as question-answering scenarios, refer to specific contexts in which evaluation data is generated with the core objective of assessing the model's ability to perform natural language question answering based on given document content. In this scenario, the evaluation data simulates the interaction process where a user poses a natural language question about document information and requires the model to provide an accurate answer based on its understanding and reasoning of the document content.

[0054] Unstructured text data refers to text information that lacks a predefined, structured data model (such as a database table structure). Its content exists as continuous strings of natural language, lacking explicit field separators, fixed formats, or type constraints. Common forms include plain text or rich text content such as business reports, meeting minutes, explanatory documents, and communication records. In this implementation, the aforementioned execution entity performs the knowledge graph construction process in the document question-and-answer scenario as follows: First, extract triples representing the subject, predicate, and object from the data blocks of unstructured text data; Second, combine multiple triples to generate a knowledge graph in the document question-and-answer scenario.

[0055] Before parsing unstructured text data, you can perform operations such as format conversion (e.g., converting Word format to plain text format), table extraction, and text cleaning (e.g., removing invalid characters, formatting marks, and uniform encoding) to ensure the reliability and accuracy of the data.

[0056] As an example, the text is first segmented into a series of logically independent data blocks according to predefined rules (such as fixed length, semantic boundaries, or natural paragraphs). For each data block, the semantic understanding and information extraction capabilities of a large-scale artificial intelligence model are used to perform in-depth analysis of its text content. This analysis aims to identify and extract key entities explicitly mentioned or implicitly contained in the text as candidate subjects and objects, and specific semantic relationships connecting these entities as candidate predicates. Based on this analysis, a (subject, predicate, object) triple is generated for each pair of semantically related entities in the data block. This triple is a structured representation of the facts expressed in the original text segment. After processing all data blocks, a set of a large number of discrete triples is obtained. Subsequently, these triples are integrated: different string representations pointing to the same real-world entity are identified and merged, normalized into unified graph nodes; at the same time, predicates expressing the same or similar semantic relationships are classified and standardized. Finally, based on the normalized nodes and standardized relation edges, all triples are spliced ​​together into an interconnected and unified semantic network, namely a knowledge graph in the document question-answering scenario.

[0057] Similar to data query scenarios, in document question-and-answer scenarios, the knowledge graph can be displayed on the front end. Users can view the knowledge graph on the front end, perform modification operations such as deleting irrelevant nodes or adding missing relationships, and review the graph. The reviewed knowledge graph can be stored in a database, such as Neo4j.

[0058] In this implementation, semantic triples are automatically extracted and standardized from unstructured text and integrated into a unified semantic network, thereby realizing the explicit and structured representation of implicit knowledge in the text and laying a solid knowledge foundation for generating semantically accurate question-answering evaluation data.

[0059] In some optional implementations of this embodiment, before performing the first step to obtain the triples, the execution entity may also perform the following operations: First, determine the partitioning method of the unstructured text data based on the data capacity of the unstructured text data; then, partition the unstructured text data using the partitioning method to obtain multiple data blocks.

[0060] As an example, for received unstructured text data, its data volume is first analyzed, such as calculating its total number of characters. Based on the comparison between this volume and a preset threshold, the system dynamically determines the specific segmentation method. If the data volume is less than the first threshold (e.g., 10K characters), it is segmented by natural paragraphs to ensure that each data block is a semantically complete paragraph. If the data volume is between the first threshold and a higher second threshold (e.g., 50K characters), it is preferentially segmented by document chapter titles, with paragraph segmentation supplemented within each chapter to maintain the document's macro-structure. If the data volume exceeds the second threshold, a fixed-length sliding window is used for segmentation. The window length is set to ensure that the block content does not exceed the upper limit of tokens (e.g., 2000 tokens) for a single processing of the large language model, and a certain overlap area (e.g., 200 tokens) is set between consecutive windows to avoid being split in the middle of sentences or key semantic units. Regardless of the segmentation method used, the corresponding segmentation algorithm is executed to convert the original unstructured text data into a series of data blocks that meet processing constraints and maintain semantic coherence as much as possible.

[0061] In this implementation, the most suitable partitioning strategy is dynamically selected based on the data capacity. This ensures that each data block meets the processing limitations of a large model while maximizing the maintenance of the semantic integrity and structure of the text, providing an optimized input foundation for subsequent high-quality knowledge triple extraction.

[0062] Step 202: Using a large artificial intelligence model and based on the knowledge graph, generate evaluation data for the model to be evaluated.

[0063] In this embodiment, the aforementioned execution entity generates evaluation data for the model to be evaluated based on a knowledge graph using a large artificial intelligence model.

[0064] Large-scale AI models (or simply large models) specifically refer to a type of large-scale parameterized AI model (e.g., deep learning models with hundreds of billions or even trillions of parameters, such as large language models based on the Transformer architecture) that are pre-trained on massive amounts of data and possess powerful semantic understanding, logical reasoning, and content generation capabilities. It can receive complex prompts containing task instructions, background knowledge (such as structured input from knowledge graphs), and generation requirements, and based on these, autonomously reason to output natural language text or other structured data formats that conform to complex semantic and logical constraints, such as large language models and multimodal large models. The model to be evaluated refers to the model whose performance is assessed using the generated evaluation data, such as large language models, multimodal large models, or other neural network models.

[0065] As an example, firstly, the knowledge graph or its key components, in the form of structured data (such as lists of triples, graph description text, or specific serialization formats), are combined with predefined generation instructions (i.e., prompts) that specify the evaluation data format, style, difficulty, and coverage requirements. These together constitute the input to the large model. Then, based on its pre-trained general knowledge and deep understanding of the specific knowledge graph content driven by the input instructions, the large model performs reasoning and creation, directly outputting evaluation data (such as natural language questions and their corresponding standard answers or query statements) that conform to the instruction requirements and are in the correct format.

[0066] As another example, firstly, based on the structure or content of the knowledge graph, a set of "prototypes" or "materials" for evaluation data are initially generated through rules or algorithms (e.g., a core question skeleton reflecting the relationships within the graph, a query logic framework, or key answer points). Then, these prototypes / materials, along with transformation instructions designed to control the final output form and quality (e.g., requiring the skeleton to be expanded into fluent and natural sentences, the logic framework to be instantiated into specific queries, or the key answer points to be professionally expressed), are input into the large model. Following these instructions, the large model, leveraging its powerful natural language understanding, logical analysis capabilities, and domain knowledge, refines, polishes, instantiates, or transforms the format of these input prototypes / materials, ultimately producing high-quality evaluation data that aligns with the business scenario.

[0067] In some optional implementations of this embodiment, the execution entity performs step 202 as follows to generate evaluation data for the data query scenario: The first step is to group the field entities in the description data according to business scenarios, resulting in multiple business groups. The second step is to combine these business groups based on their business relevance, resulting in business group combinations. The third step is to generate evaluation data, including evaluation requests and structured query statements, based on the relationship chains in the knowledge graph corresponding to the business group combinations.

[0068] As an example, firstly, based on the business meaning of the data table entity to which each field entity belongs and the semantics of the field entity name itself, field entities belonging to the same or closely related business topics are classified into the same business group, thus obtaining multiple business groups. Each business group contains several field entities from one or more data table entities.

[0069] Secondly, the semantic relevance between these business groups is analyzed. Specifically, the identifiers of each business group (such as names generated based on the tables to which its included fields belong and the core themes) are extracted, and natural language processing techniques are used to analyze the degree of co-occurrence, similarity, or logical association of these identifiers in terms of vocabulary and semantics. Based on this, the relevance between business groups is calculated.

[0070] Subsequently, based on a preset relevance threshold or strategy (such as combining the N groups with the highest relevance), multiple business groups with strong semantic relationships are aggregated into a single business group combination. Finally, in the knowledge graph, a complete relationship chain connecting all data table entities and field entities involved within this business group combination is located and extracted. Based on the semantic relationships and path structure contained in this relationship chain, an evaluation data pair is generated, containing an evaluation request with a natural language description (corresponding to the business semantics of the relationship chain) and its corresponding executable structured query statement (corresponding to the database implementation logic of the relationship chain).

[0071] Neo4j graph database allows for specific queries on a constructed structured knowledge graph. Once a business group consisting of several semantically related field entities is identified, a graph traversal query is initiated in the Neo4j database storing the knowledge graph of the data query scenario. This query uses the field entities contained in the business group as the starting or target nodes, exploring and identifying all possible paths connecting these field entities through predefined relational edges such as "contains," "associates," and "constrained by" in the knowledge graph. These paths fully reveal how fields are interconnected through table structure, foreign key references, or business logic, forming one or more field relationship chains from the starting point to the ending point.

[0072] To improve the rationality and accuracy of business group combinations, combination screening rules can be used. These rules, implemented during the business group combination generation phase, are quality control logic used to determine whether multiple business groups can constitute a meaningful and measurable business group combination. Combination screening rules include, but are not limited to, the following aspects: Semantic relevance requires that the names or topics of the business groups within the combination be highly related in natural language semantics, forming a coherent, higher-level business issue (for example, "order information" and "customer information" can be combined into "customer order analysis").

[0073] Business rationality: The combination must conform to industry common sense and actual business processes (for example, the combination of "inventory quantity" and "purchase order" is reasonable, while the combination of "inventory quantity" and "weather data" is considered invalid).

[0074] Logical consistency is required; the implicit attributes or definitions of business groups within a combination must not contain inherent contradictions (for example, combining a business group for "fuel vehicles" with a business group for "charging records" would create a logical conflict).

[0075] By applying this rule, invalid and absurd business combinations can be automatically filtered out, retaining only valid business combinations that are semantically coherent, business-reasonable, and logically consistent. This ensures that the subsequently generated evaluation data (queries or Q&A) has high-quality business authenticity and logical reliability.

[0076] After the evaluation data generation process is completed, all evaluation data can be stored in a specified file, such as an Excel file. Single-round evaluation data and multi-round evaluation data obtained in subsequent embodiments are stored separately, and hardware execution files are stored in a database, such as a MySQL database.

[0077] In this implementation, by dynamically grouping and associating business fields based on semantic relevance, evaluation data that reflects real and complex cross-table query needs can be automatically constructed, significantly improving the breadth of coverage and logical authenticity of the evaluation data for actual business scenarios.

[0078] In some optional implementations of this embodiment, the evaluation data includes single-round evaluation data. Single-round evaluation data refers to the smallest evaluation unit consisting of an independent evaluation request and its corresponding structured query statement. Each unit simulates a complete, context-independent user interaction to evaluate the model's ability to handle isolated requests.

[0079] In this implementation, the aforementioned execution entity performs the third step described above in the following manner to generate evaluation data for the data query scenario: based on the relationship chain corresponding to the business group combination in the knowledge graph, the evaluation samples, and the query types that the single-round evaluation data needs to cover, it generates single-round evaluation data including evaluation requests and structured query statements.

[0080] First, load one or more evaluation samples, each containing a natural language question and its corresponding SQL statement. Analyze the language style, sentence structure, and core intent expression of the natural language questions in these samples.

[0081] Subsequently, based on the data entities and business semantics covered by the obtained relationship chain, and referring to the style analysis results of the evaluation samples, a set (n) of natural language evaluation requests similar to the evaluation samples in terms of question style are generated. Then, a predefined list of query types is obtained, which enumerates various query logic categories that need to be covered (such as single attribute, multiple attribute, interval time, sorting query, negation, time inference, maximum / minimum value, mean, percentage, comparison, distribution, trend, etc.). Based on the query logic supported by the current relationship chain, for each query type in the list, one or more natural language evaluation requests conforming to the standard pattern of that type and their corresponding structured query statements are generated. Finally, the system integrates the two parts (styled requests and standard typed requests) of the generated "evaluation request-SQL statement" pairs to form the final single-round evaluation dataset.

[0082] For example, the evaluation request in a single round of evaluation data is "query the license plate numbers of trucks that passed through the East Second Ring Road today". The structured query statement is "SELECT v.license plate number FROM vehicle daily table v JOIN passage record t ON v.ID=t.vehicle ID WHERE t.route='East Second Ring Road' AND v.vehicle type='truck' AND t.date=CURDATE()".

[0083] In this implementation, by introducing evaluation samples to guide stylized generation and full coverage of standard query types, the generated single-round evaluation data has both the ability to imitate real user queries and complete coverage of various query logics, thereby effectively evaluating the model's performance under different expression styles and query intentions.

[0084] In some optional implementations of this embodiment, the execution entity performs the single-round evaluation data generation process in the following manner: First, it determines the chain type of the relationship chain, wherein the chain type includes a single-table chain that represents entities in the relationship chain being covered by a data table and a multi-table chain that is covered by a combination of multiple data tables; then, it generates single-round evaluation data according to the relationship chain, evaluation sample, query type and the evaluation data generation rules corresponding to the chain type.

[0085] First, analyze all entities involved in the relationship chain, identifying their respective table entities by querying the knowledge graph or data tables. Based on the uniqueness of these table entities, determine the chain type as either a single-table chain (involving only one table) or a multi-table chain (involving multiple tables). Next, based on the determined chain type, activate the corresponding evaluation data generation rule set. The core difference between these rule sets lies in the structural constraints on the generated structured query statements. For example, the single-table chain rule set prohibits generating any cross-table join operations, while the multi-table chain rule set requires or allows the generation of cross-table operations containing JOINs or subqueries.

[0086] Then, under the constraints of the selected rule set, the specific generation of evaluation data is performed: on the one hand, the expression style and sentence structure of the natural language questions in the evaluation samples are analyzed, and based on the semantic content of the current relationship chain, several natural language evaluation requests that imitate the style of the samples are generated; on the other hand, the system traverses the query type list (such as sorting, aggregation, multi-condition filtering, etc.), and for each type supported by the current relationship chain, generates a natural language evaluation request that conforms to the standard pattern of that type. For each natural language evaluation request generated above, the system constructs a structured query statement that precisely matches the specific fields and tables mapped by the relationship chain and strictly adheres to the SQL constraints of the current chain type rule set. Finally, all the generated "evaluation request-SQL statement" pairs together constitute the single-round evaluation data.

[0087] In this implementation, by combining the semantics of the relational chain, the style of the evaluation samples, the logic of the standard query type, and the SQL structural constraints of the chain type, evaluation data with semantic accuracy, expressive diversity, logical completeness, and syntactic rationality can be generated.

[0088] During the generation of single-round evaluation data, the large AI model, based on the evaluation data generation rules corresponding to the relationship chain, evaluation examples, query type, and chain type as input, can also combine time type, complexity control, and the naturalness requirements of the evaluation request expression to generate single-round evaluation data. Time type: includes dynamic time (such as today, yesterday, this week) and fixed time (such as June, 2024).

[0089] In some optional implementations of this embodiment, the evaluation data includes multi-turn evaluation data. Multi-turn evaluation data in a data query scenario refers to a sequence of evaluation units consisting of a series of semantically or logically coherent and context-dependent evaluation requests and their corresponding SQL statements. This sequence simulates a multi-turn continuous dialogue, where the semantic understanding and correct processing of subsequent requests depend on the context of previous interactions, and is used to evaluate the model's dialogue state management, context understanding, and cross-turn reasoning capabilities.

[0090] In this implementation, the aforementioned execution entity can generate multi-round evaluation data for document question-and-answer scenarios in the following way: based on the relationship chain corresponding to the business group combination in the knowledge graph and the multi-round query strategy, multi-round evaluation data including evaluation requests and structured query statements is generated.

[0091] Multi-round query strategies represent the query logic between multiple rounds of evaluation requests, and one or more multi-round query strategies can be flexibly set according to the actual situation.

[0092] As an example, firstly, seed single-round evaluation data is generated using the single-round evaluation data generation method described in the above implementation. Then, based on the logical evolution paradigm defined by the selected multi-round query strategy, multiple evaluation requests are selected or combined from the seed single-round evaluation data, and contextual dependencies between them are constructed. For example, under a stepwise filtering strategy, a broad full query request is first generated. Then, based on the structure or semantics of the query results, one or more subsequent requests are automatically generated, progressively adding specific filtering conditions to the original query conditions. The SQL statement of each subsequent request contains the constraint logic of the preceding request. Under a comparative analysis strategy, an overall statistical query is first generated. Subsequent requests are then built around comparing, drilling down, or delving deeper into different subsets or dimensions, with their SQL statements built on the result set or grouping of the preceding query. Under an aggregate statistical strategy, a request sequence is generated, progressing from overall statistics to detailed statistics by different dimensions, and then to cross-analysis. Finally, this series of evaluation requests and SQL statement pairs, which are semantically and logically coherent and dependent, are sequentially combined into a multi-round evaluation data instance.

[0093] As another example, firstly, following the generation method of single-round evaluation data, a first-round evaluation data is generated as the starting point, containing the first natural language evaluation request and its corresponding structured query statement (SQL1). Subsequently, the system analyzes the request semantics in the first-round evaluation data and the query logic and potential result set of SQL1 according to the logical evolution rules defined by the selected multi-round query strategy. Based on this analysis, subsequent rounds of evaluation data are directly derived and generated. For example, under a stepwise filtering strategy, the system analyzes the query conditions of SQL1, generates a subsequent natural language request that semantically requests "further narrowing the scope," and generates a new SQL statement (SQL2) that adds one or more filtering conditions to the WHERE clause of SQL1. Under a comparative analysis strategy, based on the aggregated query results of SQL1, a subsequent question requesting "comparing results from different dimensions" is generated, and SQL2 is generated to implement the comparison logic by adding a GROUPBY clause or splitting the query conditions based on SQL1. Under an aggregated statistical strategy, the system may generate subsequent questions requesting "drill-down analysis" based on the overall statistical results of SQL1, and generate SQL2 that introduces more granular grouping or cross-calculation. This process can be iterated to generate more rounds. Finally, this series of evaluation data generated round by round based on the output of the previous round is sequentially combined into a logically coherent multi-round evaluation data instance.

[0094] Examples of multi-round evaluation data are as follows: Round 1: Query (evaluation request): "How many operational vehicles experienced breakdowns today?" SQL (Structured Query Language): SELECT COUNT(DISTINCT CLID) FROM dwd_clzt_core_da WHERE RQ = CURDATE() AND CLYT = 'Operating Vehicle' AND GZLC_NUM IS NOT NULL AND GZLC_NUM > 0; Round 2: Query: "What is the total number of breakdowns for these vehicles?" SQL: SELECT SUM(GZLC_NUM) FROM dwd_clzt_core_da WHERE RQ = CURDATE()AND CLYT = 'Operation Car' AND GZLC_NUM IS NOT NULL AND GZLC_NUM > 0; For example, the multi-round evaluation data under the stepwise filtering strategy is: Round 1 "Query orders" (all orders) → Round 2 "Completed orders among these orders" (add status filter) → Round 3 "Orders with an amount exceeding 1000" (add amount filter).

[0095] For example, the multi-round evaluation data under the comparative analysis strategy is as follows: Round 1 "Statistical analysis of the traffic volume of various types of vehicles" (statistics) → Round 2 "Which is more common, trucks or buses?" (comparison) → Round 3 "Why is the traffic volume of trucks higher?" (analysis).

[0096] For example, the multi-round evaluation data under the aggregated statistical strategy is as follows: Round 1 "Total number of orders" (overall) → Round 2 "Statistics by region" (segmentation) → Round 3 "Average order amount in each region" (cross).

[0097] In this implementation, a sequence with logical evolution is generated by guiding a predefined multi-round query strategy, thereby generating evaluation data that can simulate complex and continuous data exploration scenarios, effectively evaluating the model's context preservation and progressive reasoning capabilities in interactive queries.

[0098] In some optional implementations of this embodiment, the execution entity performs step 202 as follows to generate evaluation data for the document question-and-answer scenario: The first step is to divide the multiple triples into multiple scenario groups based on the business scenarios to which they belong. The second step is to determine the priority of the business scenarios corresponding to each scenario group based on the relevance of the scenario group to the unstructured text data. The third step is to generate evaluation data, including evaluation requests and standard answers, based on the knowledge graph and the priority.

[0099] As an example, firstly, the semantic attributes of the entities in each triple and the type of relationship they represent are analyzed. Based on a set of predefined or learned business scenario classification rules, all triples are divided into different scenario groups, each corresponding to a specific business theme or activity domain. Next, the thematic relevance of each scenario group to the overall original unstructured text data is evaluated. This evaluation is based on the matching degree between the core semantics reflected by the triples within a scenario group and the global theme of the text, and a relevance quantification index is calculated for each scenario group accordingly. Based on this index, all scenario groups are ranked to determine the priority order of the business scenarios corresponding to each scenario group; high-priority business scenarios are considered the core content of the text. Finally, based on the structured knowledge provided by the knowledge graph and prioritizing high-priority business scenarios, subsequent evaluation data containing assessment requests and standard answers is generated, ensuring that the generated data is more focused on the core theme of the text.

[0100] In this implementation, by organizing and prioritizing the knowledge triples according to business scenarios, the final generated question-and-answer evaluation data can focus on the core and key content of the text, thereby improving the relevance and effectiveness of the evaluation and better assessing the model's understanding ability in key business scenarios.

[0101] In some optional implementations of this embodiment, the execution entity performs the first step described above by dividing the multiple triples into multiple scene groups in the following manner: First, candidate keywords are selected from the entities in multiple triples based on their frequency of occurrence. Then, multiple business scenarios are determined based on the candidate keywords, the domains of the multiple triples and the unstructured text data. Finally, the multiple triples are divided into multiple scenario groups based on their correlation with the multiple business scenarios.

[0102] As an example, firstly, the total frequency of each entity (including subject and object) in all triples is counted, and the triples are sorted from highest to lowest frequency. The top 20 entities are then selected as candidate keywords. Next, the list of candidate keywords, sampled triples from multiple triples, and domain information from unstructured text data are input into a large-scale AI model. Based on this input, the model performs semantic analysis and clustering reasoning to identify and output multiple fine-grained business scenarios. Then, for each triple, its relevance score to each of the multiple business scenarios is calculated. During the calculation, the subject is given a higher weight, followed by the object, and then the relation. Each triple is assigned to the scenario group corresponding to the business scenario with the highest relevance score. If the large-scale AI model cannot effectively identify a business scenario, a degradation mechanism is activated, grouping candidates based on co-occurrence relationships or simply by the domain of the entities.

[0103] In this implementation, by combining entity frequency statistics with large-scale model semantic analysis, fine-grained business scenarios in text are automatically identified and knowledge triples are accurately categorized, providing a structured knowledge organization foundation for the subsequent generation of scenario-based and targeted question-and-answer evaluation data.

[0104] In some optional implementations of this embodiment, the execution entity performs the second step described above to determine the priority of the business scenarios corresponding to each of the multiple scenario groups in the following manner: First, for the multiple scenario groups, the scenario keywords corresponding to the scenario groups are determined from the candidate keywords to obtain a set of scenario keywords, and a set of scenario entities is obtained based on the entities in the scenario groups; then, a first overlap is determined between the set of scenario keywords and the set of core keywords representing the topic of unstructured text data; then, a second overlap is determined between the set of scenario entities and the set of core keywords; then, the relevance between the business scenarios corresponding to the scenario groups and the unstructured text data is determined based on the first overlap and the second overlap; finally, the priority among the multiple business scenarios is determined based on the relevance corresponding to each of the multiple business scenarios.

[0105] As an example, firstly, for each scene group, a subset of keywords highly consistent with the semantic description of that scene group is selected from the candidate keyword list to form the scene keyword set for that scene group. Simultaneously, all subject and object entities are extracted from all triples contained in that scene group, and after deduplication, they form the scene entity set for that scene group. Furthermore, by analyzing the entire unstructured text data, its core thematic words are extracted to form a core keyword set.

[0106] Next, calculate the first overlap between the scene keyword set and the core keyword set of the scene group (e.g., calculate the number of intersection elements of the two sets). At the same time, calculate the second overlap between the scene entity set and the core keyword set of the scene group.

[0107] Then, based on a preset formula (e.g., the sum of the first and second overlaps divided by the number of keywords in the core keyword set), the relevance score between the business scenario and the unstructured text data corresponding to the scenario group is calculated. Finally, the relevance scores calculated for all business scenarios are compared, and the priority among multiple business scenarios is determined according to the order of scores from high to low. For example, after calculating the relevance score for each business scenario, the system classifies them according to preset thresholds: scenario groups with scores greater than or equal to 0.7 are classified as highly relevant, scenario groups with scores between 0.4 (inclusive) and 0.7 are classified as moderately relevant, and scenario groups with scores less than 0.4 are classified as lowly relevant. This classification result is directly used to determine the final priority order of scenarios and guides the resource allocation for scenarios of different importance when generating evaluation data subsequently.

[0108] In this implementation, the correlation between each business scenario and the core theme of the document is accurately measured by quantitative overlap calculation, and priority is established accordingly, so that the allocation of evaluation resources and the generation of question and answer data can focus on the most relevant and important content areas.

[0109] In some optional implementations of this embodiment, the evaluation data includes single-round evaluation data. In a document question-and-answer scenario, single-round evaluation data refers to the smallest evaluation unit consisting of an independent evaluation request and its corresponding standard answer.

[0110] In this implementation, the aforementioned execution entity performs the third step described above in the following manner to generate evaluation data for the document question-and-answer scenario: based on the knowledge graph, priority, the types of questions that the evaluation data needs to cover, and the preset generation strategy for the evaluation data, it generates single-round evaluation data including evaluation requests and standard answers.

[0111] Question type refers to the semantic and intent category of natural language questions (assessment requests) in the evaluation data. It is used to ensure the evaluation coverage model's ability to handle different question types, including but not limited to: factual questions, which ask about specific facts, attributes, or details; yes / no questions, which require confirmation or denial of a statement or possibility; comparison questions, which require comparison of the similarities and differences between two or more entities, concepts, or solutions; causal questions, which ask about the reasons behind a phenomenon, result, or decision; and open-ended questions, which require comprehensive suggestions, solutions, or analyses and usually do not have a single standard answer.

[0112] Preset generation strategies are rules or patterns that guide large AI models on how to construct specific question formats based on knowledge graph content. These include, but are not limited to: conditional judgment type, guiding the model to generate questions in the form of "If XX condition, how should YY be done?"; comparison type, guiding the model to identify multiple parallel entities in the knowledge graph and generate comparative questions; global overview type, guiding the model to integrate three or more related entities in the knowledge graph and generate comprehensive, global questions; causal analysis type, guiding the model to identify causal links in the knowledge graph and generate questions asking for reasons; spatiotemporal condition type, guiding the model to combine specific time, location, and other scene elements to generate questions with limited conditions; multi-entity association type, mandating that the generated questions must contain two or more related entities to avoid question fragmentation; and element merging type, guiding the model to automatically merge multiple elements such as location, time, and materials into one question.

[0113] As an example, firstly, based on priority, high-priority business scenarios are selected as the focus. For each selected business scenario, relevant entities and relationships are extracted from its corresponding knowledge graph subgraph as generation materials. Then, the list of question types is traversed, and for each question type, one or more strategies suitable for that question type are selected from the preset generation strategy set (e.g., selecting the "causal analysis" strategy for "cause-based" questions). Subsequently, the knowledge graph materials of the current business scenario, the target question type, the selected generation strategy, and the generation instructions are combined to form a prompt message, which is then input into the AI ​​model. The AI ​​model uses this prompt message to reason and create, generating a natural language assessment request that meets the requirements of the specified question type and strategy, and generating the corresponding standard answer based on the provided knowledge graph materials. This process is iterated, covering different business scenarios (in priority order) and different question types, ultimately summarizing to generate a set of single-round evaluation data.

[0114] During the generation of evaluation data based on a large model, quality control rules used to constrain and guide the generation of "assessment requests" (i.e., questions) and "standard answers" can be input into the large model as part of the prompt information. This aims to ensure that the generated evaluation data is relevant to real-world scenarios, specific and clear, natural and fluent, and that the answers are accurate. Quality control rules include, but are not limited to: The scenario description requires that each generated question must include clear background information, such as time, location, triggering conditions, or key stakeholders. This ensures that the questions simulate real needs in specific situations, rather than being general questions.

[0115] Specific anchor binding requires that the question must explicitly mention at least one specific entity, action, location, or time. This prevents the question from being too vague or abstract, forcing the question to establish a direct association with specific nodes (facts) in the knowledge graph.

[0116] Meta-level questions are prohibited, as are statements in questions that refer to the document itself rather than its content, such as "The document mentions..." or "This document states...". This ensures that questions concern the knowledge described in the document, conforming to the questioning habits of real users (who would not knowingly be reading a document).

[0117] Templated expressions are prohibited. No unreplaced placeholders or obvious template markers should be retained in the generated questions or answers. This ensures that the final output is natural, complete, and ready for evaluation.

[0118] The answers must come from the original text. The standard answers generated for the questions must be strictly based on and faithful to the original text content, without any interpretation, extension, or fabrication. This ensures the objectivity and verifiability of the evaluation criteria.

[0119] In this implementation, by combining knowledge graph content, scenario priority, diverse question types, and structured generation strategies, the large model is guided to generate high-quality single-round question-and-answer evaluation data that is comprehensive, highlights key points, and is diverse in form, thereby achieving a refined and multi-dimensional evaluation of the model's comprehensive question-and-answer capabilities.

[0120] In some optional implementations of this embodiment, the above-mentioned execution entity generates single-round evaluation data in the document question-and-answer scenario in the following manner: First, based on the knowledge graph, priority, question type and preset generation strategy, multiple initial evaluation data including evaluation requests and standard answers are generated.

[0121] In this implementation, the aforementioned execution entity can use the above example to generate initial evaluation data based on the knowledge graph, priority, question type, and preset generation strategy, which will not be elaborated further here.

[0122] Then, integrity checks and semantic deduplication are performed on multiple initial evaluation data to obtain multiple filtered evaluation data. Finally, the quality of the multiple filtered evaluation data is evaluated from multiple preset evaluation dimensions, and the single-round evaluation data is determined from the multiple filtered evaluation data based on the evaluation results.

[0123] As an example, completeness checks and semantic deduplication operations are performed sequentially on this batch of initial evaluation data.

[0124] The integrity verification process specifically includes: using a large-scale artificial intelligence model to perform batch analysis on the assessment requests in each initial assessment data set to detect any unclear or ambiguous statements; if detected, attempting to rewrite them using suggested keywords or specific themes; if rewriting is not possible, deleting the assessment data set. Simultaneously, a routine check is performed on each initial assessment data set to examine for unreplaced template tags, missing clear subjects in the assessment requests, overly general statements, inappropriate citation formats, meaningless questions and answers, and anchor point verification, i.e., verifying whether the assessment requests or standard answers contain at least one scene entity or relationship extracted from the knowledge graph.

[0125] The semantic deduplication operation is as follows: using a pre-trained model under the sentence-transformers framework, the text semantic embedding vector of the evaluation request in all evaluation data that passed the integrity check is calculated. The semantic repetition problem is identified by comparing the similarity between the vectors, and the evaluation data corresponding to the identified repetition problem is deleted to obtain multiple filtered evaluation data.

[0126] To further illustrate the integrity check and semantic deduplication operations, the following example is provided: Example 1, Clean up template tags: Input: "Group 1: How to detour during construction?" Output: "How to detour during construction?" Note: Remove template prefixes such as "Group 1:" Example 2, General Problem Check: Input: "What is transportation?" Result: Discarded (The problem was too general, as is well known). Note: Generating common-sense questions is prohibited. Example 3, Article number reference check: Input: "What is the relationship between Articles 115 and 116?" Result: Discarded (only the article number is cited, but the specific content is missing). Example 4, Ambiguity Detection and Modification: Input: "What are some alternative routes?" (Missing context) Test results: Unclear. Suggested keywords: "during construction". Output: "What detour routes are available during construction?" Note: Ambiguity testing is used to supplement the context, making the problem clearer. Next, the quality of this batch of screened evaluation data is assessed from multiple preset evaluation dimensions. These preset evaluation dimensions include, but are not limited to: 1. Core Keyword Coverage (0-40 points): This dimension quantifies the relevance of the evaluation data to the core theme of the document. It begins with a set of core keywords extracted from unstructured text data. During evaluation, for each evaluation data point (including the evaluation request and standard answer), two parallel methods are used to calculate the number of matches with the core keywords: one is direct substring matching, checking whether each core keyword appears as a substring in the evaluation text; the other is word segmentation matching, first extracting the keywords from the evaluation data, then calculating the number of intersections between them and the core keyword set. The maximum match count obtained from the two methods is taken. The scoring formula is: Score = min((Number of Matches / Total Number of Core Keywords)) 40,40). To ensure basic quality, a base score (e.g., 10 points) will be assigned even if the number of matches is zero.

[0127] 2. User Core Question Similarity (0-30 points): This dimension assesses how closely the generated assessment request resembles the core questions that real users might ask. A list of user core questions is pre-set (e.g., "How to navigate around a car taller than 2.3 meters?"). During assessment, keywords are extracted from both the assessment request and each user core question, and the intersection ratio (similarity) of their keyword sets is calculated. The maximum similarity score among all user core questions is taken. The scoring formula is: Score = Maximum Similarity. 30. For example, if the maximum similarity is 0.85, the score is 25.5 points.

[0128] 3. Knowledge Link Length (0-15 points): This dimension measures the complexity and depth of the knowledge examined in the evaluation data, represented by the number of nodes in the underlying knowledge graph relationship chain. It traces the knowledge link on which the evaluation data is based (e.g., "vehicle height restriction -> underpass auxiliary road -> detour route -> East Second Ring Road"). The scoring method is: count the number of nodes in the link, score = min(number of nodes). (3,15). The more nodes there are, the more entities and relationships the problem involves, resulting in higher complexity.

[0129] 4. Question Type Weighting (0-10 points): This dimension assigns different basic weights based on the question type of the assessment request, aiming to encourage the generation of more complex questions with greater reasoning value. A pre-defined weight mapping table is used, for example: conditional judgment (10 points), comparison (9 points), causal analysis (9 points), global integration (8 points), factual (7 points), open-ended (6 points), yes / no (6 points), etc. During evaluation, the corresponding score is obtained directly based on the question type determined for the assessment request.

[0130] 5. Answer Completeness (0-5 points): This dimension assesses the sufficiency and completeness of the standard answer's reproduction of the original text information. The generated answer will be compared with the knowledge graph and the original text fragment to check whether the answer covers all the key points asked in the question and whether there are any omissions or incomplete expressions. Based on the review results, a score between 0 and 5 will be assigned.

[0131] 6. Conditional Reasoning Bonus (0-5 points): This dimension is a reward metric used to identify and reward assessment requests that contain conditional logical reasoning. If an assessment request is found to contain obvious conditional assumption structures (e.g., using phrases like "if...then..." or "when...how should..."), an additional fixed score (e.g., 5 points) will be added to the scores of other dimensions.

[0132] 7. Core Scene Recognition (0-20 points): This dimension evaluates the fit between the evaluation data and high-priority core scenes. Based on the scene priority ranking, high-priority (core) scenes and their corresponding core scene keywords with TF-IDF (Term Frequency – Inverse Document Frequency) weights are obtained. During evaluation, it is checked whether the evaluation data text contains these core scene keywords, and the sum of the TF-IDF weights of the matched keywords is calculated (or the maximum weight is taken). The scoring formula is: Score = max(TF-IDF weight of matched keywords) 20). This ensures that data most closely related to the most important scenarios receives higher ratings.

[0133] After independently scoring the seven dimensions mentioned above, a total score is calculated for each evaluation data point. A base score guarantee mechanism is in place (e.g., a total score of no less than 30 points). Subsequently, based on preset threshold ranges, the evaluation data are divided into different importance levels, for example: a total score ≥50 points is "Core," ≥35 points is "Important," ≥25 points is "General," and <25 points is "Supplementary." Finally, the evaluation data set for the final evaluation is determined based on the assessment results (e.g., ranking by level and score).

[0134] In this implementation, a multi-stage automated quality inspection and multi-dimensional quantitative evaluation screening pipeline is introduced to systematically eliminate low-quality, ambiguous, and repetitive generated results, and select high-quality data based on complex rules closely related to the evaluation objectives, thereby ensuring the high reliability, diversity, and evaluation relevance of the final evaluation dataset.

[0135] In some optional implementations of this embodiment, the evaluation data includes multi-round evaluation data. Multi-round evaluation data in a document question-answering scenario refers to a sequence of evaluation units consisting of a series of semantically or logically coherent and context-dependent evaluation requests and their corresponding answers.

[0136] In this implementation, the aforementioned execution entity can perform the third step described above to generate multi-round evaluation data for a document question-and-answer scenario as follows: Based on the knowledge graph, priority, and multi-round question-and-answer strategy, multi-round evaluation data including evaluation requests and standard answers is generated. The multi-round question-and-answer strategy represents the question-and-answer logic between multi-round evaluation requests.

[0137] As an example, firstly, based on priority, a high-priority business scenario is selected as the initial focus. Then, based on the knowledge graph subgraph within that scenario, a first-round evaluation data set containing the evaluation request and standard answer is generated. Specifically, the first-round evaluation data can also be generated by referring to the single-round evaluation data generation process described above. Subsequently, based on the logical evolution paradigm defined by the selected multi-round question-answering strategy, the evaluation requests, standard answers, and related knowledge graph content in the first-round evaluation data are analyzed. Using this as context, subsequent rounds of evaluation requests are generated. For example, under a progressively deeper strategy, a follow-up question is generated based on the first-round overview, delving into a specific detail of the previous round's answer; under a comparative extension strategy, a follow-up question might be generated based on the first-round comparison results, asking for the reasons for the comparison differences or making a selection accordingly; under a scenario expansion strategy, a subsequent question is generated based on the scenario discussed in the first round, introducing another related scenario and requiring comprehensive analysis. For each subsequent round of evaluation requests, the system generates a corresponding standard answer based on the updated context and knowledge graph. This process can be iterated to generate multi-round evaluation data containing multiple rounds of questions and answers with coherent logic.

[0138] Multi-round evaluation data in a document question-and-answer scenario is as follows: Under the progressive strategy, round 1 "How to pass through during construction?" (overview) → round 2 "What are the restrictions on the East Second Ring Road?" (details) → round 3 "What about oversized vehicles?" (condition judgment).

[0139] Under the comparative extension strategy, round 1 "What are the differences between East Second Ring Road and Xianning West Road?" (comparison) → round 2 "Why are different restrictions applied?" (cause and effect) → round 3 "Which road should I choose during the morning rush hour?" (choice).

[0140] Under the scenario-expanding strategy, round 1 "How to detour during construction?" (Scenario 1) → round 2 "What about during non-construction periods?" (Scenario 2) → round 3 "What are the common requirements?" (Comprehensive).

[0141] During the generation of multi-round evaluation data, repeated information in subsequent questions can be replaced with pronouns. For example, the original question "What are the restrictions on the East Second Ring Road among these detour routes?" can be rewritten as "What are the restrictions on the East Second Ring Road?" because the "detour routes" have been clearly defined in the previous round.

[0142] In this implementation, a predefined multi-turn question-and-answer strategy guides the natural evolution of the dialogue logic, generating evaluation data that simulates real complex interactions. This enables the effective evaluation of the model's contextual understanding, state maintenance, and progressive reasoning capabilities in continuous dialogue.

[0143] See also Figure 8 , Figure 8This is a schematic diagram of an application scenario 800 of the evaluation data generation method according to this embodiment. Target user 801 sends business-native data 804 to server 803 via terminal device 802. The server first performs structured processing on the business-native data, generating a knowledge graph 805 representing the relationships between entities in the business-native data; then, using an artificial intelligence big data model 806, it generates evaluation data for the model to be evaluated based on the knowledge graph. Evaluation data may include, for example, "Q: "Query the license plate numbers of trucks passing through the East Second Ring Road today"; SQL: SELECT v.license plate number FROM vehicle_day_table v JOIN passage_record_t ON v.ID=t.vehicle_ID WHERE t.route='East Second Ring Road' AND v.type_of_truck' AND t.date=CURDATE().

[0144] This embodiment provides a method for generating evaluation data. By structuring the original business data, a knowledge graph representing the relationships between entities in the original business data is generated. Based on the knowledge graph, evaluation data for the model to be evaluated is generated using an artificial intelligence big data model. Thus, based on the structuring advantages of the knowledge graph, it helps to solve problems such as entity confusion, logical breaks, and ambiguous relationships. Furthermore, the efficient generation capability of the artificial intelligence big data model is combined to improve the quality, credibility, and efficiency of the generated evaluation data.

[0145] Continue to refer to Figure 9 This illustrates an illustrative flow 900 of yet another embodiment of the evaluation data generation method according to the present disclosure. In flow 900, the evaluation data generation process in a data query scenario includes the following steps: Step 901: Extract entities at multiple levels and the relationships between entities at multiple levels from the description data of the database structure.

[0146] This includes multiple levels such as data table entities, field entities within data tables, field value entities under fields, and constraint rule entities under fields. Relationships include those between data table entities and field entities, between field entities and constraint rule entities, between field entities and field value entities, and between data table entities themselves.

[0147] Step 902: Generate multiple triples based on the association relationships.

[0148] Step 903: Merge triples that contain the same entity from multiple triples, and generate a knowledge graph for the data query scenario with the data table entity as the first-level node, the field entity as the second-level node, and the constraint rule entity and field value entity as the third-level nodes.

[0149] Step 904: Group the field entities in the description data according to the business scenario to obtain multiple business groups.

[0150] Step 905: Based on the business relevance between multiple business groups, combine multiple business groups to obtain a business group combination.

[0151] Step 906: Based on the relationship chains corresponding to the business group combinations in the knowledge graph, the evaluation samples, the query types that the single-round evaluation data needs to cover, and the evaluation data generation rules corresponding to the chain types of the relationship chains, generate single-round evaluation data including evaluation requests and structured query statements.

[0152] Chain types include single-table chains, which represent entities in a relational chain that are covered by a single data table, and multi-table chains, which are covered by a combination of multiple data tables.

[0153] Step 907: Generate multi-round evaluation data, including evaluation requests and structured query statements, based on the multi-round query strategy.

[0154] Multi-round query strategy represents the query logic between multiple rounds of evaluation requests.

[0155] The evaluation data generation method in this embodiment, process 900, compared to process 200 above, specifically describes the evaluation data generation process in a data query scenario. Based on the structured advantages of knowledge graphs, it helps to solve problems such as entity confusion, logical breaks, and ambiguous relationships. Furthermore, it combines the efficient generation capabilities of large artificial intelligence models, thereby improving the generation quality, credibility, and efficiency of evaluation data in a data query scenario.

[0156] Continue to refer to Figure 10 The illustration shows a schematic flow 1000 of another embodiment of the evaluation data generation method according to the present disclosure. In flow 1000, the evaluation data generation process in a document question-and-answer scenario includes the following steps: Step 1001: Extract triples representing the subject, predicate, and object from the data blocks of unstructured text data.

[0157] Step 1002: Combine multiple triples to generate a knowledge graph for the document question-and-answer scenario.

[0158] Step 1003: Divide the multiple triples into multiple scenario groups according to the business scenarios to which they belong.

[0159] Step 1004: For multiple scenario groups, determine the priority of the business scenarios corresponding to each scenario group based on the correlation between the scenario group and the unstructured text data.

[0160] Step 1005: Based on the knowledge graph, priority, the types of questions that the evaluation data needs to cover, and the preset generation strategy for the evaluation data, generate single-round evaluation data including evaluation requests and standard answers.

[0161] Step 1006: Generate multi-round evaluation data, including assessment requests and standard answers, based on the knowledge graph, priorities, and multi-round question-answering strategy.

[0162] Among them, the multi-turn question-answering strategy represents the question-answering logic between multiple rounds of evaluation requests.

[0163] The evaluation data generation method in this embodiment, process 1000, compared to process 200 above, specifically describes the evaluation data generation process in the document question-and-answer scenario. Based on the structured advantages of knowledge graphs, it helps to solve problems such as entity confusion, logical breaks, and ambiguous relationships. Furthermore, it combines the efficient generation capabilities of large artificial intelligence models, thereby improving the generation quality, credibility, and efficiency of evaluation data in the document question-and-answer scenario.

[0164] Continue to refer to Figure 11 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of an evaluation data generation apparatus, which is similar to... Figure 2 Corresponding to the method embodiments shown, the system can be specifically applied to various electronic devices.

[0165] like Figure 11 As shown, the evaluation data generation device 1100 includes: a graph generation unit 1101, configured to process business native data in a structured manner and generate a knowledge graph representing the relationships between entities in the business native data; and a data generation unit 1102, configured to generate evaluation data for the model to be evaluated based on the knowledge graph using an artificial intelligence big data model.

[0166] In some optional implementations of this embodiment, the graph generation unit 1101 is further configured to: use the processing method corresponding to the generation scenario of the evaluation data to structure the original business data and generate a knowledge graph representing the relationship between entities in the original business data.

[0167] In some optional implementations of this embodiment, the generation scenario includes a data query scenario, the business native data is description data with a database structure, and the graph generation unit 1101 is further configured to: determine the association relationship between entities at multiple levels in the description data, generate multiple triples, wherein the triples represent the association relationship between two entities; and combine the multiple triples to generate a knowledge graph in the data query scenario.

[0168] In some optional implementations of this embodiment, the graph generation unit 1101 is further configured to: extract entities at multiple levels and the relationships between entities at multiple levels in the description data, wherein the multiple levels include data table entities, field entities in the data table, field value entities under the fields, and constraint rule entities under the fields; and generate multiple triples based on the relationships.

[0169] In some optional implementations of this embodiment, the association relationships include the association relationships between data table entities and field entities, the association relationships between field entities and constraint rule entities, the association relationships between field entities and field value entities, and the association relationships between data table entities.

[0170] In some optional implementations of this embodiment, the graph generation unit 1101 is further configured to: merge triples that include the same entity from multiple triples, and generate a knowledge graph for the data query scenario with the data table entity as the first-level node, the field entity as the second-level node, and the constraint rule entity and the field value entity as the third-level node.

[0171] In some optional implementations of this embodiment, the data generation unit 1102 is further configured to: group the field entities in the description data according to the business scenario to obtain multiple business groups; combine the multiple business groups according to the business correlation between them to obtain a business group combination; and generate evaluation data including evaluation requests and structured query statements according to the relationship chain in the knowledge graph corresponding to the business group combination.

[0172] In some optional implementations of this embodiment, the evaluation data includes single-round evaluation data, and the data generation unit 1102 is further configured to generate single-round evaluation data including evaluation requests and structured query statements based on the relationship chain corresponding to the business group combination in the knowledge graph, the evaluation samples, and the query types that the single-round evaluation data needs to cover.

[0173] In some optional implementations of this embodiment, the data generation unit 1102 is further configured to: determine the chain type of the relationship chain, wherein the chain type includes a single-table chain representing entities in the relationship chain that are covered by a data table and a multi-table chain that is covered by a combination of multiple data tables; and generate single-round evaluation data according to the relationship chain, evaluation sample, query type and the evaluation data generation rules corresponding to the chain type.

[0174] In some optional implementations of this embodiment, the evaluation data includes multi-round evaluation data, and the data generation unit 1102 is further configured to generate multi-round evaluation data including evaluation requests and structured query statements based on the relationship chain corresponding to the business group combination in the knowledge graph and the multi-round query strategy, wherein the multi-round query strategy represents the query logic between the multi-round evaluation requests.

[0175] In some optional implementations of this embodiment, the generation scenario includes a document question-and-answer scenario, the original business data is unstructured text data, and the graph generation unit 1102 is further configured to: extract triples representing the subject, predicate and object from the data blocks of unstructured text data; and combine multiple triples to generate a knowledge graph in the document question-and-answer scenario.

[0176] In some optional implementations of this embodiment, the above apparatus further includes: a data partitioning unit (not shown in the figure), configured to: determine the partitioning method of the unstructured text data according to the data capacity of the unstructured text data; and partition the unstructured text data using the partitioning method to obtain multiple data blocks.

[0177] In some optional implementations of this embodiment, the data generation unit 1102 is further configured to: divide the multiple triples into multiple scenario groups according to the business scenarios to which each triple belongs; for the multiple scenario groups, determine the priority of the business scenarios corresponding to each of the multiple scenario groups according to the correlation between the scenario groups and the unstructured text data; and generate evaluation data including evaluation requests and standard answers according to the knowledge graph and the priority.

[0178] In some optional implementations of this embodiment, the data generation unit 1102 is further configured to: filter candidate keywords from the entities in the multiple triples according to the frequency of occurrence of the entities in the multiple triples; determine multiple business scenarios according to the candidate keywords, the multiple triples and the domain to which the unstructured text data belongs; and divide the multiple triples into multiple scenario groups according to the correlation between the triples and the multiple business scenarios.

[0179] In some optional implementations of this embodiment, the data generation unit 1102 is further configured to: for multiple scenario groups, determine the scenario keywords corresponding to the scenario groups from the candidate keywords to obtain a set of scenario keywords, and obtain a set of scenario entities based on the entities in the scenario groups; determine a first degree of overlap between the set of scenario keywords and the set of core keywords representing the topic of unstructured text data; determine a second degree of overlap between the set of scenario entities and the set of core keywords; determine the relevance between the business scenarios corresponding to the scenario groups and the unstructured text data based on the first degree of overlap and the second degree of overlap; and determine the priority among the multiple business scenarios based on the relevance of each of the multiple business scenarios.

[0180] In some optional implementations of this embodiment, the evaluation data includes single-round evaluation data, and the data generation unit 1102 is further configured to generate single-round evaluation data including evaluation requests and standard answers based on the knowledge graph, priority, the question types to be covered by the evaluation data, and the preset generation strategy of the evaluation data.

[0181] In some optional implementations of this embodiment, the data generation unit 1102 is further configured to: generate multiple initial evaluation data including assessment requests and standard answers based on the knowledge graph, priority, question type and preset generation strategy; perform integrity verification and semantic deduplication operations on the multiple initial evaluation data to obtain multiple filtered evaluation data; perform quality assessment on the multiple filtered evaluation data from multiple preset evaluation dimensions, and determine single-round evaluation data from the multiple filtered evaluation data based on the assessment results.

[0182] In some optional implementations of this embodiment, the evaluation data includes multi-round evaluation data, and the data generation unit 1102 is further configured to generate multi-round evaluation data including evaluation requests and standard answers based on the knowledge graph, priority and multi-round question answering strategy, wherein the multi-round question answering strategy represents the question answering logic between multi-round evaluation requests.

[0183] In this embodiment, an evaluation data generation device is provided. The graph generation unit in the evaluation data generation device performs structured processing of the original business data and generates a knowledge graph that represents the relationship between entities in the original business data. The data generation unit generates evaluation data for the model to be evaluated based on the knowledge graph through an artificial intelligence big data model. Thus, based on the structured advantages of the knowledge graph, it helps to solve problems such as entity confusion, logical breaks, and ambiguous relationships. Furthermore, the efficient generation capability of the artificial intelligence big data model is superimposed, which improves the generation quality, credibility, and generation efficiency of the evaluation data.

[0184] According to embodiments of this disclosure, this disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to implement the evaluation data generation method described in any of the above embodiments.

[0185] According to embodiments of this disclosure, this disclosure also provides a readable storage medium storing computer instructions that enable a computer to implement the evaluation data generation method described in any of the above embodiments when executed.

[0186] This disclosure provides a computer program product that, when executed by a processor, can implement the evaluation data generation method described in any of the above embodiments.

[0187] Figure 12 A schematic block diagram of an example electronic device 1200 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0188] like Figure 12 As shown, device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1202 or a computer program loaded from storage unit 1208 into random access memory (RAM) 1203. The RAM 1203 may also store various programs and data required for the operation of device 1200. The computing unit 1201, ROM 1202, and RAM 1203 are interconnected via bus 1204. Input / output (I / O) interface 1205 is also connected to bus 1204.

[0189] Multiple components in device 1200 are connected to I / O interface 1205, including: input unit 1206, such as keyboard, mouse, etc.; output unit 1207, such as various types of monitors, speakers, etc.; storage unit 1208, such as disk, optical disk, etc.; and communication unit 1209, such as network card, modem, wireless transceiver, etc. Communication unit 1209 allows device 1200 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0190] The computing unit 1201 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 performs the various methods and processes described above, such as the evaluation data generation method. For example, in some embodiments, the evaluation data generation method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1208. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1200 via ROM 1202 and / or communication unit 1209. When the computer program is loaded into RAM 1203 and executed by the computing unit 1201, one or more steps of the evaluation data generation method described above may be performed. Alternatively, in other embodiments, the computing unit 1201 may be configured to perform a method for generating evaluation data by any other suitable means (e.g., by means of firmware).

[0191] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0192] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to the processor or controller of a general-purpose computer, special-purpose computer, or other programmable test data generation device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0193] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0194] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0195] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0196] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, also known as cloud computing servers or cloud hosts, which are hosting products within the cloud computing service system to address the management difficulties and weak business scalability inherent in traditional physical hosts and Virtual Private Servers (VPS) services; they can also be servers for distributed systems or servers incorporating blockchain technology.

[0197] According to the technical solution of the embodiments of this disclosure, a method and apparatus for generating evaluation data are provided. By structuring business native data, a knowledge graph representing the relationship between entities in the business native data is generated. Based on the knowledge graph, evaluation data of the model to be evaluated is generated by an artificial intelligence big data model. Thus, based on the structuring advantages of the knowledge graph, it helps to solve problems such as entity confusion, logical breakage, and ambiguous relationships. Furthermore, the efficient generation capability of the artificial intelligence big data model is superimposed, which improves the generation quality, credibility, and generation efficiency of the evaluation data.

[0198] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution provided in this disclosure can be achieved, and this is not limited herein.

[0199] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for generating evaluation data, comprising: Structured processing of raw business data generates a knowledge graph representing the relationships between entities in the raw business data; Using a large-scale artificial intelligence model, evaluation data for the model to be evaluated is generated based on the knowledge graph.

2. The method according to claim 1, wherein, The structured processing of the original business data generates a knowledge graph representing the relationships between entities in the original business data, including: Using the processing method corresponding to the generation scenario of the evaluation data, the original business data is structured and processed to generate a knowledge graph representing the relationship between entities in the original business data.

3. The method according to claim 2, wherein, The generated scenarios include data query scenarios, where the native business data is descriptive data of the database structure, and The processing method corresponding to the generation scenario of the evaluation data is used to structure the original business data and generate a knowledge graph representing the relationships between entities in the original business data, including: Determine the relationships between entities at multiple levels in the description data and generate multiple triples, wherein the triples represent the relationships between two entities; By combining multiple triples, a knowledge graph for the data query scenario is generated.

4. The method according to claim 3, wherein, The step of determining the relationships between entities at multiple levels in the description data and generating multiple triples includes: Extract entities of multiple levels and the relationships between entities of multiple levels in the description data, wherein the multiple levels include data table entities, field entities in the data table, field value entities under the field, and constraint rule entities under the field; Based on the aforementioned association, multiple triples are generated.

5. The method according to claim 4, wherein, The relationships include the relationships between the data table entities and the field entities, the relationships between the field entities and the constraint rule entities, the relationships between the field entities and the field value entities, and the relationships between the data table entities.

6. The method according to claim 4, wherein, The step of combining multiple triples to generate a knowledge graph for the data query scenario includes: Merge multiple triples that include the same entity, and use the data table entity as the first-level node, the field entity as the second-level node, and the constraint rule entity and the field value entity as the third-level nodes to generate a knowledge graph for the data query scenario.

7. The method according to any one of claims 3-6, wherein, The step of generating evaluation data for the model to be evaluated based on the knowledge graph includes: The fields in the description data are grouped according to business scenarios to obtain multiple business groups; Based on the business relevance between the multiple business groups, the multiple business groups are combined to obtain a business group combination; Based on the relationship chains in the knowledge graph corresponding to the business group combination, evaluation data including evaluation requests and structured query statements is generated.

8. The method according to claim 7, wherein, The evaluation data includes single-round evaluation data, and The step of generating evaluation data, including evaluation requests and structured query statements, based on the relationship chains in the knowledge graph corresponding to the business group combination, includes: Based on the relationship chains corresponding to the business group combination in the knowledge graph, the evaluation samples, and the query types that the single-round evaluation data needs to cover, single-round evaluation data including evaluation requests and structured query statements is generated.

9. The method according to claim 8, wherein, The step of generating single-round evaluation data, including evaluation requests and structured query statements, based on the relationship chains corresponding to the business group combination in the knowledge graph, evaluation samples, and the query types that the single-round evaluation data needs to cover, includes: Determine the chain type of the relationship chain, wherein the chain type includes a single-table chain that represents entities in the relationship chain being covered by a single data table and a multi-table chain that is covered by a combination of multiple data tables; The single-round evaluation data is generated based on the relationship chain, the evaluation sample, the query type, and the evaluation data generation rules corresponding to the chain type.

10. The method according to any one of claims 7-9, wherein, The evaluation data includes data from multiple rounds of evaluation, and The step of generating evaluation data, including evaluation requests and structured query statements, based on the relationship chains in the knowledge graph corresponding to the business group combination, includes: Based on the relationship chains and multi-round query strategies corresponding to the business group combinations in the knowledge graph, multi-round evaluation data including evaluation requests and structured query statements are generated, wherein the multi-round query strategy represents the query logic between multi-round evaluation requests.

11. The method according to claim 2, wherein, The generation scenarios include document question-and-answer scenarios, and the native business data is unstructured text data. The processing method corresponding to the generation scenario of the evaluation data is used to structure the original business data and generate a knowledge graph representing the relationships between entities in the original business data, including: Extract triples representing the subject, predicate, and object from the data blocks of the unstructured text data; By combining multiple triples, a knowledge graph for the document question-and-answer scenario is generated.

12. The method according to claim 11, wherein, Before extracting the triples representing the subject, predicate, and object from the data blocks of the unstructured text data, the method further includes: Based on the data capacity of the unstructured text data, determine the method for dividing the unstructured text data; The unstructured text data is divided using the aforementioned partitioning method to obtain multiple data blocks.

13. The method according to claim 11, wherein, The step of generating evaluation data for the model to be evaluated based on the knowledge graph includes: Based on the business scenario to which each of the triples belongs, the triples are divided into multiple scenario groups; For multiple scenario groups, the priority of the business scenario corresponding to each scenario group is determined based on the correlation between the scenario group and the unstructured text data; Based on the knowledge graph and the priority, evaluation data including assessment requests and standard answers is generated.

14. The method according to claim 13, wherein, The step of dividing the multiple triples into multiple scenario groups according to the business scenario to which each triple belongs includes: Candidate keywords are selected from the entities in the triples based on the frequency of occurrence of the entities in the triples. Based on the candidate keywords, the multiple triples, and the domain to which the unstructured text data belongs, multiple business scenarios are determined; Based on the correlation between the triples and the multiple business scenarios, the multiple triples are divided into multiple scenario groups.

15. The method according to claim 13, wherein, For multiple scenario groups, the priority of the business scenario corresponding to each scenario group is determined based on the correlation between the scenario group and the unstructured text data, including: For multiple scene groups, the scene keywords corresponding to the scene groups are determined from the candidate keywords to obtain a set of scene keywords, and a set of scene entities is obtained based on the entities in the scene groups; Determine the first degree of overlap between the set of scenario keywords and the set of core keywords representing the topic of the unstructured text data; Determine the second degree of overlap between the set of scene entities and the set of core keywords; Based on the first overlap and the second overlap, the correlation between the business scenario corresponding to the scenario group and the unstructured text data is determined; The priority among the multiple business scenarios is determined based on their respective correlations.

16. The method according to any one of claims 13-15, wherein, The evaluation data includes single-round evaluation data, and The step of generating assessment data, including assessment requests and standard answers, based on the knowledge graph and the priority, includes: Based on the knowledge graph, the priority, the types of questions the evaluation data needs to cover, and the preset generation strategy of the evaluation data, single-round evaluation data including evaluation requests and standard answers is generated.

17. The method according to claim 16, wherein, The step of generating single-round evaluation data, including evaluation requests and standard answers, based on the knowledge graph, the priority, the question types that the evaluation data needs to cover, and the preset generation strategy of the evaluation data, includes: Based on the knowledge graph, the priority, the question type, and the preset generation strategy, multiple initial evaluation data, including evaluation requests and standard answers, are generated. Perform integrity checks and semantic deduplication operations on multiple initial evaluation data to obtain multiple filtered evaluation data; The quality of multiple filtered evaluation data is evaluated from multiple preset evaluation dimensions, and the single-round evaluation data is determined from the multiple filtered evaluation data based on the evaluation results.

18. The method according to claim 13, wherein, The evaluation data includes data from multiple rounds of evaluation, and The step of generating assessment data, including assessment requests and standard answers, based on the knowledge graph and the priority, includes: Based on the knowledge graph, the priority, and the multi-round question-answering strategy, multi-round evaluation data including evaluation requests and standard answers is generated, wherein the multi-round question-answering strategy represents the question-answering logic between multi-round evaluation requests.

19. An apparatus for generating evaluation data, comprising: The graph generation unit is configured to process business native data in a structured manner and generate a knowledge graph that represents the relationships between entities in the business native data. The data generation unit is configured to generate evaluation data for the model to be evaluated based on the knowledge graph using a large artificial intelligence model.

20. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-18.

21. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-18.

22. A computer program product, comprising: A computer program that, when executed by a processor, implements the method according to any one of claims 1-18.