Vertical domain AI application evaluation data distribution semi-automatic fitting method and system

By constructing a two-layer directory structure and a semi-automated approach, combined with historical data analysis and expert review, the problems of distribution bias and insufficient integration of business knowledge in general AI evaluation datasets in vertical fields were solved. This enabled the construction of efficient and accurate evaluation datasets, improving the performance of models in real-world business scenarios.

CN120851216BActive Publication Date: 2026-01-02传申弘安智能(深圳)有限公司 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511344655.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2026-01-02
Estimated Expiration
2045-09-19

AI Technical Summary

Technical Problem

Existing general AI evaluation datasets suffer from distribution bias, insufficient integration of business knowledge, and difficulty in balancing automated construction and manual review in vertical applications, resulting in poor model performance in actual deployments.

Method used

By acquiring business specification documents to construct a two-level directory structure, combining historical data analysis to identify high-frequency or core attribute combinations, identifying test points and question type ratios, and using OpenAI knowledge ability assessment and Charles Sanders Peirce reasoning ternary classification method to design knowledge-based and logic-based questions, a semi-automated assessment set is constructed.

Benefits of technology

It achieves high business relevance and accuracy in the evaluation dataset, reduces construction costs and time consumption, adapts to rapidly changing business needs, and ensures the accuracy and economy of model performance evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120851216B_ABST
    Figure CN120851216B_ABST
Patent Text Reader

Abstract

The application discloses a vertical AI application evaluation data distribution semi-automatic fitting method and system. The method comprises the following steps: obtaining a business specification document, and constructing a double-layer directory structure according to the business specification document; obtaining historical data; based on the double-layer directory structure and the historical data, high-frequency or core attribute combinations are counted to obtain test points; the types of test questions in the historical data are analyzed to determine the proportion of the types of test questions; according to the test points and the proportion of the types of test questions, a corresponding number of knowledge-based questions and logic-based questions are extracted to form a preliminary evaluation set; the preliminary evaluation set is audited, and the preliminary evaluation set is optimized and adjusted according to the audit result. By implementing the method of the application, the utility problem of the existing evaluation data set in the application of the vertical field can be solved, and the model performance evaluation can accurately reflect the actual business performance and has cost effectiveness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to data fitting methods, and more specifically to a semi-automatic fitting method and system for the distribution of evaluation data in vertical AI applications. Background Technology

[0002] Currently, mainstream general-purpose AI vendors often reveal a mismatch between the evaluation datasets they construct and the actual application needs of vertical domain AI when applying these datasets to evaluate AI performance. This phenomenon is mainly attributed to the utility problem commonly encountered by general data analysis methods: although the test datasets may achieve high scores, their actual performance after deployment falls far short of expectations, preventing high performance in closed environments from translating into real-world productivity. Specifically, this mismatch can be summarized into three major technical flaws:

[0003] First, the distribution of the test dataset deviates significantly from that of the data in the actual production environment, making it difficult for the test results to accurately reflect the model's performance in real business scenarios. Second, in vertical domain practices, there is a lack of effective integration between business knowledge and test data, manifested in a disconnect between the labeling system and business rules, as well as insufficient coverage of test point combinations. This fragmentation not only leads to violations of basic industry standards or results that are meaningless or illogical, but also overlooks the evaluation of key observation object attribute combinations, thus affecting the model's performance on key business metrics. Finally, striking a balance between automated construction of evaluation datasets and manual review is difficult: on the one hand, relying entirely on experts for manual annotation is both expensive and time-consuming; on the other hand, while purely automated construction methods are more efficient, they are prone to deviating from real business needs, increasing risks.

[0004] These issues combined limit the effectiveness of general AI evaluation methods in vertical industries. Therefore, to improve the practical effectiveness of AI applications in vertical industries, it is necessary to address the aforementioned technical challenges in a targeted manner.

[0005] Therefore, it is necessary to design a new method to address the utility problem of existing evaluation datasets in vertical applications, ensuring that model performance evaluation accurately reflects actual business performance while also being cost-effective. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a semi-automatic fitting method and system for the distribution of evaluation data for vertical AI applications.

[0007] To achieve the above objectives, the present invention adopts the following technical solution: a semi-automatic fitting method for the distribution of vertical AI application evaluation data, comprising:

[0008] Obtain the business specification document and construct a two-level directory structure based on the business specification document;

[0009] Obtain historical data;

[0010] Based on the aforementioned two-layer directory structure and the aforementioned historical data, high-frequency or core attribute combinations are statistically analyzed to obtain the test points;

[0011] Analyze the question types in the historical data to determine the proportion of each question type;

[0012] Based on the test points and question types, a corresponding number of knowledge-based questions and logic-based questions were selected to form a preliminary evaluation set.

[0013] The preliminary evaluation set is reviewed, and the preliminary evaluation set is optimized and adjusted based on the review results.

[0014] The further technical solution is as follows: the step of obtaining the business specification document and constructing a two-level directory structure based on the business specification document includes:

[0015] Obtain the business specification document;

[0016] The business specification document is analyzed to obtain core elements, which include business scenarios, observation objects and their attributes;

[0017] Based on the aforementioned core elements, a multi-level nested structure is constructed to obtain a preliminary two-level directory;

[0018] The preliminary two-level catalog is reviewed and adjusted to obtain the adjusted two-level catalog;

[0019] The adjusted two-level directory is verified and improved using historical business data to obtain a corrected two-level directory.

[0020] The corrected two-level directory is then subject to final review to obtain the two-level directory structure.

[0021] Its further technical solution is: the method of obtaining test points by combining the two-level directory structure with historical data statistics of high-frequency or core attribute combinations includes:

[0022] By using historical data to perform statistical analysis on the attributes of the observed objects, high-frequency or core attribute combinations can be identified to obtain the test points.

[0023] The further technical solution is as follows: the analysis of the question types in the historical data to determine the proportion of question types includes:

[0024] Analyze the question types in the historical data, identify and calculate the proportion and coverage of knowledge-based and logic-based questions in different business scenarios, so as to obtain the question type ratio.

[0025] The further technical solution is as follows: based on the ratio of test points and question types, a corresponding number of knowledge-based questions and logic-based questions are extracted to form a preliminary evaluation set, including:

[0026] Based on the test points and question types, the number of test questions for each test point is allocated according to a certain ratio, and the proportion of knowledge-based questions and logic-based questions in the evaluation set is determined to obtain the preliminary evaluation set.

[0027] The further technical solution is as follows: Based on the ratio of test points and question types, the number of test questions for each test point is allocated according to a certain proportion, and the proportion of knowledge-based questions and logic-based questions in the evaluation set is determined to obtain a preliminary evaluation set, including:

[0028] Based on the aforementioned test points and question types, a corresponding number of knowledge-based and logic-based questions are randomly sampled using a strategy of random sampling by test point and random sampling by question type to obtain a preliminary evaluation set.

[0029] The further technical solution is as follows: the knowledge-based questions adopt OpenAI's knowledge ability assessment method, focusing on short knowledge questions to examine direct information retrieval and understanding abilities, and classifying long knowledge questions involving multi-step reasoning as questions for logical reasoning tasks;

[0030] The logic questions are based on Charles Sanders Peirce's triadic classification of reasoning, which subdivides logical ability assessment into three core reasoning skills: abduction, induction, and deduction, in order to comprehensively evaluate the logical processing capabilities of the AI ​​system.

[0031] This invention also provides a semi-automatic fitting system for the distribution of evaluation data for vertical AI applications, including:

[0032] A building unit is used to obtain business specification documents and build a two-level directory structure based on the business specification documents;

[0033] Historical data acquisition unit, used to acquire historical data;

[0034] The test point determination unit is used to obtain test points by statistically analyzing high-frequency or core attribute combinations based on the two-level directory structure and the historical data.

[0035] The question type ratio determination unit is used to analyze the question types in the historical data and determine the question type ratio.

[0036] The preliminary assessment set generation unit is used to extract a corresponding number of knowledge-based questions and logic-based questions to form a preliminary assessment set based on the test points and question type ratios.

[0037] The review unit is used to review the preliminary evaluation set and optimize and adjust the preliminary evaluation set based on the review results.

[0038] The present invention also provides a computer device, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the above-described method.

[0039] The present invention also provides a storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0040] The advantages of this invention compared to existing technologies are as follows: This invention acquires and analyzes business specification documents and historical data to construct a two-layer directory structure that covers complete business scenarios and combinations of observed object attributes. Based on this, it determines high-frequency or core test points and the proportion of question types. Subsequently, based on this information, it extracts a corresponding number of questions from knowledge-based and logic-based questions to form a preliminary evaluation set. Expert review and dynamic adjustments ensure the accuracy and business representativeness of the evaluation set. This method solves the utility problem of existing evaluation datasets in vertical domain applications, ensuring that model performance evaluation not only accurately reflects actual business performance but also significantly reduces construction costs and time consumption through human-machine collaboration, maximizing cost-effectiveness. This ensures that the evaluation dataset has high business relevance and accuracy while effectively improving construction efficiency and adapting to rapidly changing business needs.

[0041] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. Attached Figure Description

[0042] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0043] Figure 1 A schematic diagram illustrating the application scenario of the semi-automatic fitting method for vertical domain AI application evaluation data distribution provided in this embodiment of the invention;

[0044] Figure 2 A flowchart illustrating the semi-automatic fitting method for vertical AI application evaluation data distribution provided in this embodiment of the invention;

[0045] Figure 3 A schematic diagram of a sub-process of the semi-automatic fitting method for vertical domain AI application evaluation data distribution provided in an embodiment of the present invention;

[0046] Figure 4 A schematic diagram of a semi-automatic fitting method for the distribution of evaluation data for vertical AI applications provided in an embodiment of the present invention;

[0047] Figure 5 A schematic block diagram of a semi-automatic fitting system for vertical domain AI application evaluation data distribution provided in an embodiment of the present invention;

[0048] Figure 6 A schematic block diagram of the building units of the semi-automatic fitting system for vertical AI application evaluation data distribution provided in this embodiment of the invention;

[0049] Figure 7 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation

[0050] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0051] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0052] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0053] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0054] Please see Figure 1 and Figure 2 , Figure 1 This is a schematic diagram illustrating an application scenario of the semi-automatic fitting method for vertical domain AI application evaluation data distribution provided in this embodiment of the invention. Figure 2This is a schematic flowchart illustrating a semi-automated fitting method for vertical AI application evaluation data distribution provided in this embodiment of the invention. This method is applied to a server. The server interacts with the terminal, starting by acquiring business specification documents to construct a two-layer directory structure. It then analyzes historical data to determine high-frequency or core attribute combinations to identify test points and analyzes question types to determine the proportion of each type. A suitable number of knowledge-based and logic-based questions are then extracted to form a preliminary evaluation set. After review and optimization, the final evaluation dataset is formed. This method particularly emphasizes designing knowledge-based and logic-based questions based on OpenAI knowledge ability assessment and the Charles Sanders Peirce reasoning ternary classification method. This ensures that the evaluation accurately reflects the actual business performance of the AI ​​system in a specific vertical domain, while also improving efficiency and reducing costs through semi-automated process optimization. This effectively solves the utility problem of existing evaluation datasets in vertical domain applications, while ensuring the accuracy and economy of model performance evaluation.

[0055] Figure 2 This is a flowchart illustrating the semi-automatic fitting method for the distribution of vertical AI application evaluation data provided in this embodiment of the invention. Figure 2 As shown, the method includes the following steps S110 to S160.

[0056] S110. Obtain the business specification document and construct a two-level directory structure based on the business specification document.

[0057] In this embodiment, the two-layer directory structure refers to a classification framework based on a specific domain's business knowledge system, used to systematically organize and manage the knowledge points and logical reasoning questions required for assessment. This structure consists of two layers: the first layer comprises business scenarios, and the second layer comprises observation objects and their attributes. First, relevant business specification documents need to be collected from relevant departments or public resources. These documents may include policy documents, operation manuals, industry standards, regulatory rules, etc. These documents are key information sources for understanding the business operation mechanisms within a specific vertical domain.

[0058] Next, the collected business specification documents are analyzed in depth to extract core elements. This step aims to identify all business scenarios directly related to the evaluation, as well as the key observation objects and their attributes within each scenario. For example, in power AI application scenarios, business scenarios might include grid fault detection and load forecasting; while observation objects might include voltage levels and current intensity.

[0059] Based on the above analysis, a preliminary two-tier directory structure was constructed. This process involves organizing the identified core elements (i.e., business scenarios and observed objects and their attributes) according to certain logical relationships. For example, for a specific business scenario "power grid fault detection," it may contain multiple observed objects such as "transformer status" and "line load," and each observed object has its own attribute values ​​such as "temperature" and "humidity."

[0060] After completing the initial two-tiered catalog, experts in relevant fields should be invited to review it. Based on the feedback, the catalog structure should be adjusted as necessary to ensure that it accurately reflects actual business needs and covers all important business scenarios and combinations of observed object attributes.

[0061] The two-level directory is validated using historical business data to check for any missing important observation objects or attributes, and the directory structure is further improved based on this, until the final version of the two-level directory is formed.

[0062] By following the steps above, a comprehensive and detailed two-tiered directory structure can be effectively established, providing a solid foundation for subsequent work such as test point distribution analysis and test question design. This method not only helps improve the quality and accuracy of the evaluation dataset but also ensures that it truly reflects the actual business situation within the vertical field.

[0063] In one embodiment, please refer to Figure 2 The above step S110 may include steps S111 to S116.

[0064] S111. Obtain the business specification document.

[0065] In this embodiment, the business specification documents include business operation manuals, business process specifications, business policy documents, and business-related form templates.

[0066] Specifically, step S111 involves "obtaining business specification documents," which is the first step in building a two-tiered directory structure based on power AI application scenarios. In essence, this stage involves collecting all necessary documents related to the target business domain, which will serve as the foundation for subsequent analysis and development.

[0067] Business specification documents mainly include, but are not limited to, the following categories:

[0068] Operations Manual: This manual details the processes and steps of daily operations, including the specific execution method for each step, the expected results, and potential problems and their solutions. For example, in the power industry, this might include how to perform grid maintenance and troubleshooting.

[0069] Business process specifications define the standard procedures for completing specific business tasks, emphasizing the logical relationships and sequence between processes. For power systems, this may involve operational specifications for generation, transmission, and distribution.

[0070] Business policy documents: These documents cover all internal regulations regarding business management, such as safety standards and quality control measures. These documents ensure that business operations comply with industry and company requirements.

[0071] Business-related form templates: These provide standard formats for recording and reporting business activities, such as work logs, equipment checklists, and incident report forms. These forms help standardize the data collection process and provide structured input for subsequent data analysis.

[0072] S112. Analyze the business specification document to obtain core elements, wherein the core elements include business scenarios, observation objects and their attributes.

[0073] In this embodiment, core elements refer to the key information that directly reflects the essence of the business and supports the subsequent construction of a two-tier directory. Specifically:

[0074] Business scenario: This refers to the specific environment or background in which the technology or method is applied. For example, in the power system, "large vehicle inspection" and "power grid maintenance" are specific business scenarios.

[0075] Observational objects: These refer to objects that need to be monitored or evaluated in a specific business scenario. These objects can be physical entities (such as cranes or tower cranes) or abstract phenomena (such as light intensity or equipment status).

[0076] Observational object attributes: Feature dimensions or attribute values ​​related to the observed object, which can be qualitative (such as presence or absence) or quantitative (such as the specific value of light intensity).

[0077] Specifically, the business specification document is input into a pre-trained large model, and the pre-trained large model is guided by specific prompts to extract structured information. Cross-validation is then performed in conjunction with the experience of business experts to identify and correct any deviations, thereby obtaining the core elements.

[0078] In this embodiment, all relevant business specification documents are collected, including but not limited to business operation manuals, business process specifications, and business policy documents.

[0079] The collected documents are digitized to ensure uniform formatting, facilitating subsequent large-scale model analysis. This involves technologies such as text conversion and image OCR recognition.

[0080] The preprocessed business specification document is input into the pre-trained language model.

[0081] Specific prompts guide the LLM (Large Language Model) to extract structured information. For example, for a business specification document about "large vehicle inspection," the prompts might be designed as: "Please identify and classify all business scenarios, observation objects, and their related attributes based on the following."

[0082] LLM will attempt to understand and parse the corresponding structured information based on the provided document content and prompts, such as a list of business scenarios, categories of observed objects, and attribute descriptions of each object.

[0083] While LLM possesses powerful natural language processing capabilities, its output may still contain certain biases or inaccuracies. Therefore, in this step, it is necessary to involve business experts in the field to review the preliminary results.

[0084] Experts can not only verify the correctness of the model's output, but also provide necessary correction suggestions, especially for complex or ambiguous situations. For example, certain business terms may have specific meanings that require expert explanation and confirmation.

[0085] Based on feedback from business experts, the model output is adjusted and improved to form a more accurate set of core elements. This process may require multiple iterations until all key elements are accurately identified and their relationships are clearly defined.

[0086] Finally, the verified and revised core elements were compiled into a detailed report. This report not only included the finalized business scenarios, observation objects and their attributes, but also recorded the decision-making basis and reasons for revisions throughout the process, becoming an important basis for the next step of building a two-tiered catalog.

[0087] By following the steps above, the core elements crucial for building a two-tiered directory can be efficiently and accurately extracted from complex business specification documents, thus laying a solid foundation for achieving the goal of comprehensive and detailed coverage of business scenarios. This method not only improves work efficiency but also enhances the reliability and practicality of the results.

[0088] S113. Construct a multi-level nested structure based on the core elements to obtain a preliminary two-level directory.

[0089] In this embodiment, the initial two-layer directory refers to a structured information organization method that organically combines business scenarios with observed objects and their attributes to form a multi-level nested system. The multi-level nested structure includes a four-level nested structure: business scenario - observed object layer - observed object attribute layer - attribute value layer.

[0090] Specifically, in this embodiment, this multi-level nested structure includes the following four levels:

[0091] Business Scenario Layer: Located at the top level, it represents the specific environment or background of the applied technology or method.

[0092] Observation object layer: Located in the second layer, it refers to the objects that need to be monitored or evaluated in specific business scenarios.

[0093] Observation object attribute layer: Located in the third layer, it describes the relevant feature dimensions or attribute values ​​of the observed object.

[0094] Attribute value layer: Located at the bottom layer, it provides the specific value or status of each attribute.

[0095] This four-level nested structure helps to record and classify the observed objects and their detailed attribute information in each business scenario, which facilitates subsequent data retrieval, analysis, and decision support.

[0096] Specifically, taking business scenarios as the core, and targeting the observation objects and various attributes under each business scenario, a two-level directory tag system with the observation object attributes as the core is constructed to obtain a preliminary two-level directory.

[0097] In this embodiment, based on the core elements extracted in the previous step (S120), all relevant business scenarios are first identified. For example, "power grid maintenance" and "large vehicle inspection".

[0098] For each business scenario, identify all the observation objects contained therein. These objects can be physical entities such as cranes and tower cranes, or abstract concepts such as light intensity and equipment status.

[0099] For each observed object, its related attributes are further analyzed. For example, for a "crane," attributes such as maximum lifting capacity and operating radius may be involved.

[0100] Specify a specific numerical value or state range for each attribute. For example, the attribute value for "maximum lifting capacity" can range from 5 tons to 50 tons.

[0101] In this step, a two-level directory tag system is built around the attributes of the observed objects. This means that each observed object has related attributes as sub-items, and each attribute has corresponding attribute values ​​as more granular information units.

[0102] By integrating the information from each of these levels, a unified four-level nested structure is formed. In this structure, the business scenario serves as the top-level node, with the observed object, observed object attributes, and attribute values ​​arranged sequentially downwards. This structure not only clearly reflects the business logic but also facilitates data management and querying.

[0103] After completing the initial two-tiered directory, it needs to be validated through real-world case studies or simulation tests to ensure that the structure accurately reflects business needs. Based on feedback, necessary adjustments and optimizations should be made to improve its usability and applicability.

[0104] By following the steps described above, a preliminary two-layer directory with a multi-level nested structure can be systematically constructed, starting from the core elements. This directory not only provides a comprehensive understanding of the complex relationships between business scenarios, observed objects, and their attributes, but also lays a solid foundation for subsequent in-depth analysis, model training, and application scenario evaluation. This method greatly improves the efficiency and accuracy of information organization, and is conducive to promoting the in-depth development of power AI application scenario evaluation.

[0105] S114. Review and adjust the preliminary two-level directory to obtain the adjusted two-level directory.

[0106] In this embodiment, the adjusted two-tier directory refers to the information organization structure after detailed review and optimization. It not only includes comprehensive and detailed descriptions of business scenarios, observation objects, and their attributes, but all content strictly conforms to business meaning, ensuring that the final directory system can effectively support subsequent data analysis, model training, and application scenario evaluation.

[0107] Specifically, the preliminary two-layer directory is reviewed and adjusted to determine whether it is comprehensive in terms of business scenarios, whether the attributes of the observed objects are comprehensive, and whether it conforms to business significance, so as to obtain the adjusted two-layer directory.

[0108] To obtain a more complete and practical revised two-level directory from the initial one, the following are the specific implementation steps:

[0109] Organize a review team composed of domain experts, technical personnel, and business personnel.

[0110] Prepare review materials, including a preliminary two-tiered catalog document, relevant business descriptions, historical data samples, etc.

[0111] Business scenario review: Check whether all important business scenarios are covered. If any important scenarios are found to be missing, they need to be added to the catalog.

[0112] Observation object review: Conduct a comprehensive review of the observation objects for each business scenario to ensure that no key objects are missed.

[0113] Attribute comprehensiveness review: For each observed object, review whether its attribute list is sufficiently detailed and can comprehensively reflect the key characteristics of the object. If important attributes are found to be missing, they should be added.

[0114] For each business scenario, observation object, and its attributes, assess whether they align with real-world business processes and operational habits. For example, if certain attributes are technically feasible but not frequently used or significant in actual business operations, consider removing or simplifying them.

[0115] Ensure that all descriptions in the directory structure are clear and easy to understand, avoiding vague or overly technical terms that could lead to comprehension difficulties.

[0116] Based on the review results, the initial two-tiered catalog will be adjusted accordingly. This may involve adding new business scenarios or observation objects, supplementing missing attributes, and correcting inaccurate or unreasonable descriptions.

[0117] During the adjustment process, maintain close communication with the business team to ensure that every change made receives full support and approval.

[0118] After the initial adjustment, the adjusted two-level directory was tested and verified using some real data to observe its performance in practical applications.

[0119] Collect user feedback and fine-tune the directory structure based on the feedback until the optimal state is achieved.

[0120] The finalized and adjusted two-level directory was compiled into a formal document, detailing the content of each level and their interrelationships.

[0121] The revised two-tier catalog was officially released, and training and support were provided to relevant personnel to ensure they could correctly understand and use this new framework.

[0122] Through the rigorous review and adjustment process described above, the quality of the initial two-tiered directory can be significantly improved, making it more closely aligned with actual business needs and laying a solid foundation for subsequent work. Furthermore, this iterative improvement approach also helps to continuously optimize the directory structure, adapting to the ever-changing business environment and technological advancements.

[0123] S115. Utilize historical business data to verify and improve the adjusted two-level directory to obtain a corrected two-level directory.

[0124] In this embodiment, the corrected two-tier directory refers to an information organization structure that has been supplemented and optimized after detailed comparison with historical business data. It not only inherits the advantages of the adjusted two-tier directory—comprehensive coverage of important business scenarios, detailed description of observed objects and their attributes, and strict adherence to business logic—but also verifies its accuracy and completeness through specific data examples, ensuring that the final directory system can more accurately support subsequent work such as data analysis, model training, and application scenario evaluation.

[0125] The adjusted two-level directory is compared with the historical business data. If the adjusted two-level directory is found to be inadequate or missing, it is supplemented and improved based on the historical business data to obtain the corrected two-level directory.

[0126] After reviewing and adjusting the initial two-tier directory (S140), a relatively complete adjusted two-tier directory was obtained. However, whether this directory structure truly meets actual business needs and accurately reflects all key business scenarios and observed object attributes still needs further verification and improvement through comparison with real historical business data. Therefore, in this step (S150), historical business data will be used as a benchmark to comprehensively verify the adjusted two-tier directory, and supplements and improvements will be made based on the problems found.

[0127] Specifically, collect and organize relevant historical business data, including but not limited to transaction records, customer information, product details, and marketing activity reports.

[0128] Ensure that the selected data is representative and covers as many business scenarios and observation object types as possible.

[0129] Based on the revised two-tier directory, construct a framework or template for comparison. This framework should clearly list the business scenarios, observation objects, and their attributes for each level.

[0130] Design a set of evaluation metrics or rules to measure the degree of matching between catalog items and actual data, such as coverage and accuracy.

[0131] Import historical business data into the aforementioned comparison framework and check one by one whether each item can be found in the dataset.

[0132] For each business scenario, observation object, and its attributes, analyze its performance in actual data and identify any deficiencies or omissions in the catalog.

[0133] When it is found that the adjusted two-tier directory fails to cover important business scenarios, or does not mention key observation objects or attributes, these issues should be recorded.

[0134] For each problem, corresponding supplementary measures or modification suggestions are formulated based on the specific characteristics of historical business data.

[0135] Based on the established plan, necessary supplements and improvements will be made to the adjusted two-tier directory. This may involve adding new business scenarios or observation objects, expanding the existing attribute list, and correcting inaccurate or outdated descriptions.

[0136] Throughout this process, maintain good communication with business experts and the technical team to ensure that every change made receives full support and approval.

[0137] After the initial improvements are completed, the corrected two-tier directory is tested and verified again using historical business data to observe its improvement effect.

[0138] If there are still unresolved issues or new discoveries, continue repeating the above steps until the expected goal is achieved.

[0139] The finalized corrected two-level directory was compiled into a formal document, detailing the changes made in each iteration and the reasons for them.

[0140] Provide necessary training and support to relevant personnel to help them understand the changes in the new version of the catalog and the significance behind them.

[0141] Through the above series of verification and improvement efforts based on historical business data, the accuracy and usability of the adjusted two-tier directory can be significantly improved, making it more aligned with actual business needs and providing strong support for subsequent data-driven decision-making. Furthermore, this approach helps to continuously optimize the directory structure, ensuring its adaptability to the ever-changing business environment.

[0142] S116. Perform a final review on the corrected two-level directory to obtain the two-level directory structure.

[0143] In this embodiment, the two-tier directory structure refers to a structure built upon a corrected two-tier directory, rigorously reviewed to ensure it fully covers all necessary business scenarios, accurately reflects the actual recording of various business specifications and data, and guarantees the entire directory's operability. Specifically, the two-tier directory structure should meet the following conditions:

[0144] Complete scenario coverage: Covers all key business processes and activities within an enterprise or organization, leaving nothing out.

[0145] Accurate attribute descriptions: For each observed object and its attributes, precise and detailed descriptions are provided, which not only meet business specifications but also remain faithful to historical data records.

[0146] The catalog is highly operable: it is well-designed, easy to understand and implement, and helps improve work efficiency and decision-making quality.

[0147] Specifically, the corrected two-layer directory is subject to a final review to determine whether the scenario fully covers the specifications and actual business, whether the attributes simultaneously meet the business specifications and actual data records, and whether the directory is operable, in order to obtain the two-layer directory structure.

[0148] After detailed verification and optimization in step S150, a more accurate and practical corrected two-level directory was obtained. However, to ensure that the directory not only conforms to business specifications but also meets the needs of actual operation, a comprehensive and meticulous final review (S160) is required. This step aims to comprehensively evaluate the corrected two-level directory from multiple dimensions—including the completeness of the business scenario, the accuracy of attribute descriptions, and the operability of the directory structure—to derive a two-level directory structure that can effectively support subsequent work.

[0149] Collect and organize the latest standard documents, policy documents, and technical manuals and other reference materials related to current business.

[0150] Prepare a sample of historical business data for comparative analysis, ensuring that the data is broad and representative.

[0151] The review process brought together experts from various fields, such as business analysts, data scientists, IT professionals, and frontline employee representatives.

[0152] Clearly define the roles and responsibilities of each member to ensure the professionalism and comprehensiveness of the review process.

[0153] Scenario integrity review: Check one by one whether the corrected two-level directory contains all important business scenarios, paying special attention to those niche or marginalized scenarios that are easily overlooked.

[0154] Attribute accuracy verification: Thoroughly verify whether the definition of each observed object and its attributes is clear and accurate, whether it fully considers various situations that may occur in actual operation, and compare and confirm with historical data.

[0155] Feasibility assessment: Test the actual application effect of the catalog, examine whether it is easy to understand and use, and whether it can effectively guide daily work practices.

[0156] If any discrepancies are found during the review process, they should be immediately recorded, and solutions should be discussed with the relevant responsible persons.

[0157] Specific modification suggestions were made for the problems found, and some contents of the catalog were readjusted if necessary until the review standards were met.

[0158] Based on all the review comments, a detailed evaluation report was prepared, clearly pointing out the advantages, disadvantages and improvement suggestions of the corrected two-level catalog.

[0159] Determine the final two-level directory structure as the basis for the next steps.

[0160] The evaluation results will be promptly fed back to relevant departments and individuals to ensure that everyone understands the new catalog system and the logic behind it.

[0161] Specialized training courses will be arranged as needed to help employees better master the application skills of the new catalog and improve overall work efficiency.

[0162] By following the steps above, the corrected two-tier directory can be effectively transformed into a truly meaningful two-tier directory structure, providing strong support for the company's strategic planning, operational management, and technology development. Furthermore, this rigorous review mechanism helps to continuously optimize the directory structure, keeping it in optimal condition to adapt to ever-changing market environments and internal needs.

[0163] A two-tiered directory refers to a business scenario extracted from business specifications and the observed objects within it. Based on the actual application of these business scenarios, the attributes and attribute values ​​of the observed objects are further extracted, thus forming a comprehensive many-to-many two-tiered directory covering all business scenarios. For example... Figure 3 As shown.

[0164] Two-level directory format: Adopting a four-level nested structure: "Business Scenario – Observation Object Layer – Observation Object Attributes – Attribute Value Layer". Each level is detailed below:

[0165] Business Scenario: Defines the specific business areas for knowledge application, such as "large vehicle inspection".

[0166] Observation objects: Identify and classify key targets in business scenarios, such as "cranes" and "tower cranes" in "large vehicle inspection".

[0167] Observation object attributes: describe the key features or dimensions of the observed object, such as the "existence" attribute used to determine whether a specific observed object exists.

[0168] Attribute values: provide quantifiable information in various forms, such as discrete values ​​and continuous intervals.

[0169] By employing a hierarchical mapping approach, the fragmentation problem inherent in traditional question-and-answer models is addressed, improving the clarity of business logic and the diversity of reasoning. An LLM model is used for intelligent attribute mining, combined with manual review, enabling efficient and accurate dataset construction. This not only comprehensively covers business scenarios but also demonstrates the correlation between observed object attribute combinations and business logic, ensuring the business significance of the evaluation results. A semi-automated process of "human intervention + large-scale model" is utilized to rapidly process large amounts of data while ensuring accuracy and reducing the subjectivity inherent in traditional manual definitions.

[0170] Suppose we want to build a two-level directory for power grid equipment maintenance for a certain power company:

[0171] Business scenario: Power grid equipment maintenance;

[0172] Observation objects: transformers, cables, substations, etc.;

[0173] Attributes of the observed object: temperature, voltage level, operating status, etc.;

[0174] Attribute values: Normal, Abnormal, High temperature, Low temperature, etc.

[0175] By following the steps above, a detailed two-tiered directory can be built, which not only helps technical personnel better understand and perform maintenance tasks, but also provides decision support for management. Furthermore, through continuous historical data analysis and expert feedback, this directory can be continuously optimized to better align with actual business needs.

[0176] By systematically processing business specification documents to identify core elements, including business scenarios, observation objects, and their attributes, a multi-level nested structure is constructed based on these elements. Through review, adjustment, and verification, a two-layer evaluation directory tightly integrated with business logic is ultimately formed. This method effectively solves three major problems of traditional knowledge graphs in vertical business scenarios: First, it clarifies the hierarchical relationship by clearly defining the multi-level nested structure, avoiding ambiguity; second, it emphasizes starting from business specifications, ensuring a shared semantic framework between business and technical teams, enhancing communication and collaboration efficiency; finally, it utilizes semi-automated methods combined with historical data for verification and improvement, significantly reducing manual costs and achieving an efficient and low-cost knowledge graph construction process. This not only improves construction speed and accuracy but also ensures that the final two-layer evaluation directory closely aligns with actual business needs, possessing high practicality and operability.

[0177] S120, Obtain historical data.

[0178] In this embodiment, historical data refers to a collection of past records and information related to business operations in a specific vertical domain. This data typically contains detailed information about business processes, such as the status of observed objects, operational results, and user feedback, which are crucial for constructing evaluation sets.

[0179] First, it's necessary to identify which data sources can provide valuable historical data. These data sources may include, but are not limited to:

[0180] Internal database: Various business records accumulated within a company or organization.

[0181] Publicly available external resources: published industry reports, data analysis from market research institutions, etc.

[0182] Third-party platforms: Data provided by other companies or platforms that cooperate with the business.

[0183] From the aforementioned data sources, select data directly relevant to the current business scenario and the attributes of the observed objects. For example, in a power AI application scenario, if the focus is on power grid fault detection, historical records involving parameters such as voltage, current, and transformer status need to be collected.

[0184] Because the raw data may contain missing values, outliers, etc., it needs to be cleaned and preprocessed. Specific measures include:

[0185] Remove duplicates: Ensure that each record is unique.

[0186] Imputing missing values: Use statistical methods (such as mean imputation) or other algorithms to estimate missing values.

[0187] Correct outliers: Identify and correct illogical data points.

[0188] The cleaned data undergoes further processing to extract key features that are helpful for analyzing the distribution of test points. This step typically involves complex calculations and algorithmic applications, such as using machine learning models to identify which attribute combinations best reflect key points in actual business operations.

[0189] S130. Based on the two-level directory structure and the historical data, high-frequency or core attribute combinations are statistically analyzed to obtain the test points.

[0190] In this embodiment, "test points" refer to combinations of observed object attributes that frequently occur during business operations or have a significant impact on business outcomes. These attribute combinations are not only the core content of the evaluation system but also key indicators for assessing the performance of AI models. By analyzing historical data and combining it with a two-tiered directory architecture, these key attribute combinations can be identified and included as test points in the evaluation set.

[0191] Specifically, historical data is used to perform statistical analysis on the attributes of the observed objects to identify high-frequency or core attribute combinations in order to obtain the test points.

[0192] Frequency statistics are performed on the attributes of each observed object to identify which attribute combinations most frequently appear in historical data. For example, in power grid maintenance scenarios, the association between certain voltage fluctuation patterns and equipment failures may be high-frequency attribute combinations.

[0193] Given the importance of the business, it's crucial to consider not only the frequency of occurrence but also the degree of impact of the attribute combination on business objectives. For example, even if a certain attribute combination appears infrequently, it should be considered a core assessment point if it directly affects the safe operation of critical equipment.

[0194] Based on the above analysis, we have determined which attribute combinations should be selected as test points. These test points will be used to design subsequent test questions to ensure that the assessment system comprehensively covers key areas of business knowledge.

[0195] The selection of test sites should also take into account their representativeness and general applicability, so as to facilitate their application across different business scenarios.

[0196] For example, suppose we are developing an AI system for power company grid maintenance. In this case:

[0197] Two-tier directory: Business scenarios: power grid maintenance, fault detection, load management, etc.

[0198] Observational objects: voltage, current, temperature, humidity, transformer status, etc.

[0199] Historical data analysis: Filter all data on the power grid's operating status from log files from the past few years.

[0200] After cleaning the data, extract the relevant parameters (such as abnormal voltage or excessively high temperature) for each fault occurrence.

[0201] Attribute combination statistics: It was found that when voltage fluctuations exceed a certain threshold and the temperature is higher than the normal range, transformers are prone to failure. This specific attribute combination (i.e., abnormal voltage + high temperature) is considered a frequently tested topic.

[0202] Identify the test points: Set the combination of "voltage fluctuation + temperature abnormality" as one of the test points, and design a series of test questions around this test point to test whether the AI ​​system can accurately predict potential faults and provide effective maintenance suggestions.

[0203] In this way, step S130 not only identifies key points in business operations but also provides a solid foundation for the design of subsequent evaluation sets. This approach ensures that the evaluation system reflects actual business needs while effectively assessing the performance of the AI ​​system.

[0204] S140. Analyze the question types in the historical data and determine the proportion of each question type.

[0205] In this embodiment, the question type ratio refers to the proportion and coverage of knowledge-based and logic-based test questions based on historical data statistics in a specific business scenario. Determining the question type ratio not only helps in understanding the composition of the current assessment system but also provides a basis for subsequent optimization of assessment content.

[0206] Specifically, the question types in the historical data are analyzed, and the proportion and coverage of knowledge-based and logic-based questions in different business scenarios are identified and calculated to obtain the question type ratio.

[0207] Among them, the knowledge-based questions adopt OpenAI's knowledge ability assessment method, focusing on short knowledge questions to test direct information retrieval and understanding abilities, and classifying long knowledge questions involving multi-step reasoning as questions of logical reasoning tasks;

[0208] The logic questions are based on Charles Sanders Peirce's triadic classification of reasoning, which subdivides logical ability assessment into three core reasoning skills: abduction, induction, and deduction, in order to comprehensively evaluate the logical processing capabilities of the AI ​​system.

[0209] Knowledge-based questions: These questions primarily test the test taker's memory and understanding of basic concepts, principles, and operational procedures. For example, in power AI application scenarios, questions about transformer working principles and power grid maintenance standards fall into this category.

[0210] Logical reasoning questions: These questions focus on assessing the test taker's analytical, reasoning, and problem-solving abilities. For example, given a set of power grid operating parameters, the question asks for the determination of whether there are potential faults and the proposal of corresponding solutions.

[0211] Extract all relevant exam question records from historical data. Ensure this data covers various exam question instances across different business scenarios.

[0212] The selected data is cleaned to remove incomplete or irrelevant information. Then, each question is categorized based on predefined classification criteria (such as knowledge-based and logic-based).

[0213] For each business scenario, the total number of knowledge-based and logic-based questions is counted separately.

[0214] Calculate the percentage of each question type in the total number of questions. For example, in a power grid maintenance scenario, if there are 50 questions, 30 of which are knowledge-based questions and 20 are logic-based questions, then the percentage of knowledge-based questions is 60% and the percentage of logic-based questions is 40%.

[0215] The assessment should determine whether the different types of questions comprehensively cover the key knowledge and skills required for the business area. For example, in power grid maintenance, there should be both sufficient knowledge-based questions about the working principles of equipment and sufficient logic-based questions about fault diagnosis and handling.

[0216] Check that each question type covers different difficulty levels to ensure a comprehensive assessment of the test taker's level. For knowledge-based questions, this might mean both simple memorization questions and complex application questions; for logic-based questions, it should include a range of difficulty levels from basic analysis to advanced reasoning.

[0217] Based on the above analysis, and in conjunction with business needs and assessment objectives, the final question type ratios should be adjusted and determined. If it is found that the proportion of a certain type of question is too high or too low, affecting the balance or effectiveness of the assessment, the number of questions of that type can be increased or decreased.

[0218] At the same time, attention should be paid to maintaining a reasonable combination of question types so that the assessment can not only comprehensively reflect the test taker's knowledge level, but also accurately measure their practical skills and problem-solving abilities.

[0219] Specifically, the knowledge-based question construction method adopts the question construction and evaluation standards for short-form objective knowledge questions in OpenAI's knowledge ability evaluation method. Its core requirement is that each question must have a unique, clear, and unambiguous objective answer (e.g., specifying the scope by limiting it to "which city?" or "which year?"), and the answer should be time-invariant, based on facts that do not change over time (such as historical events or scientific constants). If the content involves potentially changing information, the time point must be clearly stated (e.g., "up to 2023"). Simultaneously, the answer to each question must be directly supported by authoritative sources (such as official documents) to ensure verifiable evidence. Furthermore, questions must be concise and clear, and answers must be short texts (such as words, phrases, numbers, and dates), focusing on short knowledge questions to assess direct information retrieval and matching abilities. Longer knowledge questions involving multi-step reasoning are categorized as logical reasoning tasks. The logical questions are based on Charles Sanders Peirce's triadic classification of reasoning, with logical ability assessment subdivided into three core reasoning skills: abduction, induction, and deduction, to comprehensively evaluate the logical processing capabilities of the AI ​​system. We do not directly use OpenAI's SimpleQA or other mentioned benchmark sets as our question bank. These benchmark sets, in patent documents, serve to define and exemplify the design principles and evaluation criteria for "knowledge-based questions" and "logic-based questions," rather than serving as ready-made data sources. Our actual method for obtaining corresponding questions is based on our unique business domain, generated from scratch through a semi-automated process. The core of this process lies in transforming business knowledge into specific exam questions. Specifically, the process begins by analyzing business specification documents to construct a structured two-layer syllabus consisting of business scenarios, observed objects, and their attributes, clarifying the scope of question generation. Then, combining historical business data statistical analysis or expert experience, we extract high-frequency or key attribute combinations as test points from this syllabus. Next, based on the short question-and-answer and factually unique answer principles advocated by OpenAI's SimpleQA benchmark, we develop question generation prompts adapted to specific business scenarios in our vertical domain, transforming test points into knowledge-based questions, or constructing multi-dimensional logic-based questions by referring to Peirce's three reasoning classification methods. Finally, the generated questions are reviewed and optimized by business experts to ensure their business accuracy and standardized question types, thus forming the final domain evaluation set.

[0220] S150. Based on the test points and question type ratios, select a corresponding number of knowledge-based questions and logic-based questions to form a preliminary evaluation set.

[0221] In this embodiment, the preliminary assessment set refers to a collection of knowledge-based and logic-based questions selected based on the analyzed test points and their corresponding question type ratios, according to a certain allocation strategy (such as random sampling by test point and random sampling by question type). This set aims to comprehensively assess the test taker's knowledge mastery and practical skills in different business scenarios.

[0222] Based on the test points and question types, the number of test questions for each test point is allocated according to a certain ratio, and the proportion of knowledge-based questions and logic-based questions in the evaluation set is determined to obtain the preliminary evaluation set.

[0223] Based on the aforementioned test points and question types, a corresponding number of knowledge-based and logic-based questions are randomly sampled using a strategy of random sampling by test point and random sampling by question type to obtain a preliminary evaluation set.

[0224] Based on the results obtained from step S140, clarify the main test points and the proportion of various question types (knowledge-based and logic-based) in each business scenario.

[0225] Determine the specific knowledge points or skills points that need to be covered under each type of test point, and set the distribution of the number of questions within each test point accordingly.

[0226] Random sampling by test center: For each test center, a corresponding number of questions are randomly selected from the relevant historical dataset based on its importance and the predetermined number of questions. This ensures that the selected questions fully reflect the core content of the test center.

[0227] Random sampling by question type: Based on the required proportion of knowledge-based and logic-based questions, the question types are further subdivided within the selected test points, and a certain number of knowledge-based and logic-based questions are randomly selected from each.

[0228] First, calculate the total number of questions that the assessment set needs to include. Then, allocate the specific number of questions according to the importance of each test point, and determine the specific number of knowledge-based and logic-based questions by considering the proportion of question types.

[0229] Using computer algorithms or other effective methods, knowledge-based and logic-based questions for each test point are randomly selected from historical datasets until the predetermined quantity requirement is met.

[0230] All the questions obtained through the above steps are compiled to form a preliminary evaluation set.

[0231] Check whether the content of the assessment set meets the expected coverage of test points and the required proportion of question types. If necessary, adjust the selection of certain test points or question types to optimize the overall balance of the assessment set.

[0232] S160. Review the preliminary evaluation set and optimize and adjust it based on the review results.

[0233] In this embodiment, after the initial evaluation set is constructed (i.e., step S150), a comprehensive review of the evaluation set is required to ensure the quality and applicability of its content. This process not only helps to identify potential problems or deficiencies, but also allows for necessary optimization and adjustments based on specific feedback, thereby improving the effectiveness and accuracy of the evaluation set.

[0234] Ensure the questions are accurate, unambiguous, and accurately reflect the core requirements of the test points. Verify that the assessment set covers all predetermined test points and question types to ensure comprehensiveness. Assess whether the difficulty level of the questions is appropriate for the target group, avoiding questions that are too difficult or too easy, which could negatively impact the assessment results. Ensure that the questions are designed fairly and impartially for all test takers, without any bias or prejudice.

[0235] The review panel should consist of individuals with relevant expertise and extensive experience. Members may include education experts, industry veterans, and technology consultants.

[0236] Clearly define the standards and guidelines for review, such as the accuracy, clarity, logical consistency, and relevance of the questions to the test points.

[0237] Define the scoring rules and set specific scoring criteria for each audit dimension.

[0238] Each question in the evaluation set is reviewed independently by members of the review team, and any problems or suggestions found are recorded.

[0239] All issues raised during the initial review were discussed collectively, and a consensus was reached.

[0240] Based on the results of the group discussion, the evaluation set was revised as necessary and then reviewed again to ensure that all issues were properly resolved.

[0241] You can invite a small target group to conduct trial tests and obtain first-hand user feedback directly.

[0242] Analyze the test data to understand the actual application effect of the evaluation set, such as the average score rate and common error types.

[0243] Based on the review results and feedback, the preliminary evaluation set was optimized and adjusted accordingly:

[0244] Correcting errors: Correcting any factual errors, unclear statements, or misleading information found.

[0245] Supplementary content: If certain test points are not fully covered, or the number of questions of a certain type is insufficient, then corresponding content needs to be added.

[0246] Adjusting the difficulty: Based on the actual test results, adjust the difficulty of the questions appropriately to make them more suitable for the ability level of the target audience.

[0247] Optimize the structure: Improve the overall structure and layout of the evaluation set, such as rearranging the order of questions to make the transition between knowledge-based questions and logic-based questions more natural and smooth.

[0248] Please see Figure 4 This embodiment's method integrates a two-tiered directory and historical data analysis of test point distribution. After constructing the two-tiered directory based on business specifications, it combines historical data to statistically analyze the attribute combinations of observed objects, thus deriving the test point distribution. Specifically, it first filters information matching the attributes of observed objects from historical data, then statistically analyzes the attribute combinations of observed objects appearing in the historical data, and determines high-frequency or core attribute combinations as test points based on the frequency and business importance of the attribute combinations.

[0249] Based on historical data, a distribution analysis of exam question types was conducted to determine the distribution of two main question types: knowledge and logic. The proportion and scenario coverage of these two question types in historical exams were calculated to clarify their distribution characteristics across different business scenarios.

[0250] Based on the distribution sampling of test points and test questions, and by fitting the distribution of test point distribution and test question type distribution, a preliminary evaluation set is generated. First, based on the frequency of test point distribution, the number of test questions for each test point is allocated proportionally. Simultaneously, combined with the question type distribution table, the proportion of knowledge-based and logic-based questions in the evaluation set is determined. Based on these proportions, test questions are randomly sampled by test point and by question type. The sampling results are then summarized to form a preliminary evaluation set.

[0251] The data distribution is dynamically adjusted based on the review process. Business experts are submitted to review the rationality of the distribution, and dynamic adjustments are made based on the feedback.

[0252] The syllabus is a two-tiered directory, derived from business specifications (such as policy documents, operation manuals, industry standards, and regulatory rules). It comprises two layers: business scenarios and observation objects. Observation objects can be further subdivided into observation object attributes and attribute values, and observation objects can be reused across multiple scenarios. Ultimately, the two-tiered directory is formed through the structure of business scenarios and observation objects. The syllabus defines the scope of business knowledge to be tested, ensuring that test cases have business significance and meet actual business needs.

[0253] The test points are the combinations of observation objects and attributes in the two-level directory. They are based on the discovery and summarization of common mistakes and key points in the business process. Specifically, they are the combinations of multiple observation objects and attributes in the two-level directory.

[0254] The test questions are constructed using a "1+3 question method," with one type of short knowledge-based questions to test knowledge ability, and three types of questions—abductive reasoning, inductive reasoning, and deductive reasoning—to assess logical ability, thus achieving complete coverage of knowledge reserves and logical reasoning ability for large models.

[0255] The knowledge and ability test questions refer to OpenAI's knowledge and ability evaluation methods (applied to benchmark sets such as openai-SimpleQA, openai-SimpleVQA, and Alibaba-Chinese SimpleQA). The knowledge and ability assessment is divided into short knowledge questions and long knowledge questions. This evaluation system will focus on the in-depth analysis of short knowledge questions, while long knowledge questions rely on multi-step deduction and are essentially classified as logical reasoning tasks.

[0256] The logic test questions are based on the triadic classification of reasoning proposed by Charles Sanders Peirce (this classification method has been adopted and developed in the field of artificial intelligence, such as the MME-Reasoning benchmark study case). In terms of AI logic ability, it can be further subdivided into abduction, induction, deduction and other abilities.

[0257] Compared to existing technologies, the method in this embodiment employs a semi-automated data distribution analysis approach that fits the production distribution. It addresses three major challenges: test dataset distribution deviating from actual business operations, test data being disconnected from business knowledge, and the inability to simultaneously balance efficiency and accuracy in test data construction. The method optimizes these challenges through distribution fitting and semi-automation, offering the following advantages:

[0258] This method uses dual business pillars (business knowledge and business data) to fit the business distribution. The core objective of vertical domain AI application evaluation is to ensure that the performance of AI applications on real business systems is consistent with the evaluation results. Therefore, the data distribution of the test set needs to be consistent with the data distribution of the business system, i.e., it must be representative of the business. To achieve this goal, this embodiment proposes a test set data analysis process supported by dual business pillars. Based on business knowledge, it performs distribution analysis on all business data from the perspective of business applications, achieving full coverage of business scenarios and observation object attributes, ensuring that the test set comprehensively and accurately reflects the distribution of business data.

[0259] This paper proposes a semi-automated data analysis method combining expert and large-scale model construction to improve human-machine collaboration efficiency. Vertical AI applications often require data analysis of massive amounts of business data, and there are significant professional barriers to vertical business knowledge. Therefore, the efficiency of the evaluation data construction method and the high quality of the evaluation dataset are extremely important. To achieve this goal, this embodiment proposes a semi-automated data analysis method based on a combination of large-scale models and human business expert review. The large-scale model improves the initial data analysis efficiency, while the human business expert review enhances the accuracy of the dataset and the output quality. This achieves synergy between expert and large-scale model technologies, reduces the investment cost of business experts, and achieves evaluation data analysis that balances efficiency and quality.

[0260] Therefore, this embodiment studies a data analysis method for fitting production distribution to address the discrepancy between the test dataset distribution and the actual production distribution. The goal is to ensure that the model's performance on test data accurately reflects its performance in the production environment.

[0261] Explore methods for integrating and mapping business knowledge with test data to overcome the disconnect between the two. By establishing a mapping relationship between a tagging system and business rules, ensure that the test point combination covers the core areas of business concern and complies with basic industry rules, thereby using test data to accurately evaluate the model's performance on key business test points.

[0262] This paper develops a method that combines automated data collection with manual review to address the problems of high cost and time consumption associated with fully manual annotation, and the tendency for fully automated data collection to deviate from actual business applications. This method aims to balance cost-effectiveness and accuracy, promoting the effective construction of datasets for evaluating the effectiveness of AI applications in vertical industries.

[0263] The aforementioned semi-automated fitting method for vertical domain AI application evaluation data distribution acquires and analyzes business specification documents and historical data to construct a two-layer directory structure that covers complete business scenarios and observed object attribute combinations. Based on this, it determines high-frequency or core test points and question type ratios. Subsequently, based on this information, it extracts a corresponding number of questions from knowledge-based and logic-based questions to form an initial evaluation set. Expert review and dynamic adjustments ensure the accuracy and business representativeness of the evaluation set. This method solves the utility problem of existing evaluation datasets in vertical domain applications, ensuring that model performance evaluation not only accurately reflects actual business performance but also significantly reduces construction costs and time consumption through human-machine collaboration, maximizing cost-effectiveness. This approach guarantees the evaluation dataset has high business relevance and accuracy while effectively improving construction efficiency and adapting to rapidly changing business needs.

[0264] Figure 5 This is a schematic block diagram of a semi-automatic fitting system 300 for evaluating the distribution of vertical AI application data, provided in an embodiment of the present invention. Figure 5As shown, corresponding to the above-described semi-automatic fitting method for vertical AI application evaluation data distribution, this invention also provides a semi-automatic fitting system 300 for vertical AI application evaluation data distribution. This semi-automatic fitting system 300 includes a unit for executing the above-described semi-automatic fitting method for vertical AI application evaluation data distribution, and the system can be configured in a server. Specifically, please refer to... Figure 5 The semi-automatic fitting system 300 for the vertical domain AI application evaluation data distribution includes a construction unit 301, a historical data acquisition unit 302, a test point determination unit 303, a question type ratio determination unit 304, a preliminary evaluation set generation unit 305, and an auditing unit 306.

[0265] The system comprises: a construction unit 301 for acquiring business specification documents and constructing a two-level directory structure based on these documents; a historical data acquisition unit 302 for acquiring historical data; a test point determination unit 303 for statistically analyzing high-frequency or core attribute combinations based on the two-level directory structure and the historical data to obtain test points; a question type ratio determination unit 304 for analyzing the question types in the historical data and determining the question type ratio; a preliminary evaluation set generation unit 305 for extracting a corresponding number of knowledge-based and logic-based questions to form a preliminary evaluation set based on the test points and question type ratio; and an auditing unit 306 for auditing the preliminary evaluation set and optimizing and adjusting it based on the audit results.

[0266] In one embodiment, such as Figure 6 As shown, the construction unit 301 includes a document acquisition subunit 3011, an analysis subunit 3012, a structure construction subunit 3013, a review subunit 3014, an improvement subunit 3015, and an audit subunit 3016.

[0267] The document acquisition subunit 3011 is used to acquire business specification documents; the analysis subunit 3012 is used to analyze the business specification documents to obtain core elements, wherein the core elements include business scenarios, observation objects and their attributes; the structure construction subunit 3013 is used to construct a multi-level nested structure based on the core elements to obtain a preliminary two-level directory; the review subunit 3014 is used to review and adjust the preliminary two-level directory to obtain an adjusted two-level directory; the improvement subunit 3015 is used to use historical business data to verify and improve the adjusted two-level directory to obtain a corrected two-level directory; and the audit subunit 3016 is used to conduct a final audit of the corrected two-level directory to obtain the two-level directory structure.

[0268] In one embodiment, the test point determination unit 303 is used to perform statistical analysis on the attributes of the observed object using historical data, identify high-frequency or core attribute combinations, and obtain test points.

[0269] In one embodiment, the question type ratio determination unit 304 is used to analyze the question types in the historical data, identify and calculate the proportion and coverage of knowledge-based and logic-based questions in different business scenarios, so as to obtain the question type ratio.

[0270] In one embodiment, the preliminary evaluation set generation unit 305 is used to allocate the number of test questions for each test point according to a certain ratio based on the test point and question type ratio, and to determine the proportion of knowledge-based questions and logic-based questions in the evaluation set, so as to obtain a preliminary evaluation set.

[0271] In one embodiment, the preliminary evaluation set generation unit 305 is used to extract a corresponding number of knowledge-based questions and logic-based questions according to the ratio of test points and question types, using a strategy of random sampling by test points and random sampling by question types, in order to obtain a preliminary evaluation set.

[0272] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned semi-automatic fitting system 300 for vertical AI application evaluation data distribution and each unit can be referred to the corresponding description in the aforementioned method embodiments. For the sake of convenience and brevity, it will not be repeated here.

[0273] The aforementioned semi-automatic fitting system 300 for evaluating the distribution of vertical AI application evaluation data can be implemented as a computer program, which can, for example... Figure 7 It runs on the computer device shown.

[0274] Please see Figure 7 , Figure 7 This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a server, wherein the server can be a standalone server or a server cluster composed of multiple servers.

[0275] See Figure 7 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.

[0276] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, cause the processor 502 to perform a semi-automatic fitting method for the distribution of evaluation data for vertical AI applications.

[0277] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.

[0278] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a semi-automatic fitting method for the distribution of evaluation data for vertical AI applications.

[0279] This network interface 505 is used for network communication with other devices. Those skilled in the art will understand that... Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0280] The processor 502 is used to run a computer program 5032 stored in the memory to perform the following steps:

[0281] Obtain business specification documents and construct a two-level directory structure based on the business specification documents; obtain historical data; based on the two-level directory structure and the historical data, statistically analyze high-frequency or core attribute combinations to obtain test points; analyze the test question types in the historical data and determine the question type ratio; according to the test points and question type ratio, extract a corresponding number of knowledge-based questions and logic-based questions to form a preliminary evaluation set; review the preliminary evaluation set and optimize and adjust the preliminary evaluation set based on the review results.

[0282] The knowledge-based questions are assessed using OpenAI's knowledge ability evaluation method, focusing on short knowledge questions to test direct information retrieval and comprehension abilities, while long knowledge questions involving multi-step reasoning are classified as logical reasoning tasks.

[0283] The logic questions are based on Charles Sanders Peirce's triadic classification of reasoning, which subdivides logical ability assessment into three core reasoning skills: abduction, induction, and deduction, in order to comprehensively evaluate the logical processing capabilities of the AI ​​system.

[0284] In one embodiment, when the processor 502 implements the step of obtaining the business specification document and constructing a two-level directory structure based on the business specification document, it specifically implements the following steps:

[0285] Obtain the business specification document; analyze the business specification document to obtain core elements, wherein the core elements include business scenarios, observation objects and their attributes; construct a multi-level nested structure based on the core elements to obtain a preliminary two-level directory; review and adjust the preliminary two-level directory to obtain an adjusted two-level directory; verify and improve the adjusted two-level directory using historical business data to obtain a corrected two-level directory; conduct a final review of the corrected two-level directory to obtain the two-level directory structure.

[0286] In one embodiment, when the processor 502 implements the step of combining the two-level directory structure with historical data to statistically analyze high-frequency or core attribute combinations to obtain test points, the specific implementation steps are as follows:

[0287] By using historical data to perform statistical analysis on the attributes of the observed objects, high-frequency or core attribute combinations can be identified to obtain the test points.

[0288] In one embodiment, when the processor 502 performs the step of analyzing the question types in the historical data and determining the proportion of question types, it specifically implements the following steps:

[0289] Analyze the question types in the historical data, identify and calculate the proportion and coverage of knowledge-based and logic-based questions in different business scenarios, so as to obtain the question type ratio.

[0290] In one embodiment, when the processor 502 implements the step of extracting a corresponding number of knowledge-based questions and logic-based questions to form a preliminary evaluation set based on the ratio of test points and question types, the specific implementation steps are as follows:

[0291] Based on the test points and question types, the number of test questions for each test point is allocated according to a certain ratio, and the proportion of knowledge-based questions and logic-based questions in the evaluation set is determined to obtain the preliminary evaluation set.

[0292] In one embodiment, when the processor 502 implements the step of allocating the number of test questions for each test point according to a certain ratio based on the test point and question type ratio, and determining the proportion of knowledge-based questions and logic-based questions in the evaluation set to obtain a preliminary evaluation set, the processor 502 specifically implements the following steps:

[0293] Based on the aforementioned test points and question types, a corresponding number of knowledge-based and logic-based questions are randomly sampled using a strategy of random sampling by test point and random sampling by question type to obtain a preliminary evaluation set.

[0294] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0295] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.

[0296] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein when executed by a processor, the computer program causes the processor to perform the following steps:

[0297] Obtain business specification documents and construct a two-level directory structure based on the business specification documents; obtain historical data; based on the two-level directory structure and the historical data, statistically analyze high-frequency or core attribute combinations to obtain test points; analyze the test question types in the historical data and determine the question type ratio; according to the test points and question type ratio, extract a corresponding number of knowledge-based questions and logic-based questions to form a preliminary evaluation set; review the preliminary evaluation set and optimize and adjust the preliminary evaluation set based on the review results.

[0298] The knowledge-based questions are assessed using OpenAI's knowledge ability evaluation method, focusing on short knowledge questions to test direct information retrieval and comprehension abilities, while long knowledge questions involving multi-step reasoning are classified as logical reasoning tasks.

[0299] The logic questions are based on Charles Sanders Peirce's triadic classification of reasoning, which subdivides logical ability assessment into three core reasoning skills: abduction, induction, and deduction, in order to comprehensively evaluate the logical processing capabilities of the AI ​​system.

[0300] In one embodiment, when the processor executes the computer program to implement the steps of obtaining the business specification document and constructing a two-level directory structure based on the business specification document, the processor specifically implements the following steps:

[0301] Obtain the business specification document; analyze the business specification document to obtain core elements, wherein the core elements include business scenarios, observation objects and their attributes; construct a multi-level nested structure based on the core elements to obtain a preliminary two-level directory; review and adjust the preliminary two-level directory to obtain an adjusted two-level directory; verify and improve the adjusted two-level directory using historical business data to obtain a corrected two-level directory; conduct a final review of the corrected two-level directory to obtain the two-level directory structure.

[0302] In one embodiment, when the processor executes the computer program to implement the step of statistically analyzing high-frequency or core attribute combinations based on the two-level directory structure and historical data to obtain test points, the specific implementation steps are as follows:

[0303] By using historical data to perform statistical analysis on the attributes of the observed objects, high-frequency or core attribute combinations can be identified to obtain the test points.

[0304] In one embodiment, when the processor executes the computer program to analyze the question types in the historical data and determine the proportion of question types, it specifically implements the following steps:

[0305] Analyze the question types in the historical data, identify and calculate the proportion and coverage of knowledge-based and logic-based questions in different business scenarios, so as to obtain the question type ratio.

[0306] In one embodiment, when the processor executes the computer program to implement the step of extracting a corresponding number of knowledge-based questions and logic-based questions to form a preliminary evaluation set according to the ratio of test points and question types, the specific implementation steps are as follows:

[0307] Based on the test points and question types, the number of test questions for each test point is allocated according to a certain ratio, and the proportion of knowledge-based questions and logic-based questions in the evaluation set is determined to obtain the preliminary evaluation set.

[0308] In one embodiment, when the processor executes the computer program to implement the steps of allocating the number of test questions for each test point according to a certain ratio based on the test point and question type ratio, and determining the proportion of knowledge-based questions and logic-based questions in the evaluation set to obtain a preliminary evaluation set, the processor specifically implements the following steps:

[0309] Based on the aforementioned test points and question types, a corresponding number of knowledge-based and logic-based questions are randomly sampled using a strategy of random sampling by test point and random sampling by question type to obtain a preliminary evaluation set.

[0310] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.

[0311] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0312] In the embodiments provided by this invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of each unit is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0313] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the system of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0314] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0315] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A semi-automatic fitting method for the distribution of evaluation data in vertical AI applications, characterized in that, include: Obtain the business specification document and construct a two-level directory structure based on the business specification document; Obtain historical data; Based on the aforementioned two-layer directory structure and the aforementioned historical data, high-frequency or core attribute combinations are statistically analyzed to obtain the test points; Analyze the question types in the historical data to determine the proportion of each question type; Based on the test points and question types, a corresponding number of knowledge-based questions and logic-based questions were selected to form a preliminary evaluation set. The preliminary evaluation set is reviewed, and the preliminary evaluation set is optimized and adjusted based on the review results; The step of obtaining the business specification document and constructing a two-level directory structure based on the business specification document includes: Obtain the business specification document; The business specification document is analyzed to obtain core elements, which include business scenarios, observation objects and their attributes; Based on the core elements, a multi-level nested structure is constructed to obtain a preliminary two-level directory. The multi-level nested structure includes a business scenario, an observation object layer, observation object attributes, and an attribute value layer. The preliminary two-level catalog is reviewed and adjusted to obtain the adjusted two-level catalog; The adjusted two-level directory is verified and improved using historical business data to obtain a corrected two-level directory. The corrected two-level directory is then subject to final review to obtain the two-level directory structure.

2. The semi-automatic fitting method for vertical domain AI application evaluation data distribution according to claim 1, characterized in that, The method of combining the two-level directory structure with historical data to statistically analyze high-frequency or core attribute combinations to obtain test points includes: By using historical data to perform statistical analysis on the attributes of the observed objects, high-frequency or core attribute combinations can be identified to obtain the test points.

3. The semi-automatic fitting method for vertical domain AI application evaluation data distribution according to claim 1, characterized in that, The analysis of the question types in the historical data to determine the proportion of each question type includes: Analyze the question types in the historical data, identify and calculate the proportion and coverage of knowledge-based and logic-based questions in different business scenarios, so as to obtain the question type ratio.

4. The semi-automatic fitting method for vertical domain AI application evaluation data distribution according to claim 1, characterized in that, Based on the aforementioned test points and question type ratios, a corresponding number of knowledge-based and logic-based questions are selected to form a preliminary assessment set, including: Based on the test points and question types, the number of test questions for each test point is allocated according to a certain ratio, and the proportion of knowledge-based questions and logic-based questions in the evaluation set is determined to obtain the preliminary evaluation set.

5. The semi-automatic fitting method for vertical AI application evaluation data distribution according to claim 4, characterized in that, The method involves allocating the number of questions for each test point according to a certain ratio based on the test points and question types, and determining the proportion of knowledge-based questions and logic-based questions in the assessment set to obtain a preliminary assessment set, including: Based on the aforementioned test points and question types, a corresponding number of knowledge-based and logic-based questions are randomly sampled using a strategy of random sampling by test point and random sampling by question type to obtain a preliminary evaluation set.

6. The semi-automatic fitting method for vertical domain AI application evaluation data distribution according to claim 1, characterized in that, The knowledge-based questions are assessed using OpenAI's knowledge ability evaluation method, focusing on short knowledge questions to test direct information retrieval and comprehension abilities, while long knowledge questions involving multi-step reasoning are classified as logical reasoning tasks. The logic questions are based on Charles Sanders Peirce's triadic classification of reasoning, which subdivides logical ability assessment into three core reasoning skills: abduction, induction, and deduction, in order to comprehensively evaluate the logical processing capabilities of the AI ​​system.

7. A semi-automatic fitting system for vertical AI application evaluation data distribution, wherein, during operation, the semi-automatic fitting method for vertical AI application evaluation data distribution as described in any one of claims 1-6 is characterized in that, include: A building unit is used to obtain business specification documents and build a two-level directory structure based on the business specification documents; Historical data acquisition unit, used to acquire historical data; The test point determination unit is used to obtain test points by statistically analyzing high-frequency or core attribute combinations based on the two-level directory structure and the historical data. The question type ratio determination unit is used to analyze the question types in the historical data and determine the question type ratio. The preliminary assessment set generation unit is used to extract a corresponding number of knowledge-based questions and logic-based questions to form a preliminary assessment set based on the test points and question type ratios. The review unit is used to review the preliminary evaluation set and optimize and adjust the preliminary evaluation set based on the review results.

8. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method as described in any one of claims 1 to 6.

9. A storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Test paper generation method and device

    CN112699283A

  • Method for automatically generating multi-task objective question evaluation set in vertical field of large language model

    CN119227818A