Knowledge base construction methods, devices, equipment, storage media and products

By preprocessing the original business materials and parsing the large model, and combining multiple correction strategies to build a high-quality business process knowledge base, the problem of knowledge base construction in the initial stage of intelligent systems is solved, and efficient initialization of intelligent systems is achieved.

CN122088643APending Publication Date: 2026-05-26CHINA UNICOM (SHANGHAI) IND INTERNET CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNICOM (SHANGHAI) IND INTERNET CO LTD
Filing Date
2026-02-09
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies cannot effectively utilize unstructured business knowledge to build a knowledge base in the initial stage of intelligent system construction, leading to difficulties in the initial deployment of intelligent systems.

Method used

By preprocessing the original business materials, using a large model for semantic parsing to generate candidate business data, and making multiple corrections when the preset indicators do not meet the conditions, until the semantic accuracy, semantic novelty and semantic stability indicators are met, a business process knowledge base is constructed.

Benefits of technology

It enables the rapid and accurate extraction of relevant business knowledge from unstructured business materials, builds a high-quality business process knowledge base, and supports the initial deployment of intelligent systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122088643A_ABST
    Figure CN122088643A_ABST
Patent Text Reader

Abstract

This invention discloses a method, apparatus, device, storage medium, and product for constructing a knowledge base. The method includes: preprocessing original business materials to obtain initial business data; performing semantic parsing on the initial business data using a large model to generate candidate business data; if preset indicators of the candidate business data do not meet preset conditions, revising the candidate business data multiple times according to a preset revision strategy until the preset conditions are met; wherein the preset indicators include at least one of semantic accuracy, semantic novelty, and semantic stability indicators of the candidate business data; determining target business data based on the revision results, and constructing a business process knowledge base based on the target business data. This invention can construct a business process knowledge base based on original business materials in the initial stage of intelligent system construction, thereby supporting the initial deployment of the intelligent system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent cognitive technology, and in particular to methods, apparatus, devices, storage media and products for knowledge base construction. Background Technology

[0002] With the widespread application of large-scale models and intelligent systems in government, finance, and manufacturing industries, enterprises' demand for building intelligent question answering, process automation, and business decision-making systems continues to grow. However, during the system construction process, business knowledge in specialized fields often exists in unstructured materials or expert experience. This knowledge cannot be directly used for system construction or training. Therefore, it is necessary to build a knowledge base to provide knowledge sources for intelligent systems.

[0003] Currently, existing technologies for building knowledge bases largely rely on pre-set expert rule templates, structured tag samples, and historical question-and-answer data. However, these methods are suitable for enhancing or fine-tuning the system when existing knowledge has been accumulated, but not for building a knowledge base under cold-start conditions in the initial stage of system construction. Therefore, there is an urgent need for a knowledge base construction method to support the initial deployment of intelligent systems. Summary of the Invention

[0004] This invention provides a knowledge base construction method, apparatus, device, storage medium, and product, enabling the rapid and accurate extraction of relevant business knowledge from unstructured business materials during the initial stage of intelligent system construction to generate a business knowledge base, thereby supporting the initial deployment of the intelligent system.

[0005] According to one aspect of the present invention, a method for constructing a knowledge base is provided, comprising: Preprocess the original business materials to obtain initial business data; The initial business data is semantically parsed using a large model, and candidate business data is generated. If the preset indicators of the candidate business data do not meet the preset conditions, the candidate business data shall be corrected multiple times according to the preset correction strategy until the preset conditions are met; wherein, the preset indicators include at least one of the semantic accuracy indicator, semantic novelty indicator, and semantic stability indicator of the candidate business data. The target business data is determined based on the correction results, and a business process knowledge base is constructed based on the target business data.

[0006] According to another aspect of the present invention, a knowledge base construction apparatus is provided, comprising: The preprocessing module is used to preprocess the original business materials to obtain initial business data; The candidate business data generation module is used to perform semantic parsing on the initial business data using a large model and generate candidate business data. The correction module is used to correct the candidate business data multiple times according to a preset correction strategy when the preset indicators of the candidate business data do not meet the preset conditions, until the preset conditions are met; wherein, the preset indicators include at least one of the semantic accuracy indicator, semantic novelty indicator, and semantic stability indicator of the candidate business data. The business process knowledge base construction module is used to determine the target business data based on the correction results, and to construct the business process knowledge base based on the target business data.

[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the knowledge base construction method according to any embodiment of the present invention.

[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the knowledge base construction method according to any embodiment of the present invention.

[0009] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the knowledge base construction method described in any embodiment of the present invention.

[0010] The technical solution of this invention preprocesses the original business materials to obtain initial business data, providing a foundation for subsequent extraction of business knowledge using a large model. By using the large model to perform semantic parsing on the initial business data and generating candidate business data, it enables rapid mining of potential semantic information within the initial business data. If the preset indicators of the candidate business data do not meet preset conditions, the candidate business data is corrected multiple times according to a preset correction strategy until the preset conditions are met. The preset indicators include at least one of the semantic accuracy, semantic novelty, and semantic stability indicators of the candidate business data. These steps ensure that the candidate business data is semantically accurate, novel, and stable, resulting in higher data quality for the subsequently obtained target business data. By determining the target business data based on the correction results and constructing a business process knowledge base based on the target business data, it is possible to ensure that the intelligent system deployed based on this business process knowledge base has more comprehensive and accurate knowledge support.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart of a knowledge base construction method provided in Embodiment 1 of the present invention; Figure 2 This is a flowchart of a knowledge base construction method provided in Embodiment 2 of the present invention; Figure 3 This is a flowchart of a knowledge base construction method provided in Embodiment 3 of the present invention; Figure 4 This is a schematic diagram of the structure of a knowledge base construction device according to Embodiment 4 of the present invention; Figure 5 This is a schematic diagram of the structure of an electronic device that implements the knowledge base construction method of Embodiment 5 of the present invention. Detailed Implementation

[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0015] It should be noted that the terms "first," "second," "initial," and "target," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0016] Example 1 Figure 1 This is a flowchart illustrating a knowledge base construction method provided in Embodiment 1 of the present invention. This embodiment is applicable to the construction of business knowledge graphs or business knowledge bases. The method can be executed by a knowledge base construction device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes: S101. Preprocess the original business materials to obtain initial business data.

[0017] In this embodiment, raw business materials can be unstructured original records generated by an enterprise in a certain field (e.g., the power industry or the automotive industry) during its daily operations. Examples include Standard Operating Procedure (SOP) documents, business rule documents, historical meeting minutes, historical interview records, product specifications, and enterprise policy-related documents. Raw business materials can include various forms of multimodal materials related to the business, such as text, tables, and images. Initial business data can be understood as structured data that provides a semantic foundation for large models to understand the enterprise's business knowledge.

[0018] For example, the categories of original business materials are determined based on file type and business theme; the text content of the original business materials is cleaned and semantically segmented, for example, incomplete text content is removed, long texts are partially segmented according to preset identifiers (e.g., pre-defined symbols) and / or preset vocabulary, and semantically repetitive text content is removed; all original business materials are unified to the same format (e.g., JSON or CSV format), and the metadata of each original business material is extracted to add a structured index to each original business material, thereby obtaining initial business data; wherein, the structured index may include source information, business theme, and time information corresponding to the initial business data. Through the above preprocessing, the originally scattered and unstructured original business materials can be transformed into structured data that can be used by large models.

[0019] S102. Use the large model to perform semantic parsing on the initial business data and generate candidate business data.

[0020] In this embodiment, the large model can be a Large Language Model (LLM). LLMs have massive amounts of sample data and model parameters, for example, billions or even hundreds of billions of parameters. They are trained on large-scale text data or other modal data and can be used for various natural language processing tasks, exhibiting excellent text generation capabilities. For example, each initial business data item can be input into the large language model, which then outputs corresponding candidate business data. This candidate business data may include correct business knowledge or incorrect business knowledge. Each candidate business data item generally includes business scenario content, triggering condition content, and business action or conclusion content. For example, "A customer's purchase order exceeds 10,000 yuan and requires manager approval before goods can be prepared." Here, "customer purchase order" can be understood as business scenario content, "customer's purchase order exceeds 10,000 yuan" as triggering condition content, and "requires manager approval before goods can be prepared" as business action or conclusion content.

[0021] S103. If the preset indicators of the candidate business data do not meet the preset conditions, the candidate business data shall be corrected multiple times according to the preset correction strategy until the preset conditions are met; wherein, the preset indicators include at least one of the semantic accuracy indicator, semantic novelty indicator, and semantic stability indicator of the candidate business data.

[0022] In this embodiment, the preset correction strategy can be understood as a pre-defined strategy for correcting candidate business data that does not meet preset conditions. For example, adjusting the prompt information of the large model so that the large model can output more accurate candidate business data.

[0023] When determining the semantic accuracy metric of candidate business data, it is possible to judge whether the candidate business data violates preset constraint rules, conflicts with existing rules, and contains logically contradictory descriptions. When determining the semantic novelty metric of candidate business data, it is possible to judge whether the candidate business data introduces new business scenario content, supplements previously uncovered trigger condition content and corresponding business action or conclusion content, and whether the similarity with existing samples is lower than a preset similarity threshold; where existing samples can be previously generated candidate business data that meets preset conditions. When determining the semantic stability metric of candidate business data, it is possible to calculate the degree of difference between candidate business data and historical candidate business data (which can be candidate business data generated by the large model in historical rounds), determine whether unexpected drift occurs in the key decision path (unexpected drift can be understood as a mismatch between trigger condition content and business action or conclusion content; the key decision path can be identified by pre-defined keywords), and calculate the proportion of content with consistent semantics in the business action or conclusion content of the candidate business data when trigger condition content contains the same or similar semantics (e.g., different descriptions of the same term can be trigger condition content with similar semantics).

[0024] The preset conditions can be set according to preset indicators. For example, the above indicators can be quantified into numerical scores to obtain the score corresponding to each candidate business data. Score thresholds can be set for each indicator, and then the preset conditions corresponding to each indicator can be determined according to the score thresholds corresponding to each indicator. Alternatively, a total score threshold can be set based on the sum of the scores of each indicator, and the preset conditions can be determined according to the total score threshold. This invention does not impose any specific limitations.

[0025] For example, the semantic accuracy, semantic novelty, and semantic stability scores of the candidate business data are calculated respectively, and it is determined whether each score is greater than or equal to the corresponding score threshold. If there are cases where the scores are less than the score threshold, the prompt information of the large model can be adjusted, and the candidate business data can be regenerated using the large model until the scores of the semantic accuracy, semantic novelty, and semantic stability of the newly generated candidate business data are greater than or equal to the corresponding score threshold.

[0026] S104. Determine the target business data based on the correction results, and construct a business process knowledge base based on the target business data.

[0027] In this embodiment, the correction result can be understood as the latest candidate business data obtained after the candidate business data is corrected for the last time according to the preset correction strategy, and all preset indicators of the latest candidate business data meet the preset conditions. The target business data can be data in a structured form, such as data in the form of "condition-conclusion". The latest candidate business data can be extracted into rules to obtain the target business data; and the target business data can be structurally decomposed to build a business process knowledge base. The rule extraction can include: statistically analyzing frequently occurring business condition combinations and triggering events, constructing a condition-conclusion rule expression based on the business condition combinations and triggering events, summarizing stable and recurring decision paths and business patterns based on the constructed rule expressions, and generating reusable business rules and process descriptions.

[0028] A business process knowledge base can be understood as a collection of business knowledge that can be retrieved and used by large models or intelligent systems (e.g., intelligent question-answering systems). The constructed business process knowledge base can serve as a knowledge base for these large models, ensuring that they possess more specialized knowledge and higher efficiency when subsequently parsing initial business data in the same domain. It can also be used to provide business knowledge support for other intelligent systems (e.g., intelligent business question-answering systems).

[0029] Optionally, the method further includes: when the preset indicators of the candidate business data meet preset conditions, marking the candidate business data as high-quality business data, and the high-quality business data can be used to determine the target business data.

[0030] For example, rules are extracted from high-quality business data to identify target business data.

[0031] This invention provides a knowledge base construction method. It preprocesses original business materials to obtain initial business data, laying the foundation for subsequent extraction of business knowledge using a large model. By using the large model to perform semantic parsing on the initial business data and generating candidate business data, it enables rapid mining of potential semantic information within the initial business data. If the preset indicators of the candidate business data do not meet preset conditions, the candidate business data is corrected multiple times according to a preset correction strategy until the preset conditions are met. The preset indicators include at least one of semantic accuracy, semantic novelty, and semantic stability indicators of the candidate business data. These steps ensure that the candidate business data is semantically accurate, novel, and stable, resulting in higher data quality for the subsequently obtained target business data. By determining the target business data based on the correction results and constructing a business process knowledge base based on the target business data, it is possible to ensure that the intelligent system deployed based on this business process knowledge base has more comprehensive and accurate knowledge support.

[0032] Example 2 Figure 2 This is a flowchart of a knowledge base construction method provided in Embodiment 2 of the present invention. This embodiment is a refinement based on the above embodiments. Figure 2 As shown, the method includes: S201. Preprocess the original business materials to obtain initial business data.

[0033] S202. Based on preset prompts, the ICL paradigm of context learning guides the large model to perform semantic parsing on the initial business data and generate candidate business data.

[0034] The preset prompt information includes the generation strategy corresponding to the candidate business data; the generation strategy is used to guide the large model to extract standard business processes, abnormal business processes, business rule boundary conditions, and different expressions of business terms from the initial business data, so as to generate candidate business data of various business types.

[0035] In this embodiment, the preset prompt information can be constructed based on the prompt template corresponding to the large model; the prompt template may include constraint columns such as business objectives, role constraints, terminology conventions, and logical restrictions, and specific content can be filled in under each constraint column to obtain the preset prompt information.

[0036] For example, a preset number (e.g., 1 to 2) of example samples are embedded in the prompt template of the large model. These include at least one example sample from standard business process examples, abnormal business process examples, business rule boundary condition examples, and different expressions of business terms. Using the few-shot in-context learning (ICL) method, the large model is guided to perform semantic understanding of the initial business data. According to the generation strategy specified by the preset prompt information in the prompt template, standard business processes, abnormal business processes, business rule boundary conditions, and different expressions of business terms are extracted from the initial business data. Candidate business data is generated based on the above extraction results.

[0037] S203. If the preset indicators of the candidate business data do not meet the preset conditions, the candidate business data shall be corrected multiple times according to the preset correction strategy until the preset conditions are met; wherein, the preset indicators include at least one of the semantic accuracy indicator, semantic novelty indicator, and semantic stability indicator of the candidate business data.

[0038] Optionally, the preset conditions include a first preset condition corresponding to the semantic accuracy index, a second preset condition corresponding to the semantic novelty index, and a third preset condition corresponding to the semantic stability index; wherein, when the preset indicators of the candidate service data do not meet the preset conditions, the candidate service data is corrected multiple times according to a preset correction strategy, including at least one of the following A1 to A3: A1. If the semantic accuracy index does not meet the first preset condition, increase the proportion of rule-constrained prompt words in the preset prompt information of the large model, and regenerate candidate business data.

[0039] In this embodiment, the first preset condition can be determined based on a first score threshold preset for the semantic accuracy indicator. For example, if the score of the first indicator corresponding to the semantic accuracy indicator is lower than the first score threshold, more rule constraint-type prompt words are added to the prompt template of the large model to obtain the first preset prompt information, and the first preset prompt information is used to guide the large model to regenerate candidate business data.

[0040] A2. If the semantic novelty index does not meet the second preset condition, add a preset novel example to the preset prompt information of the large model and regenerate the candidate business data; wherein the similarity between the preset novel example and the candidate business data is lower than the preset similarity threshold.

[0041] In this embodiment, the second preset condition can be determined based on a second score threshold preset for the semantic novelty index. For example, if the score of the second index corresponding to the semantic novelty index is lower than the second score threshold, a preset novel example is added to the prompt template of the large model to obtain second preset prompt information, which is then used to guide the large model to regenerate candidate business data. The similarity between the preset novel example and the candidate business data can be determined based on their semantic similarity.

[0042] A3. If the semantic stability index does not meet the third preset condition, add standard examples and / or reverse examples to the preset prompt information of the large model, and regenerate candidate business data; wherein, the standard examples are examples that conform to the preset specifications; and the reverse examples are examples that do not conform to the preset specifications.

[0043] In this embodiment, the third preset condition can be determined based on a third score threshold preset for the semantic stability index; the preset specification can be pre-defined based on the logic of the business process.

[0044] For example, if the score of the third indicator corresponding to the semantic stability index is lower than the third score threshold, standard examples and / or reverse examples are added to the prompt template of the large model, and corresponding identification information is added to the standard examples and reverse examples respectively to obtain the third preset prompt information. Then, the large model is guided to regenerate candidate business data using the third preset prompt information. Furthermore, the standard examples can also be determined based on candidate business data that meet preset conditions.

[0045] By following the steps above, the preset prompts of the large model can be optimized and adjusted in a targeted manner according to the characteristics of the candidate business data, so as to ensure that the large model can quickly and efficiently generate high-quality candidate business data in the next iteration.

[0046] S204. Determine the target business data based on the correction results, and construct a business process knowledge base based on the target business data.

[0047] This invention, through its embodiments, guides a large model to perform semantic parsing on initial business data using the Context Learning (ICL) paradigm based on preset prompts, and generates candidate business data. This can efficiently generate high-quality candidate business data without adjusting the parameters of the large model. Furthermore, the generation strategy in the preset prompts provides clear rule constraints for the generated candidate business data, improving the semantic accuracy, semantic novelty, and semantic stability of the generated candidate business data, thus laying the foundation for the subsequent construction of a business process knowledge base.

[0048] Example 3 Figure 3 This is a flowchart of a knowledge base construction method provided in Embodiment 2 of the present invention. This embodiment is a refinement based on the above embodiments. Figure 3 As shown, the method includes: S301. Preprocess the original business materials to obtain initial business data.

[0049] S302. Use the large model to perform semantic parsing on the initial business data and generate candidate business data.

[0050] S303. If the preset indicators of the candidate business data do not meet the preset conditions, the candidate business data shall be corrected multiple times according to the preset correction strategy until the preset conditions are met; wherein, the preset indicators include at least one of the semantic accuracy indicator, semantic novelty indicator, and semantic stability indicator of the candidate business data.

[0051] In this embodiment, the process of correcting candidate business data can also be assisted by a third-party system (which can be an intelligent model with domain knowledge). For example, if the index score corresponding to the preset index of some candidate business data is lower than the preset score threshold, or if it still does not meet the preset conditions after multiple iterations of the large model, it means that these candidate business data are low-quality business data. The preset index information corresponding to these candidate business data, as well as the historical generation trajectory and conflict points, can be submitted to the third-party system for review. The third-party system corrects these candidate business data into new candidate business data that meet the preset conditions and uses it to determine the target business data structured data in the future. The third-party system corrects the data in the following ways: confirming whether the judgment results of each preset index are consistent with the actual business understanding, providing supplementary explanations on at least one of the rule boundaries, priorities, and implicit conditions, and marking the types of business rules that need to be added or revised.

[0052] Optionally, the process of correcting candidate business data can also be carried out through human-computer interaction. For example, low-quality business data can be handed over to domain experts for correction until it meets the preset conditions, and the target business data can be determined based on the final correction result.

[0053] S304. Determine the target business data based on the correction results, and split the target business data to obtain business entities and business rules.

[0054] For example, each piece of target business data is broken down into business entities and business rules; whereby business rules include triggering conditions, decision conclusions or processing actions, constraint relationships, and priority information. Words or phrases in the target business data containing semantics such as objects, roles, resources, and states can be broken down into business entities; words or phrases in the target business data containing semantics such as combinations of business conditions and threshold constraints can be broken down into triggering conditions.

[0055] S305. Map business entities to entities in triples, map the relationship between business entities and business rules to relationships in triples, map business rules to rules in triples, and add the triples to the business rule knowledge graph.

[0056] For example, the splitting results corresponding to the target business data can be converted into the structure of nodes and relationships in a knowledge graph. Specifically, mapping rules can be pre-defined, and business entities can be mapped to triple entities according to the mapping rules, the association between business entities and business rules can be mapped to triple relationships, and business rules can be mapped to triple rules, thereby obtaining triples, and the triples can be added to the business rule knowledge graph.

[0057] Optionally, before adding the triple to the business rule knowledge graph, the method further includes: verifying the triple corresponding to the target business data using the graph structure of the business rule knowledge graph based on preset verification conditions; if the verification fails, revising the target business data multiple times according to the preset correction strategy until the target triple corresponding to the revised target business data satisfies the preset verification conditions; wherein, adding the triple to the business rule knowledge graph includes: adding the target triple to the business rule knowledge graph.

[0058] In this embodiment, the preset verification conditions may include at least one of the following: whether it produces a conflicting conclusion with an existing rule under the same conditions, whether it forms a circular dependency or a logical closed loop, and whether there are isolated rules or nodes without triggering paths.

[0059] For example, based on preset verification conditions, the graph structure of the business rule knowledge graph is used to verify the triples corresponding to the target business data. Taking the preset verification condition of whether it produces a conflicting conclusion with an existing rule under the same conditions as an example, the rule link with the same conditions as the triples corresponding to the current target business data is retrieved in the business rule knowledge graph. The conclusion under the condition is obtained according to the rule link, and the conclusion is semantically compared with the conclusion in the triples corresponding to the current target business data. If there is a semantic conflict, the triples corresponding to the current target business data do not meet the preset verification conditions.

[0060] If the verification fails, the target business data can be corrected multiple times according to the preset correction strategy until the target triple corresponding to the corrected target business data meets the preset verification conditions. For example, the prompt template corresponding to the large model can be adjusted according to the preset correction strategy to guide the large model to regenerate candidate business data that meets the preset conditions. Based on the candidate business data, the corresponding target business data (i.e., the corrected target business data) can be re-determined. The target business data can be converted into triples through the above steps, and the triples can be verified again until the verification passes, indicating that the triple (i.e., the target triple) meets the preset verification conditions. Then, the target triple is added to the business rule knowledge graph. The above steps utilize the graph structure of the business rule knowledge graph to efficiently verify the triples corresponding to each target business data, ensuring that the generated business rule knowledge graph has higher reliability.

[0061] S306. Extract the triggering relationships and conditional dependencies between different rules from the business rule knowledge graph, and perform structured processing on the triggering relationships and conditional dependencies to generate a business process knowledge base.

[0062] For example, triggering relationships and conditional dependencies between different rules can be extracted from a business rule knowledge graph. Specifically, for a given entity node in the business rule knowledge graph, multiple rule nodes that are related to it can be retrieved using the graph structure. The triggering relationships and conditional dependencies are then structured to generate a business process knowledge base. Specifically, rule links are constructed based on the association edges between entity nodes and rule nodes, thereby generating the business process knowledge base based on these rule links. These rule links can include standard business processing flows, exception and fallback handling flows, multi-condition branches and decision paths, and multi-role collaborative execution sequences, etc.

[0063] This invention constructs a business rule knowledge graph by splitting target business data and converting the splitting results into triples. This transforms fragmented and poorly correlated business data into a highly correlated and reasonable graph structure. Based on the characteristics of the graph structure, it accurately mines the triggering and dependency relationships between rules and transforms them into a business process knowledge base, providing knowledge support for the subsequent deployment of intelligent systems.

[0064] In some embodiments, the method further includes: receiving a business request; filtering target knowledge related to the business request in the business process knowledge base; using an agent to reason about the business request based on the target knowledge to obtain a reasoning result corresponding to the business request; obtaining target association data corresponding to the reasoning result; wherein the target association data is used to characterize the correctness of the reasoning result and the reasoning process; filtering abnormal rules in the business process knowledge base according to the target association data; adjusting the preset prompt information of the large model based on the abnormal rules to update the business process knowledge base. The advantage of this setup is that it achieves a closed-loop update of the business process knowledge base based on business requests; by filtering abnormal rules according to target association data, it achieves accurate positioning of abnormal rules based on the reasoning result and reasoning process; by adjusting the preset prompt information of the large model based on the abnormal rules, it achieves continuous optimization of the large model and ensures that the constructed business process knowledge base can be dynamically updated.

[0065] For example, an intelligent question-answering system can be deployed based on a business process knowledge base. For instance, if an employee asks the system, "What processes are required for a purchase order exceeding 10,000 yuan?", the system can locate the question in the business rule knowledge graph and business process knowledge base based on the question's task type and input conditions. This allows it to match the corresponding business entity and rule subset, available process paths, and strong constraint rules. Utilizing a Retrieval-Augmented Generation (RAG) mechanism, the matched rules, processes, and example samples are used as context input to the integrated agent. The example samples serve as context examples, guiding the agent to follow business logic during reasoning, thereby obtaining the reasoning result corresponding to the question. Optionally, the intelligent question-answering system can integrate multiple intelligent agents, each undertaking different tasks. For example, it can integrate a first intelligent agent for verifying rule consistency, a second intelligent agent for reasoning about process paths, and a third intelligent agent for assessing risks and anomalies. Each intelligent agent can share the business rule knowledge graph and the business process knowledge base, perform independent reasoning, and cross-validate their respective reasoning results. If there are conflicts in the validation results, the reasoning can be re-performed, thereby ensuring the accuracy and stability of the reasoning results.

[0066] The target-related data corresponding to the reasoning results can include: user feedback and correction behavior on the reasoning results, quality indicators of the reasoning results (such as consistency scores and confidence distribution), execution trajectory of actual business processes, and newly emerging real business examples and abnormal cases. This target-related data can be used to characterize the correctness of the reasoning results and reflect the system's reasoning process. The intelligent question answering system can filter abnormal rules in the business process knowledge base based on the acquired target-related data. Specifically, the intelligent question answering system can periodically calculate preset evaluation indicators based on the target-related data, such as the triggering frequency and accuracy of each business rule, the consistency between the process path and the actual execution path, and the degree of deviation between the reasoning results and the business rules. Based on the above indicator values, abnormal rules that have not been triggered for a long time, are frequently rejected, and cause reasoning errors can be filtered out. In order to ensure that the reasoning results of the intelligent question answering system have higher reliability, the business process knowledge base can be updated in real time. For example, the preset prompt information of the large model can be adjusted based on the abnormal rules, and the score thresholds corresponding to each preset indicator of the candidate business data can be changed.

[0067] Example 4 Figure 4 This is a schematic diagram of a knowledge base construction device provided in Embodiment 4 of the present invention. Figure 4As shown, the device includes: a preprocessing module 401, a candidate business data generation module 402, a correction module 403, and a business process knowledge base construction module 404.

[0068] The preprocessing module is used to preprocess the original business materials to obtain initial business data; The candidate business data generation module is used to perform semantic parsing on the initial business data using a large model and generate candidate business data. The correction module is used to correct the candidate business data multiple times according to a preset correction strategy when the preset indicators of the candidate business data do not meet the preset conditions, until the preset conditions are met; wherein, the preset indicators include at least one of the semantic accuracy indicator, semantic novelty indicator, and semantic stability indicator of the candidate business data. The business process knowledge base construction module is used to determine the target business data based on the correction results, and to construct the business process knowledge base based on the target business data.

[0069] This invention provides a knowledge base construction apparatus. It preprocesses original business materials to obtain initial business data, laying the foundation for subsequent extraction of business knowledge using a large model. By using the large model to perform semantic parsing on the initial business data and generating candidate business data, it enables rapid mining of potential semantic information within the initial business data. If the preset indicators of the candidate business data do not meet preset conditions, the candidate business data is corrected multiple times according to a preset correction strategy until the preset conditions are met. The preset indicators include at least one of semantic accuracy, semantic novelty, and semantic stability indicators of the candidate business data. These steps ensure that the candidate business data is semantically accurate, novel, and stable, resulting in higher data quality for the subsequently obtained target business data. By determining the target business data based on the correction results and constructing a business process knowledge base based on the target business data, it is possible to ensure that the intelligent system deployed based on this business process knowledge base has more comprehensive and accurate knowledge support.

[0070] Optionally, the candidate business data generation module is specifically used for: Based on preset prompts, the ICL (Internal Context Learning) paradigm guides the large model to perform semantic parsing on the initial business data and generate candidate business data. The preset prompts include the generation strategy corresponding to the candidate business data. The generation strategy is used to guide the large model to extract different expressions of standard business processes, abnormal business processes, business rule boundary conditions, and business terms from the initial business data to generate candidate business data of various business types.

[0071] Optionally, the preset conditions include a first preset condition corresponding to the semantic accuracy index, a second preset condition corresponding to the semantic novelty index, and a third preset condition corresponding to the semantic stability index; the correction module includes at least one of the following correction units: The first correction unit is used to increase the proportion of rule-constrained prompt words in the preset prompt information of the large model and regenerate candidate business data when the semantic accuracy index does not meet the first preset condition. The second correction unit is used to add a preset novel example to the preset prompt information of the large model and regenerate candidate business data when the semantic novelty index does not meet the second preset condition, until the second preset condition is met; wherein the similarity between the preset novel example and the candidate business data is lower than a preset similarity threshold. The third correction unit is used to add standard examples and / or reverse examples to the preset prompt information of the large model and regenerate candidate business data when the semantic stability index does not meet the third preset condition, until the third preset condition is met; wherein, the standard examples are examples that conform to the preset specifications; and the reverse examples are examples that do not conform to the preset specifications.

[0072] Optionally, the business process knowledge base construction module includes: The target business data determination unit is used to determine the target business data based on the correction results. The target business data splitting unit is used to split the target business data to obtain business entities and business rules; The triple mapping unit is used to map the business entity to a triple entity, map the association between the business entity and the business rule to a triple relationship, map the business rule to a triple rule, and add the triple to the business rule knowledge graph. The structured processing unit is used to extract the triggering relationships and conditional dependencies between different rules from the business rule knowledge graph, and to perform structured processing on the triggering relationships and conditional dependencies to generate a business process knowledge base. Optionally, the business process knowledge base construction module further includes: The triplet verification unit is used to verify the triplet corresponding to the target business data based on preset verification conditions and the graph structure of the business rule knowledge graph before adding the triplet to the business rule knowledge graph. If the verification fails, the target business data is modified multiple times according to the preset correction strategy until the target triplet corresponding to the modified target business data meets the preset verification conditions. The step of adding the triplet to the business rule knowledge graph includes adding the target triplet to the business rule knowledge graph.

[0073] Optionally, the device further includes: The reasoning module is used to receive a business request, filter target knowledge related to the business request from the business process knowledge base, use an intelligent agent to reason about the business request based on the target knowledge, and obtain the reasoning result corresponding to the business request. The target association data acquisition module is used to acquire the target association data corresponding to the reasoning result; wherein, the target association data is used to characterize the correctness of the reasoning result and the reasoning process; The business process knowledge base update module is used to filter the abnormal rules in the business process knowledge base according to the target associated data, and adjust the preset prompt information of the large model based on the abnormal rules in order to update the business process knowledge base.

[0074] The knowledge base construction apparatus provided in this embodiment of the invention can execute the knowledge base construction method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0075] Example 5 Figure 5 A schematic diagram of an electronic device 500 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0076] like Figure 5As shown, the electronic device 500 includes at least one processor 501 and a memory, such as a read-only memory (ROM) 502 and a random access memory (RAM) 503, communicatively connected to the at least one processor 501. The memory stores computer programs executable by the at least one processor. The processor 501 can perform various appropriate actions and processes based on the computer program stored in the ROM 502 or loaded into the RAM 503 from storage unit 508. The RAM 503 can also store various programs and data required for the operation of the electronic device 500. The processor 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0077] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows electronic device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0078] Processor 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 501 performs the various methods and processes described above, such as knowledge base construction methods.

[0079] In some embodiments, the knowledge base construction method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by processor 501, one or more steps of the knowledge base construction method described above may be performed. Alternatively, in other embodiments, processor 501 may be configured to execute the knowledge base construction method by any other suitable means (e.g., by means of firmware).

[0080] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0081] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0082] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0083] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0084] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0085] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0086] This disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the knowledge base construction method provided in the above embodiments.

[0087] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0088] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for constructing a knowledge base, characterized in that, include: Preprocess the original business materials to obtain initial business data; The initial business data is semantically parsed using a large model, and candidate business data is generated. If the preset indicators of the candidate business data do not meet the preset conditions, the candidate business data shall be corrected multiple times according to the preset correction strategy until the preset conditions are met; wherein, the preset indicators include at least one of the semantic accuracy indicator, semantic novelty indicator, and semantic stability indicator of the candidate business data. The target business data is determined based on the correction results, and a business process knowledge base is constructed based on the target business data.

2. The knowledge base construction method according to claim 1, characterized in that, The step of using a large model to perform semantic parsing on the initial business data and generating candidate business data includes: Based on preset prompts, the ICL (Internal Context Learning) paradigm guides the large model to perform semantic parsing on the initial business data and generate candidate business data. The preset prompts include the generation strategy corresponding to the candidate business data. The generation strategy is used to guide the large model to extract different expressions of standard business processes, abnormal business processes, business rule boundary conditions, and business terms from the initial business data to generate candidate business data of various business types.

3. The knowledge base construction method according to claim 1, characterized in that, The preset conditions include a first preset condition corresponding to the semantic accuracy index, a second preset condition corresponding to the semantic novelty index, and a third preset condition corresponding to the semantic stability index. Wherein, when the preset indicators of the candidate service data do not meet the preset conditions, the candidate service data is corrected multiple times according to a preset correction strategy, including at least one of the following: If the semantic accuracy index does not meet the first preset condition, increase the proportion of rule-constrained prompt words in the preset prompt information of the large model, and regenerate candidate business data. If the semantic novelty index does not meet the second preset condition, a preset novel example is added to the preset prompt information of the large model, and candidate business data is regenerated; wherein the similarity between the preset novel example and the candidate business data is lower than a preset similarity threshold. If the semantic stability index does not meet the third preset condition, standard examples and / or reverse examples are added to the preset prompt information of the large model, and candidate business data is regenerated; wherein, the standard examples are examples that conform to the preset specifications; and the reverse examples are examples that do not conform to the preset specifications.

4. The knowledge base construction method according to claim 1, characterized in that, The step of constructing a business process knowledge base based on the target business data includes: The target business data is split to obtain business entities and business rules; The business entity is mapped to a triple entity, the association between the business entity and the business rule is mapped to a triple relationship, the business rule is mapped to a triple rule, and the triple is added to the business rule knowledge graph. The triggering relationships and conditional dependencies between different rules are extracted from the business rule knowledge graph, and the triggering relationships and conditional dependencies are structured to generate a business process knowledge base.

5. The knowledge base construction method according to claim 4, characterized in that, Before adding the triples to the business rule knowledge graph, the following is also included: Based on preset verification conditions, the triples corresponding to the target business data are verified using the graph structure of the business rule knowledge graph. If the verification fails, the target business data is corrected multiple times according to the preset correction strategy until the target triples corresponding to the corrected target business data meet the preset verification conditions. The step of adding the triples to the business rule knowledge graph includes: The target triple is added to the business rule knowledge graph.

6. The knowledge base construction method according to any one of claims 1-5, characterized in that, The method further includes: Receive a business request, filter target knowledge related to the business request from the business process knowledge base, use an intelligent agent to reason about the business request based on the target knowledge, and obtain the reasoning result corresponding to the business request; Obtain the target association data corresponding to the reasoning result; wherein, the target association data is used to characterize the correctness of the reasoning result and the reasoning process; Based on the target-related data, the abnormal rules in the business process knowledge base are filtered, and the preset prompt information of the large model is adjusted based on the abnormal rules in order to update the business process knowledge base.

7. A knowledge base construction apparatus, characterized in that, include: The preprocessing module is used to preprocess the original business materials to obtain initial business data; The candidate business data generation module is used to perform semantic parsing on the initial business data using a large model and generate candidate business data. The correction module is used to correct the candidate business data multiple times according to a preset correction strategy when the preset indicators of the candidate business data do not meet the preset conditions, until the preset conditions are met; wherein, the preset indicators include at least one of the semantic accuracy indicator, semantic novelty indicator, and semantic stability indicator of the candidate business data. The business process knowledge base construction module is used to determine the target business data based on the correction results, and to construct the business process knowledge base based on the target business data.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the knowledge base construction method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the knowledge base construction method according to any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the knowledge base construction method according to any one of claims 1-6.