Knowledge slice data processing method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202610785300.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]然而,现有方案侧重于静态全量知识图谱的构建,在为具体岗位生成系统提示词时,将整张本体或大规模子图直接序列化注入,容易导致大模型的上下文窗口被无关概念稀释,从而导致占用模型注意力以造成推理成本较大的情况
[0018]本申请实施例提出的知识切片数据处理方法、装置、电子设备及存储介质,方法包括:首先,获取针对目标智能代理的种子上下文对象,种子上下文对象包括待抽取的业务载荷以及关联的归属元数据;然后,基于业务载荷进行候选知识抽取,得到多个候选知识;之后,基于预设的租户层知识库和/或共享层知识库,根据归属元数据将多个候选知识进行跨层级匹配,得到每个候选知识对应的匹配状态结果,并基于匹配状态结果生成目标知识切片;最后,当目标智能代理运行时,基于目标知识切片中的激活项目进行局部序列化处理,得到业务知识提示词,并将业务知识提示词添加至目标智能代理的系统提示词中。本申请实施例通过引入归属元数据与业务载荷协同驱动的抽取及跨层级匹配机制,实现了从全量知识库中按需构建针对特定岗位实例的目标知识切片,通过对切片内激活项目进行局部序列化处理,改变了以往将静态全量本体或大规模子图直接注入大模型的方式,有效缓解了冗余信息对上下文窗口的稀释作用,减小了无关概念对模型注意力的占用,从而在降低词元消耗及推理成本的同时,增强了业务知识与运行时语境的适配性,并利用租户层与共享层的逻辑隔离机制,以提升多租户场景下知识边界治理的安全性与注入内容的可追溯性。
Smart Images

Figure CN122840184A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to methods, apparatuses, electronic devices and storage media for processing knowledge slice data. Background Technology
[0002] In recent years, intelligent agents (AI agents) built on large language models have been widely deployed in business roles within enterprises. To enable general-purpose large models to understand specific industries and business scenarios, related technologies generally adopt an approach that combines ontology or knowledge graph construction with retrieval generation. At runtime, industry knowledge, internal enterprise knowledge, and business rules specific to each role are injected into the large model through system prompts in a templated manner, thereby achieving business context-assisted reasoning.
[0003] However, existing solutions focus on the construction of static full knowledge graphs. When generating system prompts for specific positions, the entire ontology or a large-scale subgraph is directly serialized and injected. This can easily lead to the dilution of the context window of a large model by irrelevant concepts, resulting in a high inference cost due to the occupation of model attention. Summary of the Invention
[0004] This application provides a knowledge slice data processing method, apparatus, electronic device, and storage medium that can solve the above-mentioned technical problems.
[0005] To achieve the above objectives, a first aspect of this application proposes a knowledge slice data processing method, the method comprising: Obtain a seed context object for the target intelligent agent, the seed context object including the business payload to be extracted and the associated attribution metadata; Based on the aforementioned business payload, candidate knowledge is extracted to obtain multiple candidate knowledge items; Based on a preset tenant layer knowledge base and / or shared layer knowledge base, multiple candidate knowledges are matched across layers according to the attribution metadata to obtain a matching status result corresponding to each candidate knowledge, and a target knowledge slice is generated based on the matching status result. When the target intelligent agent runs, it performs local serialization processing based on the activated items in the target knowledge slice to obtain business knowledge prompts, and adds the business knowledge prompts to the system prompts of the target intelligent agent.
[0006] In some embodiments, obtaining the seed context object for the target intelligent agent includes: Obtain structured dialogue data generated by the construction end for the job template, and / or obtain unstructured business documents input by the user end for the target intelligent agent; Based on the attribution metadata, the structured dialogue data and / or the unstructured business documents are uniformly encapsulated to obtain the seed context object.
[0007] In some embodiments, the extraction of candidate knowledge based on the service payload to obtain multiple candidate knowledge items includes: Based on the business payload, candidate knowledge is extracted to obtain multiple initial candidate knowledge; Based on the normalized form and candidate type of the initial candidate knowledge, the multiple initial candidate knowledge are subjected to precise deduplication processing to obtain deduplicated candidate knowledge. The candidate names, candidate types, and attribute key names of the deduplicated candidate knowledge are concatenated to obtain text representation data, and the text representation data is input into the pre-trained embedding model to obtain the corresponding candidate knowledge representation vector. Calculate the pairwise similarity between two candidate knowledge representation vectors, and perform clustering processing on the deduplicated candidate knowledge based on the numerical comparison relationship between the pairwise similarity and the preset deduplication threshold to obtain multiple candidate knowledge clusters; For each candidate knowledge cluster, the attribute set and evidence fragments of the candidate knowledge within the cluster are merged to obtain a merged candidate set, which includes multiple candidate knowledge clusters.
[0008] In some embodiments, the step of performing cross-level matching of multiple candidate knowledge bases based on a preset tenant layer knowledge base and / or shared layer knowledge base according to the attribution metadata to obtain a matching status result corresponding to each candidate knowledge base, and generating a target knowledge slice based on the matching status result, includes: Based on the tenant ID in the attribution metadata, the target tenant layer knowledge base is determined from multiple tenant layer knowledge bases; Based on preset concept type constraints, the text representation vector corresponding to each candidate knowledge is retrieved and matched in the target tenant layer knowledge base and / or shared layer knowledge base to determine the target node; Determine the overall similarity between each candidate knowledge and the target node; Based on the numerical comparison relationship between the comprehensive similarity and the preset matching threshold, the matching status result corresponding to each candidate knowledge is obtained, and a target knowledge slice is generated based on the matching status result.
[0009] In some embodiments, the step of retrieving and matching the text representation vector corresponding to each candidate knowledge in the target tenant layer knowledge base and / or shared layer knowledge base based on preset concept type constraints to determine the target node includes: For each candidate knowledge, based on the concept type constraint corresponding to the candidate knowledge, filtering is performed in the target tenant layer knowledge base and / or the shared layer knowledge base to obtain at least one candidate knowledge node whose type matches the concept type constraint; Determine the vector similarity between the text representation vector of the candidate knowledge and the representation vectors corresponding to all candidate knowledge nodes; Based on the vector similarity, a nearest neighbor search is performed on all the candidate knowledge nodes to obtain the target node.
[0010] In some embodiments, determining the comprehensive similarity between each candidate knowledge and the target node includes: For each candidate knowledge, the name dimension similarity is determined based on the vector matching degree between the candidate knowledge and the target node in name semantics and the character matching degree in name string. Based on the relative logical distance between the candidate knowledge and the target node in the preset concept inheritance topology, the type dimension similarity is determined; The similarity of the attribute dimension is determined based on the intersection-union ratio of the candidate knowledge and the set of attribute key names contained in the target node, and the consistency weighting of the attribute values corresponding to the same attribute key names. The comprehensive similarity is obtained by weighting the similarity based on the name dimension, the type dimension, and the attribute dimension.
[0011] In some embodiments, the preset matching threshold includes a full matching threshold and a partial matching threshold. The step of obtaining the matching state result corresponding to each candidate knowledge based on the numerical comparison relationship between the comprehensive similarity and the preset matching threshold, and generating a target knowledge slice based on the matching state result, includes: When the overall similarity is greater than or equal to the full matching threshold, a matching status result representing the matched items is generated, and the corresponding candidate knowledge is confirmed as the activated item. When the overall similarity is less than the full matching threshold and greater than or equal to the partial matching threshold, a matching state result representing the partial matching is generated, and a difference description object between the candidate knowledge and the target node is generated. When the overall similarity is less than the partial matching threshold, a matching state result representing a new candidate is generated, and the corresponding candidate knowledge is marked as a ghost state item. The target knowledge slice is obtained based on at least one of the activated item, the difference description object, and the ghost state item.
[0012] In some embodiments, after obtaining the target knowledge slice, the method further includes: When the attribution type identifier in the attribution metadata represents the construction end, the target knowledge slice is presented in a graph view format that includes knowledge nodes and related relationship edges; When the attribution type identifier represents the user end, the target knowledge slice is presented in a business confirmation list view format that hides the graph structure; Obtain user confirmation instructions, and perform status updates on the difference description object and the ghost state item based on the user confirmation instructions.
[0013] In some embodiments, the step of performing local serialization processing based on activated items in the target knowledge slice to obtain business knowledge prompts includes: Based on the project type of the activated project in the target knowledge slice, obtain the corresponding target concept information; Based on a preset prompt word assembly template, the target concept information is converted into a text format to obtain the business knowledge prompt words that include business concept paragraphs and business relationship paragraphs.
[0014] In some embodiments, the method further includes: Add the ghost-state project to the pending review queue corresponding to the target tenant layer knowledge base; The pre-defined change policy checker is used to perform gating verification on the ghost state items in the queue to be reviewed, and the verification result is obtained. When the verification result is passed, the candidate knowledge corresponding to the ghost state project is updated to the officially activated node in the target tenant layer knowledge base; Retrieve historical knowledge slices that reference the candidate knowledge in the target tenant layer knowledge base, and synchronously update the ghost state items contained in all the historical knowledge slices to the active items.
[0015] To achieve the above objectives, a second aspect of this application provides a knowledge slice data processing apparatus, the apparatus comprising: The acquisition module is used to acquire a seed context object for the target intelligent agent. The seed context object includes the business payload to be extracted and the associated attribution metadata. The candidate knowledge determination module is used to extract candidate knowledge based on the business payload to obtain multiple candidate knowledge. The knowledge slice determination module is used to perform cross-level matching of multiple candidate knowledge based on the preset tenant layer knowledge base and / or shared layer knowledge base, according to the attribution metadata, to obtain the matching status result corresponding to each candidate knowledge, and to generate a target knowledge slice based on the matching status result. The prompt word update module is used to perform local serialization processing on the activated items in the target knowledge slice when the target intelligent agent is running, to obtain business knowledge prompt words, and to add the business knowledge prompt words to the system prompt words of the target intelligent agent.
[0016] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the knowledge slicing data processing method as described in the first aspect.
[0017] To achieve the above objectives, a fourth aspect of the present application provides a storage medium, which is a computer-readable storage medium storing a computer program that, when executed by a processor, implements the knowledge slice data processing method as described in the first aspect.
[0018] The knowledge slice data processing method, apparatus, electronic device, and storage medium proposed in this application include: First, obtaining a seed context object for a target intelligent agent, the seed context object including the business payload to be extracted and the associated attribution metadata; then, extracting candidate knowledge based on the business payload to obtain multiple candidate knowledge; next, based on a preset tenant layer knowledge base and / or shared layer knowledge base, performing cross-level matching of the multiple candidate knowledge according to the attribution metadata to obtain the matching status result corresponding to each candidate knowledge, and generating a target knowledge slice based on the matching status result; finally, when the target intelligent agent runs, performing local serialization processing based on the activated items in the target knowledge slice to obtain business knowledge prompts, and adding the business knowledge prompts to the system prompts of the target intelligent agent. This application's embodiments introduce an extraction and cross-level matching mechanism driven by attribution metadata and business payload, enabling the on-demand construction of target knowledge slices for specific job instances from the full knowledge base. By performing local serialization processing on activated items within the slices, it changes the previous method of directly injecting static full ontology or large-scale subgraphs into large models, effectively alleviating the dilution effect of redundant information on the context window and reducing the occupation of irrelevant concepts on model attention. This reduces lexical consumption and inference costs while enhancing the adaptability of business knowledge to runtime context. Furthermore, it utilizes the logical isolation mechanism between the tenant layer and the shared layer to improve the security of knowledge boundary governance and the traceability of injected content in multi-tenant scenarios.
[0019] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description
[0020] Figure 1 This is a flowchart of a knowledge slice data processing method provided in an embodiment of this application.
[0021] Figure 2 yes Figure 1 The flowchart for step 101.
[0022] Figure 3 This is a schematic diagram of a three-layer knowledge data model and attribution relationship provided in another embodiment of this application.
[0023] Figure 4 yes Figure 1 The flowchart for step 102.
[0024] Figure 5 yes Figure 1 The flowchart for step 103.
[0025] Figure 6 yes Figure 5 The flowchart for step 502.
[0026] Figure 7 yes Figure 5 The flowchart for step 503.
[0027] Figure 8 yes Figure 5 The flowchart for step 504.
[0028] Figure 9 This is another embodiment of the present application, providing a flowchart for the visualization generation of knowledge slice data.
[0029] Figure 10 yes Figure 1 The flowchart for step 104.
[0030] Figure 11 This is a schematic diagram of a process for slicing and extracting the main waterline, provided in another embodiment of this application.
[0031] Figure 12 This is another embodiment of the present application, providing a flowchart for updating ghost-state project data.
[0032] Figure 13 This is a flowchart illustrating a novel concept of backflow closed loop, dual gating, and automatic cross-instance hit, provided in another embodiment of this application.
[0033] Figure 14 This is a schematic diagram of the structure of a knowledge slice data processing device provided in another embodiment of this application.
[0034] Figure 15This is a schematic diagram of the hardware structure of an electronic device provided in another embodiment of this application. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0036] It should be noted that although functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart.
[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0038] In recent years, intelligent agents (AI agents) built on large language models have been widely deployed in enterprise business roles. To enable general-purpose large models to understand specific industries and business scenarios, related technologies generally adopt an approach that combines ontology or knowledge graph construction with retrieval-enhanced generation. At runtime, industry knowledge, internal enterprise knowledge, and business rules specific to each role are injected into the large model through the templated assembly of system prompts, thereby achieving business context-assisted reasoning. However, existing solutions focus on the construction of static full knowledge graphs. When generating system prompts for specific roles, the entire ontology or a large sub-graph is often directly serialized and injected. This can easily lead to the dilution of the large model's context window by irrelevant concepts, thus occupying the model's attention and resulting in high reasoning costs. Furthermore, this full injection method may also lead to the injection of sensitive concepts that the current role should not be aware of (such as financial data), posing a risk of overstepping authority across role contexts. Moreover, the injected content is difficult to trace throughout the conversation, making it impossible to accurately locate the trigger source when hallucinations occur.
[0039] Furthermore, in multi-tenant enterprise application scenarios, existing solutions typically construct only a single-layer domain ontology. When the same platform needs to provide intelligent agent services to multiple enterprise tenants, platform-level general knowledge, enterprise-private internal knowledge, and specific business rules for particular job instances are often mixed on the same ontology. Once enterprise-private concepts are directly merged into the domain ontology, tenant boundaries are easily lost, leading to potential cross-tenant information leakage risks and redundant maintenance issues. Moreover, in handling heterogeneous input, existing technologies often unify structured and unstructured inputs into triples for graph construction, making it difficult to perceive user differences at the input end, resulting in the inability to differentiate the return path of new concepts.
[0040] On the other hand, during the knowledge update and extraction phases of intelligent agent operation, existing technologies mostly adopt a single-layer, automatic update method that generates candidates from large models and directly merges them into the ontology, lacking a backflow mechanism subject to dual gating by change policies and manual review. Simultaneously, the extraction process typically only retains or discards cases with similar names or partially overlapping attributes—"approximate but not perfect matches"—without an independent intermediate channel for manual confirmation. This single-step merging method, where each update overwrites the previous one, also results in the extraction products lacking version snapshot and baseline differentiation capabilities. This makes it difficult for the platform to accurately merge with enterprise-customized private content after upgrading to an industry-standard ontology, and the business context accumulated during instance operation cannot be retained as a portable and inheritable asset.
[0041] To address the aforementioned technical challenges, this application's embodiments introduce an extraction and cross-level matching mechanism driven by attribution metadata and business payload. This enables the on-demand construction of target knowledge slices for specific job instances from the full knowledge base. By performing local serialization processing on activated items within the slices, it changes the previous method of directly injecting static full ontology or large-scale subgraphs into large models. This effectively alleviates the dilution effect of redundant information on the context window and reduces the attentional burden of irrelevant concepts on the model. Consequently, while reducing lexical consumption and inference costs, it enhances the adaptability of business knowledge to runtime context. Furthermore, by utilizing the logical isolation mechanism between the tenant layer and the shared layer, it improves the security of knowledge boundary governance and the traceability of injected content in multi-tenant scenarios.
[0042] The knowledge slice data processing method, apparatus, electronic device, and storage medium provided in the embodiments of this application will be further described below.
[0043] Reference Figure 1 This is an optional flowchart of the knowledge slice data processing method provided in the embodiments of this application. Figure 1 The method described may include, but is not limited to, steps 101 to 104. It is also understood that this embodiment... Figure 1The order of steps 101 to 104 is not specifically limited; the order of steps can be adjusted or certain steps can be added or removed according to actual needs. The knowledge slice data processing method provided in this application can be applied to any processing system with computing resources (such as servers, smart terminals, computing processors, etc.).
[0044] Step 101: Obtain the seed context object for the target intelligent agent. The seed context object includes the business payload to be extracted and the associated attribution metadata.
[0045] Step 101 will be described in detail below.
[0046] In step 101 of some embodiments, in actual business scenarios, the sources of knowledge differ. For example, there may be structured dialogue data input by technical personnel on the construction side, or unstructured business documents such as standard operating procedures uploaded by business personnel on the user side. The solution of this application abstracts and encapsulates these two different types of input into a seed context object. The seed context object contains the valid business payload to be analyzed and the attribution metadata used to identify the seed type, attribution type and tenant number. This shields the source differences at the front end of the pipeline, so that the subsequent extraction process does not need to be aware of the input differences, as described below.
[0047] Reference Figure 2 Obtaining the seed context object for the target smart agent includes the following steps 201 to 202.
[0048] Step 201: Obtain the structured dialogue data generated by the builder for the job template, and / or obtain the unstructured business documents input by the user for the target intelligent agent.
[0049] Step 202: Based on the attribution metadata, the structured dialogue data and / or unstructured business documents are uniformly encapsulated to obtain the seed context object.
[0050] Steps 201 to 202 are described in detail below.
[0051] In step 201 of some embodiments, the data input format of the intelligent agent is classified and acquired according to different application stages. At the construction end used to design job templates, the system usually guides the answer through structured dialogue tools (prompting the user to answer questions such as "What problem does this job solve? / What is the benchmark for a real job? / What business concepts need to be understood? / What capabilities are required? / What systems need to be integrated?"), thereby acquiring the structured dialogue data input by technical personnel. At the user end used to actually hire a smart agent, the system receives unstructured business documents uploaded by business personnel (including but not limited to standard operating procedures, rules and regulations, business descriptions, organizational structures, application programming interface documents, database structure definition files, etc.), such as standard operating procedures, rules and regulations, or business descriptions.
[0052] In step 202 of some embodiments, in order to enable the subsequent extraction process to use the same pipeline for isomorphic processing, the system uniformly encapsulates structured dialogue data and / or unstructured business documents into seed context objects. In this encapsulation operation, the system uses attribution metadata (such as seed type, payload, attribution type, attribution number, tenant number, etc.) to structurally associate them, thereby shielding the differences in data presentation format at the input end to obtain the seed context object.
[0053] Through steps 201 to 202 above, by introducing a dual-source collection and encapsulation mechanism at the front end of the extraction pipeline, the structured dialogue at the construction end and the unstructured document at the user end are uniformly abstracted into a standard seed context object containing attribution metadata. This eliminates the format differences of different business roles on the input carrier, enabling the knowledge extraction engine to achieve reuse of the underlying backbone process. At the same time, by retaining and transmitting attribution metadata, accurate routing basis can be provided for downstream branch links without changing the core extraction logic.
[0054] In addition, before obtaining the seed context object of the target intelligent agent, the proposed solution also needs to construct a three-layer knowledge data model, which includes the following steps.
[0055] Define a concept node object: its fields must include at least: concept number, concept name, concept type, layer identifier (value can be "industry layer" or "enterprise layer"), attribute set (including the name, data type, value range and constraints of each attribute), status, source identifier, and corresponding industry layer concept number (can be empty, used for enterprise layer nodes to be associated with industry layer nodes to support inheritance).
[0056] Define a concept relation edge object: its fields must include at least: relation ID, source concept ID, target concept ID, relation type, and relation metadata. The relation type supports five semantic constraints: symmetric, transitive, reflexive, antisymmetric, and functional.
[0057] Define a slice object: its fields should include at least the following: slice number, attribution type (value is "template" or "instance"), attribution number (points to a job template or a smart agent instance), source layer combination (identifies which category each item in this slice comes from, such as industry layer, enterprise layer, or newly generated category), status (draft, activated, archived, pending review, etc.), and generation source identifier (points to the seed document or conversation session that triggered the generation of this slice).
[0058] Define a slice project object: its fields must include at least: slice number, project type (value is "concept item" or "relation item"), referenced object number, and source tag (value is "template baseline", "tenant increment" or "instance increment", corresponding to the three sources of ownership respectively).
[0059] The above four types of objects are stored in the database and assigned permissions according to "shared layer / tenant layer / instance layer": Shared layer (industry layer concept, template slice): all tenants can read, only the platform can write; Tenant layer (enterprise layer concept, new concept pending approval within a tenant): completely isolated between enterprises, enterprise administrators can write; Instance layer (instance slice, instance slice project): bound to a specific smart agent instance, different instances within the same enterprise are not visible to each other by default.
[0060] Establish monotonic visibility constraints—lower-level slices can reference upper-level concept nodes, but the reverse is not allowed; new concepts are only visible in the "ghost state" project of the slice itself before being approved, and cannot be written into the concept master table of its upper level (industry level or enterprise level).
[0061] Reference Figure 3 This is a schematic diagram of a three-layer knowledge data model and attribution relationship architecture provided in an embodiment of this application. For example... Figure 3 As shown, the data model is logically and explicitly divided into a shared layer, a tenant layer, and an instance layer. The shared layer, maintained uniformly by the platform, mainly includes an industry-level concept master table and a template slice master table, granting read access to all tenants but allowing only the platform to write to it. The tenant layer achieves absolute data isolation between different enterprises (such as tenant A and tenant B), mainly including an enterprise-level concept master table and a queue of new concepts awaiting approval, subject to dual gating by change policies and administrator review, granting enterprise administrators write permissions within the private domain.
[0062] Furthermore, the instance layer is directly bound to specific intelligent agent instances (such as customer service specialist instance A1, financial assistant instance A2, etc., belonging to tenant A), and includes instance slices and their corresponding slice project tables for each instance. Different instances within the same enterprise maintain isolation by default. In this three-tier architecture, monotonic visibility constraints are strictly followed between layers: lower-level instance slices can reference enterprise-private concepts in the tenant layer or industry-wide concepts in the shared layer, but upward references from upper layers to lower layers are strictly prohibited. Simultaneously, instance slices can be generated through forking from the template slice master table. Moreover, any newly extracted concept, before formal approval, can only be seen in a "ghost state" within its own slice, thus completely eliminating cross-tenant information leakage and the contamination of shared knowledge by private data at the underlying mechanism level.
[0063] Based on this pre-built three-layer knowledge data model, the next step will be to generate subsequent knowledge slices.
[0064] Step 102: Extract candidate knowledge based on business payload to obtain multiple candidate knowledge.
[0065] Step 102 is described in detail below.
[0066] In step 102 of some embodiments, the business payload is segmented into text, and then named entity recognition technology using large language model or non-large language model approach is used to identify elements such as business concepts, actions, resources, constraints or roles from the text segments to form multiple candidate knowledge. Each candidate knowledge typically includes a candidate name, type, key attributes and source evidence fragments in the original text.
[0067] In this application, the effective business payload of the seed context is segmented: if it is an unstructured document, it is slidably segmented according to paragraph title, paragraph separator and maximum character length (e.g. 800–1500 characters) to obtain a set of text segments T = {t_1, t_2, …, t_n}, and each segment retains its start and end character positions in the original text; if it is a structured dialogue, each round of question and answer is a text segment, and the round number is retained.
[0068] Next, preliminary candidate concept generation is performed (using a large model + post-processing verification, and supporting alternative implementations).
[0069] For each text segment t_k, construct structured prompt words and feed them into the large model. The prompt word template contains at least three parts: Task description: "Please identify the business concepts, actions, resources, constraints, and roles from the following text and return them in JSON array format. Each item contains the fields: name (concept name), type (concept type), attributes (list of key attribute values), and evidence (key phrases in the original text)"; Output format constraints: limit the return value to a JSON array, prohibit free text output, and provide 1–3 domain-independent examples; Context: the input text segment t_k.
[0070] After the large model is returned, post-processing is performed: the output results are formatted using a JSON parser; for each evidence field, a character-level search is performed in the original text t_k, and if no match is found, the item is discarded (to prevent illusions); the validated items are merged with the position information of t_k to obtain the candidate item set C_k for that segment.
[0071] In addition, in the embodiments of this application, a non-large model route of "Chinese named entity recognition model (such as BERT-CRF, UIE information extraction model) + relation classification model" can also be used to produce C_k. The solution of this application does not limit the specific model selection.
[0072] Reference Figure 4 Based on the business payload, candidate knowledge is extracted to obtain multiple candidate knowledge, including the following steps 401 to 405.
[0073] Step 401: Extract candidate knowledge based on the business payload to obtain multiple initial candidate knowledge.
[0074] Step 402: Based on the normalized form and candidate type of the initial candidate knowledge, perform precise deduplication processing on multiple initial candidate knowledge to obtain deduplicated candidate knowledge.
[0075] Step 403: Concatenate the candidate names, candidate types, and attribute key names of the deduplicated candidate knowledge to obtain text representation data, and input the text representation data into the pre-trained embedding model to obtain the corresponding candidate knowledge representation vector.
[0076] Step 404: Calculate the pairwise similarity between the two candidate knowledge representation vectors, and perform clustering processing on the deduplicated candidate knowledge based on the numerical comparison relationship between the pairwise similarity and the preset deduplication threshold to obtain multiple candidate knowledge clusters.
[0077] Step 405: For each candidate knowledge cluster, merge the attribute set and evidence fragment of the candidate knowledge within the cluster to obtain a merged candidate set, which includes multiple candidate knowledge clusters.
[0078] Steps 401 to 405 are described in detail below.
[0079] In step 401 of some embodiments, the received business payload is first subjected to preliminary text analysis and information extraction to identify and output initial candidate knowledge C_k, which includes elements such as business concepts, actions, and resources. Since business documents may be lengthy or contain multiple rounds of dialogue, the initial results directly extracted at this stage often contain a large number of duplicates or redundant expressions. In step 402 of some embodiments, the initial candidate knowledge C_k is merged into C = C_k is used to normalize the names (including removing whitespace, unifying capitalization, converting between simplified and traditional Chinese characters, and removing common suffixes such as "system" and "process"). Then, the normalized form and candidate type are used as joint judgment conditions to accurately deduplicate completely identical items, resulting in a set of deduplication candidate knowledge C' with preliminary noise reduction.
[0080] In step 403 of some embodiments, in order for the model to understand the comprehensive characteristics of the business entity, the system constructs structured text representation data by concatenating the candidate name, candidate type, and attribute key names contained in the concept for each deduplicated candidate knowledge c_i in the C' set, i.e., text_i = "name:<candidate name>| type:<candidate type>| attrs:<attribute key name concatenation>". Subsequently, this text representation data is input into a pre-trained embedding model (such as BGE / GTE / E5 general Chinese embedding or industry-adapted embedding) and mapped to a high-dimensional candidate knowledge representation vector v_i ∈ ^d, where dimension d is typically 768 or 1024, is used to capture its deep semantic features.
[0081] In step 404 of some embodiments, the system calculates the pairwise similarity (e.g., cosine similarity) between different candidate knowledge representation vectors, i.e., sim(c_i, c_j) = (v_i · v_j) / (||v_i|| · ||v_j||). Next, a preset deduplication threshold is introduced to perform clustering processing (e.g., greedy clustering based on the threshold θ_dedup). When the pairwise similarity of two candidate knowledge vectors is greater than or equal to the deduplication threshold, and their concept types are consistent, the system determines that they point to the same entity in the business context, and then merges them into the same cluster, thereby obtaining multiple semantically similar candidate knowledge clusters.
[0082] In the proposed scheme, during the clustering process: First, each candidate is initialized as an independent cluster; then, candidate pairs (c_i, c_j) are traversed from largest to smallest according to sim: if sim(c_i, c_j) ≥ θ_dedup (e.g., 0.86–0.92) and the two are of the same type, the clusters are merged; if more robustness is desired, DBSCAN or hierarchical clustering can be used instead. The proposed scheme does not limit the specific clustering algorithm.
[0083] In step 405 of some embodiments, the clustered candidate knowledge clusters are integrated with their internal information and representative items are generated. For each cluster, the most representative name is selected, and the attribute sets of all candidate knowledge within the cluster are merged. When encountering the same attribute key name, the value with the most evidence is retained. Simultaneously, the system aggregates all source evidence fragments within the cluster into a list, ultimately outputting a cohesive and complete merged candidate set C''. The candidate knowledge in this set then enters the downstream cross-layer matching stage as a high-purity data source.
[0084] In this application's scheme, during the cluster merging process, for multiple candidate options {c_{i1}, c_{i2}, ...} within each cluster: Representative name selection: The normalized name with the highest frequency of occurrence within the cluster is chosen; if tied, the one with the largest sum of evidence fragment lengths is selected; Attribute set merging: The attribute sets of all candidates within the cluster are joined, and for the same key name, the value with the highest number of evidence is selected; Evidence fragment merging: All evidence fragments within the cluster are aggregated into a list for easy traceability in subsequent audits; After cluster merging is completed, a merged candidate set C'' is obtained, which is the final candidate knowledge.
[0085] Through steps 401 to 405 above, a two-stage merging and deduplication mechanism combining literal precise deduplication and deep semantic clustering is constructed. This effectively compensates for the shortcomings of single precise deduplication in dealing with the high-frequency synonyms and homonyms in multi-source business texts. By fusing candidate names, types, and attribute keys to construct text representations, and using a pre-trained embedding model to calculate pairwise vector similarity for clustering, redundant concepts with similar semantics and pointing to the same actual business rules can be pre-fused with high quality before cross-level library matching, which consumes significant system resources. This reduces the computational power consumption of downstream cross-level knowledge retrieval and effectively integrates the attributes and evidence fragments of various synonymous entities, improving the richness and traceability of business concepts entering the target knowledge slice.
[0086] Step 103: Based on the preset tenant layer knowledge base and / or shared layer knowledge base, perform cross-level matching of multiple candidate knowledge according to the attribution metadata, obtain the matching status result corresponding to each candidate knowledge, and generate target knowledge slices based on the matching status results.
[0087] Step 103 will be described in detail below.
[0088] In step 103 of some embodiments, since the underlying data architecture of this solution explicitly splits knowledge into a shared layer (industry-wide knowledge) and a tenant layer (enterprise-owned knowledge), the system, based on the attribution metadata, allows candidate knowledge to be queried and compared sequentially in the tenant layer knowledge base and the shared layer knowledge base of the tenant to which it belongs; and through similarity calculation, the system determines the matching status of each candidate knowledge (e.g., divided into "matched", "partially matched" or "new candidate"), and selects the corresponding concept node references and relation edge references accordingly, and combines them to generate a target knowledge slice for the job template or instance. The target knowledge slice refers to the managed cognitive subset that is cut out from the upper-layer knowledge ontology as needed for a specific job template or specific intelligent agent instance, composed of concept node references and relation edge references, and used for local serialization injection into the intelligent agent at runtime, as described below.
[0089] Reference Figure 5 Based on a preset tenant layer knowledge base and / or shared layer knowledge base, multiple candidate knowledge are matched across layers according to the attribution metadata to obtain the matching status result corresponding to each candidate knowledge, and a target knowledge slice is generated based on the matching status result, including the following steps 501 to 504.
[0090] Step 501: Determine the target tenant layer knowledge base from multiple tenant layer knowledge bases based on the tenant number in the attribution metadata.
[0091] Step 502: Based on the preset concept type constraints, the text representation vector corresponding to each candidate knowledge is retrieved and matched in the target tenant layer knowledge base and / or shared layer knowledge base to determine the target node.
[0092] Steps 501 to 502 are described in detail below.
[0093] In step 501 of some embodiments, for each candidate knowledge c in C'', the tenant number carried in the seed context object is first used to accurately map and locate the specific enterprise layer (i.e., target tenant layer) knowledge base to which the current business belongs in a logically isolated multi-tenant environment to establish the ownership of knowledge data, thereby ensuring that subsequent query matching actions are carried out independently within the tenant's private domain, preventing cross-tenant data overreach or information leakage.
[0094] In step 502 of some embodiments, for each candidate knowledge c extracted in the early stage, the system performs preliminary pre-screening constraints based on its candidate type (such as business concept, action, resource, etc.), and then performs vector nearest neighbor retrieval (i.e. Top-K) in the target tenant layer knowledge base and the shared industry layer knowledge base in turn to efficiently and specifically recall potential matching target nodes, as described below.
[0095] Reference Figure 6 Based on preset concept type constraints, the text representation vector corresponding to each candidate knowledge is retrieved and matched in the target tenant layer knowledge base and / or shared layer knowledge base to determine the target node, including the following steps 601 to 603.
[0096] Step 601: For each candidate knowledge, based on the concept type constraint corresponding to the candidate knowledge, filter in the target tenant layer knowledge base and / or shared layer knowledge base to obtain at least one candidate knowledge node whose type matches the concept type constraint.
[0097] Step 602: Determine the vector similarity between the text representation vector of the candidate knowledge and the representation vectors corresponding to all candidate knowledge nodes.
[0098] Step 603: Based on vector similarity, perform nearest neighbor retrieval among all candidate knowledge nodes to obtain the target node.
[0099] Steps 601 to 603 are described in detail below.
[0100] In step 601 of some embodiments, for each extracted candidate knowledge, the system does not blindly calculate the distance directly in the large-scale underlying knowledge base. Instead, it first performs a pre-screening operation in the target tenant layer and / or shared layer knowledge base based on the concept type (such as business concept, role, action, resource or constraint) of the candidate knowledge as a constraint condition, so as to extract those nodes in the knowledge base whose type is consistent with the current candidate knowledge, forming a set of candidate knowledge nodes, thereby effectively reducing the search space for subsequent calculations.
[0101] In step 602 of some embodiments, after the type coarse screening is completed, text vectorization technology is introduced to measure the semantic proximity between knowledge points. The name and related features of the candidate knowledge are converted into a specific text representation vector (e.g., a high-dimensional vector processed by a pre-trained embedding model). Then, the spatial distance or similarity (e.g., cosine similarity) between the text representation vector of the candidate knowledge and the representation vectors corresponding to each candidate knowledge node selected above is calculated.
[0102] In step 603 of some embodiments, based on the vector similarity value, a nearest neighbor search is performed within the set of candidate knowledge nodes of the same type, such as a Top-K search sorted by vector similarity, so as to recall one or more nodes that are most closely related to the current candidate knowledge in terms of deep semantics from the candidate range, and establish them as target nodes t for subsequent comprehensive similarity and threshold determination.
[0103] Through steps 601 to 603 above, introducing concept type constraints in the early stage of retrieval can quickly filter out a large number of noisy nodes that do not match the business scope (such as avoiding invalid comparisons between "role" and "action"), reducing the computing power overhead in the background; on the other hand, using representation vectors for similarity retrieval within nodes of the same type overcomes the limitations of traditional simple character matching, which is easily affected by synonyms and synonyms, and ensures high quality of target node recall in the deep semantic dimension.
[0104] Step 503: Determine the overall similarity between each candidate knowledge and the target node.
[0105] Step 503 will be described in detail below.
[0106] In step 503 of some embodiments, for the target node to be recalled, the comprehensive similarity between the candidate knowledge and the target node is further refined. This comprehensive similarity abandons the single text literal comparison, but integrates the weighted calculation results of name similarity (vector and character-level matching degree), type similarity (concept inheritance relationship) and attribute similarity (intersection over union ratio and consistency weighting of attribute key-value pairs), so as to more comprehensively and realistically measure the degree of fit between the new candidate knowledge and the existing nodes in the database in the business context, as described below.
[0107] Reference Figure 7 The comprehensive similarity between each candidate knowledge and the target node is determined, including the following steps 701 to 704.
[0108] Step 701: For each candidate knowledge, determine the name dimension similarity based on the vector matching degree between the candidate knowledge and the target node in name semantics and the character matching degree in name string.
[0109] Step 702: Determine the type dimension similarity based on the relative logical distance between candidate knowledge and target node in the preset concept inheritance topology.
[0110] Step 703: Determine the attribute dimension similarity based on the intersection-union ratio of the set of attribute key names contained in the candidate knowledge and the target node, and the consistency weighting of the attribute values corresponding to the same attribute key names.
[0111] Step 704: Perform weighted processing based on name dimension similarity, type dimension similarity, and attribute dimension similarity to obtain the comprehensive similarity.
[0112] Steps 701 to 704 are described in detail below.
[0113] In step 701 of some embodiments, since conventional text representation vectors are often not sensitive enough to changes in local characters when processing short business concept names, the solution of this application introduces character-level matching based on the literal meaning of the name string, in addition to calculating the semantic matching degree (such as cosine similarity) between embedded vectors. In specific implementations, the system usually takes the maximum value of these two as the final name dimension similarity, thereby compensating for the deficiency of single vector retrieval in being insensitive to changes in short names.
[0114] In this application, the name dimension similarity sim_name(c, t) is calculated by taking the maximum value of the embedding cosine similarity cos(emb(c.name), emb(t.name)) and the character-level Jaro-Winkler similarity (to compensate for the problem that the embedding is not sensitive to changes in short names).
[0115] In step 702 of some embodiments, business concepts usually have an inherent classification system. Therefore, the present application scheme places candidate knowledge and target nodes in a preset concept type inheritance topology (such as a concept type inheritance table) for evaluation. If the two are completely identical in type, a higher logical proximity is assigned; if the two have a subordinate relationship such as "parent-child (is-a)", a lower logical distance score is assigned; if they are completely unrelated, the score is lower, thereby obtaining the type dimension similarity.
[0116] In this application, the type dimension similarity sim_type(c, t) is calculated as follows: if c.type == t.type, it is 1.0; if c.type and t.type have an is-a parent-child relationship in the concept type inheritance table, it is 0.7; otherwise, it is 0.0.
[0117] In step 703 of some embodiments, the attribute key name sets contained in the candidate knowledge and the target node are first extracted, and the intersection-union ratio (such as the Jaccard coefficient) of the two sets is calculated to measure the degree of overlap of the attribute structure. Then, the specific attribute values corresponding to the same attribute key names are extracted, the proportion of equal value strings is calculated, and this is used as a consistency weighting coefficient to correct and weight the aforementioned intersection-union ratio, thereby obtaining the attribute dimension similarity.
[0118] In this application, the attribute dimension similarity sim_attr(c, t) is calculated as follows: based on the Jaccard coefficient of the attribute name sets on both sides, the values of the same attribute name are weighted by character equivalence — sim_attr = J(A_c, A_t) × (1 + α · agree_ratio), where J is the Jaccard coefficient, agree_ratio is the proportion of equal string values under the same attribute name, α is 0.2, and the result does not exceed 1.0.
[0119] In step 704 of some embodiments, in actual complex business contexts, the contribution of name, type, and attribute to determining whether two concepts are equivalent usually differs. Therefore, according to a pre-set weight allocation principle (e.g., name has a higher weight, attribute has a lower weight, and type is used as an auxiliary factor), the similarity of the name dimension, the similarity of the type dimension, and the similarity of the attribute dimension are weighted and summed (default weights 0.45 / 0.20 / 0.35), and a comprehensive similarity value that can objectively reflect the overall matching degree of the two is obtained as shown in the following formula.
[0120] score(c, t) = w_name · sim_name + w_type · sim_type + w_attr ·sim_attr Through steps 701 to 704 above, a comprehensive similarity evaluation model interwoven with three dimensions—name, type, and attribute—is constructed. This overcomes the one-sidedness of evaluation caused by relying solely on character matching or local semantic comparison in traditional retrieval schemes. The name dimension takes into account the semantics of long texts and variations of phrase characters, the type dimension introduces the logical inclusiveness of hierarchical topology, and the attribute dimension delves into business details for value weighting verification. This provides a solid and reliable data quantification basis for subsequent cross-level three-state matching, thereby improving the accuracy of automatic deduplication and merging of business knowledge in complex contexts.
[0121] Step 504: Based on the numerical comparison relationship between the comprehensive similarity and the preset matching threshold, obtain the matching status result corresponding to each candidate knowledge, and generate the target knowledge slice based on the matching status result.
[0122] Step 504 will be described in detail below.
[0123] In step 504 of some embodiments, the system compares the calculated comprehensive similarity score(c, t) with a pre-set matching threshold. Based on the range in which the comparison falls, the candidate knowledge is precisely defined as a specific matching state such as "matched," "partially matched," or "new candidate." Furthermore, based on these clear matching state results, the system categorizes and includes the corresponding items (e.g., directly including matched items and including partially matched items in a pending confirmation state), thereby generating a controlled draft of the target knowledge slice.
[0124] The following section will further describe how to obtain the matching status results.
[0125] Reference Figure 8 The preset matching thresholds include a full matching threshold θ_match and a partial matching threshold θ_partial. Based on the numerical comparison between the comprehensive similarity and the preset matching thresholds, the matching status result corresponding to each candidate knowledge is obtained, and the target knowledge slice is generated based on the matching status result, including the following steps 801 to 804.
[0126] Step 801: When the overall similarity is greater than or equal to the full matching threshold, generate a matching status result representing the matched items, and confirm the corresponding candidate knowledge as the activated items.
[0127] Step 802: When the overall similarity is less than the full matching threshold and greater than or equal to the partial matching threshold, generate a matching status result representing the partial matching, and generate a difference description object between the candidate knowledge and the target node.
[0128] Step 803: When the overall similarity is less than the partial matching threshold, generate a matching state result representing a new candidate and mark the corresponding candidate knowledge as a ghost state item.
[0129] Step 804: Obtain the target knowledge slice based on at least one of the activated project, the difference description object, and the ghost state project.
[0130] Steps 801 to 804 are described in detail below.
[0131] In step 801 of some embodiments, the system uses a preset full-match threshold θ_match (such as any value between 0.90 and 0.95) to determine the high similarity of candidate knowledge. When the calculated comprehensive similarity reaches or exceeds the full-match threshold, i.e., score ≥ θ_match, it indicates that the currently extracted candidate knowledge is highly consistent with the target node in the underlying knowledge base. At this time, the system generates a state result representing "matched" and directly confirms the candidate knowledge as a usable activation item, includes it in the knowledge slice draft, and references the existing concept number, without additional manual intervention.
[0132] In step 802 of some embodiments, intermediate situations such as "similar names, consistent types, or partially overlapping attributes" that are common in the actual context of enterprises are processed. By introducing a partial matching threshold θ_partial (such as any value between 0.55 and 0.70) that is lower than the full matching threshold θ_match, when the overall similarity falls within the range between these two thresholds, i.e. θ_partial ≤ score < θ_match, the system determines it as "partial match". In order to support subsequent human-machine collaborative decision-making, the system not only outputs the matching status, but also extracts the specific differences between the two (such as attribute differences, relationship differences, name differences, etc.), and then generates a unique difference description object so that it can be directed to the manual confirmation channel.
[0133] In step 803 of some embodiments, when the overall similarity is lower than the partial matching threshold, i.e., score < θ_partial, it indicates that the candidate knowledge has not found sufficiently similar reference nodes in the current target tenant layer and / or shared layer. Based on this, the system generates a matching status result for a "new candidate". To prevent unapproved new concepts from directly participating in the injection of prompt words into the large model or polluting the ontology library, the system marks them as "ghost" items, placing them in an isolated transitional state, waiting to enter the policy-gated return queue later.
[0134] In step 804 of some embodiments, based on the three types of data objects generated by the comparison judgment, namely, the activated items confirmed for reuse, the items with attached difference descriptions pending manual confirmation, and the ghost state items in the transition period, the system aggregates these three (or one / more of them) to construct a structured target knowledge slice, namely {matched list, partially matched list, new candidate list}. Each element carries evidence fragments and matching / unmatched hierarchical labels. In this way, data of various matching states are classified and contained in the same slice object.
[0135] In this application, items in the "Matched" list are directly included in the draft as slice items. The "Source Tag" is marked according to the combination of the attribution type and the hit layer (when the attribution type of the build end is "Template", it is marked "Template Baseline" or "Tenant Increment"; when the attribution type of the user end is "Instance", it is marked "Tenant Increment" or "Instance Increment"). The relationship edges between the selected concepts are automatically pulled from the concept relationship edge master table and included as relationship items. Items in the partially matched list are included in the draft in the "Pending Confirmation" state with a difference description, and do not participate in the subsequent system prompt word injection until confirmation. Items in the new candidate list are included in the draft in the "Ghost State", do not participate in the system prompt word injection, and enter the return channel.
[0136] Through steps 801 to 804 above, by introducing a three-state matching mechanism with dual threshold grading, a separate manual confirmation channel is opened for the near-but not completely matching situations that frequently occur in enterprise scenarios by generating difference description objects; at the same time, new candidates are identified as isolated "ghost states" to avoid polluting the existing graph or large model prompt words before verification by unapproved new concepts, thereby improving the business fit of knowledge extraction and reuse, and effectively ensuring the security and controllability of system knowledge evolution.
[0137] Through steps 501 to 504 above, by performing type-constrained vector retrieval in a multi-level knowledge base and combining it with a multi-dimensional comprehensive similarity formula for threshold comparison, a cross-level tri-state matching mechanism with clear boundaries is realized. This overcomes the coarse-grained defect of binary judgment in existing technologies, which can only retain or discard candidate concepts that are approximately but not completely matched. It can not only accurately reuse existing target nodes, but also provide an independent intermediate processing channel for concepts with partially overlapping attributes or names, thereby improving the accuracy of knowledge slice content generation and the standardization of enterprise knowledge graph iteration.
[0138] Reference Figure 9 After obtaining the target knowledge slice, the knowledge slice data processing method further includes the following steps 901 to 903.
[0139] Step 901: When the attribution type identifier in the attribution metadata represents the construction end, the target knowledge slice is presented in a graph view format containing knowledge nodes and related relationship edges.
[0140] Step 902: When the attribution type identifier represents the user end, the target knowledge slices are presented in a business confirmation list view format that masks the graph structure.
[0141] Step 903: Obtain user confirmation instructions and perform status updates on the difference description object and ghost state items based on user confirmation instructions.
[0142] Steps 901 to 903 are described in detail below.
[0143] In step 901 of some embodiments, data is presented from the perspective of the technical personnel designing the job template. Based on the seed context object generated in the aforementioned steps, the attribution metadata is read. When the attribution type of the data is identified as "construction end", the system determines that the current operating entity has the ability to maintain the graph. At this time, the target knowledge slice draft is rendered into an intuitive concept graph visualization view (such as a multi-column display of nodes + node edges + attribute tables + source evidence fragments), which fully displays the knowledge nodes, relationship edges, and corresponding attribute tables and source evidence to the technical personnel, so that the technical personnel can directly perform in-depth maintenance operations such as editing, adding, or merging on the graph structure.
[0144] In step 902 of some embodiments, an adaptive transformation of the display layer is performed from the perspective of the business personnel who actually employ intelligent agents. When the attribution type is identified as "user end", the system determines that the current operating entity is more concerned with business fit than the underlying data structure. At this time, the same slice draft data is transcribed and rendered into a common business confirmation list view. In this view, the complex graph topology and field details are deliberately hidden. For example, "The system identifies the concepts involved in this position as: return policy, after-sales work order, customer level. Please confirm whether each item corresponds to your company's actual business." The graph structure and field details are not exposed, thereby significantly reducing the reading and review threshold for business personnel.
[0145] Understandably, both views share the same slice draft data, with the only difference being in the front-end rendering layer.
[0146] In step 903 of some embodiments, after the system completes the presentation of the differentiated view on the front end, it enters the confirmation and feedback stage of human-computer collaboration. Whether it is a technical staff member or a business staff member, after they check or annotate the "partially matched" items (corresponding to the objects of difference description) and "new candidates" items (corresponding to ghost state items) left over from the previous cross-layer matching on their respective views, the system will uniformly obtain these user confirmation instructions. Subsequently, the system will perform status updates of the underlying data according to the specific content of the instructions, such as confirming the partially matched items as adopting existing nodes or upgrading them to new concepts, and marking ghost state items as submitted for business review, etc., ultimately promoting the flow of knowledge slices from the draft state to the active state.
[0147] In this application scheme, for matched projects: users can choose to keep / delete them; for partially matched projects: users can choose to "adopt the target node (automatically cover differences)," "establish a new concept (upgrade to a completely new candidate)," or "mark it as needing further discussion (keep it pending confirmation)"; for ghost-state completely new candidates: on the user side, users can mark "the business really needs it, please submit for review"; on the build side, technical personnel can directly complete and write it; the user's confirmation action is recorded in the source_label and status fields of the slice project.
[0148] In addition, after a user submits a submission, the confirmed slice is written to the slice slice table, and the slice master record status is set to "active".
[0149] Through steps 901 to 903 above, by introducing a differentiated dual-view presentation and collaborative confirmation mechanism driven by attribution type, the same knowledge slice draft can be flexibly and adaptively transformed into a graph view that is easy for technical personnel to edit in depth, or a simple business list that is easy for business personnel to quickly select, depending on the different operating terminals. This eliminates the cognitive barriers for business personnel when facing complex graph structures, effectively lowers the threshold for using the system, and enables technical and business teams to collaborate seamlessly on the same set of data assets, improving the processing efficiency of manual confirmation and status transition for fuzzy concepts and new knowledge.
[0150] Step 104: When the target intelligent agent runs, it performs local serialization processing based on the activated items in the target knowledge slice to obtain business knowledge prompts, and adds the business knowledge prompts to the system prompts of the target intelligent agent.
[0151] Step 104 will be described in detail below.
[0152] In step 104 of some embodiments, when the target intelligent agent receives a user request and starts inference, the system retrieves the corresponding target knowledge slice based on the agent's affiliation number. The system extracts only the active concept items and relation items within the slice, converts their names, definitions, attributes, and relation edges into text paragraphs (i.e., business knowledge prompts) according to a preset template, and then assembles them into the intelligent agent's system prompts before sending them to the large model to perform the inference task.
[0153] Reference Figure 10 Based on the activated items in the target knowledge slice, local serialization processing is performed to obtain business knowledge prompt words, including the following steps 1001 to 1002.
[0154] Step 1001: Based on the project type of the activated project in the target knowledge slice, obtain the corresponding target concept information.
[0155] Step 1002: Based on the preset prompt word assembly template, the target concept information is converted into text format to obtain business knowledge prompt words containing business concept paragraphs and business relationship paragraphs.
[0156] Steps 1001 to 1002 are described in detail below.
[0157] In step 1001 of some embodiments, the local data extraction operation for a specific proxy instance limits the extraction scope to the effective target knowledge slice. At this time, according to the "project type" of each activated item in the slice, the corresponding target concept information is extracted. For items marked as "concept items", the corresponding concept name, definition and key attribute information are extracted in a targeted manner; for items marked as "relationship items", the system will extract the corresponding relationship edge information.
[0158] In step 1002 of some embodiments, the extracted structured knowledge graph data is converted into a text form that the large model can understand. At this time, a preset prompt word assembly template is called to serialize and transcribe the acquired target concept information. For example, the relevant information of the "concept item" is serialized according to the template requirements to generate the "business concept" paragraph in the system prompt words. At the same time, the information of the "relationship item" is converted into the "business relationship" paragraph according to a fixed format such as "<source concept>--[relationship type]--><target concept>". Then, these paragraphs are spliced together to generate complete business knowledge prompt words, so that they can be assembled with other modules such as job personality prompt words and sent to the large model for inference.
[0159] In this application's scheme, when the intelligent agent instance receives a user request, it retrieves the active slice according to its ownership number: extracts the concept name, definition, and key attributes corresponding to all "concept items" in the slice, and serializes them into a "business concept" paragraph with system prompt words according to a preset template; extracts the relationship edges corresponding to all "relationship items" in the slice, and serializes them into a "business relationship" paragraph according to the format "<source concept>--[relationship type]--><target concept>"; only the content within the scope of this slice is injected, and concepts outside this intelligent agent slice are not injected, nor are slices of other job instances in the enterprise injected; the serialization result is assembled with job personality prompt words, skill call constraints, and other modules and then sent to the large model for inference.
[0160] Through steps 1001 to 1002 above, by performing local serialization and template assembly based on the activated items within the target knowledge slice range, it is ensured that only the concept names, key attributes, and relation edges required by the current intelligent agent are injected into the large model at runtime. This effectively reduces the dilution of the context window by irrelevant concepts and lowers the cost of model attention being occupied. It also physically isolates sensitive concepts that the current position should not perceive, reducing the risk of cross-position context overreach. This makes the knowledge injection process of the large model both business accuracy and data security.
[0161] In one example, when a company hires a job template from the market to create a new smart agent instance, the new instance slice is initialized as a reference copy of the baseline slice corresponding to that template version (i.e., the "template baseline" layer). Then, following the above process, the company's uploaded business data generates two layers of "tenant increment" and "instance increment", forming a three-layer slice of "template baseline + tenant increment + instance increment".
[0162] Reference Figure 11 This is a schematic diagram of a process for slicing and extracting the main waterline according to an embodiment of this application. Figure 11 As shown, this pipeline supports input from both ends with the same source. Whether it's structured dialogue data generated by the build end (technicians) or business documents uploaded by the user end (business personnel), both are uniformly encapsulated at the front end into a seed context object containing fields such as seed type, payload, attribution type, and tenant. Subsequently, the system sequentially extracts candidate concepts through a large model combined with JSON post-processing (or using named entity recognition and relation classification models), and performs semantic-level deduplication and merging using embedding vectors and threshold clustering techniques, thereby providing a structured and deduplicated candidate knowledge set for downstream comparison.
[0163] After extraction and deduplication, the process enters the core cross-layer three-state matching stage. The system prioritizes searching candidate knowledge at the enterprise level, followed by the industry level, and calculates the comprehensive similarity. Based on the comparison result of this comprehensive similarity with preset dual thresholds, candidate knowledge is precisely divided into three state paths: "Matched" items with sufficient similarity are directly included in the slice draft; "Partially Matched" items with moderate similarity generate difference descriptions and are transferred to the channel requiring manual confirmation; "New Candidates" with low similarity are treated as ghost state items and enter a return queue subject to dual gating mechanisms. This diversion mechanism effectively avoids the simplistic binary discarding problem in traditional knowledge construction.
[0164] Based on the slice drafts generated from the above three types of status items, the system renders differentiated collaborative views according to the user's affiliation. For the build side, a graph view containing details of nodes, relationships, and attributes is presented; for the user side, a business confirmation checklist view that masks the complex graph structure is presented. After receiving the corresponding review and confirmation instructions, the target knowledge slice is officially activated. Finally, the system extracts only the activated knowledge items within the slice, performs local serialization processing, converts them into text format, and injects them into the system prompts in the runtime large model, thus completing a controlled business context injection loop.
[0165] Through steps 101 to 104 above, by unifying and abstracting heterogeneous inputs from multiple terminals into seed context objects with attribution metadata, and combining them with a controlled cross-level comparison and judgment mechanism, target knowledge slices composed of concepts and relationships can be cut out as needed from massive multi-tenant data. During the intelligent agent operation phase, only local knowledge nodes within the slice are serialized and injected into the large model, effectively avoiding the problems of the large model context window being diluted by irrelevant concepts and the overreach of authority in sensitive cross-position contexts caused by directly importing the full static graph. While controlling lexical consumption and inference attention, the business knowledge injected into the large model has a high degree of job relevance and process traceability.
[0166] In addition, refer to Figure 12 The knowledge slice data processing method provided in this application also includes the following steps 1201 to 1204.
[0167] Step 1201: Add the ghost-state project to the pending review queue corresponding to the target tenant layer knowledge base.
[0168] Step 1202: Use the preset change strategy checker to perform gating verification on ghost state items in the audit queue and obtain the verification results.
[0169] Step 1203: When the verification result representation passes, update the candidate knowledge corresponding to the ghost state project to the formal activation node in the target tenant layer knowledge base.
[0170] Step 1204: Retrieve historical knowledge slices that reference candidate knowledge in the target tenant layer knowledge base, and synchronously update all ghost state items contained in all historical knowledge slices to active items.
[0171] Steps 1201 to 1204 are described in detail below.
[0172] In step 1201 of some embodiments, when a business user triggers the generation of a new candidate on the user end, these items marked as ghost state are added to the pending review queue under the target tenant layer knowledge base, thus serving as a buffer before the subsequent manual and policy gating processes.
[0173] In step 1202 of some embodiments, a preset change policy checker is used to read the change policy configured by the current tenant (including rule items such as concept type sensitivity classification, single batch quantity limit, blacklist keywords and source document credibility). Through these compliance rules of business dimensions, gating verification is performed on ghost projects in the pending review queue, thereby obtaining judgment results such as "pass", "block" or "transfer to manual discretion".
[0174] In step 1203 of some embodiments, when the ghost state project successfully passes the verification of the aforementioned change policy checker and is approved by the tenant administrator's review queue, the system formally writes the candidate knowledge into the enterprise layer concept master table, and at the same time changes its status from the isolated "ghost state" to the formal "active" state.
[0175] In step 1204 of some embodiments, once the new concept is upgraded to a formal activation node, the system performs a reverse scan using the candidate number to retrieve all historical slice projects within the enterprise that reference the ghost candidate. Subsequently, the system synchronously upgrades these retrieved historical projects to activation projects, enabling them to immediately participate in the injection of system prompt words.
[0176] In this application, if the candidate belongs to the "template" category and the operator is the platform / technical personnel, the candidate concept can be directly written into the main table of industry or enterprise level concepts after being reviewed by the technical personnel. The writing level is specified by the "target level" field when the user submits the application.
[0177] If a candidate belongs to the "Instance" category and the operator is a business user, the candidate concept is not directly written to the enterprise-level concept master table. Instead, it enters two gates: The first gate: Change Policy Inspector: This reads the change policy object configured for the tenant. The change policy must include at least the following rules: concept type sensitivity classification (e.g., sensitive categories such as "involving personal information," "involving compliance standards," and "involving financial accounting" require dual review); maximum number of new concepts per batch (to prevent business users from submitting a large amount of noise at once); blacklisted keywords (e.g., containing "password," "bank card," etc.); source document credibility (metadata verification of self-uploaded documents, such as issuing department and document version). The Change Policy Inspector outputs three judgments for each candidate: "Pass," "Block," and "Transfer to Manual Judgment." The second gate: Tenant Administrator Queue: Candidates that pass the policy check enter the enterprise administrator's pending review queue. Queue entries must include at least: the full text of the candidate concept, evidence fragments, submitter, submitted instance, related slices, and the policy inspector's judgment result. The administrator performs four operations for each entry: "Accept," "Accept after modification," "Reject," and "Return to submitter."
[0178] Once a candidate concept is adopted by the administrator: it is written to the enterprise-level concept master table, and its status is upgraded from "ghost" to "active"; the system performs a reverse scan of all slice projects that reference this ghost candidate (searched by candidate number), upgrading their status from "ghost" to "active," making the project immediately participate in system prompt word injection. Once a candidate concept is upgraded to a formal enterprise-level concept, other intelligent agent instances within the enterprise can automatically match this concept as "matched" during the next cross-level matching, avoiding the repeated rediscovery of the same concept on multiple instances. This mechanism constitutes a closed loop of knowledge accumulation within the enterprise.
[0179] Reference Figure 13 This is a flowchart illustrating a novel concept of backflow closed loop, dual gating, and cross-instance automatic hit provided in an embodiment of this application. Figure 13 As shown, when an instance slice on the user side extracts a completely new, unmatched candidate concept during business operations, these concepts are first isolated as "ghost" projects and are not directly written into the enterprise-level concept master table. Subsequently, these ghost projects enter the first gate, namely the change policy checker. This checker performs pre-emptive automated verification of candidate concepts based on preset policy rules (such as concept sensitivity classification, single batch quantity limit, blacklist keywords, and source document credibility), and decides whether to block and return the submission to the submitter, or approve and transfer it to the manual stage, based on the verification results.
[0180] Candidate concepts that pass the policy check then enter the second gate, namely the tenant administrator's review queue. In this queue, the system presents the administrator with the full text of the candidate concept, related evidence fragments, submitter and instance information, related slices, and the policy checker's judgment results, allowing the administrator to make a comprehensive judgment to perform operations such as adoption, modified and adopted, or rejection and return. When a phantom-state candidate concept is confirmed as adopted by the administrator, the system officially writes it into the enterprise-level concept master table and changes its status from "phantom state" to "active state".
[0181] After a candidate concept is upgraded to an active state, the system will trigger the underlying data synchronization and reuse mechanism. On one hand, the system reverse-scans all historical slice projects that referenced the ghost candidate, synchronizing and upgrading their status to "active," making the relevant slice projects effective immediately. On the other hand, this closed-loop mechanism enables the internal knowledge accumulation and sharing within the enterprise, namely, other intelligent agent instances within the enterprise (such as...) Figure 13 When instances X, Y, and Z in the system perform the next cross-layer knowledge matching, they can automatically match the upgraded concept. This "one-time review, enterprise-wide effect" approach avoids the repeated extraction and manual confirmation of the same concept across multiple business instances, improving the efficiency of enterprise-wide private knowledge updates and transfer.
[0182] Through steps 1201 to 1204 above, a new concept backflow closed loop is constructed, subject to dual gating by change policies and a pending review queue. This ensures that new knowledge generated by business personnel must undergo strict compliance and quality checks before being formally entered into the database, preventing business noise or sensitive unauthorized data from polluting the enterprise's private knowledge base. In addition, in conjunction with the reverse synchronization mechanism, once a ghost-state project is approved, it can not only immediately activate the local injection of the original slice, but also cause other intelligent agent instances within the same enterprise to automatically hit the upgraded concept in the next matching. This forms a closed loop of enterprise knowledge accumulation that allows for one-time review and cross-instance reuse, improving the business efficiency of knowledge backflow while ensuring the security of multi-tenant knowledge graph evolution.
[0183] In addition, in this application, a version snapshot record will be generated when any of the following three events occur: the slice becomes effective for the first time from a draft; a slice project status change is triggered by a concept upgrade approved by the administrator; the platform releases a new version of the job template, and the template baseline slice is updated. Each snapshot record shall include at least: snapshot number, slice number, snapshot payload (full project copy of the slice), parent snapshot number (forming a version chain), generator, generation time, and triggering event type.
[0184] In addition, the differential calculation process between the instance slice and the template baseline slice is as follows: given the instance slice S_inst (the latest active snapshot) and the template baseline slice S_base (the snapshot when the template version was released) corresponding to its referenced template version, it includes the following steps.
[0185] Step 1: Concept Item Differentiation: Using "Reference Object Number" (the referenced concept number) as the key, compare the concept item sets on both sides: Common = {x | x ∈ S_inst.concepts ∩ S_base.concepts}; InstAdd = {x | x ∈ S_inst.concepts \ S_base.concepts}; InstDel = {x | x ∈ S_base.concepts \ S_inst.concepts}.
[0186] Step 2: Relationship Item Difference: Using the (source concept ID, relation type, target concept ID) triple as the key, perform the same set difference on both sides of the relation item set as above.
[0187] Step 3: Attribute-level Differentiation: For each pair in Common (instance-side x_i, baseline-side x_b), compare the attributes_json field attribute by attribute: If the attribute key sets on both sides are different or the same key has different values, output the Modified item and the reason for the change (attribute addition / attribute deletion / value modification).
[0188] Step 4: Source Tag Layering and Categorization: Based on the "Source Tag" field of each project, categorize the difference results into: "Tenant Additions" set ← InstAdd ∪ Modified for projects with the source tag "Tenant Additions"; "Instance Additions" set ← InstAdd ∪ Modified for projects with the source tag "Instance Additions"; "Baseline Removals" set ← InstDel; "Baseline Consistency" set ← Other Common items.
[0189] When the platform releases a new version of the job template (upgrading the baseline slice from S_base_old to S_base_new), a three-way merge is performed on each instance slice S_inst that is already bound to the template: Input: S_inst (old on the instance side), S_base_old (old on the template side), S_base_new (new on the template side); Output: the merged S_inst_new and the conflict list Conflicts; The corresponding algorithm is as follows.
[0190] BaseDiff = diff(S_base_old, S_base_new) gets the set of new / deleted / modified items on the template side; InstDiff = diff(S_base_old, S_inst) obtains the "tenant increment" and "instance increment" on the instance side; S_inst_new = S_base_new + InstDiff (First, the instance-side incremental is shifted onto the new template baseline). If the same item is modified in both BaseDiff and InstDiff and the modifications are inconsistent, add Conflicts for the user to decide.
[0191] When a smart agent instance is archived, its slices are not physically deleted; only the "Status" field is changed to "Archived". When an enterprise changes job templates or a person changes for the same job, the migration assistant calculates the TenantAdditions ∪ InstanceAdditions of the old instance and uses it as the initialization input for the new instance to run through the slice generation pipeline again. This makes the "business context taught" for a certain job a portable and inheritable asset.
[0192] In addition, the concept references in the slice establish cross-references with the "required concept references" field of the "skill (capability pack)" object, so that one of the prerequisites for "this skill can be activated on a certain instance" is that all the concepts it declares have been included in the slice; when the slice is shrunk (a concept is removed), the related skills that reference the concept are automatically put into the pending activation state.
[0193] Slice differences also serve as input for the automatic generation of evaluation use cases: for a new concept entering a slice, the evaluation use case generator is triggered to produce an "input-expected output" pair for that concept, which is then included in the instance evaluation set.
[0194] This application also provides a knowledge slice data processing apparatus, which can implement the above-described knowledge slice data processing method, as described above. Figure 14 The device 1400 includes: The acquisition module 1410 is used to acquire the seed context object for the target intelligent agent. The seed context object includes the business payload to be extracted and the associated attribution metadata. The candidate knowledge determination module 1420 is used to extract candidate knowledge based on the business payload to obtain multiple candidate knowledge. The knowledge slice determination module 1430 is used to perform cross-level matching of multiple candidate knowledge based on the preset tenant layer knowledge base and / or shared layer knowledge base and according to the attribution metadata, to obtain the matching status result corresponding to each candidate knowledge, and to generate the target knowledge slice based on the matching status result. The prompt word update module 1440 is used to perform local serialization processing on the activated items in the target knowledge slice when the target intelligent agent is running, to obtain business knowledge prompt words, and to add the business knowledge prompt words to the system prompt words of the target intelligent agent.
[0195] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, the specific implementation of the knowledge slice data processing device is basically the same as the specific implementation of the knowledge slice data processing method described above, and will not be repeated here.
[0196] This application also provides an electronic device, including: At least one memory; At least one processor; At least one program; The program is stored in a memory, and the processor executes the at least one program to implement the knowledge slice data processing method described above in this application. The electronic device can be any smart terminal, including mobile phones, tablets, personal digital assistants (PDAs), and in-vehicle computers.
[0197] Please see Figure 15 , Figure 15 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 1501 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 1502 can be implemented in the form of ROM (Read Only Memory), static storage device, dynamic storage device, or RAM (Random Access Memory). The memory 1502 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1502 and is called and executed by the processor 1501 using the knowledge slice data processing method of the embodiments of this application. The input / output interface 1503 is used to implement information input and output; The communication interface 1504 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1505 transmits information between various components of the device (e.g., processor 1501, memory 1502, input / output interface 1503, and communication interface 1504); The processor 1501, memory 1502, input / output interface 1503 and communication interface 1504 are connected to each other within the device via bus 1505.
[0198] This application also provides a storage medium, which is a computer-readable storage medium, storing a computer program that, when executed by a processor, implements the above-described knowledge slice data processing method.
[0199] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0200] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0201] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0202] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0203] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0204] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0205] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0206] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. The coupling or direct coupling or communication connection between the shown or discussed units may be through some interfaces, or indirect coupling or communication connection between the apparatus or units, and may be electrical, mechanical, or other forms.
[0207] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0208] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0209] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0210] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A knowledge slice data processing method, characterized in that, The method includes: Obtain a seed context object for the target intelligent agent, the seed context object including the business payload to be extracted and the associated attribution metadata; Based on the aforementioned business payload, candidate knowledge is extracted to obtain multiple candidate knowledge items; Based on a preset tenant layer knowledge base and / or shared layer knowledge base, multiple candidate knowledges are matched across layers according to the attribution metadata to obtain a matching status result corresponding to each candidate knowledge, and a target knowledge slice is generated based on the matching status result. When the target intelligent agent runs, it performs local serialization processing based on the activated items in the target knowledge slice to obtain business knowledge prompts, and adds the business knowledge prompts to the system prompts of the target intelligent agent.
2. The knowledge slice data processing method according to claim 1, characterized in that, The step of obtaining the seed context object for the target intelligent agent includes: Obtain structured dialogue data generated by the construction end for the job template, and / or obtain unstructured business documents input by the user end for the target intelligent agent; Based on the attribution metadata, the structured dialogue data and / or the unstructured business documents are uniformly encapsulated to obtain the seed context object.
3. The knowledge slice data processing method according to claim 1, characterized in that, The process of extracting candidate knowledge based on the service payload yields multiple candidate knowledge items, including: Based on the business payload, candidate knowledge is extracted to obtain multiple initial candidate knowledge; Based on the normalized form and candidate type of the initial candidate knowledge, the multiple initial candidate knowledge are subjected to precise deduplication processing to obtain deduplicated candidate knowledge. The candidate names, candidate types, and attribute key names of the deduplicated candidate knowledge are concatenated to obtain text representation data, and the text representation data is input into the pre-trained embedding model to obtain the corresponding candidate knowledge representation vector. Calculate the pairwise similarity between two candidate knowledge representation vectors, and perform clustering processing on the deduplicated candidate knowledge based on the numerical comparison relationship between the pairwise similarity and the preset deduplication threshold to obtain multiple candidate knowledge clusters; For each candidate knowledge cluster, the attribute set and evidence fragments of the candidate knowledge within the cluster are merged to obtain a merged candidate set, which includes multiple candidate knowledge clusters.
4. The knowledge slice data processing method according to claim 1, characterized in that, The method, based on a preset tenant-layer knowledge base and / or shared-layer knowledge base, performs cross-layer matching on multiple candidate knowledge items according to the attribution metadata, obtains a matching status result corresponding to each candidate knowledge item, and generates a target knowledge slice based on the matching status result, including: Based on the tenant ID in the attribution metadata, the target tenant layer knowledge base is determined from multiple tenant layer knowledge bases; Based on preset concept type constraints, the text representation vector corresponding to each candidate knowledge is retrieved and matched in the target tenant layer knowledge base and / or shared layer knowledge base to determine the target node; Determine the overall similarity between each candidate knowledge and the target node; Based on the numerical comparison relationship between the comprehensive similarity and the preset matching threshold, the matching status result corresponding to each candidate knowledge is obtained, and a target knowledge slice is generated based on the matching status result.
5. The knowledge slice data processing method according to claim 4, characterized in that, The step of determining the target node by searching and matching the text representation vector corresponding to each candidate knowledge in the target tenant layer knowledge base and / or shared layer knowledge base based on preset concept type constraints includes: For each candidate knowledge, based on the concept type constraint corresponding to the candidate knowledge, filtering is performed in the target tenant layer knowledge base and / or the shared layer knowledge base to obtain at least one candidate knowledge node whose type matches the concept type constraint; Determine the vector similarity between the text representation vector of the candidate knowledge and the representation vectors corresponding to all candidate knowledge nodes; Based on the vector similarity, a nearest neighbor search is performed on all the candidate knowledge nodes to obtain the target node.
6. The knowledge slice data processing method according to claim 4, characterized in that, Determining the comprehensive similarity between each candidate knowledge and the target node includes: For each candidate knowledge, the name dimension similarity is determined based on the vector matching degree between the candidate knowledge and the target node in name semantics and the character matching degree in name string. Based on the relative logical distance between the candidate knowledge and the target node in the preset concept inheritance topology, the type dimension similarity is determined; The similarity of the attribute dimension is determined based on the intersection-union ratio of the candidate knowledge and the set of attribute key names contained in the target node, and the consistency weighting of the attribute values corresponding to the same attribute key names. The comprehensive similarity is obtained by weighting the similarity based on the name dimension, the type dimension, and the attribute dimension.
7. The knowledge slice data processing method according to claim 4, characterized in that, The preset matching threshold includes a full matching threshold and a partial matching threshold. The matching status result corresponding to each candidate knowledge is obtained based on the numerical comparison relationship between the comprehensive similarity and the preset matching threshold, and a target knowledge slice is generated based on the matching status result, including: When the overall similarity is greater than or equal to the full matching threshold, a matching status result representing the matched items is generated, and the corresponding candidate knowledge is confirmed as the activated item. When the overall similarity is less than the full matching threshold and greater than or equal to the partial matching threshold, a matching state result representing the partial matching is generated, and a difference description object between the candidate knowledge and the target node is generated. When the overall similarity is less than the partial matching threshold, a matching state result representing a new candidate is generated, and the corresponding candidate knowledge is marked as a ghost state item. The target knowledge slice is obtained based on at least one of the activated item, the difference description object, and the ghost state item.
8. The knowledge slice data processing method according to claim 7, characterized in that, After obtaining the target knowledge slice, the method further includes: When the attribution type identifier in the attribution metadata represents the construction end, the target knowledge slice is presented in a graph view format that includes knowledge nodes and related relationship edges; When the attribution type identifier represents the user end, the target knowledge slice is presented in a business confirmation list view format that hides the graph structure; Obtain user confirmation instructions, and perform status updates on the difference description object and the ghost state item based on the user confirmation instructions.
9. The knowledge slice data processing method according to claim 1, characterized in that, The process of performing local serialization on activated items in the target knowledge slice to obtain business knowledge prompts includes: Based on the project type of the activated project in the target knowledge slice, obtain the corresponding target concept information; Based on a preset prompt word assembly template, the target concept information is converted into a text format to obtain the business knowledge prompt words that include business concept paragraphs and business relationship paragraphs.
10. The knowledge slice data processing method according to claim 7, characterized in that, The method further includes: Add the ghost-state project to the pending review queue corresponding to the target tenant layer knowledge base; The pre-defined change policy checker is used to perform gating verification on the ghost state items in the queue to be reviewed, and the verification result is obtained. When the verification result is passed, the candidate knowledge corresponding to the ghost state project is updated to the officially activated node in the target tenant layer knowledge base; Retrieve historical knowledge slices that reference the candidate knowledge in the target tenant layer knowledge base, and synchronously update the ghost state items contained in all the historical knowledge slices to the active items.
11. A knowledge slice data processing device, characterized in that, The device includes: The acquisition module is used to acquire a seed context object for the target intelligent agent. The seed context object includes the business payload to be extracted and the associated attribution metadata. The candidate knowledge determination module is used to extract candidate knowledge based on the business payload to obtain multiple candidate knowledge. The knowledge slice determination module is used to perform cross-level matching of multiple candidate knowledge based on the preset tenant layer knowledge base and / or shared layer knowledge base, according to the attribution metadata, to obtain the matching status result corresponding to each candidate knowledge, and to generate a target knowledge slice based on the matching status result. The prompt word update module is used to perform local serialization processing on the activated items in the target knowledge slice when the target intelligent agent is running, to obtain business knowledge prompt words, and to add the business knowledge prompt words to the system prompt words of the target intelligent agent.
12. An electronic device, characterized in that, The device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the knowledge slice data processing method according to any one of claims 1 to 10.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the knowledge slice data processing method according to any one of claims 1 to 10.