Quality evaluation method and device of knowledge graph, computing equipment and storage medium

By processing the target triples and reference examples of the knowledge graph with a large model, the problem of quantifying contradictory knowledge in the knowledge graph is solved, and the objective quantitative evaluation of the knowledge graph quality and the improvement of relationship confidence are achieved.

CN120632115APending Publication Date: 2025-09-12HUAWEI TECH CO LTD

Patent Information

Application Number
CN202410289646.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-11
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies make it difficult to reliably quantitatively evaluate contradictory knowledge or conflicts in knowledge graphs, resulting in the inability to quantify knowledge quality.

Method used

A large model is used to process the target triples and reference examples in the knowledge graph, and the relationship confidence of the target triples is obtained through computing equipment. The multi-domain knowledge and powerful semantic understanding capabilities of the large model are used for quantitative evaluation.

Benefits of technology

It achieves objective quantitative evaluation of knowledge graph quality, improves the prediction performance of target triple relationship confidence, and improves the reliability and accuracy of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632115A_ABST
    Figure CN120632115A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge graph quality evaluation method and device, computing equipment and a storage medium, and relates to the technical field of machine learning. Based on the fact that a large model has multi-domain knowledge and strong semantic understanding ability, a computing device adopts the large model to process reference examples of a target triple and a target triple, the relation confidence degree of knowledge described by the target triple is directly estimated, and objective quantitative evaluation of the knowledge graph quality is achieved. Moreover, the input of the large model not only comprises the target triad, but also comprises the reference example of the target triad, and the relationship attribute of the relationship in the reference triad included in the reference example is matched with the relationship attribute of the relationship in the target triad. And the prediction performance of a large model on the relation confidence coefficient in the target triad based on the relations in the reference triads can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of machine learning technology, and in particular to a method, apparatus, computing device, and storage medium for evaluating the quality of a knowledge graph. Background Art

[0002] A knowledge graph is a graphical model used to represent and store knowledge. By organizing entities, relationships, and attributes into nodes, edges, and labels, it provides a structured, systematic, and standardized description of all knowledge in the world. Knowledge graph technology offers many advantages when used to represent knowledge, including clear semantics, good interpretability, information traceability, and strong domain expertise. However, due to the limited input information in knowledge graphs and the diverse knowledge represented by different input information, this leads to contradictory knowledge and conflicts in the knowledge graph, making it difficult for users to reliably quantify and effectively evaluate the quality of different knowledge in the knowledge graph. Summary of the Invention

[0003] The present application provides a knowledge graph quality assessment method, apparatus, computing device and storage medium, which solves the problem that the quality of knowledge cannot be quantified due to contradictory knowledge or conflicts in the knowledge graph.

[0004] This application adopts the following technical solution.

[0005] In the first aspect, the present application provides a quality assessment method for a knowledge graph. The quality assessment method is executed by a computing device or a chip (such as a processor) in the computing device, and the quality assessment method includes: the computing device obtains a quality assessment request, and the quality assessment request indicates that the quality of the knowledge graph is quantified; and the computing device obtains the target triples and reference examples of the target triples in the knowledge graph according to the quality assessment request. Among them, the target triples include: a head entity, a tail entity, the relationship between the head entity and the tail entity, and the attributes of the head entity and the attributes of the tail entity; the aforementioned reference examples include: multiple reference triples, and the relationship attributes of the relationships in the reference triples match the relationship attributes of the relationships in the target triples. Finally, the computing device processes the target triples and the reference examples based on the large model to obtain the relationship confidence of the target triples.

[0006] In this application, since the large model has multi-domain knowledge and powerful semantic understanding capabilities, the computing device uses the large model to process the target triples and reference examples of the target triples, directly estimating the confidence of the relationships of the knowledge described by the target triples, and achieving an objective quantitative evaluation of the quality of the knowledge graph. Moreover, since the input of the large model includes not only the target triples but also the reference examples of the target triples, the relational attributes of the relationships in the reference triples included in the reference examples match the relational attributes of the relationships in the target triples, which is conducive to improving the prediction performance of the large model for the confidence of the relationships in the target triples based on the relationships in these reference triples.

[0007] In conjunction with the quality assessment method provided in the first aspect, in an optional implementation, the plurality of reference triples include: positive triples and negative triples whose relational attributes are consistent with those of the relation in the target triple. The positive triples indicate true knowledge, and the negative triples indicate false knowledge.

[0008] In the first possible case, true knowledge means that the confidence of the entities, relationships and attributes in the triples determined by the computing device is greater than or equal to the confidence threshold, and false knowledge means that the confidence of the entities, relationships and attributes in the triples determined by the computing device is less than the aforementioned confidence threshold.

[0009] In the second possible case, true knowledge refers to the knowledge that the entities, relations and attributes in the triples confirmed by the user conform to the laws of nature, and false knowledge refers to the knowledge that the relations between the entities in the triples confirmed by the user do not conform to the laws of nature.

[0010] In combination with the quality assessment method provided in the first aspect, in an optional implementation, a computing device obtains a target triple and a reference example of the target triple in a knowledge graph, including: the computing device obtains the target triple from the knowledge graph, and determines a reference triple pool from multiple triple pools based on the relational attributes of the relations in the target triple; the relational attributes of the relations in all triples in the reference triple pool are consistent with the relational attributes of the relations in the target triple. Furthermore, the computing device selects M positive triples and N negative triples from the reference triple pool, and obtains a reference example of the target triple based on the selected M positive triples and N negative triples. M and N are both positive integers.

[0011] In this application, since the computing device divides different triple example pools according to relationship attributes, and extracts positive example triples and negative example triples with corresponding relationship attributes as reference contexts (reference examples), the target triples and the reference examples have the same relationship attributes, which can assist the large model to better understand the relationship between entities in the target triples, thereby improving the large model's prediction performance on the relationship confidence of the target triples.

[0012] In conjunction with the quality assessment method provided in the first aspect, in an optional implementation, before the computing device obtains triples and reference examples of triples in the knowledge graph, the quality assessment method provided in this application further includes: the computing device obtains one or more historical knowledge graphs, and divides the triples in the one or more historical knowledge graphs into multiple triple example pools according to relationship attributes. In particular, the relationship attributes of the relationships in all triples included in a triple example pool are consistent, and the multiple triple example pools include a reference triple example pool.

[0013] Optionally, the multiple relationship attributes corresponding to the multiple triple example pools include two or more of the following: causal relationship, temporal relationship, hierarchical relationship, interpersonal relationship, affiliation relationship or conditional relationship.

[0014] In conjunction with the quality assessment method provided in the first aspect, in an optional implementation, the quality assessment method provided in this application further includes: if the relationship confidence of the target triple is greater than or equal to a first threshold, or if the relationship confidence of the target triple is less than or equal to a second threshold, then the computing device updates the reference triple example pool based on the target triple. Wherein, the second threshold is less than or equal to the aforementioned first threshold.

[0015] In this application, the computing device updates the reference triple example pool according to the confidence interval to which the relationship confidence of the target triple belongs, which is conducive to improving the reliability of the reference triple example pool, so that when the large model determines the reference examples of other triples based on the reference triple example pool, the prediction performance of the large model for the relationship confidence of other triples is improved.

[0016] Exemplarily, when the relationship confidence of the target triple is greater than or equal to the first threshold, the computing device uses the target triple as a positive example triple to update the reference triple example pool.

[0017] As another example, when the relationship confidence of the target triple is less than or equal to the second threshold, the computing device uses the target triple as a counterexample triple to update the reference triple example pool.

[0018] In conjunction with the quality assessment method provided in the first aspect, in an optional implementation, a computing device processes a target triple and a reference example based on a large model to obtain a relationship confidence score for the target triple, including: generating a sequence text including the target triple and the reference example, and generating a first question based on the sequence text; and inputting the first question into the large model to obtain a relationship confidence score for the target triple.

[0019] In combination with the quality assessment method provided in the first aspect, in an optional implementation, a computing device generates a sequence text including a target triple and a reference example, including: the computing device inserts the target triple and the reference example into a reference question template to obtain the sequence text.

[0020] In combination with the quality assessment method provided in the first aspect, in an optional implementation method, a computing device generates a sequence text including a target triple and a reference example, including: the computing device inputs the target triple and the reference example into a set text conversion model to obtain a sequence text, and the text conversion model is used to convert the input information into a sequence text.

[0021] In combination with the quality assessment method provided in the first aspect, in an optional implementation, a computing device generates a sequence text including a target triple and a reference example, including: the computing device arranges the contents of the target triple and the reference example in order to obtain the sequence text.

[0022] In this application, the computing device can choose one of a variety of methods for generating sequence text, so that when the quality assessment method of the knowledge graph provided in this application is applied on different computing devices, a sequence text generation method that is most suitable for the computing device or user needs can be adopted, avoiding the problem of reduced quantification effect caused by the limited resources of the computing device (such as computing resources, storage resources and network resources), and improving the quantification effect of the computing device on the knowledge graph.

[0023] In conjunction with the quality assessment method provided in the first aspect, in an optional implementation, a computing device inputs a first question into a large model to obtain a relationship confidence score for a target triple, including: the computing device inputs the first question into the large model and outputs a text to be processed. The computing device uses a regular expression to extract the location of the prediction result from the text to be processed and obtains a probability vector output by the large model at that location; thereby, the computing device determines a predicted probability for the target triple from the probability vector and normalizes the predicted probability to obtain the relationship confidence score for the target triple.

[0024] In conjunction with the quality assessment method provided in the first aspect, in an optional implementation, a computing device inputs a first question into a large model to obtain a relationship confidence score of a target triple, including: the computing device inputs the first question into the large model and outputs a text to be processed; the computing device obtains a probability vector output by the large model at the first position in the text to be processed; and the computing device determines a predicted probability of the target triple from the probability vector and normalizes the predicted probability to obtain the relationship confidence score of the target triple.

[0025] The above two ways for the computing device to obtain the relationship confidence of the target triple are only embodiments provided in this application. The computing device combines the obtained sequence texts to obtain the question input (the first question), so that the large model can evaluate the confidence of the target triple. Exemplarily, there are multiple schemes for the computing device to combine the sequence texts to obtain the question input, for example: question = "{reference triple sequence text 1}, is true knowledge; ...; {reference triple sequence text K+1}, is false knowledge; ...; {input triple sequence text}, is...", and finally the computing device inputs the question into the large model to obtain the corresponding output.

[0026] In combination with the quality assessment method provided in the first aspect, in an optional implementation, after the computing device obtains the quality assessment request, the quality assessment method provided in this application also includes: the computing device obtains the target tuple and the reference example of the target tuple in the knowledge graph according to the quality assessment request; wherein the target tuple includes: the first entity, the second entity, the relationship between the first entity and the second entity, the attribute of the first entity, the attribute of the second entity and the extension item, and the extension item includes one or more of time, knowledge features or knowledge attributes; the reference example of the target tuple includes: multiple reference tuples, and the relationship attributes of the relationship in the reference tuple match the relationship attributes of the relationship in the target tuple. And, the computing device processes the target tuple and the reference example of the target tuple based on the large model to obtain the relationship confidence of the target tuple. The target tuple provided in this application can refer to a quadruple, a quintuple or other, etc. Taking a quadruple as an example, in addition to the content contained in the triple, the quadruple can also include one of time, knowledge features or knowledge attributes.

[0027] In a second aspect, the present application provides a knowledge graph quality assessment device. The quality assessment device can be applied to a computing device or a chip (such as a processor) in a computing device, and the quality assessment device includes: a module or software unit for implementing the method of the first aspect or any optional implementation of the first aspect.

[0028] Exemplarily, the quality assessment device of the knowledge graph provided by the present application includes: a first acquisition module, a second acquisition module and a quality assessment module. The first acquisition module is used to obtain a quality assessment request, and the quality assessment request indicates the quantification of the quality of the knowledge graph. The second acquisition module is used to obtain the target triples and reference examples of the target triples in the knowledge graph according to the quality assessment request; wherein the target triples include: a head entity, a tail entity, the relationship between the head entity and the tail entity, and the attributes of the head entity and the attributes of the tail entity; the reference examples include: multiple reference triples, and the relationship attributes of the relationships in the reference triples match the relationship attributes of the relationships in the target triples. The quality assessment module is used to: process the target triples and reference examples based on the large model to obtain the relationship confidence of the target triples.

[0029] In a third aspect, the present application provides a chip. The chip includes an interface circuit and a control circuit. The interface circuit is configured to obtain a quality assessment request and, in conjunction with the control circuit, implement the method of the first aspect or any optional implementation of the first aspect.

[0030] In a fourth aspect, the present application provides a computing device comprising a memory and a processor, wherein the memory stores program code, and when the processor executes the program code, the processor is configured to implement the method of the first aspect or any optional implementation of the first aspect.

[0031] In a fifth aspect, the present application provides a computer-readable storage medium. The computer-readable storage medium includes computer instructions. When the computer instructions are executed in a computing device, the computing device is used to implement the method of the first aspect or any optional implementation of the first aspect.

[0032] In a sixth aspect, the present application provides a computer program product. When the computer program product is executed in a computing device, the computing device is configured to implement the method in the first aspect or any optional implementation of the first aspect.

[0033] In the seventh aspect, the present application provides a knowledge graph quantification system. The knowledge graph quantification system provided by the present application includes: a graph generation device and a computing device provided in the fourth aspect. The graph generation device and the computing device communicate through a wired connection or a wireless connection. The graph generation device is used to generate a knowledge graph including multiple triples, and the computing device is used to quantify the quality of the knowledge graph generated by the graph generation device. For example, the computing device can adopt the above first aspect or any optional implementation method in the first aspect to quantify the quality of the knowledge graph, which will not be elaborated here.

[0034] Regarding the beneficial effects of the technical solutions provided in aspects 2 to 7, reference may be made to the description of aspect 1 or any optional implementation of aspect 1, and no further description is given here. Based on the implementations provided in the above aspects, this application can also be further combined to provide more implementations. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 A schematic diagram of the architecture of a knowledge graph quantization system provided in this application;

[0036] Figure 2 A schematic diagram of the structure of a chip provided in this application;

[0037] Figure 3 Schematic diagram of the process of a knowledge graph quality assessment method provided for this application Figure 1 ;

[0038] Figure 4 Schematic diagram of the process of a knowledge graph quality assessment method provided for this application Figure 2 ;

[0039] Figure 5 Schematic diagram of the process of a knowledge graph quality assessment method provided for this application Figure 3 ;

[0040] Figure 6 Schematic diagram of the process of a knowledge graph quality assessment method provided for this application Figure 4 ;

[0041] Figure 7 Schematic diagram of the process of a knowledge graph quality assessment method provided for this application Figure 5 ;

[0042] Figure 8 Schematic diagram of the process of a knowledge graph quality assessment method provided for this application Figure 6 ;

[0043] Figure 9 A schematic diagram of the structure of a knowledge graph quality assessment device provided in this application. DETAILED DESCRIPTION

[0044] The present application provides a method for evaluating the quality of a knowledge graph. Based on the large model's multi-domain knowledge and powerful semantic understanding capabilities, a computing device uses the large model to process target triples and reference examples of target triples, directly estimating the confidence of the relationships of the knowledge described by the target triples, and achieving an objective quantitative evaluation of the quality of the knowledge graph. Moreover, since the input of the large model includes not only target triples but also reference examples of target triples, the relational attributes of the relationships in the reference triples included in the reference examples match the relational attributes of the relationships in the target triples, which is conducive to improving the prediction performance of the large model for the confidence of the relationships in the target triples based on the relationships in these reference triples.

[0045] This application can be applied not only to existing machine learning technology, artificial intelligence (AI) technology, and knowledge graph quantification scenarios, but also to future machine learning technology, AI technology, and knowledge graph quantification scenarios. The terms used in the implementation methods of this application are only used to explain the specific embodiments of this application and are not intended to limit this application. The following is a brief introduction to some concepts that may be involved in this application.

[0046] Knowledge Graph: The knowledge graph, powered by machine learning, uses natural language processing to build a comprehensive view of entities, relationships, and attributes through a semantic enrichment process. When ingesting data, this process enables the knowledge graph to identify individual objects (entities) and understand the relationships between different objects (entities). This working knowledge is then compared and integrated with other related and similar data sets. Once the knowledge graph is complete, question-answering and search systems are able to retrieve and reuse comprehensive answers to given queries. While consumer-oriented products have demonstrated their time-saving capabilities, the same system can also be applied in business environments, thereby avoiding manual data collection and integration work and supporting business decision-making.

[0047] Entity: The basic element in the knowledge graph, which is a node in the knowledge graph.

[0048] Relationship: Represents the relationship between entities and is an edge in the knowledge graph.

[0049] Attribute: refers to the characteristics or features of an entity, representing the attributes of the entity, and is the value or label of a node in the knowledge graph.

[0050] Relationship attributes: the type of relationship, such as causal relationship, temporal relationship, hierarchical relationship, interpersonal relationship, affiliation relationship or conditional relationship, etc.

[0051] Triple: A common representation of knowledge graphs, such as (head, relation, tail), where head represents the head entity, tail represents the tail entity, and relation represents the relationship between the head entity and the tail entity. It is the basis for building knowledge graphs.

[0052] Relation confidence: The probability that the actual relationship between entities is the relationship recorded in the triple.

[0053] Large models: A machine learning model with a large number of parameters and computing resources. The main feature of large models is the large number of parameters, which can improve the model's generalization ability and performance through large amounts of data. Large models are more complex, with deeper and more complex network structures, and can capture richer features and relationships, thereby improving the expressive power of large models. Large models can be divided into sparse large models and dense large models. Among them, sparse large models have a large number of sparse parameters and are commonly used for tasks such as search, recommendation, and advertising. Sparse large models are characterized by massive samples and large-scale sparse parameters, and are suitable for training using CPU / graphics processing unit (GPU) parameter server mode. Dense large models have most parameters with non-zero values ​​and no obvious sparsity characteristics. They are commonly used for computer vision (CV) and natural language processing (NLP) tasks. Dense large models are characterized by regular sample data and large-scale dense parameters, and are suitable for training using pure GPU collective communication mode.

[0054] Large Language Models (LLMs): Leveraging their powerful computing power and sophisticated algorithms, they can effectively process these massive amounts of data, providing users with efficient and accurate information processing and analysis services. Large language models not only excel in understanding and generating human language, but also demonstrate tremendous potential in solving complex problems and tasks. For example, large language models are widely used in automated question-answering systems, text summarization, machine translation, and language generation, significantly improving efficiency and accuracy. Especially when processing large datasets, large language models can uncover deep patterns and connections to support decision-making. Furthermore, the self-learning capabilities of large language models enable them to continuously evolve, continuously improving their performance and intelligence through continuous learning from new data.

[0055] Commonly used large language models include the Bidirectional Encoder Representation from Transformer (BERT) model, which can pre-train deep bidirectional representations using unlabeled text by jointly adjusting the left and right contexts in all layers.

[0056] It is worth noting that in some optional situations, the large language model can also be referred to as a large model. The large model provided in the embodiment of the present application can refer not only to a large language model, but also to a model whose model parameters reach a certain level. For example, according to the changes in the field in which the model is applied, the large model can also refer to a model with various functions such as image processing functions, human-computer interaction functions, semantic search and dialogue functions. This application does not limit the fields in which the large model can be applied and the specific name. In this article, for the sake of simplicity of description, it is named as a large model, but this should not be understood as a limitation of this application and will not be repeated later.

[0057] The following is an illustrative description of the scenarios in which the embodiments of the present application can be applied, with reference to the accompanying drawings.

[0058] Figure 1 This is a schematic diagram of the architecture of a knowledge graph quantization system provided in this application. The knowledge graph quantization system includes a computing device 110, an acceleration device 115, and a client device 120. The computing device 110 is a common computer device. The user can input the knowledge graph into the computing device 110 through the client device 120, and the computing device 110 quantifies the quality of the knowledge graph. The computing device 110 also outputs the quantization result of the knowledge graph (the relationship confidence of each triple) to the client device 120. The client device 120 is a terminal device, including but not limited to a personal computer, a server, a mobile phone, a tablet computer, or a smart car.

[0059] The computing devices provided in the embodiments of the present application may include, but are not limited to, electronic devices with computing capabilities or virtual devices with computing capabilities. For example, when the computing device is an electronic device with computing capabilities, the computing device may be, but is not limited to, a server, a personal computer, a tablet computer, or other electronic device. For another example, when the computing device is a virtual device with computing capabilities, the computing device may be a virtual machine, a microservice, or a cloud service, and may be provided to users through a cloud service subscription model, where users may select different subscription tiers based on their needs.

[0060] The computing device 110 includes an input / output (I / O) interface 114, a processor 111, and a memory 112. The I / O interface 114 is used to communicate with devices external to the computing device 110. For example, the client device 120 inputs data and sends quantization tasks to the computing device 110 through the I / O interface 114. After the computing device 110 processes the input data (such as quantifying the quality of the knowledge graph), it sends the processed output results to the client device 120 through the I / O interface 114.

[0061] The processor 111 is the computing core and control core of the computing device 110. It may include: a central processing unit (CPU), a specific integrated circuit, other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. In actual applications, the computing device 110 may also include multiple processors. The processor 111 may include one or more processor cores. An operating system and other software programs are installed in the processor 111, so that the processor 111 can access the memory 112 and various peripheral component interconnect express (PCIe) devices.

[0062] The processor 111 is connected to the memory 112 via a double data rate (DDR) bus or other type of bus. The memory 112 is the main memory of the computing device 110. Memory 112 is typically used to store various running software in the operating system, input data received from the client device 120, and output results to be sent to the client device 120 in the future. To improve the access speed of the processor 111, the memory 112 needs to have a fast access speed. In traditional computer devices, dynamic random access memory (DRAM) is generally used as the memory 112. In addition to DRAM, the memory 112 can also be other random access memories, such as static random access memory (SRAM). In addition, the memory 112 can also be read-only memory (ROM). For example, the read-only memory can be programmable read-only memory (PROM) or erasable programmable read-only memory (EPROM). This embodiment does not limit the number and type of memory 112.

[0063] Optionally, in order to store data persistently, the knowledge graph quantification system is further provided with a data storage system 113, which may be located outside the computing device 110 (e.g., Figure 1 ), exchanges data with the computing device 110 via a network. Optionally, the data storage system 113 may also be located inside the host, such as exchanging data with the processor 111 via a bus 116. In this case, the data storage system 113 is a hard disk.

[0064] The accelerator 115 is used to perform the quantization task of the knowledge graph. The processor 111 sends the received AI task and input data to the accelerator 115, and the accelerator 115 completes the AI ​​task according to the input data and sends the processing result to the processor 111. Figure 1 As shown, the acceleration device 115 can be directly inserted into a card slot on the motherboard of the computing device 110 and exchange data with the processor 111 through the bus 116. Figure 1 The bus 116 in the embodiment may also be replaced by a bus acceleration device 115 of a compute express link (CXL), a universal serial bus (USB) protocol or other protocols for data transmission.

[0065] In addition, the above-mentioned acceleration device 115 may not be directly inserted into the card slot on the motherboard of the computing device 110, but may be located in the acceleration device. For example, the acceleration device is a device independent of the computing device 110, such as an acceleration card. In this case, the computing device 110 can be connected to the acceleration device 115 through a wired network such as a network cable, or through a wireless hotspot or a wireless network such as Bluetooth. If the acceleration device 115 is used to process quantization tasks, such as quantizing the quality of the knowledge graph, the acceleration device can be implemented by one or more chips. For example, the chip includes any one of a CPU, a graphics processing unit (GPU), a neural-network processing unit (NPU), a tensor processing unit (TPU), an FPGA, and an ASIC. Among them, the GPU, also known as a display core, a visual processor, or a display chip, is a microprocessor specifically used for image computing on personal computers, workstations, game consoles, and some mobile devices (such as tablets, smartphones, etc.). The NPU simulates human neurons and synapses at the circuit level and uses a deep learning instruction set to directly process large numbers of neurons and synapses. A single instruction completes the processing of a group of neurons. ASICs are suitable for integrated circuit products with a single purpose.

[0066] For example, Figure 1 The processor 111 in the embodiment may be implemented by a chip, such as Figure 2 As shown, Figure 2 This is a schematic diagram of the structure of a chip provided in the present application. For example, the chip 200 includes a core 201 , a CPU 202 , a system buffer 203 and a DDR 206 .

[0067] Among them, CPU 202 is used to accept AI tasks (such as compression tasks, decompression tasks, neural network computing tasks, knowledge graph quantization tasks, etc.), and call core 201 to perform the task. In the case where the chip 200 has multiple cores 201, CPU 202 is also used to undertake scheduling tasks. For example, CPU 202 can be implemented by an ARM processor, which is small in size, low in power consumption, uses a 32-bit reduced instruction set, and has simple and flexible addressing. Of course, in some embodiments, CPU 202 can also be implemented by other processors.

[0068] Core 201 is used to provide the computing power required for the quantization task of the knowledge graph. In an optional scenario, core 201 includes a load / store unit (LSU), a cube computing unit, a scalar computing unit, a vector computing unit and a buffer. Among them, the LSU is used to load data to be processed and store processed data. It can also be used for reading and writing management of internal data in the core between different buffers, and to complete some format conversion operations. The cube computing unit is used to provide the core computing power of matrix multiplication. The scalar computing unit is a single instruction single data (SISD) processor. This type of processor only processes one piece of data (usually an integer or floating point number) at the same time. The vector computing unit, also known as an array processor, is a processor that can directly operate a set of arrays or vectors for calculations. The number of buffers may be one or more. For example, the buffer primarily refers to the level 1 cache (L1 buffer). The buffer is used to temporarily store data that core 201 repeatedly uses, thereby reducing bus read and write times. Furthermore, certain data format conversion functions require the source data to be located in the buffer. In this embodiment, since the buffer is located in the core, the distance between the core's cube computing units and the data storage area is shortened, reducing the cube computing units' access to DDR 206, thereby reducing data access latency and core data processing latency.

[0069] The system buffer 203 mainly refers to a level 2 buffer (L1 buffer or L2 cache), which is used to temporarily store input data, intermediate results or final results passing through the chip.

[0070] DDR 206 is an off-chip memory that can be replaced by high bandwidth memory (HBM) or other off-chip memory. DDR 206 is located between the chip and the external memory, overcoming the access speed limitation of the shared memory when computing resources read and write.

[0071] The input / output (I / O) device 205 included in chip 200 refers to the hardware that performs data transmission and can also be understood as the device that interfaces with the I / O interface. Common I / O devices include network cards, printers, keyboards, and mice. All external storage devices, such as hard drives, floppy disks, and optical disks, can also serve as I / O devices.

[0072] In some application scenarios, data encoding or decoding is required. Therefore, chip 200 may also include a codec 204 and an I / O device 205. Codec 204 is used to encode or decode data. It should be understood that in some optional scenarios, codec 204 may also be designed as a codec unit (software module) and integrated into core 201.

[0073] The core 201, the CPU 202, the system buffer 203, the encoder / decoder 204, the I / O device 205, and the DDR 206 are connected via a bus. The bus may include a path for transmitting information between the above components (such as the CPU 202 and the system buffer 203). In addition to the data bus, the bus may also include a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus may be a PCIe bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. For example, the core 201 can access these I / O devices 205 via the PCIe bus. The core 201 is connected to the system buffer 203 via the DDR bus. Here, different system buffers 203 may use different data buses to communicate with the core 201. Therefore, the DDR bus can also be replaced with other types of data buses. The embodiment of the present application does not limit the bus type.

[0074] For example, after CPU 202 loads the data to be processed by an AI task (such as a knowledge graph to be quantized) into DDR 206, the LSU in core 201 reads (loads) the data from DDR 206, quantifies the quality of the knowledge graph, and then obtains the target data. Once the processing results are obtained, the LSU loads (stores) the results into DDR 206. The network interface card then sends the inference results to the client device 120 or to the data storage system 113 for persistent storage.

[0075] It is worth noting that Figure 1 The acceleration device 115 shown can also be Figure 2 The chip 200 shown is implemented, and this application is not limited to this.

[0076] It should be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the computing device. In other embodiments, the computing device and chip may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0077] In an embodiment of the present application, a large language model (large model or LLM) may be deployed in a computing device or a processor (or chip) in the computing device, and other neural network models or algorithm models with quality quantification functions of a knowledge graph may also be deployed without limitation.

[0078] The following combination Figure 1 and Figure 2 The content shown provides a detailed description of the quality assessment method of the knowledge graph provided in this application.

[0079] Figure 3 Schematic diagram of the process of a knowledge graph quality assessment method provided for this application Figure 1 The quality assessment method of the knowledge graph can be executed by a computing device or a chip or processor in the computing device. The quality assessment method of the knowledge graph can be executed by a computing device, which can be Figure 1 The computing device 110, client device 120, or acceleration device 115 shown, or Figure 2 The chip 200 shown in FIG. 200 and the like. For hardware implementation of the computing device, please refer to the aforementioned Figure 1 and Figure 2 In some optional examples, the quality assessment method of the knowledge graph can also be performed by other computing devices. For the hardware implementation of the computing device, please refer to the aforementioned Figure 1 and Figure 2 The description is not repeated here.

[0080] Here, we take the computing device executing the quality assessment method of the knowledge graph provided in this embodiment as an example. Figure 3 As shown, the knowledge graph quality assessment method provided in this embodiment includes the following steps S310 to S330.

[0081] S310: The computing device obtains a quality assessment request.

[0082] The quality assessment request indicates the quantification of the quality of the knowledge graph.

[0083] In a possible example, the quality assessment request is generated during the execution of a task by the computing device.

[0084] In another possible example, the quality assessment request is received by the computing device from another device, which may be, for example, Figure 1 For example, the quality assessment request carries a knowledge graph and an identifier indicating the quality quantification of the knowledge graph; for another example, a computing device stores a knowledge graph, and the quality assessment request carries an identifier of the knowledge graph and an identifier indicating the quality quantification of the knowledge graph.

[0085] The above two possible examples are merely methods for obtaining quality assessment requests provided in this embodiment and should not be construed as limitations on this application.

[0086] S320. The computing device obtains the target triples and reference examples of the target triples in the knowledge graph according to the quality assessment request.

[0087] The target triplet includes a head entity (head), a tail entity (tail), the relationship between the head and tail entities, and the attributes of the head and tail entities. The head entity (head) refers to the starting entity in the triplet and is usually marked with an "h"; the tail entity (tail) refers to the ending entity in the triplet and is usually marked with a "t".

[0088] In the embodiments of the present application, the head entity and the tail entity refer to different objects; the relationship attributes of a triple refer to the type of relationship stored in the triple. For example, the relationship between the head entity and the tail entity may include, but is not limited to, a causal relationship, a temporal relationship, a hierarchical relationship, an interpersonal relationship, an affiliation relationship, or a conditional relationship. The attributes of the head entity include the value or label of the head entity, and the attributes of the tail entity include the value or label of the tail entity.

[0089] Exemplarily, the causal relationship includes: the head entity is the cause and the tail entity is the effect. The temporal relationship includes: the execution timing of the head entity comes first and the execution timing of the tail entity comes later. The hierarchical relationship includes: the scope covered by the head entity includes the scope covered by the tail entity. Interpersonal relationships include: the interpersonal relationship between the head entity and the tail entity, such as father and son, father and daughter, husband and wife, mother and son, mother and daughter, brothers, sisters, friends, lovers or others. The belonging relationship includes: the tail entity belongs to the head entity. The conditional relationship includes: when the head entity meets certain conditions, the operation corresponding to the tail entity is executed or the value of the tail entity is output, etc.

[0090] The relationships between the above entities are merely examples provided in this embodiment and should not be construed as limitations on this application. Entities may also have custom relationships or more diverse relationships, which will not be elaborated here.

[0091] The reference examples of the target triples acquired by the computing device in S320 include: a plurality of reference triples, and the relationship attributes of the relations in the reference triples match the relationship attributes of the relations in the target triples.

[0092] Optionally, the multiple reference triples include: positive example triples and negative example triples that are consistent with the relationship attributes of the relationship in the target triple, the content indicated by the positive example triples is true knowledge, and the content indicated by the negative example triples is false knowledge.

[0093] In the first possible example, true knowledge refers to the confidence of the entities, relationships and attributes in the triples determined by the computing device being greater than or equal to the confidence threshold, and false knowledge refers to the confidence of the entities, relationships and attributes in the triples determined by the computing device being less than the aforementioned confidence threshold. For example, the confidence threshold can be determined by the computing device based on the recognition accuracy of true knowledge and false knowledge of the large model, such as the confidence threshold being 80%. For another example, the confidence threshold can also be user-defined, such as the confidence threshold being 60%. For another example, the confidence threshold can also be a model parameter value preset in the large model, such as the confidence threshold being 30%. This application does not limit the specific value of the confidence threshold.

[0094] In the second possible example, true knowledge refers to the knowledge that the entities, relations, and attributes in the triples confirmed by the user conform to the laws of nature, and false knowledge refers to the knowledge that the relations between the entities in the triples confirmed by the user do not conform to the laws of nature.

[0095] The above two possible examples are merely optional methods provided in this embodiment and should not be understood as limitations of this application.

[0096] In an optional implementation, the reference examples of the target triples can be extracted by the computing device from a pre-built pool of triple examples. Exemplarily, before the computing device obtains the triples and reference examples of the triples in the knowledge graph, the knowledge graph quality assessment method provided in the embodiment of the present application further includes: the computing device constructing a pool of triple examples containing different relationship attributes.

[0097] As a feasible example, the process of a computing device constructing a triple example pool may include: the computing device obtains one or more historical knowledge graphs, and divides the triples in the one or more historical knowledge graphs into multiple triple example pools according to relationship attributes. Among them, the relationship attributes of the relationships in all triples included in a triple example pool are consistent. Among them, the multiple relationship attributes corresponding to the multiple triple example pools include two or more of the following: causal relationship, temporal relationship, hierarchical relationship, interpersonal relationship, affiliation relationship or conditional relationship, etc. The description of each relationship attribute can be referred to the aforementioned embodiment and will not be repeated here.

[0098] For example, the stage of building an example pool by a computing device may include two steps: dividing relation attributes and generating positive and negative example triplets.

[0099] During the step of assigning relationship attributes, the computing device refers to the concept of entity attributes and defines relationship attributes, such as "causality," "time sequence," "hierarchy," and "condition." Relationship connections with the same relationship attributes share similar characteristics. For example, relationships with the "causality" attribute include "cause," "influence," and "treatment."

[0100] In the step of generating positive and negative example triples, the computing device generates triples according to different relation attributes to form a relation pool.

[0101] There are multiple approaches to generating positive triples. For example, in Option 1, a computing device extracts all triples from an existing or historical knowledge graph. Experts then eliminate low-quality triples and classify the remaining triples according to the definition of relationship attributes to construct a pool of triples. Alternatively, in Option 2, a computing device leverages existing large-scale models to generate knowledge graphs, inputs specific relationships, and automatically generates a large number of positive triples with corresponding relationships—that is, real knowledge.

[0102] There are also many ways for computing devices to generate counterexample triples. For example, the computing device randomly replaces the inter-entity relationship of a triple with other relationships that are not attributes of the current relationship, and in this way constructs counterexample triples, that is, false knowledge.

[0103] In an optional scenario, if a triple example pool is pre-built, the specific process of the computing device obtaining the target triples and reference examples of the target triples in the knowledge graph is as follows: Figure 4 Provides a feasible implementation method. Figure 4 Schematic diagram of the process of a knowledge graph quality assessment method provided for this application Figure 2 The above-mentioned S320 may include the following S320a to S320c.

[0104] S320a. The computing device obtains the target triple from the knowledge graph.

[0105] For example, the target triple here can refer to any triple in the knowledge graph, or it can be a specified triple, which is not limited in this application.

[0106] S320b: The computing device determines a reference triple example pool from multiple triple example pools according to the relation attributes of the relations in the target triple.

[0107] The multiple triple example pools may include triple example pool 1 to triple example pool n, where n is a positive integer greater than or equal to 2. The relationship attributes of the relations in all triples in the reference triple example pool are consistent with the relationship attributes of the relations in the target triples.

[0108] Exemplarily, the relational attribute of the relationship in the target triple is “causal relationship”, and the relational attribute of the relationship of the triples included in the reference triple pool determined by the computing device is also “causal relationship”.

[0109] For example, the computing device determines the relational attributes of the current target triple based on the division of relational attributes in multiple triple example pools, and then finds a triple cluster with corresponding relational attributes from multiple pre-generated triple example pools (reference triple example pools).

[0110] S320c: The computing device selects M positive example triplets and N negative example triplets from the reference triple pool, and obtains a reference example of the target triplet based on the selected M positive example triplets and N negative example triplets.

[0111] M and N in S320c are both positive integers.

[0112] For example, the computing device randomly extracts M positive example triplets and N negative example triplets from the reference triple pool as reference context. Since the extracted reference triples (reference examples) have the same relational attributes as the input target triples, using the reference triples (reference examples) as input to the large model can help the large model understand the relationship between entities in the target triples and improve the large model's estimation performance of the relationship confidence.

[0113] It is worth noting that M and N can be different or the same, such as M=N=K, where K is a positive integer representing the number of positive example triplets or negative example triplets.

[0114] In an embodiment of the present application, since the computing device divides different triple example pools according to relationship attributes, and extracts positive example triples and negative example triples with corresponding relationship attributes as reference contexts (reference examples), the target triples and the reference examples have the same relationship attributes, which can assist the large model to better understand the relationship between entities in the target triples, thereby improving the large model's prediction performance on the relationship confidence of the target triples.

[0115] Please continue to refer to Figure 3 The quality assessment method of the knowledge graph provided in the embodiment of the present application also includes the following S330.

[0116] S330: The computing device processes the target triples and reference examples based on the large model to obtain the relationship confidence of the target triples.

[0117] The large model provided in this embodiment can be a large language model such as BERT or eDataMate, or it can refer to other large language models. These large language models have multi-domain knowledge and powerful semantic understanding capabilities, and this application does not limit this.

[0118] In the embodiment of the present application, since the large model has multi-domain knowledge and powerful semantic understanding capabilities, the computing device uses the large model to process the target triples and reference examples of the target triples, directly estimates the relationship confidence of the knowledge described by the target triples, and realizes the objective quantitative evaluation of the quality of the knowledge graph.

[0119] Moreover, since the input of the large model includes not only the target triples, but also the reference examples of the target triples, the relational attributes of the relations in the reference triples included in the reference examples match the relational attributes of the relations in the target triples, which is conducive to improving the prediction performance of the large model on the confidence of the relations in the target triples based on the relations in these reference triples.

[0120] Optionally, for the process of computing the relationship confidence of the target triplet, the following is combined with Figure 5 A possible example is provided. Figure 5 Schematic diagram of the process of a knowledge graph quality assessment method provided for this application Figure 3 The above-mentioned S330 may include the following S510 to S530.

[0121] S510: The computing device generates a sequence text including a target triple and a reference example.

[0122] Sequence text (sequence, seq) refers to text data formed by arranging various types of data in a sequence.

[0123] Optionally, the computing device may generate the above-mentioned sequence text using one or more of the following optional generation methods.

[0124] The computing device defines the target triple as (h, r, t), where h represents the head entity, r represents the relationship between entities, and t represents the tail entity. There are multiple ways to convert the target triple into sequence text. The following provides three optional methods for generating sequence text.

[0125] The first optional method of generating sequence text: the computing device inserts the target triple and the reference example into the reference question template to obtain the sequence text. In this implementation, the reference question template refers to a question template with a fixed format, which provides blank information for insertion. These blank information are arranged in a certain order as a sequence text that can be recognized by the computing device. After the computing device obtains the triple, it fills the position of the blank information with the information in the triple, thereby determining the sequence text of the triple. For example, the computing device inserts the entities and relationships in the target triple into the vacant positions in the template according to a pre-defined template related to the relationship (reference question template), seq = Template (h, r, t).

[0126] The second optional method of generating sequence text: the computing device inputs the target triple and the reference example into the set text conversion model to obtain the sequence text, and the text conversion model is used to convert the input information into sequence text. For example, the computing device uses the text conversion model to directly convert the triple into the corresponding sequence text, such as the T5 model. seq = T5 (h, r, t). The text conversion model can be trained by the computing device, or it can be installed in the computing device by other devices. For example, the text conversion model can be but not limited to a deep neural network or other algorithm model.

[0127] A third optional method for generating sequence text is to sequentially arrange the target triple and the contents of the reference example to generate the sequence text. For example, the computing device sequentially arranges the entities and relations in the target triple without adding any additional text, such as seq = "h[#SEP]r[#SEP]t", where [#SEP] is a delimiter, such as a space or comma.

[0128] The above three methods of generating sequence texts are only optional examples provided in this embodiment. It should not be understood that this application can only use one or more of the three generation methods to generate sequence text. The way in which computing devices generate sequence texts can continue to change with the development of technology, and this application does not limit this.

[0129] S520: The computing device generates a first question according to the sequence text.

[0130] Exemplarily, in the question generation step, the computing device combines the sequence texts obtained in S510 to obtain a question input (question, Qes), so that the large model can evaluate the confidence of the target triple. There are multiple schemes for the computing device to combine different triples to obtain the question input, for example: Qes = "{reference triple sequence text 1}, is true knowledge; ...; {reference triple sequence text K+1}, is false knowledge; ...; {input triple sequence text}, is..." Finally, the first question (Qes) obtained by the computing device is input into the large model to obtain the corresponding output.

[0131] In this embodiment, the first question includes a sequence of texts of different reference triples and a label indicating whether each reference triple is true knowledge or false knowledge. The tail of the first question includes the target triple to be identified and a label indicating whether the target triple is true knowledge or false knowledge. When the computing device inputs the first question into the large model, the large model predicts the relationship confidence of the target triple based on the reference triples and their labels in the first question, and the target triple and its label, as shown in S530 below.

[0132] S530: The computing device inputs the first question into the large model to obtain the relationship confidence of the target triple.

[0133] Optionally, the computing device determines the relationship confidence of the target triple from the output result of the large model. Two optional methods are provided below.

[0134] In a first optional approach, the computing device obtains the relationship confidence of the target triple by: inputting the first question into a large model, outputting a text to be processed; extracting the location of the predicted result from the text to be processed using a regular expression, and obtaining the probability vector output by the large model at that location. Finally, the computing device determines the predicted probability of the target triple from the probability vector output by the large model, and normalizes the predicted probability to obtain the relationship confidence of the target triple.

[0135] Exemplarily, the computing device uses a regular expression to extract the location of the prediction result from the output text of the large model, and then obtains the probability vector output by the large model at the corresponding position (representing the probability of each candidate token at that position), extracts the prediction probabilities of the two candidate tokens representing true knowledge and false knowledge and normalizes them, and finally outputs the normalized probability of the token representing true knowledge as the relationship confidence of the target triplet, P(target triplet) = Softmax(p(token = "true"), p(token = "false")).

[0136] In a second optional approach, the computing device obtains the relationship confidence of the target triple by: inputting a first question into a large model, outputting a text to be processed; and obtaining a probability vector output by the large model for the first position in the text to be processed. Finally, the computing device determines the predicted probability of the target triple from the probability vector output by the large model, and normalizes the predicted probability to obtain the relationship confidence of the target triple.

[0137] Exemplarily, the computing device obtains the probability vector at the first position output by the large model (representing the probability of each candidate token at this position), and similarly, extracts the predicted probabilities of the two candidate tokens representing true knowledge and false knowledge and normalizes them, and finally outputs the normalized probability of the token representing true knowledge as the relationship confidence of the target triplet, P(target triplet) = Softmax(p(token = "true"), p(token = "false")).

[0138] The two methods for the computing device to obtain the relationship confidence of the target triple are merely embodiments provided in this application. The computing device combines the obtained sequence texts to obtain the question input (the first question), so that the large model can evaluate the confidence of the target triple. Since the target triple input to the large model has the same relationship attributes as the reference triple, the reference triple can assist the large model in better understanding the relationship between the entities in the target triple, thereby improving the large model's prediction performance for the relationship confidence of the target triple.

[0139] In view of the quality assessment method of the knowledge graph provided in the above embodiment, the embodiment of the present application also provides a feasible specific example, such as Figure 6 As shown, Figure 6 Schematic diagram of the process of a knowledge graph quality assessment method provided for this application Figure 4 The quality assessment method provided in this embodiment includes four stages in sequence: generating a triple pool, selecting a corresponding relationship context, constructing a question input, and outputting a confidence score.

[0140] 1. Generate a triplet example pool.

[0141] The computing device constructs different triple example pools based on the relationship attribute division, which are used to select the reference context of the target triple from them. The specific construction method of the triple example pool can be referred to the description of S320 above and will not be repeated here.

[0142] 2. Select the corresponding relationship context.

[0143] The computing device randomly selects positive and negative triples with the same or similar relational attributes from the triple pool generated in the previous stage as reference contexts (i.e., reference examples of the aforementioned target triples) based on the relational attributes of the relations in the target triples included in the current knowledge graph. Figure 4 The description is not repeated here.

[0144] 3. Construct the problem input.

[0145] The computing device converts the selected reference context and input triples (target triples) into sequence texts, and combines multiple sequence texts to construct the question input (i.e., the first question described above). The process of the computing device generating the first question can be referred to the description of S510 and S520 above and will not be repeated here.

[0146] 4. Confidence output.

[0147] The computing device extracts the confidence probability from the output result of the large model as output. The process of the computing device obtaining the relationship confidence of the target triple can refer to the description of S530 above and will not be repeated here.

[0148] In order to continuously improve the prediction performance of the large model for the relationship confidence of different triples in the knowledge graph, the embodiment of the present application provides an optional implementation method, such as Figure 7 As shown, Figure 7 Schematic diagram of the process of a knowledge graph quality assessment method provided for this application Figure 5 The quality assessment method provided in this embodiment includes Figure 6 In addition to the four stages shown, a fifth stage is included: adding triples to the triple pool or updating the triple pool.

[0149] Exemplarily, if the relationship confidence of the target triple is greater than or equal to a first threshold, or if the relationship confidence of the target triple is less than or equal to a second threshold, the computing device updates the reference triple pool based on the target triple. The second threshold is less than or equal to the first threshold. For example, the first threshold is α and the second threshold is β.

[0150] Exemplarily, when the relationship confidence of the target triple is greater than or equal to the first threshold, the computing device uses the target triple as a positive example triple to update the reference triple example pool.

[0151] As another example, when the relationship confidence of the target triple is less than or equal to the second threshold, the computing device uses the target triple as a counterexample triple to update the reference triple example pool.

[0152] It is worth noting that if the first threshold is the same as the second threshold, when the relationship confidence of the target triple is greater than or equal to the first threshold (or the second threshold), the computing device uses the target triple as a positive example triple to update the reference triple pool; when the relationship confidence of the target triple is less than the first threshold (or the second threshold), the computing device uses the target triple as a negative example triple to update the reference triple pool.

[0153] Alternatively, if the first threshold is the same as the second threshold, then when the relationship confidence of the target triple is greater than the first threshold (or the second threshold), the computing device uses the target triple as a positive example triple to update the reference triple pool; when the relationship confidence of the target triple is less than or equal to the first threshold (or the second threshold), the computing device uses the target triple as a negative example triple to update the reference triple pool.

[0154] In this application, the computing device updates the reference triple example pool according to the confidence interval to which the relationship confidence of the target triple belongs, which is conducive to improving the reliability of the reference triple example pool, so that when the large model determines the reference examples of other triples based on the reference triple example pool, the prediction performance of the large model for the relationship confidence of other triples is improved.

[0155] The above embodiments are all described with the triples in the knowledge graph as an example, but the method provided by the knowledge graph provided in this application can also be applied to other tuples in the knowledge graph, such as the other tuples can refer to quadruplets, quintuples or others. Figure 8 , and illustrate how to apply the quality assessment method to other tuples in the knowledge graph. Figure 8 Schematic diagram of the process of a knowledge graph quality assessment method provided for this application Figure 6 The quality assessment method provided in this embodiment includes the following S810 to S830.

[0156] S810: The computing device obtains a quality assessment request.

[0157] For the detailed description of S810, please refer to the above-mentioned content of S310, which will not be repeated here.

[0158] S820. The computing device obtains the target tuple and the reference example of the target tuple in the knowledge graph according to the quality assessment request.

[0159] The target tuple includes: a first entity, a second entity, a relationship between the first entity and the second entity, an attribute of the first entity, an attribute of the second entity, and an extension item, wherein the extension item includes one or more of time, knowledge features, or knowledge attributes. The reference example of the target tuple includes: multiple reference tuples, and the relationship attributes of the relationship in the reference tuple match the relationship attributes of the relationship in the target tuple.

[0160] Taking the target tuple as a quadruple as an example, in addition to the content included in the aforementioned target triple, the quadruple also includes time, knowledge characteristics, knowledge attributes or other characteristics that describe knowledge, and the reference tuple of the quadruple (reference quadruple) also includes extended items corresponding to the quadruple.

[0161] For the detailed description of S820, please refer to the above-mentioned content of S320, which will not be repeated here.

[0162] S830: The computing device processes the target tuple and a reference example of the target tuple based on the large model to obtain the relationship confidence of the target tuple.

[0163] For the relationships between entities in different tuples, combined with the multi-domain knowledge and powerful semantic understanding capabilities of the large model, the computing device uses the large model to predict the relationship confidence of unilateral relationships, and can objectively quantify the quality of knowledge described by different tuples in the knowledge graph.

[0164] In combination with the content provided in the above embodiments, the quality assessment method of the knowledge graph provided in the above embodiments of the present application can be applied to the construction and evaluation process of the knowledge graph, thereby automatically generating an objective quantitative assessment result of the knowledge quality described by triples or other tuples in the knowledge graph using large model technology, that is, the confidence level of the relationship between entities in the triples (or other tuples). In specific implementation, the quality assessment method of the knowledge graph provided in this application can be provided to users or devices as a module in the knowledge graph tool chain.

[0165] It is understood that in order to implement the functions in the above embodiments, the computing device includes hardware structures and / or software modules corresponding to the execution of each function. It should be readily apparent to those skilled in the art that, in combination with the units and method steps of each example described in the embodiments disclosed in this application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in hardware or in a manner driven by computer software depends on the specific application scenario and design constraints of the technical solution.

[0166] Combined with the above Figures 3 to 8 , describes in detail the quality assessment method of the knowledge graph provided by the embodiment of this application, and will be combined with Figure 9 , describing the quality assessment device of the knowledge graph provided according to the embodiment of the present application.

[0167] Figure 9 This is a schematic diagram of the structure of a knowledge graph quality assessment device provided by this application. These knowledge graph quality assessment devices can be used to implement the functions of the computing device in the above method embodiment, and thus can also achieve the beneficial effects of the above method embodiment. In the embodiment of this application, the knowledge graph quality assessment device can be as follows Figure 1 Any of the devices shown, such as the computing device 110, the client device 120, and the acceleration device 115, or the computing devices shown in subsequent embodiments, may also be a module (such as a chip) applied to the device.

[0168] like Figure 9As shown, the quality assessment device 900 of the knowledge graph includes: a first acquisition module 910, a second acquisition module 920 and a quality assessment module 930. The first acquisition module 910 is used to obtain a quality assessment request, and the quality assessment request indicates that the quality of the knowledge graph is quantified. The second acquisition module 920 is used to obtain the target triples and reference examples of the target triples in the knowledge graph according to the quality assessment request; wherein the target triples include: a head entity, a tail entity, the relationship between the head entity and the tail entity, and the attributes of the head entity and the attributes of the tail entity; the reference examples include: multiple reference triples, and the relationship attributes of the relationships in the reference triples match the relationship attributes of the relationships in the target triples. The quality assessment module 930 is used to: process the target triples and reference examples based on the large model to obtain the relationship confidence of the target triples.

[0169] The knowledge graph quality assessment device 900 of the embodiment of the present application can be implemented by a software module. The knowledge graph quality assessment device 900 according to the embodiment of the present application can correspond to executing the method described in the embodiment of the present application, and the above-mentioned and other operations and / or functions of each module in the knowledge graph quality assessment device 900 are respectively for implementing the method flow in the aforementioned figures, and for the sake of brevity, they are not repeated here.

[0170] It is worth noting that if the quality assessment device 900 of the knowledge graph is implemented through a software module, for example, the software module can be provided to users through a cloud service subscription model, and users can choose different subscription levels according to their needs; for example, the software module can also provide enterprise-level customized services with professional domain customization, interface personalization and extended functions according to the needs of users or enterprises.

[0171] In addition, the quality assessment device for quantifying the quality of the knowledge graph provided by this application can also be made into a value-added service and provided to users, which is not limited by this application. When the quality assessment device 900 of the knowledge graph is implemented through a software module, the quality assessment device 900 of the knowledge graph can also be embedded in the eDataMate TM Or other large language model (large model) tool chain systems.

[0172] The knowledge graph quality assessment device 900 of the embodiment of the present application can also be implemented by hardware. For example, the hardware refers to a computing device. For the specific implementation of the computing device, please refer to Figure 1 The description is not repeated here.

[0173] In addition, when the knowledge graph quality assessment device 900 is implemented by a knowledge graph quantization system, the knowledge graph quantization system may include Figure 1A computing device and a graph generation device are provided. The graph generation device and the computing device communicate via a wired or wireless connection. The graph generation device is used to generate a knowledge graph including multiple triples, and the computing device is used to quantify the quality of the knowledge graph generated by the graph generation device. For example, the computing device can be used to execute the quality assessment method provided in the aforementioned embodiment.

[0174] The method steps in this embodiment can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in RAM, flash memory, ROM, PROM, EPROM, EEPROM, registers, hard disk, mobile hard disk, CD-ROM or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an ASIC. In addition, the ASIC can be located in a computing device or an electronic device. Of course, the processor and storage medium can also exist as discrete components in a network device or a terminal device.

[0175] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the process or function described in the embodiments of the present application is performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device or other programmable device. The computer program or instruction can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instruction can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a tape; it can also be an optical medium, such as a digital video disc (DVD); it can also be a semiconductor medium, such as a solid state drive (SSD).

[0176] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A knowledge graph quality assessment method, characterized in that: The method comprises: Obtain a quality assessment request, where the quality assessment request indicates quantifying the quality of the knowledge graph; Obtaining, according to the quality assessment request, a target triple in the knowledge graph and a reference example of the target triple; The target triple includes: a head entity, a tail entity, a relationship between the head entity and the tail entity, and attributes of the head entity and the tail entity; the reference example includes: multiple reference triples, and the relationship attributes of the relationships in the reference triples match the relationship attributes of the relationships in the target triples; The target triplet and the reference example are processed based on the large model to obtain the relationship confidence of the target triplet.

2. The method according to claim 1, characterized in that The multiple reference triples include: positive example triples and negative example triples that are consistent with the relationship attributes of the relationship in the target triple, the content indicated by the positive example triples is true knowledge, and the content indicated by the negative example triples is false knowledge.

3. The method according to claim 1 or 2, characterized in that The obtaining of the target triple in the knowledge graph and a reference example of the target triple includes: Obtaining a target triple from the knowledge graph; Determining a reference triple pool from multiple triple pools according to the relational attributes of the relations in the target triples; wherein the relational attributes of the relations in all triples in the reference triple pool are consistent with the relational attributes of the relations in the target triples; M positive example triplets and N negative example triplets are selected from the reference triple pool, and reference examples of the target triplet are obtained based on the selected M positive example triplets and N negative example triplets; M and N are both positive integers.

4. The method according to any one of claims 1 to 3, characterized in that Before obtaining the triples in the knowledge graph and the reference examples of the triples, the method further includes: Obtain one or more historical knowledge graphs; Dividing the triples in the one or more historical knowledge graphs into a plurality of triple example pools according to relationship attributes; Among them, the relationship attributes of the relations in all triples included in a triple example pool are consistent, and the multiple triple example pools include the reference triple example pool.

5. The method according to claim 4, characterized in that The multiple relationship attributes corresponding to the multiple triple example pools include two or more of the following: causal relationship, temporal relationship, hierarchical relationship, interpersonal relationship, belonging relationship or conditional relationship.

6. The method according to any one of claims 3 to 5, characterized in that The method further comprises: If the relationship confidence of the target triplet is greater than or equal to a first threshold, or the relationship confidence of the target triplet is less than or equal to a second threshold, the reference triplet pool is updated according to the target triplet; the second threshold is less than or equal to the first threshold.

7. The method according to any one of claims 1 to 6, characterized in that The processing of the target triple and the reference example based on the large model to obtain the relationship confidence of the target triple includes: generating a sequence text including the target triple and the reference example; generating a first question according to the sequence text; The first question is input into the large model to obtain the relationship confidence of the target triple.

8. The method according to claim 7, characterized in that The generating of the sequence text including the target triple and the reference example includes: Inserting the target triple and the reference example into a reference question template to obtain the sequence text; Alternatively, the target triple and the reference example are input into a set text conversion model to obtain the sequence text, wherein the text conversion model is used to convert the input information into the sequence text; Alternatively, the target triples and the contents of the reference examples are arranged in order to obtain the sequence text.

9. The method according to any one of claims 1 to 8, characterized in that After obtaining the quality assessment request, the method further includes: Obtaining a target tuple and a reference example of the target tuple in the knowledge graph according to the quality assessment request; The target tuple includes: a first entity, a second entity, a relationship between the first entity and the second entity, attributes of the first entity, attributes of the second entity, and an extension item, wherein the extension item includes one or more of time, knowledge features, or knowledge attributes; the reference example of the target tuple includes: multiple reference tuples, and the relationship attributes of the relationships in the reference tuples match the relationship attributes of the relationships in the target tuple; The target tuple and a reference example of the target tuple are processed based on the large model to obtain a relation confidence of the target tuple.

10. A knowledge graph quality assessment device, characterized in that: The quality assessment device comprises: A first acquisition module is configured to acquire a quality assessment request, wherein the quality assessment request indicates quantifying the quality of the knowledge graph; A second acquisition module is configured to acquire a target triple and a reference example of the target triple in the knowledge graph according to the quality assessment request; The target triple includes: a head entity, a tail entity, a relationship between the head entity and the tail entity, and attributes of the head entity and the tail entity; the reference example includes: multiple reference triples, and the relationship attributes of the relationships in the reference triples match the relationship attributes of the relationships in the target triples; The quality assessment module is used to: process the target triple and the reference example based on the large model to obtain the relationship confidence of the target triple.

11. A chip, characterized in that: It comprises an interface circuit and a control circuit; the interface circuit is used to obtain a quality assessment request and cooperate with the control circuit to execute the method according to any one of claims 1 to 9.

12. A computing device, characterized in that The method comprises a memory and a processor, wherein the memory stores a program code, and when the processor executes the program code, the processor is configured to execute the method according to any one of claims 1 to 9.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium includes computer instructions; when the computer instructions are executed in a computing device, the computing device executes the method according to any one of claims 1 to 9.

14. A computer program product, characterized in that When the computer program product is run in a computing device, the computing device performs the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Cash flow prediction method and system based on machine learning

    CN116579482A

  • Knowledge graph quality evaluation method and device, storage medium and electronic equipment

    CN117033668A

Cited By

  • Risk conduction prediction method and system based on combined deduction of time sequence diagram and large model

    CN120851620A