Model knowledge editing method, device and equipment

By using the model ontology knowledge editing method, the prompt words based on the concept of domain branch are obtained and generated output results, the performance degradation caused by knowledge strengthening and error labeling of large language models in specific domains is solved, and better Q&A performance is achieved.

CN119990280APending Publication Date: 2025-05-13ZHEJIANG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510125434.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing large language models have difficulty achieving better Q&A performance degradation caused by knowledge enhancement and error labeling in specific fields.

Method used

The first output result of the model is obtained based on the first prompt word associated with the concept of the field branch, and a second prompt word is generated in combination with the first prompt word and model ontology knowledge to obtain the second output result of the model. When the two output results are inconsistent, the model is edited to use the model ontology knowledge to drive the self-training process and strengthen the model's knowledge in a specific field.

Benefits of technology

This method can alleviate performance degradation due to error labeling, improve the model's knowledge enhancement effect in specific fields, and thus help the model achieve better Q&A performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990280A_ABST
    Figure CN119990280A_ABST
Patent Text Reader

Abstract

One or more embodiments of the invention relate to a model knowledge editing method and apparatus, a computing device and a computer program product. The method comprises the steps that a first output result of a model is obtained based on a first cue word, and the first cue word is associated with a domain branch concept in an ontology knowledge base of the model; the method further includes obtaining a second output result of the model based on a second cue word, the second cue word being generated based on the first cue word and ontology knowledge associated with the domain branch concept. The method further includes determining that the first output result and the second output result are inconsistent. In addition, the method further comprises training a model at least based on the first cue word and the second output result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present specification generally relate to the field of computer technology, and more specifically to a model knowledge editing method, apparatus, computing device, and computer program product. Background Art

[0002] Existing domain-specific large language models (LLMs) can be divided into two categories: those trained from scratch using domain-specific corpora; and those that are continuously trained on general models. The first type of model is built from the most basic architecture and then trained entirely on domain-specific corpora. By learning and analyzing a large amount of text data in a specific domain, the model gradually grasps the language patterns, professional terms, knowledge structure and other characteristics of the domain, thereby building a language model suitable for the domain. The second type of model is based on an already trained general large language model, and then further trained or fine-tuned using a domain-specific corpus. The general model already has certain language understanding and generation capabilities. Through continuous training on domain-specific data, the model can combine these general capabilities with domain knowledge to adapt to tasks in specific fields.

[0003] Models trained from scratch using domain-specific corpora can deeply learn various details and characteristics of a specific domain because the training data comes entirely from that domain. They show strong professionalism and accuracy when handling tasks within the domain. Models continuously trained on general models can not only handle tasks within the domain after being trained with domain-specific data, but also retain adaptability to other domains to a certain extent, and have better overall performance due to the generalization ability of the general model itself.

[0004] The paradigm of self-training involves the model itself generating data and using this self-generated data for further training. Traditional self-training methods usually use a trained model to annotate the data, and then improve the model performance based on this newly annotated data. Due to its simplicity and efficiency, this training paradigm has gradually migrated to large language models. Given the high cost of manually annotating training data or using more powerful proprietary models (such as GPT-4), many works have begun to use the language model itself to synthesize training data. Summary of the invention

[0005] In view of this, one or more embodiments of the present specification provide a model knowledge editing method, apparatus, computing device and computer program product, which can obtain a first output result of the model based on a first prompt word associated with a domain branch concept, and then generate a second prompt word in combination with the first prompt word and the model ontology knowledge to obtain a second output result of the model, and then perform knowledge editing on the model when it is determined that the two output results are inconsistent, so that the model ontology knowledge can be used to drive the self-training process, strengthen the model's knowledge in a specific domain, and reduce the performance degradation caused by incorrect labels, so as to help the model achieve better question-answering performance.

[0006] In a first aspect of the present specification, a model knowledge editing method is provided. The method includes obtaining a first output result of the model based on a first prompt word, the first prompt word being associated with a domain branch concept in an ontology knowledge base of the model. The method also includes obtaining a second output result of the model based on a second prompt word, the second prompt word being generated based on the first prompt word and the ontology knowledge associated with the domain branch concept. The method also includes determining that the first output result and the second output result are inconsistent. In addition, the method also includes training the model based on at least the first prompt word and the second output result.

[0007] In a second aspect of the present specification, a model knowledge editing device is provided. The device includes a first output result acquisition unit, which is configured to acquire a first output result of the model based on a first prompt word, and the first prompt word is associated with a domain branch concept in the ontology knowledge base of the model. The device also includes a second output result acquisition unit, which is configured to acquire a second output result of the model based on a second prompt word, and the second prompt word is generated based on the first prompt word and the ontology knowledge associated with the domain branch concept. The device also includes an output result comparison unit, which is configured to determine that the first output result and the second output result are inconsistent. In addition, the device also includes a model training unit, which is configured to train the model based on at least the first prompt word and the second output result.

[0008] In a third aspect of the present specification, a computing device is provided, comprising: at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the computing device to execute the method as described in the first aspect of the present specification.

[0009] In a fourth aspect of the present specification, a computer program product is provided, comprising machine executable instructions, which, when executed by a device, cause the device to perform the method according to the first aspect of the present specification.

[0010] It should be understood that the contents described in the Summary of the Invention are not intended to limit the key or important features of the embodiments of this specification, nor are they intended to limit the scope of this disclosure. Other features of this specification will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other purposes, features and advantages of the various embodiments of the present specification will become more easily understood through the following detailed description with reference to the accompanying drawings. In the accompanying drawings, various embodiments of the present specification will be described in an exemplary and non-limiting manner, wherein:

[0012] Figure 1 A schematic diagram showing the structure of a computing device according to some embodiments of the present specification is shown;

[0013] Figure 2 A flowchart of a model knowledge editing method according to some embodiments of the present specification is shown;

[0014] Figure 3 A schematic diagram showing a knowledge editing process according to some embodiments of the present specification;

[0015] Figure 4 A schematic diagram showing the structure of ontology knowledge according to some embodiments of this specification;

[0016] Figure 5 A block diagram showing a model knowledge editing device according to some embodiments of the present specification; and

[0017] Figure 6 A block diagram of an electronic device according to some embodiments of the present specification is shown.

[0018] Throughout the drawings, the same or similar reference numbers denote the same or similar elements. DETAILED DESCRIPTION

[0019] The embodiments of the present specification will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present specification are shown in the accompanying drawings, it should be understood that the present specification can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present specification. It should be understood that the drawings and embodiments of the present specification are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0020] In the description of the embodiments of this specification, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0021] As mentioned above, models that continuously train on general models have gradually become mainstream due to their ability to utilize rich and diverse data from seed models and more efficient training processes. However, such models usually rely on a large amount of training data to adapt to their respective fields, and these raw corpora are only injected into the seed models in a fragmented manner without being systematically organized. Compared with scattered large-scale corpora, concept-level structured knowledge in ontologies can play an important role in knowledge management and semantic search, and also has the potential to enhance LLMs, while existing training methods rarely use ontology knowledge as the basic knowledge source for training corpora.

[0022] In addition, previous self-training methods usually rely on gold label to screen low-quality instruction data and focus more on improvements within a single dataset. This method not only requires increasing data acquisition costs, but also may lack reliability and comprehensiveness for some complex or cross-domain instruction data.

[0023] To this end, a model knowledge editing method is provided in an embodiment of this specification. First, a first output result of the model is obtained based on a first prompt word associated with a domain branch concept, and then a second prompt word is generated in combination with the first prompt word and the model ontology knowledge to obtain a second output result of the model, and then the model is knowledge edited when it is determined that the two output results are inconsistent. In this way, the model ontology knowledge can be used to drive the self-training process, strengthen the model's knowledge in a specific field, and reduce the performance degradation caused by incorrect labels, so as to help the model achieve better question-answering performance.

[0024] Figure 1 FIG. 1 shows a schematic diagram of a computing device 100 according to some embodiments of the present specification. Figure 1 As shown, the computing device 100 may include a seed model 102. The seed model 102 is a model that serves as a starting point and basis in the field of machine learning and deep learning, especially in the training process of a large language model. The seed model 102 is a neural network model with a specific architecture, such as a model based on a Transformer architecture. It defines the basic structure of the model, including the connection mode of neurons, the number of layers, the dimension, etc., and provides a framework for subsequent training and learning.

[0025] The seed model 102 can be pre-trained on a large-scale general corpus, and learn the general laws of language, such as the co-occurrence relationship of words, the structural pattern of sentences, the representation of semantics, etc., through self-supervised learning and other methods, to form a preliminary understanding and knowledge reserve of the language. After pre-training on large-scale data, the seed model 102 can store rich language knowledge and world knowledge, such as grammatical rules, lexical semantics, common facts, etc., providing a strong knowledge foundation for the subsequent learning of specific tasks, and can help the model understand and process related language tasks faster.

[0026] like Figure 1 As shown, the seed model 102 may include an ontology knowledge base 104. The ontology knowledge base 104 is a structured database for storing and managing knowledge related to the seed model 102. As a knowledge representation form based on ontology, the ontology knowledge base 104 formally describes and organizes the knowledge related to the seed model 102, such as concepts, relationships, and attributes, in a machine-understandable manner, and generally uses semantic web technology to implement knowledge storage and query. For example, the ontology knowledge base 104 may include conceptual-level structured knowledge related to the medical field, including but not limited to the definitions, hypernyms, and synonyms of various medical branch concepts.

[0027] The ontology knowledge base 104 can be constructed in the following way: first, key knowledge can be extracted from the relevant literature, papers, technical reports, open source code and other resources of the seed model 102, or implicit knowledge can be acquired through communication with model developers and domain experts, and then the acquired knowledge can be formally represented using ontology language (such as OWL or RDF) to convert the knowledge into a form that can be understood and processed by computers. Then, the represented knowledge can be stored in a semantic database, such as Neo4j, Jena, etc., for efficient query, retrieval and reasoning operations.

[0028] like Figure 1 As shown, the concept-level structured knowledge in the ontology knowledge base 104 can be used to edit the seed model 102, and the model 106 aligned with the ontology knowledge can be obtained to strengthen the model's knowledge in a specific field, thereby improving the model's question-answering performance. Compared with the fragmented original corpus, the concept-level structured knowledge in the ontology knowledge base 104 can play a better role in knowledge management and semantic search.

[0029] It should be understood that the architecture and functions in the computing device 100 are described for exemplary purposes only, and do not imply any limitation on the scope of the present specification. The embodiments of the present specification may also be applied to other training systems with different structures and / or functions.

[0030] The following will combine Figures 2 to 6The process according to the embodiment of this specification is described in detail. For ease of understanding, the specific data mentioned in the following description are exemplary and are not intended to limit the scope of protection of the present disclosure. It should be understood that the embodiments described below may also include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this respect.

[0031] Figure 2 2 shows a schematic flow chart of a model knowledge editing method 200 according to some embodiments of the present specification. The method 200 may be performed by, for example Figure 1 It should be understood that the method 200 may also include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this respect.

[0032] like Figure 2 As shown, in box 210, method 200 can obtain a first output result of the model based on a first prompt word, and the first prompt word is associated with a domain branch concept in the ontology knowledge base of the model. In some implementations, the model can be a generative language model or a seed model of a generative language model. The ontology knowledge base of the model can include concept-level structured knowledge associated with one or more domains, and the domain branch concept can be a concept branch of a domain involved in the ontology knowledge base, such as a concept of neurological disorder in the medical field.

[0033] In some embodiments, the first prompt word can be generated based on at least one of the following prompt word generation templates: a first prompt word generation template, used to guide the model to generate an information set (which can be called a knowledge card) associated with a domain branch concept, the information set at least including the definition of the domain branch concept, related concepts and usage examples; a second prompt word generation template, used to guide the model to generate the definition of the domain branch concept and the difference between the domain branch concept and the related concepts; and a third prompt word generation template, used to guide the model to generate the research status of the domain branch concept. Optionally, three first prompt words can be generated for each domain branch concept for model training. Taking the medical field as an example, specific examples of the above three prompt word generation templates are as follows:

[0034]

[0035]

[0036] In block 220, method 200 may obtain a second output result of the model based on a second prompt word, where the second prompt word is generated based on the first prompt word and ontology knowledge associated with the domain branch concept. In some embodiments, the ontology knowledge associated with the domain branch concept includes at least one of the following: a definition of the domain branch concept; a hypernym of the domain branch concept; and a synonym of the domain branch concept.

[0037] In some embodiments, the computing device 100 can obtain the definition, hypernym and synonym of the domain branch concept by searching the ontology knowledge base of the model. However, the ontology knowledge base may lack the definition of some concepts. In some embodiments, in response to the definition of a specific concept not being included in the ontology knowledge base, the computing device can generate the definition based on few-sample learning. The process based on few-sample learning can be performed by inputting the following example prompt words into the generative language model:

[0038]

[0039] Thus, an example of combining the first prompt word generated based on the first prompt word generation template and the second prompt word of ontology knowledge is as follows:

[0040]

[0041] At block 230, method 200 may determine that the first output result and the second output result are inconsistent. In some embodiments, computing device 100 may first determine a similarity score between the first output result and the second output result. The similarity score may be a hybrid score based on three metrics. For example, computing device 100 may determine the cosine similarity between the first output result and the second output result, the similarity based on the longest common subsequence, and the similarity based on n-gram matching, respectively, and then determine the similarity score based on the above three similarities (e.g., determined by a simple summation).

[0042] If the similarity score is lower than the threshold, the computing device 100 may determine that the first output result and the second output result are inconsistent. For multiple branch concepts in a specific field, in some embodiments, the computing device may score the comparison results obtained for each branch concept, and then select the k groups of comparison results with the lowest similarity scores as inconsistent results.

[0043] In block 240, method 200 may train a model based on at least the first prompt word and the second output result. In some embodiments, computing device 100 may train a model based on supervised fine-tuning (SFT) using the first prompt word and the second output result as training corpus, or may train a model based on direct preference optimization (DPO) using the first prompt word, the first output result, and the second output result as training corpus.

[0044] SFT uses supervised learning to fine-tune the model based on the pre-trained language model. It is based on a large amount of manually annotated data, which contains the input text and the corresponding expected output text. By minimizing the difference between the model's predicted output and the annotated output, the model's parameters are adjusted so that the model can better adapt to specific tasks or fields. DPO directly optimizes preferences. It is based on human preference information for text, such as human ranking of the quality of different generated texts. The goal of DPO is to make the text generated by the model more in line with human preferences. By optimizing an objective function related to preferences, the model learns how to generate text that is more popular with humans.

[0045] Optionally, the model may be trained in combination with SFT and DPO (eg, training in stages). For different fields, the trained model may be selected based on the final training effect.

[0046] In this way, the model ontology knowledge can be used to drive the self-training process, strengthen the model's knowledge in specific fields, and reduce the performance degradation caused by incorrect labels, so as to help the model achieve better question-answering performance.

[0047] Figure 3 FIG. 1 is a schematic diagram showing a knowledge editing process according to some embodiments of the present specification. Figure 3 As shown, the first prompt word 302 can be generated based on the template and input into the generative language model 304 to be trained to obtain the first output result 306. Then, the second prompt word 308 can be generated based on the first prompt word 302 and the relevant ontology knowledge, and input into the generative language model 304 to obtain the second output result 310.

[0048] Figure 4 FIG. 2 shows a schematic diagram of the structure of ontology knowledge according to some embodiments of the present specification. Figure 4 As shown, ontology knowledge 402 may include ontology concept definitions 404, concept hypernyms 406, and concept synonyms 408. Ontology knowledge 402 may be obtained by searching ontology knowledge base 410. For some concepts 412 lacking definitions in ontology knowledge base 410, corresponding ontology concept definitions 404 may be generated by generative language model 414 based on few-shot learning.

[0049] return Figure 3, the computing device can judge the performance and question-answering performance of the generative language model 304 based on the similarity between the first output result 306 and the second output result 310. If the first output result 306 and the second output result 310 are consistent, it indicates that the performance of the generative language model 304 is good. If they are inconsistent, it indicates that the generative language model 304 needs to be edited to improve performance. Determining whether the output results are consistent can be achieved by obtaining the similarity scores of the two. Specifically, a mixed similarity score can be obtained based on the cosine similarity of the two, the similarity based on the longest common subsequence, and the similarity based on n-gram matching.

[0050] In step 312, the computing device can select k groups of comparison results with the lowest similarity scores to filter out inconsistent results, and use the inconsistent results as training corpus to train the generative language model 304. Specifically, the first prompt word and the second output result can be used as SFT corpus 314 to train to obtain an SFT-aligned model 316, or the first prompt word, the first output result, and the second output result can be used as DPO corpus 318 to train to obtain a DPO-aligned model 320. In addition, the SFT and DPO alignment methods can be combined to obtain an SFT+DPO aligned model 322. Then, a knowledge-edited model with better results can be selected for different fields.

[0051] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment and the computing device embodiment, since they are basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0052] Figure 5 FIG. 4 shows a schematic block diagram of a model knowledge editing device according to some embodiments of the present specification. Figure 5 As shown, the device 500 includes: a first output result acquisition unit 502, a second output result acquisition unit 504, an output result comparison unit 506 and a model training unit 508.

[0053] In some embodiments, the first output result acquisition unit 502 is configured to acquire the first output result of the model based on the first prompt word, and the first prompt word is associated with the domain branch concept in the ontology knowledge base of the model. The second output result acquisition unit 504 is configured to acquire the second output result of the model based on the second prompt word, and the second prompt word is generated based on the first prompt word and the ontology knowledge associated with the domain branch concept. The output result comparison unit 506 is configured to determine that the first output result and the second output result are inconsistent. The model training unit 508 is configured to train the model based on at least the first prompt word and the second output result.

[0054] It should be noted that the reference Figures 1 to 4 Further actions or steps shown can be performed by Figure 5 For example, the device 500 may include more modules or units to implement the actions or steps described above, or Figure 5 Some of the units or modules shown may be further configured to implement the actions or steps described above, which will not be repeated here.

[0055] Figure 6 Schematic block diagram of an example device 600 that can be used to implement some embodiments of the present specification is shown. Figure 6 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 602 or computer program instructions loaded from a storage unit 606 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0056] A number of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0057] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a CPU, a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as method 200. For example, in some embodiments, the method 200 may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the method 200 described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the method 200 in any other appropriate manner (e.g., by means of firmware).

[0058] In some embodiments, the methods and processes described above may be implemented as a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present specification.

[0059] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples of computer-readable storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not to be interpreted as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.

[0060] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0061] The computer program instructions for performing the operation of each embodiment of this specification can be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data or source code or object code written in any combination of one or more programming languages, and the programming languages ​​include object-oriented programming languages, and conventional procedural programming languages. Computer-readable program instructions can be executed completely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on the remote computer, or completely on the remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer by any type of network-including local area network (LAN) or wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect by the Internet). In certain embodiments, by using the state information of computer-readable program instructions to customize electronic circuits, such as programmable logic circuits, field programmable gate arrays (FPGAs) or programmable logic arrays (PLAs), the electronic circuits can execute computer-readable program instructions, thereby realizing the various aspects of this specification.

[0062] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0063] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0064] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0065] The embodiments of the present specification have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

[0066] The above are only optional embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A model knowledge editing method, comprising: Based on a first prompt word, obtaining a first output result of the model, wherein the first prompt word is associated with a domain branch concept in an ontology knowledge base of the model; Based on a second prompt word, obtaining a second output result of the model, wherein the second prompt word is generated based on the first prompt word and ontology knowledge associated with the domain branch concept; Determining that the first output result and the second output result are inconsistent; as well as The model is trained based on at least the first prompt word and the second output result.

2. The method according to claim 1, wherein the ontology knowledge associated with the domain branch concept comprises at least one of the following: Definition of the concepts of the said field branches; Hypernyms of the domain branch concepts; and A synonym for the domain branch concept.

3. The method according to claim 2, wherein the method further comprises: The definition, hypernym and synonym of the domain branch concept are obtained by searching the ontology knowledge base of the model.

4. The method according to claim 3, wherein the method further comprises: In response to the ontology knowledge base not including the definition, the definition is generated based on few-shot learning.

5. The method according to claim 1, wherein determining that the first output result and the second output result are inconsistent comprises: determining a similarity score between the first output result and the second output result; as well as In response to the similarity score being lower than a threshold, it is determined that the first output result and the second output result are inconsistent.

6. The method of claim 5, wherein determining a similarity score between the first output result and the second output result comprises: Determining a cosine similarity between the first output result and the second output result; Determining a longest common subsequence-based similarity between the first output result and the second output result; Determining a similarity between the first output result and the second output result based on n-gram matching; as well as The similarity score is determined based on the cosine similarity, the longest common subsequence-based similarity, and the n-gram matching-based similarity.

7. The method according to claim 1, wherein training the model based on at least the first prompt word and the second output result comprises: Using the first prompt word and the second output result as training corpus, training the model based on supervised fine-tuning; or The first prompt word, the first output result and the second output result are used as training corpus, and the model is trained based on direct preference optimization.

8. The method according to claim 1, wherein the first prompt word is generated based on at least one of the following prompt word generation templates: A first prompt word generation template is used to guide the model to generate an information set associated with the domain branch concept, the information set at least including the definition, related concepts and usage examples of the domain branch concept; A second prompt word generation template is used to guide the model to generate the definition of the domain branch concept and the difference between the domain branch concept and related concepts; as well as The third prompt word generation template is used to guide the model to generate the research status of the branch concept in the field.

9. The method of claim 1, wherein the model comprises a generative language model.

10. A model knowledge editing device, comprising: A first output result acquisition unit is configured to acquire a first output result of the model based on a first prompt word, wherein the first prompt word is associated with a domain branch concept in an ontology knowledge base of the model; a second output result obtaining unit configured to obtain a second output result of the model based on a second prompt word, wherein the second prompt word is generated based on the first prompt word and ontology knowledge associated with the domain branch concept; an output result comparison unit, configured to determine that the first output result and the second output result are inconsistent; as well as A model training unit is configured to train the model based on at least the first prompt word and the second output result.

11. A computing device comprising: at least one processing unit; At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the computing device to perform the method as claimed in any one of claims 1 to 9.

12. A computer program product comprising machine executable instructions which, when executed by a device, cause the device to perform the method of any one of claims 1 to 9.