Method and device for automatically labeling analogy corpora through computing power of intelligent computing center
Through the computing power of the intelligent computing center, the analogy corpus is automatically marked and multiple rounds of analysis are used to use the large language model to solve the problem of low corpus quality in the intelligent computing center, and the performance and capabilities of the large language model are improved.
Patent Information
- Application Number
- CN202510088165.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-20
AI Technical Summary
The existing intelligent computing centers are used to train large language models with low corpus quality, resulting in unstable performance and poor performance of large language models trained.
The calculation power of the intelligent computing center automatically labels the analogous corpus. The method includes determining multiple prompt words, using a large language model to perform entity similarity analysis, structural similarity alignment and causal logical similarity analysis, and finally labeling the analogous relationship type of the analogous corpus.
The corpus quality is improved, the cognitive ability and analogical reasoning ability of large language models are improved, and the stability and excellent performance of the model are ensured.
Smart Images

Figure CN119988628A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the fields of computing power infrastructure and artificial intelligence technology, and in particular, to a method and device for automatically annotating analogy corpus through the computing power of an intelligent computing center. Background Art
[0002] With the development of artificial intelligence technology and computing power technology, the concept of intelligent computing center has emerged. "Intelligent computing center" refers to the use of large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, mainly for artificial intelligence applications (such as artificial intelligence deep learning model (such as large language model) development, model training and model reasoning and other scenarios) to provide the required computing power, data and algorithms. Intelligent computing center covers facilities, hardware, software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.
[0003] When training a large language model, the existing intelligent computing center needs to input a large amount of corpus into the large language model. The large language model learns the semantic relationship of the corpus and is used to complete specific tasks. However, how to obtain high-quality corpus is an urgent problem to be solved. If the quality of the corpus used to train the large language model is low, it is impossible to guarantee that the trained large language model has stable and excellent performance. Summary of the invention
[0004] The embodiments of the present invention provide a method and device for automatically annotating analogy corpus through the computing power of an intelligent computing center, which is used to solve the problem that the corpus quality used by the existing intelligent computing center for training large language models is low, thereby failing to ensure that the trained large language model has stable and excellent performance.
[0005] In order to solve the above-mentioned technical problems, the present invention is achieved as follows:
[0006] In a first aspect, an embodiment of the present invention provides a method for automatically annotating analogy corpus by using the computing power of an intelligent computing center, comprising:
[0007] Step S1: determining a first prompt word, wherein the first prompt word includes: a first analogy corpus pair as a case, a task and result indicating similarity analysis of entities in the first analogy corpus pair, and an unlabeled second analogy corpus pair; inputting the first prompt word into a large language model to obtain an entity similarity analysis result of the second analogy corpus pair;
[0008] Step S2: determining a second prompt word, wherein the second prompt word includes: the first analogy corpus pair, a task and a result indicating structural similarity alignment of sentences in the first analogy corpus pair, and the second analogy corpus pair; inputting the second prompt word into the large language model to obtain a sentence pair of structural similarity alignment of the second analogy corpus pair;
[0009] Step S3: determining a third prompt word, wherein the third prompt word includes: a sentence pair aligned with structural similarity of the first analogy corpus pair, a task and result indicating a causal logic similarity analysis of the sentence pair aligned with structural similarity of the first analogy corpus pair, and a sentence pair aligned with structural similarity of the second analogy corpus pair; inputting the third prompt word into the large language model to obtain a causal logic similarity analysis result of the second analogy corpus pair;
[0010] Step S4: Determine a fourth prompt word, wherein the fourth prompt word includes: the first analogy corpus pair and its entity similarity analysis results and causal logic similarity analysis results, the analogy relationship analysis results of the first analogy corpus pair, and the second analogy corpus pair and its causal logic similarity analysis results; input the fourth prompt word into the large language model to obtain the analogy relationship analysis results of the second analogy corpus pair.
[0011] Optionally, the method further includes:
[0012] Step S5: marking the analogy relationship type of the second analogy corpus pair according to whether the entities indicated in the analogy relationship analysis result of the second analogy corpus pair are similar in form, whether the first-order relationships are similar, and whether the higher-order relationships are similar.
[0013] Optionally, the analogy relationship type includes at least one of the following:
[0014] Literal similarity type, in which the analogy corpus pairs of the literal similarity type have similar entities, first-order relations, and higher-order relations;
[0015] True analogy type, in the analogy corpus pair of the true analogy type, the entities are dissimilar, some first-order relations are similar, and the higher-order relations are similar;
[0016] A pseudo-analogy type, in which the entities in the analogy corpus pairs are dissimilar, the first-order relations are similar, and the higher-order relations are dissimilar;
[0017] A surface similarity type, in which the analogy corpus pairs have similar entities, similar first-order relations, and dissimilar higher-order relations;
[0018] Only appearance similarity type, in the analogy corpus pairs of only appearance similarity type, the entities are similar, the first-order relations are not similar, and the higher-order relations are not similar;
[0019] An abnormal type, in which the analogy corpus pairs of the abnormal type have dissimilar entities, first-order relations and higher-order relations.
[0020] Optionally, the entity includes at least one of the following: background, character, plot and common vocabulary, and the common vocabulary includes identical vocabulary and synonyms.
[0021] Optionally, the first analogy corpus pair is of a pseudo-analogy type, in which the entities are dissimilar, the first-order relations are similar, and the higher-order relations are dissimilar.
[0022] Optionally, the first prompt word further includes: first task prompt information, the first task prompt information is used to prompt the large language model to perform a task and result of similarity analysis on entities in the first analogy corpus pair according to the instruction, perform similarity analysis on entities in the second analogy corpus pair, and output entity similarity analysis results of the second analogy corpus pair;
[0023] And / or, the second prompt word further includes: second task prompt information, the second task prompt information is used to prompt the large language model to perform a task and a result of structural similarity alignment of the sentences in the first analogy corpus pair according to the instruction, perform structural similarity alignment on the second analogy corpus pair, and output a sentence pair of the second analogy corpus pair aligned with structural similarity;
[0024] And / or, the third prompt word also includes: third task prompt information, the third task prompt information is used to prompt the large language model to perform causal logic similarity analysis on the sentence pairs aligned with structural similarity of the first analogy corpus according to the task and result of the instruction, perform causal logic similarity analysis on the sentence pairs aligned with structural similarity of the second analogy corpus, and output the causal logic similarity analysis result of the second analogy corpus;
[0025] And / or, the fourth prompt word also includes: fourth task prompt information, the fourth task prompt information is used to prompt the large language model to infer the analogy relationship analysis result of the first analogy corpus pair according to the first analogy corpus pair and its entity similarity analysis result and the causal logic similarity analysis result, and to infer and output the analogy relationship analysis result of the second analogy corpus pair according to the second analogy corpus pair and its causal logic similarity analysis result.
[0026] In a second aspect, an embodiment of the present invention provides a device for automatically annotating analogy corpus through the computing power of an intelligent computing center, including:
[0027] A first processing module is used to determine a first prompt word, wherein the first prompt word includes: a first analogy corpus pair as a case, a task and a result indicating similarity analysis of entities in the first analogy corpus pair, and an unlabeled second analogy corpus pair; input the first prompt word into a large language model to obtain an entity similarity analysis result of the second analogy corpus pair;
[0028] The second processing module is used to determine a second prompt word, wherein the second prompt word includes: the first analogy corpus pair, a task and a result indicating structural similarity alignment of sentences in the first analogy corpus pair, and the second analogy corpus pair; input the second prompt word into the large language model to obtain a sentence pair of structural similarity alignment of the second analogy corpus pair;
[0029] a third processing module, configured to determine a third prompt word, wherein the third prompt word includes: a sentence pair aligned with structural similarity of the first analogy corpus pair, a task and result indicating a causal logic similarity analysis of the sentence pair aligned with structural similarity of the first analogy corpus pair, and a sentence pair aligned with structural similarity of the second analogy corpus pair; input the third prompt word into the large language model to obtain a causal logic similarity analysis result of the second analogy corpus pair;
[0030] The fourth processing module is used to determine a fourth prompt word, wherein the fourth prompt word includes: the first analogy corpus pair and its entity similarity analysis results and causal logic similarity analysis results, the analogy relationship analysis results of the first analogy corpus pair, and the second analogy corpus pair and its causal logic similarity analysis results; the fourth prompt word is input into the large language model to obtain the analogy relationship analysis results of the second analogy corpus pair.
[0031] In a third aspect, an embodiment of the present invention provides an electronic device comprising: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the method for automatically annotating analogy corpus through the computing power of an intelligent computing center as described in the first aspect above.
[0032] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method for automatically annotating analogy corpus through the computing power of an intelligent computing center as described in the first aspect above are implemented.
[0033] In a fifth aspect, an embodiment of the present invention provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the steps of the method for automatically annotating analogy corpus through the computing power of an intelligent computing center as described in the first aspect above.
[0034] In an embodiment of the present invention, a prompt word is determined, the prompt word includes an analogy corpus pair and its analogy relationship analysis process as a case, and a large language model running in an intelligent computing center is called to analyze the analogy relationship of the unlabeled analogy corpus pair according to the analogy relationship analysis process of the analogy corpus pair as a case in the prompt word, and the analogy relationship analysis result of the unlabeled analogy corpus pair is obtained. The analogy relationship type of the analogy corpus pair can be labeled according to the analogy relationship analysis result, and high-quality analogy corpus pairs can be screened out, thereby improving the cognitive ability and analogy reasoning ability of the large language model trained with these high-quality analogy corpus pairs. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0036] Figure 1 One of the flowcharts of the method for automatically annotating analogy corpus by using the computing power of an intelligent computing center according to an embodiment of the present invention;
[0037] Figure 2 This is a second flow chart of a method for automatically annotating analogy corpus by using the computing power of an intelligent computing center according to an embodiment of the present invention;
[0038] Figure 3 is a schematic diagram of a first prompt word according to an embodiment of the present invention;
[0039] Figure 4 is a schematic diagram of a second prompt word according to an embodiment of the present invention;
[0040] Figure 5 is a schematic diagram of a third prompt word according to an embodiment of the present invention;
[0041] Figure 6 is a schematic diagram of a fourth prompt word according to an embodiment of the present invention;
[0042] Figure 7 It is a structural schematic diagram of an apparatus for automatically annotating analogy corpus through the computing power of an intelligent computing center according to an embodiment of the present invention;
[0043] Figure 8 Schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] First, the technical terms involved in the present invention are briefly explained below.
[0046] The "computing power" mentioned in the present invention is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to execute certain computing requirements. It is the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity. It mainly provides services to the society through computing power infrastructure.
[0047] The "computing power" (Computational Power, CP) described in the present invention is the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, super computing power and intelligent computing power. The commonly used unit of measurement is the number of floating point operations performed per second (FLOPS: Floating Point Operations Per Second, 1EFLOPS = 10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream notebooks. The calculation formula is: CP = CP 通用 +CP 智能 +CP 超级 .
[0048] The "Network Power" (NP) described in the present invention is a manifestation of the data transmission capability of computing facilities, including comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, etc. The carrying capacity involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities. In the embodiment of the present invention, the carrying capacity uses the memory bandwidth.
[0049] The "Storage Power" (SP) described in the present invention is the comprehensive ability of a data center in terms of data storage capacity, performance, safety and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and server built-in storage devices. The commonly used unit of measurement for storage capacity is exabyte (EB, 1EB = 2^60bytes), and the commonly used unit of measurement for performance is the number of reads and writes per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB). The disaster recovery ratio is an important manifestation of safety and reliability.
[0050] The "computing power infrastructure" described in the present invention is a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity. It can realize centralized computing, storage, transmission and application of information, and presents characteristics such as diversity and ubiquity, intelligence and agility, security and reliability, and green and low-carbon.
[0051] The “computing power” mentioned in the present invention includes general computing power, intelligent computing power and super computing power.
[0052] The “general computing power” mentioned in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0053] The "intelligent computing power" described in the present invention is a computing platform for large-scale deployment of special chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit) for various innovative applications of artificial intelligence, such as natural language processing and machine vision.
[0054] The "super computing power" mentioned in the present invention is mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, and gene analysis.
[0055] The "intelligent computing center" described in the present invention refers to a facility that provides the required computing power, data and algorithms for artificial intelligence applications (such as artificial intelligence deep learning model development, model training and model reasoning scenarios) by using large-scale heterogeneous computing power resources, including general computing power (CPU: Central Processing Unit) and intelligent computing power (GPU: Graphics Processing Unit, FPGA: Field Programmable Gate Array, ASIC: Application Specific Integrated Circuit, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.
[0056] The "computing resources" mentioned in the present invention refer to the technologies and facilities with information calculation, transmission, storage and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPU (Central Processing Unit), GPU (Graphics Processing Unit), network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, as well as supporting and guarantee resources such as wind, fire, water and electricity.
[0057] The “large language model” mentioned in the present invention refers to a large language model (LLM), which is a language model with a large parameter scale. It is designed to understand and generate human language. It is trained with a large amount of text data and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.
[0058] In the field of linguistics and natural language processing, the "corpus" mentioned in the present invention refers to text or voice data used for language research, analysis, teaching or technology development. In natural language processing (NLP), corpus is used to train machine learning models, such as language models, text classifiers, sentiment analyzers, etc. These models learn the statistical laws and patterns of language by analyzing a large amount of corpus, so that they can perform various language processing tasks, such as text generation, translation, summarization, question answering, etc.
[0059] The "prompt" mentioned in the present invention refers to a prompt or instruction designed based on natural language to guide the large language model to perform a specific task.
[0060] The "analogy" described in the present invention is the core of human cognition, which allows us to abstract information and understand new (unfamiliar) situations based on familiar situations. Text corpora with analogical relationships can be described as text corpora with dissimilar entities but similar logical structures. Among them, the logical structure can include first-order relations and higher-order relations. That is, text corpora with analogical relationships can also be described as text corpora with dissimilar entities, similar higher-order relations and some first-order relations.
[0061] The "analogy corpus pair" described in the present invention is composed of two text corpora, and the two text corpora have an analogy relationship; the text corpora with an analogy relationship can be described as text corpora with dissimilar entities but similar logical structures. The logical structure can include first-order relations and higher-order relations. That is, the text corpora with an analogy relationship can also be described as text corpora with dissimilar entities but similar higher-order relations and some first-order relations.
[0062] The "entity" mentioned in the present invention includes objects, characters, roles, etc. In this method, (story) background and plot are also included in the analysis scope of the entity.
[0063] The "first-order relationship" described in the present invention refers to a simple, direct connection between entities (such as objects, people, etc.). These relationships are usually intuitive and superficial associations, without involving deep causal or logical structures, and are mainly the relationship structures of the entities described in space, time, and interaction.
[0064] The "higher-order relationships" mentioned in the present invention refer to more abstract or complex associations that go beyond the direct connections between entities, usually involving causes, results, purposes, pattern recognition or deeper logical structures. In other words, higher-order relationships explore "why" such direct connections exist between entities, or how these connections work in a broader system or context.
[0065] The following example illustrates a text corpus with analogical relationships.
[0066] A pair of text corpora with an analogy relationship includes: "The planets in the solar system revolve around the sun" and "The electrons in an atom revolve around the nucleus". The analogy relationship includes: "planet" and "electron", "sun" and "nucleus".
[0067] Another pair of text corpora with an analogical relationship includes: "An employee received a seemingly harmless attachment that contained malware. The malware invaded his personal computer and stole his sensitive personal information." and "The citizens of Troy allowed access to the Trojan Horse containing Greek soldiers. The Greek soldiers captured Troy and stole their wealth." Among them, the analogical relationships include: "employee" and "citizens of Troy", "attachment" and "Trojan horse", "malware" and "Greek soldiers", "personal computer" and "Troy", and "sensitive personal information" and "wealth".
[0068] To solve the problem that the existing intelligent computing center uses low-quality corpus for training large language models, which cannot guarantee the stable and excellent performance of the trained large language models, please refer to Figure 1 , an embodiment of the present invention provides a method for automatically annotating analogy corpus by using the computing power of an intelligent computing center, comprising:
[0069] Step S1: determining a first prompt word, wherein the first prompt word includes: a first analogy corpus pair as a case, a task and result indicating similarity analysis of entities in the first analogy corpus pair, and an unlabeled second analogy corpus pair; inputting the first prompt word into a large language model to obtain an entity similarity analysis result of the second analogy corpus pair;
[0070] Please refer to Figure 2 , Step S1 can also be called the entity analysis step, that is, performing similarity analysis on the entities of the two analogy corpora in the analogy corpus pair.
[0071] The analogy corpus pair in the embodiment of the present invention is composed of two text corpora, that is, the first analogy corpus pair includes two text corpora, and the second analogy corpus pair also includes two text corpora.
[0072] Please refer to Figure 2 , the second analogy corpus pair in the embodiment of the present invention may be an analogy corpus pair in an unlabeled analogy dataset, and the unlabeled analogy dataset includes multiple unlabeled analogy corpus pairs. An unlabeled analogy corpus pair means that the analogy corpus pair has not been scored, or the analogy relationship type has not been labeled for the analogy corpus pair, or the analogy corpus pair has not been screened. There may be analogy corpus pairs with poor analogy ability in the unlabeled analogy dataset. Directly using the analogy corpus pairs in the unlabeled analogy dataset to train the large language model cannot guarantee that the trained large language model has stable and excellent performance.
[0073] In an embodiment of the present invention, the analogy corpus pairs in the unlabeled analogy data set can be generated in the following manner: obtaining an original text corpus, inputting the original text corpus into a large language model, and obtaining an analogy corpus pair generated based on the original text corpus, wherein the analogy corpus pair includes a first text corpus and a second text corpus having an analogy relationship, wherein the first text corpus is obtained by extracting information and expanding details of the original text corpus by the large language model, and the second text corpus is written by the large language model, and the second text corpus is similar in logical structure to the original text corpus but not in entity. The original text corpus in the original text corpus set can be text downloaded from the Internet (Web), or text downloaded from the Internet and processed after data cleaning.
[0074] Of course, in the embodiment of the present invention, it is not excluded to adopt other methods to obtain analogy corpus pairs in the unlabeled analogy dataset.
[0075] Please refer to Figure 3 , Figure 3 is a schematic diagram of the first prompt word of an embodiment of the present invention, from Figure 3 It can be seen that the first prompt word includes the first analogy corpus pair as a case, and the first analogy corpus pair includes two text corpora, such as Figure 3 The first prompt word also includes: indicating the task and result of similarity analysis on the entities in the first analogy corpus pair, such as Figure 3 The Question in indicates the task of performing similarity analysis on the entities in the first analogy corpus pair, such as Figure 3 The Answer in is the result of similarity analysis on the entities in the first analogy corpus pair, also referred to as the entity similarity analysis result of the first analogy corpus pair.
[0076] In the embodiment of the present application, the entity may include at least one of the following: background, role, plot and common vocabulary, wherein the common vocabulary includes identical vocabulary and synonyms. In the embodiment of the present invention, the similarity analysis of identical vocabulary is usually to analyze verbs or nouns in the analogy corpus pair.
[0077] In the embodiment of the present application, optionally, the entity similarity analysis result of the first analogy corpus pair may be whether the entities are similar, such as Figure 3 False in indicates dissimilarity, and True indicates similarity. Common words can be listed in a list.
[0078] In the embodiment of the present application, optionally, the result of performing similarity analysis on the entities in the first analogy corpus pair (such as Figure 3The answer in the question can also include the entity similarity analysis process and analysis results, such as Figure 3 The analysis process is, "In Base, the protagonists are the tortoise and the rabbit. They had a race. The rabbit took a break on the way because of its self-confidence, and finally the tortoise won the race. In Target, the protagonists are the thief and the security guard. The thief always escapes because he runs fast, but the security guard sets a trap and finally the thief is caught. Therefore, the specific backgrounds of the two (the tortoise and the rabbit race and the security guard catching the thief) are different: the fable is different from the real event, and the responsibilities of the characters (the tortoise and the rabbit, the thief and the security guard) are also different: the tortoise and the rabbit are in a competitive relationship, and the thief and the security guard are in a catching relationship. But there are certain similarities in the ups and downs of the plot development. They are all stories about failure caused by pursuit. In addition, there is no common vocabulary."
[0079] In some embodiments, optionally, the first prompt word also includes: first task prompt information, the first task prompt information is used to prompt the large language model to perform the task and results of similarity analysis on entities in the first analogy corpus pair according to the instructions, perform similarity analysis on entities in the second analogy corpus pair, and output the entity similarity analysis results of the second analogy corpus pair. Through the first task prompt information, the large language model can be more prepared to obtain the user's intention.
[0080] It should be noted that Figure 3 The unlabeled second analogy pair is not shown in the first prompt word.
[0081] Step S2: determining a second prompt word, wherein the second prompt word includes: the first analogy corpus pair, a task and a result indicating structural similarity alignment of sentences in the first analogy corpus pair, and the second analogy corpus pair; inputting the second prompt word into the large language model to obtain a sentence pair of structural similarity alignment of the second analogy corpus pair;
[0082] Please refer to Figure 2 , Step S2 can also be called a sentence mapping step, that is, comparing the sentences in the two analogy corpora in the analogy corpus pair to obtain sentences with structural similarity in the two analogy corpora.
[0083] Please refer to Figure 4 , Figure 4 is a schematic diagram of the second prompt word of an embodiment of the present invention, from Figure 4 It can be seen that the second prompt word includes the first analogy corpus pair as a case, and the first analogy corpus pair includes two text corpora, such as Figure 4 The second prompt word also includes: indicating the task and result of aligning the structural similarity of the sentences in the first analogy corpus pair, such as Figure 4The Question in indicates the task of aligning the sentences in the first analogy corpus pair by structural similarity, such as Figure 4 The Answer in is the result of structural similarity alignment of the sentences in the first analogy corpus pair, also referred to as the sentence pair aligned by structural similarity of the first analogy corpus pair.
[0084] It should be noted that, in an embodiment of the present invention, when performing structural similarity alignment on an analogy corpus pair, it is necessary to ensure that all sentences in the two text corpora in the analogy corpus pair participate in the structural similarity alignment. If a sentence in one text corpus does not have a corresponding structural similarity aligned sentence in the other text corpus, a prompt message indicating that there is no structural similarity aligned sentence is returned.
[0085] Optionally, the sentence described in the embodiment of the present invention refers to a sentence containing only one period.
[0086] In some embodiments, optionally, the second prompt word also includes: second task prompt information, the second task prompt information is used to prompt the large language model to perform the task and results of structurally aligning the sentences in the first analogy corpus pair according to the instructions, perform structurally aligning the second analogy corpus pair, and output the sentence pairs of the second analogy corpus pair aligned with structural similarity. Through the second task prompt information, the large language model can be more prepared to obtain the user's intention.
[0087] Step S3: determining a third prompt word, wherein the third prompt word includes: a sentence pair aligned with structural similarity of the first analogy corpus pair, a task and result indicating a causal logic similarity analysis of the sentence pair aligned with structural similarity of the first analogy corpus pair, and a sentence pair aligned with structural similarity of the second analogy corpus pair; inputting the third prompt word into the large language model to obtain a causal logic similarity analysis result of the second analogy corpus pair;
[0088] Please refer to Figure 2 ,Step S3 can also be called the relation alignment step, that is, causal logical similarity analysis is performed on each sentence pair of the analogy corpus.
[0089] Please refer to Figure 5 , Figure 5 is a schematic diagram of a third prompt word according to an embodiment of the present invention, Figure 5 It can be seen that the third prompt word includes the sentence pairs aligned with the structural similarity of the first analogy corpus pair. The third prompt word also includes: indicating the task and result of performing causal logic similarity analysis on the sentence pairs aligned with the structural similarity of the first analogy corpus pair, such as Figure 5The Question in indicates the task of analyzing the causal logic similarity of the sentence pairs aligned with the structural similarity of the first analogy corpus, such as Figure 5 The Answer in is the result of causal logic similarity analysis on the sentence pairs aligned with the structural similarity of the first analogy corpus pair, also referred to as the causal logic similarity analysis result of the first analogy corpus pair.
[0090] In the embodiment of the present invention, optionally, the causal logic similarity analysis result of the first analogy corpus pair may include an analysis process and an analysis result, such as Figure 5 As shown, the analysis process is:
[0091] “1. Because one party (Target) contains 'None', it is classified as an irrelevant group.
[0092] 2. The rabbit in Base is confident that he can win because of his fast speed, which corresponds to the thief in Target who can escape because he runs fast. Both are examples of confidence or success due to speed advantage and are classified into similar groups.
[0093] 3. Because one party (Target) contains 'None', it is classified as an irrelevant group.
[0094] 4. Base describes the rabbit stopping to rest because of complacency, while Target describes the guard using a trick to catch the thief. The characters described in the two do not correspond in structure and the content has no logical relevance, so they are classified as unrelated groups.
[0095] 5. Because Base describes the tortoise continuing to move forward while the rabbit is resting, while Target describes the thief being knocked unconscious by a carriage while escaping, both describe the main characters losing consciousness, but the reasons are different: the rabbit rests because it thinks the tortoise is too slow and it would not be a big deal even if it rests for a while, while the thief did not want to stop, but was knocked unconscious by an unexpected carriage while escaping. He was not confident in his personal ability and voluntarily lay down to be caught by the security guards, so they are classified as dissimilar groups.
[0096] 6. The ending of Base is that the rabbit wakes up and finds failure, while the ending of Target is that the thief wakes up and is surrounded by guards. Both describe the unfavorable situation faced by the protagonist after regaining consciousness from a coma or sleeping state, but the causes are different: the rabbit is subjectively confident, while the thief is objectively unexpected. Therefore, they are classified as dissimilar groups.
[0097] In some embodiments, optionally, the third prompt word also includes: third task prompt information, wherein the third task prompt information is used to prompt the large language model to perform the task and results of causal logic similarity analysis on the sentence pairs aligned with structural similarity of the first analogy corpus pair according to the instructions, perform causal logic similarity analysis on the sentence pairs aligned with structural similarity of the second analogy corpus pair, and output the causal logic similarity analysis results of the second analogy corpus pair. Through the third task prompt information, the large language model can be more prepared to obtain the user's intention.
[0098] Step S4: Determine a fourth prompt word, wherein the fourth prompt word includes: the first analogy corpus pair and its entity similarity analysis results and causal logic similarity analysis results, the analogy relationship analysis results of the first analogy corpus pair, and the second analogy corpus pair and its causal logic similarity analysis results; input the fourth prompt word into the large language model to obtain the analogy relationship analysis results of the second analogy corpus pair.
[0099] Please refer to Figure 2 Step S4 can also be called the analogy summary step, that is, determining the analogy relationship analysis result of the analogy corpus pair according to the entity similarity analysis result and the causal logic similarity analysis result of the analogy corpus pair.
[0100] In some embodiments, optionally, the analogy relationship analysis result includes whether the entities are similar in appearance, whether the first-order relationships are similar, and whether the higher-order relationships are similar.
[0101] Please refer to Figure 6 , Figure 6 is a schematic diagram of a fourth prompt word according to an embodiment of the present invention, Figure 6 It can be seen that the fourth prompt word includes the first analogy corpus pair, the entity similarity analysis result of the first analogy corpus pair, the causal logic similarity analysis result, and the analogy relationship analysis result of the first analogy corpus pair. The entity similarity analysis result of the first analogy corpus pair is Figure 6 In the "Base and Target story backgrounds are different (False), the characters are different (False), and the plots are similar (True)", the causal logic similarity analysis results are Figure 6 In the formula, “len(similar group) = 1, len(dissimilar group) = 2, len(irrelevant group) = 3)”, the analogy relationship analysis result of the first analogy corpus pair is Figure 6 ""entities":"dissimilar","one-order relations":"similar","higher-order relations":"dissimilar"" in .
[0102] In some embodiments, optionally, the fourth prompt word also includes: fourth task prompt information, the fourth task prompt information is used to prompt the large language model to infer the analogy relationship analysis result of the first analogy corpus pair according to the first analogy corpus pair and its entity similarity analysis result and the causal logic similarity analysis result, and to infer and output the analogy relationship analysis result of the second analogy corpus pair according to the second analogy corpus pair and its causal logic similarity analysis result. Through the fourth task prompt information, the large language model can be more prepared to obtain the user's intention.
[0103] In an embodiment of the present invention, a prompt word is determined, the prompt word includes an analogy corpus pair and its analogy relationship analysis process as a case, and a large language model running in an intelligent computing center is called to analyze the analogy relationship of the unlabeled analogy corpus pair according to the analogy relationship analysis process of the analogy corpus pair as a case in the prompt word, and the analogy relationship analysis result of the unlabeled analogy corpus pair is obtained. The analogy relationship type of the analogy corpus pair can be labeled according to the analogy relationship analysis result, and high-quality analogy corpus pairs can be screened out, thereby improving the cognitive ability and analogy reasoning ability of the large language model trained with these high-quality analogy corpus pairs.
[0104] In the embodiment of the present invention, the execution order of the above steps S1 and S2 is not limited, and step S1 can be executed first and then step S2, or step S2 can be executed first and then step S1, or step S1 and step S2 can be executed at the same time. Step S3 needs to be executed after step S2, and step S4 needs to be executed after step S1 and step S3.
[0105] In some embodiments, optionally, the method further comprises:
[0106] Step S5: marking the analogy relationship type of the second analogy corpus pair according to whether the entities indicated in the analogy relationship analysis result of the second analogy corpus pair are similar in form, whether the first-order relationships are similar, and whether the higher-order relationships are similar.
[0107] In some embodiments, optionally, when outputting the analogy relationship analysis result of the second analogy corpus pair, the large language model may also label and output the analogy relationship type of the second analogy corpus pair based on whether the entities are similar in shape, whether the first-order relationships are similar, and whether the higher-order relationships are similar as indicated in the analogy relationship analysis result of the second analogy corpus pair.
[0108] In some embodiments, optionally, the analogy relationship type includes at least one of the following:
[0109] Literal similarity type, in which the analogy corpus pairs of the literal similarity type have similar entities, first-order relations, and higher-order relations;
[0110] True analogy type, in the analogy corpus pair of the true analogy type, the entities are dissimilar, some first-order relations are similar, and the higher-order relations are similar;
[0111] A pseudo-analogy type, in which the entities in the analogy corpus pairs are dissimilar, the first-order relations are similar, and the higher-order relations are dissimilar;
[0112] A surface similarity type, in which the analogy corpus pairs have similar entities, similar first-order relations, and dissimilar higher-order relations;
[0113] Only appearance similarity type, in the analogy corpus pairs of only appearance similarity type, the entities are similar, the first-order relations are not similar, and the higher-order relations are not similar;
[0114] An abnormal type, in which the analogy corpus pairs of the abnormal type have dissimilar entities, first-order relations and higher-order relations.
[0115] In some embodiments, optionally, the first analogy corpus pair is of a pseudo-analogy type, in which the entities are not similar, the first-order relations are similar, and the higher-order relations are not similar. Since the analogy corpus pairs of the pseudo-analogy type include more comprehensive situations, it is preferred to use the analogy corpus of the pseudo-analogy type as a case.
[0116] The computing power of the intelligent computing center in the embodiment of the present invention may include a GPU, etc., which has high computing power, so that a large amount of analogy corpus can be processed as above, and a large number of annotated analogy corpus pairs can be obtained for the training of a large language model, which can further enhance the cognitive ability and analogy reasoning ability of the large language model trained with these analogy corpus pairs.
[0117] Please refer to Figure 7 The embodiment of the present invention further provides a device 10 for automatically annotating analogy corpus by using the computing power of an intelligent computing center, comprising:
[0118] The first processing module 11 is used to determine a first prompt word, wherein the first prompt word includes: a first analogy corpus pair as a case, a task and a result indicating similarity analysis of entities in the first analogy corpus pair, and an unlabeled second analogy corpus pair; the first prompt word is input into a large language model to obtain an entity similarity analysis result of the second analogy corpus pair;
[0119] The second processing module 12 is used to determine a second prompt word, wherein the second prompt word includes: the first analogy corpus pair, a task and a result indicating structural similarity alignment of sentences in the first analogy corpus pair, and the second analogy corpus pair; input the second prompt word into the large language model to obtain a sentence pair of structural similarity alignment of the second analogy corpus pair;
[0120] The third processing module 13 is used to determine a third prompt word, wherein the third prompt word includes: a sentence pair of the first analogy corpus pair aligned with structural similarity, a task and result indicating the causal logic similarity analysis of the sentence pair of the first analogy corpus pair aligned with structural similarity, and a sentence pair of the second analogy corpus pair aligned with structural similarity; the third prompt word is input into the large language model to obtain a causal logic similarity analysis result of the second analogy corpus pair;
[0121] The fourth processing module 14 is used to determine a fourth prompt word, which includes: the first analogy corpus pair and its entity similarity analysis results and causal logic similarity analysis results, the analogy relationship analysis results of the first analogy corpus pair, and the second analogy corpus pair and its causal logic similarity analysis results; the fourth prompt word is input into the large language model to obtain the analogy relationship analysis results of the second analogy corpus pair.
[0122] In an embodiment of the present invention, a prompt word is determined, the prompt word includes an analogy corpus pair and its analogy relationship analysis process as a case, and a large language model running in an intelligent computing center is called to analyze the analogy relationship of the unlabeled analogy corpus pair according to the analogy relationship analysis process of the analogy corpus pair as a case in the prompt word, and the analogy relationship analysis result of the unlabeled analogy corpus pair is obtained. The analogy relationship type of the analogy corpus pair can be labeled according to the analogy relationship analysis result, and high-quality analogy corpus pairs can be screened out, thereby improving the cognitive ability and analogy reasoning ability of the large language model trained with these high-quality analogy corpus pairs.
[0123] In some embodiments, optionally, the method further includes:
[0124] The labeling module is used to label the analogy relationship type of the second analogy corpus pair according to whether the entities are similar in shape, whether the first-order relationships are similar, and whether the higher-order relationships are similar indicated in the analogy relationship analysis result of the second analogy corpus pair.
[0125] In some embodiments, optionally, the analogy relationship type includes at least one of the following:
[0126] Literal similarity type, in which the analogy corpus pairs of the literal similarity type have similar entities, first-order relations, and higher-order relations;
[0127] True analogy type, in the analogy corpus pair of the true analogy type, the entities are dissimilar, some first-order relations are similar, and the higher-order relations are similar;
[0128] A pseudo-analogy type, in which the entities in the analogy corpus pairs are dissimilar, the first-order relations are similar, and the higher-order relations are dissimilar;
[0129] A surface similarity type, in which the analogy corpus pairs have similar entities, similar first-order relations, and dissimilar higher-order relations;
[0130] Only appearance similarity type, in the analogy corpus pairs of only appearance similarity type, the entities are similar, the first-order relations are not similar, and the higher-order relations are not similar;
[0131] An abnormal type, in which the analogy corpus pairs of the abnormal type have dissimilar entities, first-order relations and higher-order relations.
[0132] In some embodiments, optionally, the entity includes at least one of the following: background, character, plot, and common vocabulary, wherein the common vocabulary includes identical vocabulary and synonyms.
[0133] In some embodiments, optionally, the first analogy corpus pair is of a pseudo-analogy type, in which the entities are dissimilar, the first-order relations are similar, and the higher-order relations are dissimilar.
[0134] In some embodiments, optionally, the first prompt word further includes: first task prompt information, the first task prompt information being used to prompt the large language model to perform a task and result of similarity analysis on entities in the first analogy corpus pair according to the instruction, perform similarity analysis on entities in the second analogy corpus pair, and output entity similarity analysis results of the second analogy corpus pair;
[0135] And / or, the second prompt word further includes: second task prompt information, the second task prompt information is used to prompt the large language model to perform a task and a result of structural similarity alignment of the sentences in the first analogy corpus pair according to the instruction, perform structural similarity alignment on the second analogy corpus pair, and output a sentence pair of the second analogy corpus pair aligned with structural similarity;
[0136] And / or, the third prompt word also includes: third task prompt information, the third task prompt information is used to prompt the large language model to perform causal logic similarity analysis on the sentence pairs aligned with structural similarity of the first analogy corpus according to the task and result of the instruction, perform causal logic similarity analysis on the sentence pairs aligned with structural similarity of the second analogy corpus, and output the causal logic similarity analysis result of the second analogy corpus;
[0137] And / or, the fourth prompt word also includes: fourth task prompt information, the fourth task prompt information is used to prompt the large language model to infer the analogy relationship analysis result of the first analogy corpus pair according to the first analogy corpus pair and its entity similarity analysis result and the causal logic similarity analysis result, and to infer and output the analogy relationship analysis result of the second analogy corpus pair according to the second analogy corpus pair and its causal logic similarity analysis result.
[0138] Please refer to Figure 8 The embodiment of the present invention further provides an electronic device 20, including a processor 21, a memory 22, and a computer program stored in the memory 22 and executable on the processor 21. When the computer program is executed by the processor 21, each process of the method embodiment of automatically annotating analogy corpus by the computing power of an intelligent computing center is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0139] The embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, each process of the above-mentioned method embodiment for automatically annotating analogy corpus by the computing power of an intelligent computing center can be implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0140] The present application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the above Figure 1 The various processes of the method embodiment for automatically annotating analogy corpus through the computing power of an intelligent computing center are shown, and the same technical effect can be achieved. To avoid repetition, they will not be described here.
[0141] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0142] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0143] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation modes, which are merely illustrative rather than restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are within the protection of the present invention.
Claims
1. A method for automatically annotating analogy corpus by using the computing power of an intelligent computing center, characterized in that: include: Step S1: determining a first prompt word, wherein the first prompt word includes: a first analogy corpus pair as a case, a task and result indicating similarity analysis of entities in the first analogy corpus pair, and an unlabeled second analogy corpus pair; inputting the first prompt word into a large language model to obtain an entity similarity analysis result of the second analogy corpus pair; Step S2: determining a second prompt word, wherein the second prompt word includes: the first analogy corpus pair, a task and a result indicating structural similarity alignment of sentences in the first analogy corpus pair, and the second analogy corpus pair; inputting the second prompt word into the large language model to obtain a sentence pair of structural similarity alignment of the second analogy corpus pair; Step S3: determining a third prompt word, wherein the third prompt word includes: a sentence pair aligned with structural similarity of the first analogy corpus pair, a task and result indicating a causal logic similarity analysis of the sentence pair aligned with structural similarity of the first analogy corpus pair, and a sentence pair aligned with structural similarity of the second analogy corpus pair; inputting the third prompt word into the large language model to obtain a causal logic similarity analysis result of the second analogy corpus pair; Step S4: Determine a fourth prompt word, wherein the fourth prompt word includes: the first analogy corpus pair and its entity similarity analysis results and causal logic similarity analysis results, the analogy relationship analysis results of the first analogy corpus pair, and the second analogy corpus pair and its causal logic similarity analysis results; input the fourth prompt word into the large language model to obtain the analogy relationship analysis results of the second analogy corpus pair.
2. The method according to claim 1, characterized in that Also includes: Step S5: marking the analogy relationship type of the second analogy corpus pair according to whether the entities indicated in the analogy relationship analysis result of the second analogy corpus pair are similar in form, whether the first-order relationships are similar, and whether the higher-order relationships are similar.
3. The method according to claim 2, characterized in that The analogy relationship type includes at least one of the following: Literal similarity type, in which the analogy corpus pairs of the literal similarity type have similar entities, first-order relations, and higher-order relations; True analogy type, in the analogy corpus pair of the true analogy type, the entities are dissimilar, some first-order relations are similar, and the higher-order relations are similar; A pseudo-analogy type, in which the entities in the analogy corpus pairs are dissimilar, the first-order relations are similar, and the higher-order relations are dissimilar; A surface similarity type, in which the analogy corpus pairs have similar entities, similar first-order relations, and dissimilar higher-order relations; Only appearance similarity type, in the analogy corpus pairs of only appearance similarity type, the entities are similar, the first-order relations are not similar, and the higher-order relations are not similar; An abnormal type, in which the analogy corpus pairs of the abnormal type have dissimilar entities, first-order relations and higher-order relations.
4. The method according to claim 1, characterized in that The entity includes at least one of the following: background, character, plot and common vocabulary, wherein the common vocabulary includes identical vocabulary and synonyms.
5. The method according to claim 1, characterized in that The first analogy corpus pair is of a pseudo-analogy type. In the analogy corpus pair of the pseudo-analogy type, entities are not similar, first-order relations are similar, and higher-order relations are not similar.
6. The method according to claim 1, characterized in that The first prompt word also includes: first task prompt information, the first task prompt information is used to prompt the large language model to perform a task and result of similarity analysis on entities in the first analogy corpus pair according to the instruction, perform similarity analysis on entities in the second analogy corpus pair, and output entity similarity analysis results of the second analogy corpus pair; And / or, the second prompt word further includes: second task prompt information, the second task prompt information is used to prompt the large language model to perform a task and a result of structural similarity alignment of the sentences in the first analogy corpus pair according to the instruction, perform structural similarity alignment on the second analogy corpus pair, and output a sentence pair of the second analogy corpus pair aligned with structural similarity; And / or, the third prompt word also includes: third task prompt information, the third task prompt information is used to prompt the large language model to perform causal logic similarity analysis on the sentence pairs aligned with structural similarity of the first analogy corpus according to the task and result of the instruction, perform causal logic similarity analysis on the sentence pairs aligned with structural similarity of the second analogy corpus, and output the causal logic similarity analysis result of the second analogy corpus; And / or, the fourth prompt word also includes: fourth task prompt information, the fourth task prompt information is used to prompt the large language model to infer the analogy relationship analysis result of the first analogy corpus pair according to the first analogy corpus pair and its entity similarity analysis result and the causal logic similarity analysis result, and to infer and output the analogy relationship analysis result of the second analogy corpus pair according to the second analogy corpus pair and its causal logic similarity analysis result.
7. A device for automatically annotating analogy corpus by using the computing power of an intelligent computing center, characterized in that: include: A first processing module is used to determine a first prompt word, wherein the first prompt word includes: a first analogy corpus pair as a case, a task and a result indicating similarity analysis of entities in the first analogy corpus pair, and an unlabeled second analogy corpus pair; input the first prompt word into a large language model to obtain an entity similarity analysis result of the second analogy corpus pair; The second processing module is used to determine a second prompt word, wherein the second prompt word includes: the first analogy corpus pair, a task and a result indicating structural similarity alignment of sentences in the first analogy corpus pair, and the second analogy corpus pair; input the second prompt word into the large language model to obtain a sentence pair of structural similarity alignment of the second analogy corpus pair; a third processing module, configured to determine a third prompt word, wherein the third prompt word includes: a sentence pair aligned with structural similarity of the first analogy corpus pair, a task and result indicating a causal logic similarity analysis of the sentence pair aligned with structural similarity of the first analogy corpus pair, and a sentence pair aligned with structural similarity of the second analogy corpus pair; input the third prompt word into the large language model to obtain a causal logic similarity analysis result of the second analogy corpus pair; The fourth processing module is used to determine a fourth prompt word, wherein the fourth prompt word includes: the first analogy corpus pair and its entity similarity analysis results and causal logic similarity analysis results, the analogy relationship analysis results of the first analogy corpus pair, and the second analogy corpus pair and its causal logic similarity analysis results; the fourth prompt word is input into the large language model to obtain the analogy relationship analysis results of the second analogy corpus pair.
8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the method for automatically annotating analogy corpus by using the computing power of an intelligent computing center as described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for automatically annotating analogy corpus by using the computing power of an intelligent computing center as described in any one of claims 1 to 6.
10. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of the method for automatically annotating analogy corpus by using the computing power of an intelligent computing center as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Language large model training method, system and device and computer readable storage medium
CN118210895A
Method, device and equipment for constructing prompt automatic labeling medical text based on large model and medium
CN118445419A
Methods, systems and computer program products for analogy detection among entities using reciprocal similarity measures
US20080306944A1
Method, apparatus and system for consistency enhanced large language models
US20250013873A1
Data processing method and related device
WO2024255755A1