Method and device for generating analogy corpus through computing power of intelligent computing center

The computing power of the intelligent computing center generates text corpus pairs with analogical relationships, which solves the problem of the lack of analogy ability of existing corpus and improves the analogy reasoning ability of large language models.

CN119990318APending Publication Date: 2025-05-13DATACANVAS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510072369.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing intelligent computing centers used to train large language models lack analogical capabilities, resulting in poor analogical reasoning capabilities of large language models.

Method used

Through the computing power of the intelligent computing center, the original text corpus is obtained and prompt words are generated, and a large language model is used to generate text corpus pairs with analogical relationships, including the first text corpus and the second text corpus. The first text corpus is obtained through information extraction and details expansion, and the second text corpus is written by a large language model, with similar logical structures but not similar entities.

Benefits of technology

The generated analogous corpus pairs enrich the analogy ability of the training corpus and improve the cognitive ability and analogy reasoning ability of large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990318A_ABST
    Figure CN119990318A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for generating analogy corpora through computing power of an intelligent computing center, and the method comprises the steps: S1, obtaining an original text corpus set which comprises a plurality of original text corpora; s2, cue words are generated according to the original text corpus, the cue words are input into the large language model, a text corpus pair generated based on the original text corpus is obtained, the text corpus pair comprises a first text corpus and a second text corpus which have an analogy relation, and the first text corpus and the second text corpus are in an analogy relation; the first text corpus is obtained by carrying out information extraction and detail expansion on the original text corpus through the large language model, the second text corpus is compiled through the large language model, and the second text corpus and the original text corpus are similar in logic structure but not similar in entity. According to the method, the training corpus pairs with rich analogy ability can be obtained, so that the cognitive ability and analogy reasoning ability of a large language model trained by adopting the analogy corpus pairs are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the fields of computing power infrastructure and artificial intelligence technology, and in particular, to a method and device for generating analogy corpus through the computing power of an intelligent computing center. Background Art

[0002] With the development of artificial intelligence technology and computing power technology, the concept of intelligent computing center has emerged. "Intelligent computing center" refers to the use of large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, mainly for artificial intelligence applications (such as artificial intelligence deep learning model (such as large language model) development, model training and model reasoning and other scenarios) to provide the required computing power, data and algorithms. Intelligent computing center covers facilities, hardware, software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.

[0003] When training a large language model, the existing intelligent computing center needs to input a large amount of corpus into the large language model, and the large language model learns the semantic relationship of the corpus to complete specific tasks. However, the existing network corpus downloaded directly from the Internet is messy and lacks analogy ability. Therefore, the large language model cannot fully absorb the analogy reasoning learning nutrients from the existing corpus, which limits the expansion of its analogy reasoning ability. Summary of the invention

[0004] The embodiments of the present invention provide a method and device for generating analogy corpus through the computing power of an intelligent computing center, which is used to solve the problem that the corpus used by the existing intelligent computing center for training large language models lacks analogy ability, resulting in poor analogy reasoning ability of the large language models trained with these corpora.

[0005] In order to solve the above-mentioned technical problems, the present invention is achieved as follows:

[0006] In a first aspect, an embodiment of the present invention provides a method for generating analogy corpus by using the computing power of an intelligent computing center, comprising:

[0007] Step S1: obtaining an original text corpus, wherein the original text corpus includes a plurality of original text corpora;

[0008] Step S2: Generate prompt words according to the original text corpus, and input the prompt words into the large language model to obtain a text corpus pair generated based on the original text corpus, the text corpus pair includes a first text corpus and a second text corpus with an analogy relationship, the first text corpus is obtained by the large language model performing information extraction and detail expansion on the original text corpus, the second text corpus is written by the large language model, and the second text corpus is similar to the original text corpus in logical structure but not in entity.

[0009] Optionally, the prompt words include:

[0010] The original text corpus;

[0011] Instructing the large language model on a task;

[0012] Definition or generation prompt information of output content of the large language model, the output content including the first text corpus and the second text corpus;

[0013] The generation prompt information of the first text corpus includes: prompting the large language model to expand the first summary text of the original text corpus in detail to obtain the first text corpus;

[0014] The generation prompt information of the second text corpus includes: prompting the large language model to expand the second summary text in detail to obtain the second text corpus, the second summary text is a summary text written based on the logical structure in the original text corpus, and the second summary text is similar to the original text corpus in logical structure but not in entity.

[0015] Optionally, the output content also includes at least one of the following:

[0016] the first summary text;

[0017] the second summary text;

[0018] Analogy comparison data between the first text corpus and the second text corpus.

[0019] Optionally, the output content also includes a score for the text corpus pair, and the score includes at least one of the following: an innovation score, a significance score, and a logical consistency score.

[0020] Optionally, also include:

[0021] Step S3: Based on the scores of the text corpus pairs and the corresponding filtering thresholds, the final required text corpus pairs are selected from the multiple text corpus pairs output by the large language model to form a text corpus pair training set, and the text corpus pair training set is used to train the large language model.

[0022] Optionally, the prompt word includes one or more, and when multiple prompt words are included, step S2 includes:

[0023] Step S21: inputting the plurality of prompt words into the large language model in sequence in different steps, wherein the prompt words located later are determined according to the previous output content of the large language model.

[0024] In a second aspect, an embodiment of the present invention provides a device for generating analogy corpus by using the computing power of an intelligent computing center, including:

[0025] A first acquisition module is used to acquire an original text corpus, wherein the original text corpus includes a plurality of original text corpora;

[0026] The second acquisition module is used to generate prompt words according to the original text corpus, and input the prompt words into the large language model to obtain a text corpus pair generated based on the original text corpus, wherein the text corpus pair includes a first text corpus and a second text corpus having an analogy relationship, wherein the first text corpus is obtained by extracting information and expanding details of the original text corpus by the large language model, and the second text corpus is written by the large language model, and the second text corpus is similar to the original text corpus in logical structure but not in entity.

[0027] In a third aspect, an embodiment of the present invention provides an electronic device, comprising: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the method for generating analogy corpus through the computing power of an intelligent computing center as described in the first aspect above.

[0028] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method for generating analogy corpus through the computing power of an intelligent computing center as described in the first aspect above are implemented.

[0029] In a fifth aspect, an embodiment of the present invention provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the steps of the method for generating analogy corpus through the computing power of an intelligent computing center as described in the first aspect above.

[0030] In the embodiment of the present invention, by running a large language model in an intelligent computing center, analogy corpus pairs are automatically generated based on the original text corpus, so that training corpus pairs with rich analogy capabilities can be obtained for the training of the large language model, thereby improving the cognitive ability and analogy reasoning ability of the large language model trained with these analogy corpus pairs. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0032] Figure 1This is a flow chart of a method for generating analogy corpus by using the computing power of an intelligent computing center according to an embodiment of the present invention;

[0033] Figure 2 The second flowchart of the method for generating analogy corpus by using the computing power of an intelligent computing center according to an embodiment of the present invention;

[0034] Figure 3 A schematic diagram of a prompt word according to an embodiment of the present invention;

[0035] Figure 4 Schematic diagram of a text corpus pair according to an embodiment of the present invention;

[0036] Figure 5 It is a structural schematic diagram of an apparatus for generating analogy corpus through the computing power of an intelligent computing center according to an embodiment of the present invention;

[0037] Figure 6 Schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0038] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0039] First, the technical terms involved in the present invention are briefly explained below.

[0040] The "computing power" mentioned in the present invention is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to execute certain computing requirements. It is the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity. It mainly provides services to the society through computing power infrastructure.

[0041] The "computing power" (Computational Power, CP) described in the present invention is the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, super computing power and intelligent computing power. The commonly used unit of measurement is the number of floating point operations performed per second (FLOPS: Floating Point Operations Per Second, 1EFLOPS = 10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream notebooks. The calculation formula is: CP = CP 通用 +CP 智能 +CP 超级 .

[0042] The "Network Power" (NP) described in the present invention is a manifestation of the data transmission capability of computing facilities, including comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, etc. The carrying capacity involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities. In the embodiment of the present invention, the carrying capacity uses the memory bandwidth.

[0043] The "Storage Power" (SP) described in the present invention is the comprehensive ability of a data center in terms of data storage capacity, performance, safety and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and server built-in storage devices. The commonly used unit of measurement for storage capacity is exabyte (EB, 1EB = 2^60bytes), and the commonly used unit of measurement for performance is the number of reads and writes per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB). The disaster recovery ratio is an important manifestation of safety and reliability.

[0044] The "computing power infrastructure" described in the present invention is a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity. It can realize centralized computing, storage, transmission and application of information, and presents characteristics such as diversity and ubiquity, intelligence and agility, security and reliability, and green and low-carbon.

[0045] The “computing power” mentioned in the present invention includes general computing power, intelligent computing power and super computing power.

[0046] The “general computing power” mentioned in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0047] The "intelligent computing power" described in the present invention is a computing platform for large-scale deployment of special chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit) for various innovative applications of artificial intelligence, such as natural language processing and machine vision.

[0048] The "super computing power" mentioned in the present invention is mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, and gene analysis.

[0049] The "intelligent computing center" described in the present invention refers to a facility that provides the required computing power, data and algorithms for artificial intelligence applications (such as artificial intelligence deep learning model development, model training and model reasoning scenarios) by using large-scale heterogeneous computing power resources, including general computing power (CPU: Central Processing Unit) and intelligent computing power (GPU: Graphics Processing Unit, FPGA: Field Programmable Gate Array, ASIC: Application Specific Integrated Circuit, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.

[0050] The "computing resources" mentioned in the present invention refer to the technologies and facilities with information calculation, transmission, storage and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPU (Central Processing Unit), GPU (Graphics Processing Unit), network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, as well as supporting and guarantee resources such as wind, fire, water and electricity.

[0051] The “large language model” mentioned in the present invention refers to a large language model (LLM), which is a language model with a large parameter scale. It is designed to understand and generate human language. It is trained with a large amount of text data and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.

[0052] In the field of linguistics and natural language processing, the "corpus" mentioned in the present invention refers to text or voice data used for language research, analysis, teaching or technology development. In natural language processing (NLP), corpus is used to train machine learning models, such as language models, text classifiers, sentiment analyzers, etc. These models learn the statistical laws and patterns of language by analyzing a large amount of corpus, so that they can perform various language processing tasks, such as text generation, translation, summarization, question answering, etc.

[0053] The "prompt" mentioned in the present invention refers to a prompt or instruction designed based on natural language to guide the large language model to perform a specific task.

[0054] To solve the problem that the corpus used by the existing intelligent computing center to train large language models lacks analogy ability, resulting in poor analogy reasoning ability of large language models trained with these corpora, please refer to Figure 1 , an embodiment of the present invention provides a method for generating analogy corpus by using the computing power of an intelligent computing center, comprising:

[0055] Step S1: obtaining an original text corpus, wherein the original text corpus includes a plurality of original text corpora;

[0056] In the embodiment of the present invention, please refer to Figure 2 The original text corpus in the original text corpus set may be text downloaded from the Internet (Web), or may be text downloaded from the Internet and processed after data cleaning.

[0057] The original text corpus in the embodiment of the present invention may be a pan-text, which refers to texts in various fields, such as literature, news, science and technology, medicine, etc.

[0058] Step S2: Generate prompt words according to the original text corpus, and input the prompt words into the large language model to obtain a text corpus pair generated based on the original text corpus, the text corpus pair includes a first text corpus and a second text corpus with an analogy relationship, the first text corpus is obtained by the large language model performing information extraction and detail expansion on the original text corpus, the second text corpus is written by the large language model, and the second text corpus is similar to the original text corpus in logical structure but not in entity.

[0059] The concept of analogy is first explained below.

[0060] Analogy is at the core of human cognition, allowing us to abstract information and understand new (unfamiliar) situations based on familiar situations. Text corpora with analogical relations can be described as text corpora with dissimilar entities (including objects and characters) but similar logical structures. Among them, the logical structure can include first-order relations and higher-order relations. That is, text corpora with analogical relations can also be described as text corpora with dissimilar entities, similar higher-order relations and at least some first-order relations.

[0061] In the embodiment of the present invention, the entities of the second text corpus and the first text corpus are not similar, but the logical structures are similar. Further, optionally, the second text corpus and the first text corpus belong to different fields. For example, the second text corpus belongs to the field of natural sciences, and the first text corpus belongs to the field of stories.

[0062] The following example illustrates a text corpus with analogical relationships.

[0063] A pair of text corpora with an analogy relationship includes: "The planets in the solar system revolve around the sun" and "The electrons in an atom revolve around the nucleus". The analogy relationship includes: "planet" and "electron", "sun" and "nucleus".

[0064] Another pair of text corpora with an analogical relationship includes: "An employee received a seemingly harmless attachment that contained malware. The malware invaded his personal computer and stole his sensitive personal information." and "The citizens of Troy allowed access to the Trojan Horse containing Greek soldiers. The Greek soldiers captured Troy and stole their wealth." Among them, the analogical relationships include: "employee" and "citizens of Troy", "attachment" and "Trojan horse", "malware" and "Greek soldiers", "personal computer" and "Troy", and "sensitive personal information" and "wealth".

[0065] The first text corpus is obtained by extracting information and expanding details of the original text corpus by the large language model, wherein information extraction refers to structural processing of information contained in the text, wherein information extraction may include at least one of the following: entity extraction, relationship extraction, event extraction, viewpoint extraction, etc. Detail expansion, in the embodiment of the present invention, refers to enriching the text content obtained after information extraction (referred to as overview text in the embodiment of the present invention), and enriching the content, for example, is to strengthen the description of the part (concept or plot) that has an analogy relationship with the second text corpus.

[0066] In the embodiment of the present invention, by running a large language model in an intelligent computing center, analogy corpus pairs are automatically generated based on the original text corpus, so that training corpus pairs with rich analogy capabilities can be obtained for the training of the large language model, thereby improving the cognitive ability and analogy reasoning ability of the large language model trained with these analogy corpus pairs.

[0067] The computing power of the intelligent computing center in the embodiment of the present invention may include a GPU, etc., which has high computing power, so that the above-mentioned processing can be performed on a large amount of text downloaded from the Internet to obtain a large number of analogy corpus pairs for training large language models, which can further enhance the cognitive ability and analogy reasoning ability of the large language model trained with these analogy corpus pairs.

[0068] For some examples, please refer to Figure 3 , the prompt words include:

[0069] (1) the original text corpus;

[0070] Figure 3 In the prompt words shown, the original text corpus is not shown.

[0071] (2) tasks assigned to the large language model;

[0072] The tasks given to the large language model, such as Figure 3 The prompt in the text is "Your task is to read the text provided by the user, find out which fragments can "awaken similar logical structures deep in your memory", deeply analyze the deep logical structure of each paragraph, and use your strong document writing ability to sort out the base_text and recall_text with analogical relationships."

[0073] Among them, the text provided by the user is the original text corpus, base_text is the first text corpus, and recall_text is the second text corpus.

[0074] (3) definition of output content of the large language model or generation prompt information, the output content includes the first text corpus and the second text corpus; wherein the generation prompt information of the first text corpus includes: prompting the large language model to expand the first summary text of the original text corpus in detail to obtain the first text corpus; the generation prompt information of the second text corpus includes: prompting the large language model to expand the second summary text in detail to obtain the second text corpus, the second summary text is a summary text written based on the logical structure in the original text corpus, and the second summary text is similar to the logical structure of the original text corpus but not similar in entity.

[0075] In the embodiment of the present invention, the definition of the output content is an explanation of the output content. The generation prompt information of the output content is a prompt for the large language model on how to generate the output content.

[0076] In an embodiment of the present invention, the first summary text may be obtained by extracting information from the original text corpus. The first text corpus is obtained by expanding the first summary text in detail, wherein the first summary text is expanded in detail, i.e., the content of the first summary text is enriched, and the enriched content is, for example, strengthening the description of the part (concept or plot) in the first summary text that has an analogy relationship with the second text corpus. The second summary text is expanded in detail, i.e., the content of the second summary text is enriched, and the enriched content is, for example, strengthening the description of the part (concept or plot) in the second summary text that has an analogy relationship with the first text corpus.

[0077] In some embodiments, optionally, the prompt word may also include: the role of the large prediction model, such as Figure 3 The prompt in the text is "You are an AI wise man who is good at reading and has rich emotions."

[0078] In some embodiments, optionally, the output content further includes at least one of the following:

[0079] (4) the first summary text;

[0080] like Figure 3 The “base_scene” in the prompt word is obtained by extracting information from the original text corpus.

[0081] (5) the second summary text;

[0082] like Figure 3 The second summary text is a summary text written based on the logical structure in the original text corpus, and the second summary text is similar to the original text corpus in logical structure but not in entity.

[0083] (6) Analogy comparison data between the first text corpus and the second text corpus.

[0084] like Figure 3 The word “analogies” in the prompt word.

[0085] In some embodiments, optionally, the prompt word also includes: the format of the output content, such as Figure 3 The prompt "Your response should be a json list" indicates that the output content should be listed in a json list. For example, Figure 3The “base_scene,str” in the prompt word in the first overview text is in string format.

[0086] In some embodiments, the prompt word may optionally further include: prompt information during task processing, which is used to prompt matters needing attention during task processing by the large language model, such as Figure 3 The "Notes" in the prompt words in . Among them, the "base_text and recall_text are independent texts and cannot refer to each other" in the notes means that base_text and recall_text cannot contain each other's content, that is, the first text corpus and the second text corpus cannot contain each other's content. For example, if base_text is "Planets revolve around the sun", and if the content of recall_text generated by the large language model is "Electrons revolve around the nucleus. We can understand this phenomenon that electrons are like the earth, and the earth revolves around the sun." In this example, recall_text references the content of base_tex.

[0087] In some embodiments, the output content may also include a score for the text corpus pair, and the score may include at least one of the following: an innovation score, a significance score, a logical consistency score, etc. The innovation may be, for example, whether the second text corpus generated this time is similar to the second text corpus generated previously by the large language model. The significance score refers to whether the content of the second text corpus can convey information, provide knowledge, stimulate thinking, or transmit emotions. The logical consistency score refers to whether the theory or thinking mode of the second text corpus does not have any contradictions in its own language statements. Please refer to Figure 3 The scores in the prompt word are the scores in the embodiment of the present invention.

[0088] In some embodiments, optionally, the method of generating analogy corpus by using the computing power of an intelligent computing center further includes:

[0089] Step S3: Based on the scores of the text corpus pairs and the corresponding filtering thresholds, the final required text corpus pairs are selected from the multiple text corpus pairs output by the large language model to form a text corpus pair training set (AnalogyDataset), which is used to train the large language model.

[0090] Optionally, each type of score corresponds to a filtering threshold, for example, the innovation score corresponds to an innovation filtering threshold, and the innovation of the text corpus pair meets the requirements only when the innovation score of the second text corpus in the text corpus pair is greater than or equal to the innovation filtering threshold. The meaningfulness score corresponds to a meaningfulness filtering threshold, and the text corpus pair meets the requirements in terms of whether it is meaningful only when the meaningfulness score of the second text corpus in the text corpus pair is greater than or equal to the meaningfulness filtering threshold.

[0091] Optionally, only when all types of scores of the text corpus pair meet the requirements of the corresponding filtering threshold can the text corpus pair be used as the final required text corpus pair.

[0092] In the embodiment of the present invention, by setting a filtering threshold, higher quality text corpus pairs can be screened out for training a large language model, thereby further improving the cognitive ability and analogical reasoning ability of the large language model trained with these analogy corpus pairs.

[0093] In some embodiments, optionally, the prompt word includes one or more, and when multiple prompt words are included, step S2 includes:

[0094] Step S21: inputting the plurality of prompt words into the large language model in sequence in different steps, wherein the prompt words located later are determined according to the previous output content of the large language model.

[0095] For example, multiple prompt words can be divided into the following two steps and input into the large language model:

[0096] Step 1: The plurality of prompt words include a first prompt word, and the first prompt word is used to prompt the large prediction model to generate a first summary text and a second summary text according to the original text corpus. The first prompt word is input into the large language model to obtain the first summary text and the second summary text.

[0097] Step 2: The plurality of prompt words include a second prompt word, and the second prompt word is used to prompt the large prediction model to generate a first text corpus and a second text corpus according to the original text corpus, the first summary text, and the second summary text. The second prompt word is input into the large language model to obtain the first text corpus and the second text corpus.

[0098] Alternatively, multiple prompt words can be divided into the following three steps and input into the large language model:

[0099] Step 1: The plurality of prompt words include a third prompt word, and the third prompt word is used to prompt the large prediction model to generate a first summary text according to the original text corpus. The third prompt word is input into the large language model to obtain the first summary text.

[0100] Step 2: The plurality of prompt words include a fourth prompt word, and the fourth prompt word is used to prompt the large prediction model to generate a second summary text according to the original text corpus. The fourth prompt word is input into the large language model to obtain the second summary text.

[0101] Step 3: The plurality of prompt words include a fifth prompt word, and the fifth prompt word is used to prompt the large prediction model to generate a first text corpus and a second text corpus according to the original text corpus, the first summary text, and the second summary text. The fifth prompt word is input into the large language model to obtain the first text corpus and the second text corpus.

[0102] Alternatively, multiple prompt words can be divided into the following 4 steps and input into the large language model:

[0103] Step 1: The plurality of prompt words include a third prompt word, and the third prompt word is used to prompt the large prediction model to generate a first summary text according to the original text corpus. The third prompt word is input into the large language model to obtain the first summary text.

[0104] Step 2: The plurality of prompt words include a fourth prompt word, and the fourth prompt word is used to prompt the large prediction model to generate a second summary text according to the original text corpus. The fourth prompt word is input into the large language model to obtain the second summary text.

[0105] Step 3: The plurality of prompt words include a sixth prompt word, and the sixth prompt word is used to prompt the large prediction model to generate a first text corpus according to the original text corpus and the first summary text. The sixth prompt word is input into the large language model to obtain the first text corpus.

[0106] Step 4: The plurality of prompt words include a seventh prompt word, and the seventh prompt word is used to prompt the large prediction model to generate a second text corpus according to the original text corpus and the second summary text. The seventh prompt word is input into the large language model to obtain the second text corpus.

[0107] Please refer to Figure 4 , Figure 4 is a schematic diagram of a text corpus pair according to an embodiment of the present invention, Figure 4 In the example, Base_text is the first text corpus in the embodiment of the present invention, recall_text is the second text corpus in the embodiment of the present invention, and the analog comparison data of Base_text and recall_text quality inspection include:

[0108] 'Yu Tu's passion for spaceflight: Columbus's passion for exploration';

[0109] 'The details of aerospace in the play: description of Columbus's navigation technology';

[0110] 'Yu Tu's Growth and Space Connection: Columbus's Sailing Career';

[0111] 'Yu Tu's QQ nickname "Yutu Pounding Medicine": Columbus's sailing ship';

[0112] 'Dialogues between characters in the play recall the history of space travel: the voyage records in Columbus' diary';

[0113] 'Yu Tu's research on Chang'e-1: Columbus's research on new routes'.

[0114] Please refer to Figure 5 The embodiment of the present invention further provides a device 10 for generating analogy corpus by using the computing power of an intelligent computing center, comprising:

[0115] A first acquisition module 11 is used to acquire an original text corpus, wherein the original text corpus includes a plurality of original text corpora;

[0116] The second acquisition module 12 is used to generate prompt words according to the original text corpus, and input the prompt words into the large language model to obtain a text corpus pair generated based on the original text corpus, wherein the text corpus pair includes a first text corpus and a second text corpus having an analogy relationship, wherein the first text corpus is obtained by extracting information and expanding details of the original text corpus by the large language model, and the second text corpus is written by the large language model, and the second text corpus is similar to the original text corpus in logical structure but not in entity.

[0117] Optionally, the prompt words include:

[0118] The original text corpus;

[0119] Instructing the large language model on a task;

[0120] Definition or generation prompt information of output content of the large language model, the output content including the first text corpus and the second text corpus;

[0121] The generation prompt information of the first text corpus includes: prompting the large language model to expand the first summary text of the original text corpus in detail to obtain the first text corpus;

[0122] The generation prompt information of the second text corpus includes: prompting the large language model to expand the second summary text in detail to obtain the second text corpus, the second summary text is a summary text written based on the logical structure in the original text corpus, and the second summary text is similar to the original text corpus in logical structure but not in entity.

[0123] Optionally, the output content also includes at least one of the following:

[0124] the first summary text;

[0125] the second summary text;

[0126] Analogy comparison data between the first text corpus and the second text corpus.

[0127] Optionally, the output content also includes a score for the text corpus pair, and the score includes at least one of the following: an innovation score, a significance score, and a logical consistency score.

[0128] Optionally, also include:

[0129] A filtering module is used to select the final required text corpus pairs from the multiple text corpus pairs output by the large language model based on the scores of the text corpus pairs and the corresponding filtering thresholds to form a text corpus pair training set, and the text corpus pair training set is used to train the large language model.

[0130] Optionally, the prompt word includes one or more, and when it includes multiple prompt words, the second acquisition module 12 is used to input the multiple prompt words into the large language model in sequence in different steps, wherein the prompt word located later is determined according to the previous output content of the large language model.

[0131] In the embodiment of the present invention, by running a large language model in an intelligent computing center, analogy corpus pairs are automatically generated based on the original text corpus, so that training corpus pairs with rich analogy capabilities can be obtained for the training of the large language model, thereby improving the cognitive ability and analogy reasoning ability of the large language model trained with these analogy corpus pairs.

[0132] Please refer to Figure 6 An embodiment of the present invention further provides an electronic device 20, including a processor 21, a memory 22, and a computer program stored in the memory 22 and executable on the processor 21. When the computer program is executed by the processor 21, each process of the above-mentioned method embodiment for generating analogy corpus through the computing power of an intelligent computing center is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0133] The embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, each process of the above-mentioned method embodiment for generating analogy corpus by computing power of an intelligent computing center is implemented, and the same technical effect can be achieved. To avoid repetition, it is not described here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0134] The present application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the above Figure 1 The various processes of the method embodiment for generating analog corpus through the computing power of an intelligent computing center are shown, and can achieve the same technical effect. To avoid repetition, they will not be described here.

[0135] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0136] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present invention.

[0137] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation modes, which are merely illustrative rather than restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are within the protection of the present invention.

Claims

1. A method for generating analogy corpus by using the computing power of an intelligent computing center, characterized in that: include: Step S1: obtaining an original text corpus, wherein the original text corpus includes a plurality of original text corpora; Step S2: Generate prompt words according to the original text corpus, and input the prompt words into the large language model to obtain a text corpus pair generated based on the original text corpus, the text corpus pair includes a first text corpus and a second text corpus with an analogy relationship, the first text corpus is obtained by the large language model performing information extraction and detail expansion on the original text corpus, the second text corpus is written by the large language model, and the second text corpus is similar to the original text corpus in logical structure but not in entity.

2. The method according to claim 1, characterized in that The prompt words include: The original text corpus; Instructing the large language model on a task; Definition or generation prompt information of output content of the large language model, the output content including the first text corpus and the second text corpus; The generation prompt information of the first text corpus includes: prompting the large language model to expand the first summary text of the original text corpus in detail to obtain the first text corpus; The generation prompt information of the second text corpus includes: prompting the large language model to expand the second summary text in detail to obtain the second text corpus, the second summary text is a summary text written based on the logical structure in the original text corpus, and the second summary text is similar to the original text corpus in logical structure but not in entity.

3. The method according to claim 2, characterized in that The output content also includes at least one of the following: the first summary text; the second summary text; Analogy comparison data between the first text corpus and the second text corpus.

4. The method according to claim 2, characterized in that: The output content also includes a score for the text corpus pair, and the score includes at least one of the following: an innovation score, a significance score, and a logical consistency score.

5. The method according to claim 4, characterized in that Also includes: Step S3: Based on the scores of the text corpus pairs and the corresponding filtering thresholds, the final required text corpus pairs are selected from the multiple text corpus pairs output by the large language model to form a text corpus pair training set, and the text corpus pair training set is used to train the large language model.

6. The method according to claim 2, characterized in that The prompt word includes one or more, and when multiple prompt words are included, step S2 includes: Step S21: inputting the plurality of prompt words into the large language model in sequence in different steps, wherein the prompt words located later are determined according to the previous output content of the large language model.

7. A device for generating analogy corpus by using the computing power of an intelligent computing center, characterized in that: include: A first acquisition module is used to acquire an original text corpus, wherein the original text corpus includes a plurality of original text corpora; The second acquisition module is used to generate prompt words according to the original text corpus, and input the prompt words into the large language model to obtain a text corpus pair generated based on the original text corpus, wherein the text corpus pair includes a first text corpus and a second text corpus having an analogy relationship, wherein the first text corpus is obtained by extracting information and expanding details of the original text corpus by the large language model, and the second text corpus is written by the large language model, and the second text corpus is similar to the original text corpus in logical structure but not in entity.

8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the method for generating analogy corpus by using the computing power of an intelligent computing center as described in any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for generating analogy corpus through the computing power of an intelligent computing center as described in any one of claims 1 to 6.

10. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of the method for generating analogy corpus by using the computing power of an intelligent computing center as described in any one of claims 1 to 6.