Sample corpus generation method, training method and testing method for large model

By identifying the target entity that matches the demand attributes of the scene, using the large language model to generate sample corpus that matches the user's needs, the problem that the reply content of the large language model is difficult to meet the actual needs of users, and the accuracy and effect of intelligent question-and-answer is improved.

CN120196720APending Publication Date: 2025-06-24BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510330418.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In intelligent Q&A scenarios based on large language models, the reply content is difficult to accurately meet the actual needs of users, and there are often problems that reply is irrelevant to user needs or are highly repetitive.

Method used

By identifying the corpus requirements information intently, determining the target entity that matches the scene requirements attributes, and processing the target entity with a large language model to generate sample corpus to ensure that the sample corpus matches the user's business scenario and requirements.

Benefits of technology

It improves the quality and relevance of sample corpus, ensures that the responses output by the large model can more accurately meet the actual needs of users, and improves the effect of intelligent question and answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196720A_ABST
    Figure CN120196720A_ABST
Patent Text Reader

Abstract

The invention provides a sample corpus generation method, a training method and a testing method for a large model, and relates to the technical field of artificial intelligence, in particular to the technical fields of large language models, generative models, big data, knowledge maps and the like. The sample corpus generation method for the large model comprises the following steps: performing intention recognition on corpus demand information to obtain a corpus demand intention; a target entity matched with a scene demand attribute in the corpus demand intention is determined from a plurality of service entities, the service entities are extracted from service basic data, and the scene correlation attribute of the service entities represents the matching degree between the service entities and an execution condition for executing a specified service scene; the scene demand attribute represents a demand intention for the matching degree; and processing the target entity by using the large language model to obtain a sample corpus.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly to technical fields such as large language models, generative models, big data, knowledge graphs, etc. Background Art

[0002] With the rapid development of artificial intelligence technology, large language models capable of multi-round question-and-answer interactions with users are applied to scenarios such as intelligent e-commerce services and online education. Based on the relatively powerful semantic understanding ability and information generation ability of large models, the actual needs of users such as information consultation and information retrieval can be met through question-and-answer interaction methods in related scenarios. Summary of the Invention

[0003] The present disclosure provides a method for generating sample corpus for a large model, a training method, a testing method, an intelligent agent, a device, an electronic device, and a storage medium.

[0004] According to one aspect of the present disclosure, a method for generating sample corpus for a large model is provided, including: performing intent recognition on corpus requirement information to obtain a corpus requirement intent; determining a target entity that matches the scenario requirement attribute in the corpus requirement intent from multiple business entities, where the business entities are extracted from business basic data, the scenario relevance attribute of the business entity represents the matching degree between the business entity and the execution conditions of the specified business scenario, and the scenario requirement attribute represents the requirement intent for the matching degree; and processing the target entity using a large language model to obtain sample corpus.

[0005] According to another aspect of the present disclosure, a training method for a large model is provided, including: obtaining sample corpus, where the sample corpus is determined based on the method provided in the embodiments of the present disclosure; and training an initial large model based on the sample corpus to obtain a trained large model.

[0006] According to another aspect of the present disclosure, a testing method for a large model is provided, including: obtaining sample corpus, where the sample corpus is determined based on the method provided in the embodiments of the present disclosure; processing the sample corpus using the large model to obtain an output text; and determining the test result of the large model based on the output text.

[0007] According to another aspect of the present disclosure, there is provided a sample corpus generation device for a large model, including: an identification module configured to perform intent identification on corpus requirement information to obtain a corpus requirement intent; a first determination module configured to determine a target entity that matches the scenario requirement attribute in the corpus requirement intent from multiple business entities, where the business entities are extracted from business basic data, and the scenario relevance attribute of the business entity represents the matching degree between the business entity and the execution conditions of a specified business scenario, and the scenario requirement attribute represents the requirement intent for the matching degree; and a sample corpus obtaining module configured to process the target entity nodes using a large language model to obtain a sample corpus.

[0008] According to another aspect of the present disclosure, there is provided a training device for a large model, including: a first acquisition module configured to acquire a sample corpus, where the sample corpus is determined based on the method provided in the embodiments of the present disclosure; and a training module configured to train an initial large model based on the sample corpus to obtain a trained large model.

[0009] According to another aspect of the present disclosure, there is provided a testing device for a large model, including: a second acquisition module configured to acquire a sample corpus, where the sample corpus is determined based on the method provided in the embodiments of the present disclosure; an output text obtaining module configured to process the sample corpus using the large model to obtain an output text; and a second determination module configured to determine the test result of the large model based on the output text.

[0010] According to another aspect of the present disclosure, there is provided an intelligent agent for artificial intelligence, including: an input module configured to receive input information; a processing module configured to determine a target task based on the input information received by the input module, determine a large language model based on the target task, and obtain output information by calling the large language model to execute the method for generating a sample corpus for a large model provided in the embodiments of the present disclosure, or by calling the large model to execute the training method for the large model or the testing method for the large model; and an output module configured to output the output information obtained by the processing module.

[0011] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.

[0012] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the above method.

[0013] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, where the computer program, when executed by a processor, implements the above method.

[0014] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0016] Figure 1 Schematically shows an exemplary system architecture to which a method and apparatus for generating sample corpus for a large model according to an embodiment of the present disclosure can be applied;

[0017] Figure 2 Schematically shows a flowchart of a method for generating sample corpus for a large model according to an embodiment of the present disclosure;

[0018] Figure 3 Schematically shows a schematic diagram of a method for generating sample corpus for a large model according to an embodiment of the present disclosure;

[0019] Figure 4 Schematically shows a flowchart of generating a knowledge graph according to an embodiment of the present disclosure;

[0020] Figure 5 Schematically shows a flowchart of a method for generating sample corpus for a large model according to another embodiment of the present disclosure;

[0021] Figure 6 Schematically shows a schematic diagram of a method for generating sample corpus for a large model according to another embodiment of the present disclosure;

[0022] Figure 7 Schematically shows a flowchart of a method for training a large model according to an embodiment of the present disclosure;

[0023] Figure 8 Schematically shows a flowchart of a method for testing a large model according to an embodiment of the present disclosure;

[0024] Figure 9 Schematically shows a block diagram of an apparatus for generating sample corpus for a large model according to an embodiment of the present disclosure;

[0025] Figure 10 Schematically shows a block diagram of a training apparatus for a large model according to an embodiment of the present disclosure;

[0026] Figure 11 Schematically shows a block diagram of a testing apparatus for a large model according to an embodiment of the present disclosure;

[0027] Figure 12A structural block diagram of an agent of artificial intelligence according to an embodiment of the present disclosure is schematically shown; and

[0028] Figure 13 A schematic block diagram of an example electronic device for a sample corpus generation method for a large model, a training method for a large model, and a testing method for a large model, which can be used to implement the embodiments of the present disclosure, is shown. Detailed implementation manners

[0029] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0030] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and public order and good customs are not violated.

[0031] The inventors found that in the intelligent question-and-answer scenarios based on large language models (LLMs), multimodal large models, and other large models, the reply content output by the large model is relatively difficult to accurately meet the actual needs of users. For example, the reply content may be irrelevant to the actual business scenario of the user's current needs, or the repetition degree of the reply content related to the business scenario of the actual needs is relatively high, making it difficult to meet the actual needs of users.

[0032] The embodiments of the present disclosure provide a sample corpus generation method, a training method, a testing method, an agent, a device, an electronic device, and a storage medium for a large model. The sample corpus generation method for a large model includes: performing intent recognition on corpus requirement information to obtain a corpus requirement intent; determining, from a plurality of business entities, a target entity that matches the scenario requirement attribute in the corpus requirement intent, where the business entities are extracted from business basic data, and the scenario relevance attribute of the business entity represents the matching degree between the business entity and the execution conditions of the specified business scenario, and the scenario requirement attribute represents the requirement intent for the matching degree; and processing the target entity using a large language model to obtain a sample corpus.

[0033] According to an embodiment of the present disclosure, by identifying the corpus requirement intention of the corpus requirement information and determining, through the scenario requirement attributes in the corpus requirement intention, a target entity that matches the execution conditions of executing a specified business scenario from multiple business entities, the target entity can meet the user's requirement for the matching degree between the sample corpus and the execution conditions of the specified business scenario. Thus, the large language model can be used to process the target entity to generate a sample corpus with a high degree of fit to the user's business scenario and requirement level, so as to enrich the content of the sample corpus and ensure that the relevance between the sample corpus and the business scenario meets the training requirements of the training task for executing the specified business scenario or the testing requirements of the large model for the inference task of executing the specified business scenario, thereby improving the quality of the sample corpus.

[0034] Figure 1 Schematically shows an exemplary system architecture to which the method and apparatus for generating a sample corpus for a large model according to an embodiment of the present disclosure can be applied.

[0035] It should be noted that Figure 1 The illustration is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios. For example, in another embodiment, the exemplary system architecture to which the method and apparatus for generating a sample corpus for a large model can be applied may include a terminal device, but the terminal device can implement the method and apparatus for generating a sample corpus for a large model provided by the embodiments of the present disclosure without interacting with the server.

[0036] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0037] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as examples).

[0038] The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0039] Server 105 can be a server that provides various services, such as a background management server (for example only) that supports the content browsed by users using terminal devices 101, 102, and 103. The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0040] Server 105 can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). Server 105 can also be a server of a distributed system, or a server combined with a blockchain.

[0041] It should be noted that the method for generating sample corpus for a large model provided by the embodiments of the present disclosure can generally be executed by server 105. Correspondingly, the device for generating sample corpus for a large model provided by the embodiments of the present disclosure can generally be set in server 105. The method for generating sample corpus for a large model provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the device for generating sample corpus for a large model provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.

[0042] It should be understood, Figure 1 the numbers of terminal devices, networks, and servers in

[0043] Figure 2 are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.

[0044] As Figure 2 shown, the method for generating sample corpus for a large model includes operations S210 to S230.

[0045] In operation S210, the intention of the corpus requirement information is recognized to obtain the corpus requirement intention.

[0046] In operation S220, a target entity that matches the scenario requirement attributes in the corpus requirement intention is determined from multiple business entities.

[0047] In operation S230, the target entity is processed using a large language model to obtain a sample corpus.

[0048] According to an embodiment of the present disclosure, the corpus requirement information is used to indicate the requirement intention for the sample corpus used in the large model. The sample corpus may include sample texts for controlling the large model to perform data processing tasks, and the large model can be trained based on the sample corpus so that the trained large language model can meet the actual question-and-answer requirements of the specified business scenario. Alternatively, the sample corpus can also be used to test the base large model or the trained large model, so as to test the performance of the large model through the test response information output by processing the sample corpus by the base large model or the trained large model. The embodiment of the present disclosure does not limit the specific usage mode of the sample corpus, as long as it can be used in the large model.

[0049] It should be noted that the corpus requirement information can be represented based on natural language, or the corpus requirement information can also be information determined based on an interactive operation for a requirement intention option. The embodiment of the present disclosure does not limit the specific data type of the corpus requirement information.

[0050] According to an embodiment of the present disclosure, intent recognition of the corpus requirement information may include processing the corpus requirement information based on any type of information extraction algorithm such as a keyword extraction algorithm. Alternatively, the corpus requirement information can also be processed using a large language model to obtain the corpus requirement intention. The embodiment of the present disclosure does not limit the specific manner of obtaining the corpus requirement intention. The corpus requirement intention can represent the scenario relevance requirement between the required sample corpus and the specified business scenario, or can also represent requirements such as the expression mode of the sample corpus and the context semantic relevance degree. The embodiment of the present disclosure does not limit the specific requirement attributes represented by the corpus requirement intention.

[0051] In some embodiment modes, multiple business entities can be mapped based on a preset knowledge graph. The preset knowledge graph includes business entity nodes representing business entities, and edges between multiple business entity nodes. The business entity nodes represent business entities extracted from business basic data, and the business entities can have entity attribute information such as scenario relevance attributes. The scenario relevance attribute of a business entity represents the matching degree between the business entity and the execution conditions of the specified business scenario. The scenario requirement attribute represents the requirement intention for the matching degree. For example, the scenario requirement attribute can represent the matching degree between the subsequently generated sample corpus and the execution conditions of the specified business scenario.

[0052] It should be noted that the business basic data may include unstructured information such as user manuals, usage specifications, and emergency measures for hardware facilities such as devices and apparatuses in a specified business scenario. Alternatively, the business basic data may also include structured or unstructured data generated by executing a specified business scenario. For example, the business basic data may also include information such as the numbers, installation locations, and working hours of hardware devices.

[0053] According to an embodiment of the present disclosure, determining a target entity that matches the scenario requirement attributes in the corpus requirement intention from multiple business entities may include matching the scenario relevance attributes of the business entities based on the scenario requirement attributes to obtain target scenario relevance attributes that match the scenario requirement attributes, and determining the business entity with the target scenario relevance attributes as the target entity. The business node representing the target entity in the knowledge graph may be the target entity node. The target entity represented by the target entity node may have a scenario matching degree with the execution conditions of the specified business scenario, which can meet the actual needs of the user for the sample corpus.

[0054] In one example, the target entity node may be determined from the knowledge graph based on a Retrieval-Augmented Generation (RAG) tool or a Generative Business Intelligence (GBI) tool selected by the user.

[0055] According to an embodiment of the present disclosure, using a large language model to process the target entity to obtain a sample corpus may include using the large language model to process the target entity represented by the target entity node and the entity attribute information of the target entity to obtain a sample corpus expressed in natural language. Alternatively, using a large language model to process the target entity may also include using the large language model to process the target entity represented by the target entity node and the associated entities having an attribute relationship with the target entity, so as to output a sample corpus represented in natural language by utilizing the semantic understanding ability and text output ability of the large language model.

[0056] According to an embodiment of the present disclosure, the sample corpus may include a question corpus generated based on the target entity and the entity attributes of the target entity. For example, if the target entity is a "switch", the entity attributes of the target entity may include "number and installation location". The sample corpus may be "Where is the installation location of the switch numbered 0001?".

[0057] According to an embodiment of the present disclosure, the sample corpus may include one or more sentences of text corpus, and there may be a context semantic relationship between multiple sentences of text corpus. By controlling the prompt information input to the large language model, the large language model can be made to process the target entity to generate a sample corpus that matches the requirements of the prompt information, so as to meet the test requirements for testing the intelligent response performance of the large model for a specified business scenario based on the sample corpus, or to train the large model based on the sample corpus so that the trained large model can meet the intelligent response requirements of users for inquiries regarding a specified business scenario.

[0058] In one example, the unstructured business basic data may include files such as industry standard documents and industry specification documents related to the specified business. In addition, the business basic data may also include file data such as technical documents and user manuals related to products such as devices and apparatuses. In addition, the business basic data may also include historical question-and-answer data related to any industry standard documents, technical documents, etc. related to the specified business scenario.

[0059] In one example, the structured business basic data may include data generated during the execution of the specified business, such as device operation data, alarm data, monitoring data, etc. generated in the management system for executing the specified business. The structured business basic data may be stored in a storage device of a server such as a database based on a structured description schema such as a schema.

[0060] According to an embodiment of the present disclosure, business entities can be obtained by performing entity extraction on the business basic data, and the entity attribute information of the business entities can be determined based on the structured or unstructured basic business data. In addition, based on data such as indexes and data type organization methods indicated by the structured business basic data, the attribute relationships of the business entities over time can be determined. Or the attribute relationships between different business entities can also be determined by performing semantic analysis on the unstructured data. The attribute relationships between business entities can be represented based on the edges in the knowledge graph, and the business entity nodes in the knowledge graph can be used to represent the corresponding business entities.

[0061] It should be noted that the business entities may include any information related to the execution of the specified business, such as devices, events, locations, etc. in the specified business scenario, and the specific extraction method of the business entities in the embodiments of the present disclosure is not limited.

[0062] In one example, the knowledge graph can be constructed based on the user's interactive operation intention. For example, by processing the interactive operations input by the user, it can be determined that the user needs to extract business entities from structured business basic data or unstructured business basic data. For the intention that the user needs to extract entity and attribute relationships from structured business basic data, the structural level of the structured stored business basic data can be parsed to obtain each basic data field and the corresponding annotation information. By performing entity extraction on the annotation information, the business entities and the entity attribute information of the business entities can be determined. The entity attribute information can be, for example, the performance parameters of the device, etc. By analyzing the entity attribute information of different business entities to determine the attribute relationships between multiple business entities, an initial knowledge graph can be constructed based on the triples composed of business entities and attribute relationships. The knowledge graph can be obtained by adding scene relevance attributes to the business entities in the initial knowledge graph.

[0063] In one example, entity extraction and attribute relationship extraction are performed by performing syntactic structure analysis and semantic analysis on unstructured business basic data to obtain the business entities and the attribute relationships between the business entities. For example, the attribute relationships between the business entities can be obtained based on the following steps 1 to 7. For example, the unstructured business basic data can be a natural language sentence related to a specified road management scenario: "Data and details of cameras operating normally on the AB Expressway target hub - B city boundary section."

[0064] In step 1, syntactic structure and semantic analysis are performed on the unstructured business basic data sentence to determine the syntactic attributes and semantic attributes of each word in the sentence. For example, based on morphological and syntactic rules, the main component attributes such as the subject, predicate, and object of the sentence and the modifier component attributes such as the attributive, adverbial, and complement can be identified, and the complete semantic attribute information of the sentence can be understood in combination with the context.

[0065] For example, for the unstructured business basic data "Data and details of cameras operating normally on the AB Expressway target hub - B city boundary section.", it can be determined that "on the AB Expressway target hub - B city boundary section" is used as an adverbial of place to limit the scope. "Operating normally" describes the state of the camera and is used as an attributive to modify "camera". "Camera" is the object, and "data and details" are the relevant subsidiary information of the "camera" and are the specific content to be finally queried in the sentence. Overall, the semantics of the sentence is to query the relevant data and details of the cameras in a normal operating state on a specific section.

[0066] In step 2, core transactions are extracted from the unstructured business basic data as business entities. For example, entities such as people, places, organizations, and items with independent meanings in the business basic data can be extracted according to the business entity definition. For example, "AB Expressway Target Hub - Section of B City Boundary" belongs to an entity; "camera" is an item with independent meaning and also belongs to an entity.

[0067] In step 3, identify the words or characters that modify the entity part in the statement. Determine the words or characters in the statement that further describe or limit the business entity. For example, the words appearing before the business entity are used as attributives, the words appearing after the business entity are used as complements, and the words used to describe the occurrence of the action of the business entity are used as adverbials. For example, "AB Expressway Target Hub - Section of B City Boundary" limits the position of the "camera"; "operating normally" describes the operating state of the "camera", and both are used to modify the "camera".

[0068] In step 4, determine whether the words or characters that modify the entity part are the attributes or state descriptions of the business entity. The attributes of the business entity can be understood as relatively stable characteristics and properties of the business entity, and the attributes can be information that does not change rapidly with time or environment. The state description can be understood as information with timeliness, and the state description can represent the situation of the business entity at a specific moment or time period. For example, "operating normally" describes the operating situation of the camera at the current moment and will change at any time with time and the device status, so it belongs to the state description. The statements in this embodiment may not include entity attribute information for the time being. It should be noted that the entity attribute information of the business entity can be determined as a data set of attributes and state descriptions, or the entity attribute information of the business entity can also be determined as the attributes representing the relatively stable characteristics and properties of the business entity.

[0069] In step 5, determine the business entities with attribute relationships. Analyze whether there are attribute association situations such as execution action transfer, execution condition influence, and execution logic association among the business entities in the statement, and determine the attribute relationships between the business entities based on the attribute association situations. Among them, the large language model can be made to focus on the keywords representing attribute relationships such as verbs, prepositions, and conjunctions in the statement based on the enhanced information.

[0070] For example, in the above statement, there is a position relationship between the "camera" and the "AB Expressway Target Hub - Section of B City Boundary", which can be understood as the "camera" is located at the "AB Expressway Target Hub - Section of B City Boundary"; there is an ownership relationship between the "camera" and the "data and details", and the "data and details" belong to the "camera".

[0071] In step 6, determine the type of attribute relationship between business entities. According to the specific manifestation forms of the attribute relationships between business entities, the type of attribute relationship between business entities can be determined as attribute relationship types such as "ownership", "causality", "action", etc. For example, the attribute relationship type between "camera" and "AB Expressway Target Hub - Section of the B City Boundary" is a location relationship type; the attribute relationship between "camera" and "data and details" is an ownership attribute relationship type.

[0072] In step 7, based on the determined business entities and the attribute relationships between business entities, multiple "business entity - attribute relationship - business entity" triples can be determined. Based on the multiple triples, a knowledge graph can be constructed to obtain an initial knowledge graph. The knowledge graph is obtained by adding scene relevance attributes to the business entities in the initial knowledge graph.

[0073] According to an embodiment of the present disclosure, the scene relevance attribute of a business entity can be a scene relevance level attribute. There can be multiple scene relevance level attributes, and different scene relevance level attributes can represent different degrees of matching between the business entity and the execution conditions of a specified business scene. The scene relevance level attribute can be determined based on expert decision-making, or can also be determined based on other methods, such as being determined based on the frequency of occurrence of the business entity during the execution of the business scene. The embodiment of the present disclosure does not limit the specific method for determining the scene relevance attribute of a business entity, as long as it can represent the matching degree between the business entity and the execution conditions of a specified business scene.

[0074] In one example, the scene relevance attribute can include three different level attributes, P0, P1, and P2. The three different level attributes can indicate that the degree of matching with the execution conditions of the specified business decreases in sequence. By identifying that the scene requirement attribute in the corpus demand information is P0, the business entity with the scene relevance attribute of P0 is determined from the knowledge graph as the target business entity.

[0075] In one example, different business entities in the knowledge graph can be related to multiple types of business scenarios. The corpus requirement intention can also include a scenario requirement intention representing a specified business scenario. Multiple candidate business entities associated with the specified business scenario represented by the scenario requirement intention are determined from the knowledge graph through the scenario requirement intention, and a target business entity with a scenario relevance attribute matching the scenario requirement attribute in the scenario requirement intention is determined from the multiple candidate business entities. In this way, based on the business scenario type represented by the corpus requirement intention for the specified business scenario and the matching degree with the execution conditions of the specified business scenario, the target entity matching the corpus requirement intention of the user can be accurately determined. Furthermore, by using a large language model to process the target entity node, sample corpus matching the requirement intention represented by the corpus requirement information can be obtained, improving the quality of the sample corpus.

[0076] According to an embodiment of the present disclosure, the scenario relevance attribute includes a first relevance attribute, and the first relevance attribute is determined based on the following operations: extracting scenario description entities from the business scenario description information; determining a first detection result based on the semantic similarity between the scenario description entities and the business entities; and when the first detection result indicates that the semantic similarity meets a preset similarity condition, determining the scenario relevance attribute of the business entity as the first relevance attribute.

[0077] According to an embodiment of the present disclosure, the business scenario description information can be information related to the business scenario such as the execution tasks and business execution requirements for executing the business scenario. The business scenario description information can be unstructured natural language information. For example, the business scenario description information can be rescue process information, emergency process exercises, etc. for describing a specified business scenario as an emergency rescue scenario. The business scenario description information can also be structured information. For example, the business scenario description information can be emergency equipment operation status data, etc.

[0078] According to an embodiment of the present disclosure, the business scenario description information can be closely related to the execution conditions for executing the business scenario. By extracting scenario description entities from the business scenario description information, the necessary execution conditions for executing the specified business scenario can be represented by the scenario description entities. By determining the semantic similarity between the scenario description entities and the business entities, and when the first detection result indicates that the semantic similarity meets the preset similarity condition, it can be determined that the business entity and the scenario description entity represent the same semantics. Therefore, the business entity has a relatively close logical relationship with the necessary conditions of the business scenario.

[0079] According to an embodiment of the present disclosure, a business entity with a first correlation attribute can meet the user's requirement for generating sample corpora closely related to the execution conditions of a specified business scenario, thereby improving the matching degree between the sample corpora and the specified scenario of the user's needs, so that the large model for testing or training can achieve intelligent and accurate responses to the problems of the specified business scenario, and improve the reasoning performance of the large model for the problems of the specified business scenario.

[0080] Figure 3 Schematically shows the principle diagram of a method for generating sample corpora for a large model according to an embodiment of the present disclosure.

[0081] As Figure 3 shown, business basic data can be input into the entity extraction component and the relationship detection component in the entity analysis module 310 to obtain multiple business entities and the attribute relationships between different business entities, and then business entity triples are generated based on "business entity - attribute relationship - business entity". The business scenario description information is input into the entity extraction component in the entity analysis module 310 to obtain scenario description entities. The business entities and scenario description entities in the business entity triples are input into the scenario correlation attribute detection component 321, and it can be obtained that the business entity has a first correlation attribute or a second correlation attribute. Furthermore, the first correlation attribute or the second correlation attribute can be added to the entity attributes of the business entity to obtain an updated business entity triple. The updated business entity triple is input into the knowledge graph generation component, and a knowledge graph can be constructed based on the attribute relationships between different business entities, and an updated knowledge graph is output.

[0082] According to an embodiment of the present disclosure, the scenario correlation attribute includes a second correlation attribute, and the second correlation attribute is determined based on the following operations: using a large language model to process the scenario description information, business entities, and sub - condition prompt templates to obtain the second correlation attribute of the business entity.

[0083] According to an embodiment of the present disclosure, the sub - condition prompt template is used to control the large language model to judge whether a business entity meets the sub - conditions for executing a business process indicated by the scenario description information. The execution conditions include multiple sub - conditions, and the second correlation attribute indicates that the business entity meets a preset number of the multiple sub - conditions.

[0084] According to an embodiment of the present disclosure, the sub - conditions for executing a business process can represent necessary condition information such as execution parameters and execution devices for executing a business process.

[0085] For example, the automatic payment process for executing the automatic payment scenario requires multiple devices such as cameras, lighting fixtures, and smart gateways as multiple sub-conditions. The sub-condition prompt template can be used to prompt the large language model to determine whether the business entity represented by the business entity node representation of the business entity represents a device required for executing the automatic payment process. When it is determined that the business entity "camera failure event" matches the device required for executing the automatic payment process and the preset number is the quantity "1", it is determined that the business entity "camera failure event" has a second relevance attribute.

[0086] According to an embodiment of the present disclosure, when the semantic similarity between the business entity and the scenario description entity extracted from the business scenario description information does not meet the preset similarity condition, the large language model can be used to process the scenario description information, the business entity, and the sub-condition prompt template to understand the influence degree of the business entity on the execution of the business process. Thus, it can be determined that the business entity with the second relevance attribute represents that the business entity meets a preset number of sub-conditions, and further determine that the business entity with the second relevance attribute is closely related to the specified business scenario. In this way, based on the scenario requirement intention representing the second relevance attribute in the corpus requirement intention, the number of obtained target entities can be further expanded, avoiding the omission of corpus entities closely related to the specified business scenario, thereby improving the richness of the sample corpus while improving the information integrity of the sample corpus, and further improving the accuracy of the sample corpus for answering questions for executing the specified business scenario.

[0087] According to an embodiment of the present disclosure, the sub-condition may include at least one of a first sub-condition and a second sub-condition.

[0088] The first sub-condition represents that the entity attribute of the business entity is used for at least one process task in the execution of the business process.

[0089] In one example, the process task may include executing a business decision task in the business scenario. For example, the scenario description information is the description content of the device management business scenario, and the business decision tasks of the device management business scenario may include device procurement decision tasks, device maintenance decision tasks, device layout adjustment decision tasks, etc. The sub-condition prompt template can construct a prompt word based on the following content: when the scenario description information does not involve the service life of the device, in order to execute the device procurement decision task or the device maintenance decision task, it is necessary to obtain the service life distribution of the devices in the system to execute the device procurement decision task or the device maintenance decision task, so as to avoid device aging failures and losses.

[0090] According to an embodiment of the present disclosure, by setting a sub-condition prompt template for controlling whether a large model determines that a business entity meets the sub-conditions of a business decision task, the large language model can determine whether a business entity meets any one of multiple first sub-conditions in an equipment management business scenario through the data processing tasks indicated by the sub-condition prompt template, and further determine whether the business entity can meet the second scenario relevance condition. In this way, it can be determined whether a business entity can be used to execute any one or more business decision tasks, so that the user can determine a large number of target entities with a relatively high degree of relevance to the specified business scenario from the knowledge graph through the scenario requirement attributes of the corpus demand intention, improving the richness of the sample corpus and the scenario coverage.

[0091] In one example, the process tasks can be multiple tasks with an execution logical relationship in a business scenario. For example, executing the equipment distribution adjustment process may include an equipment type query task, an equipment location query task, and an equipment quantity query task. The equipment number is not involved in the business scenario description information, but the equipment location corresponding to the equipment number can be determined based on the equipment number entity. Therefore, the equipment location query task can be executed based on the equipment number entity, and the equipment number entity can be determined to meet the first sub-condition. Thus, the sub-condition prompt template can be used to control the large language model to understand the execution logic of the equipment process tasks. Furthermore, under the condition that the relevant business entity is not involved in the business scenario description information, the large language model can more accurately determine the business entity related to the execution of the process tasks, supplement the integrity of multiple process tasks for executing the business process, make the execution of multiple process tasks of the business process smooth, and thereby improve the diversity of target entities and the diversity of sample corpus.

[0092] The second sub-condition indicates that the execution status data associated with the business entity is used to execute the business process. The execution status data can include data generated during the execution process such as the running duration of the equipment, the number of alarms, and the location distribution. The execution status data can be stored as variable status data in the entity attribute storage area related to the business entity, so as to generate sample corpus through the target business entity, enabling the sample corpus to real-time represent the attribute change situation of the target entity that matches the task requirements of the current specified business scenario, improving the accuracy of the sample corpus for describing or querying the equipment or events in the specified business scenario, and thereby improving the corpus quality of the sample corpus.

[0093] In one example, the business scenario description information does not involve the equipment brand. The intention of the business scenario description information is to analyze the operating status of the charging system, and the operating performance of equipment of different equipment brands in the system has a greater impact on the operating status. Therefore, the large language model can be prompted based on the sub-condition prompt template to understand the distribution of equipment brands and equipment corresponding to equipment brands, which will affect the operating status of the charging system. The large language model can thus determine that the business entities "Brand A" and "Brand B" meet the second sub-condition related to the specified charging system evaluation scenario.

[0094] It should be noted that the first sub-condition or the second sub-condition involved in the embodiments of the present disclosure may include multiple ones, and the sub-condition prompt template can be used to prompt the large language model to determine whether the business entity satisfies which sub-conditions involved in executing the business process in the business scenario description information, and determine whether the business entity has the second relevance attribute by determining the number of sub-conditions that the business entity satisfies for executing the business process.

[0095] Figure 4 A flowchart for generating a knowledge graph according to an embodiment of the present disclosure is schematically shown.

[0096] like Figure 4 As shown, an indication map may be generated based on operations S401 to S405.

[0097] In operation S401, business entities may be acquired from an initial knowledge graph, and scenario description entities may be extracted from business scenario description information.

[0098] In operation S402, it is determined whether the business entity and the scene description entity meet the preset similarity condition. For example, a text matching algorithm can be used to match the business entity and the scene description entity to obtain a first detection result that represents whether the semantic similarity condition is met. Alternatively, a semantic similarity judgment can be performed on the business entity and the scene description entity based on a deep learning model for processing natural language to obtain a first detection result that represents whether the semantic similarity condition is met. In the case where the first detection result indicates that the semantic similarity condition is met, it can be determined that the judgment result of operation S402 is yes. In the case where the first detection result indicates that the semantic similarity condition is not met, it can be determined that the judgment result of operation S402 is no.

[0099] If the result of operation S402 is yes, operation S403 may be performed to determine whether the business description entity has a first relevance attribute. If the result of operation S402 is no, operation S404 may be performed to process the scenario description information, the business entity, and the sub-condition prompt template using a large language model to determine whether the business entity satisfies a preset number of sub-conditions for executing the business process, and obtain a second relevance attribute of the business entity.

[0100] In operation S405, based on the first relevance attribute or the second relevance attribute corresponding to each of the multiple business entities as a new entity attribute, the entity attributes of the corresponding business entity nodes in the initial knowledge graph are updated, so that the entity attributes of the business entity nodes in the updated knowledge graph include scenario relevance attributes, so as to determine the target entity nodes that meet the corpus requirement intent from the knowledge graph according to the scenario requirement attributes in the corpus requirement intent, and achieve strong correlation attributes between the sample corpus and the execution conditions and execution processes of the specified business scenarios involved in the corpus requirement intent, thereby improving the quality of the sample corpus.

[0101] According to an embodiment of the present disclosure, the preset knowledge graph also includes edges between business entity nodes, and the edges represent attribute relationships between different business entities. Based on the knowledge graph, the target entity is processed using a large language model to obtain a sample corpus, including: using a large language model to process the target entity, the associated entity, and the corpus constraints in the corpus requirement intent to obtain the sample corpus.

[0102] According to an embodiment of the present disclosure, an associated entity node in a knowledge graph represents the associated entity, and the associated entity node has an edge relationship with a target entity node representing the target entity. A corpus entity node may be a target entity or an associated entity.

[0103] According to an embodiment of the present disclosure, an associated entity node having an edge relationship with a target entity node may represent a business entity having an attribute relationship such as "attribution" or "location" with the target entity. Alternatively, the associated entity node may also include a business entity node having a multi-level edge relationship with the target entity node, and the associated entity may be a business entity having a multi-level attribute relationship with the target entity. By utilizing a large language model to process target entities and associated entities having an edge relationship, the business entities with a high degree of relevance to the specified scenario of the corpus demand intent and the attribute relationship between business entities can be understood based on the semantic understanding ability of the large language model to perform natural language generation tasks, so that the sample corpus can more accurately represent the corpus demand intent, and the semantic attributes of the generated sample corpus are richer, thereby improving the corpus quality of the sample corpus.

[0104] According to an embodiment of the present disclosure, corpus constraint conditions are used to constrain the semantic expression modes of corpus entities in a sample corpus. The corpus entities are target entities represented by target entity nodes or associated entities represented by associated entity nodes. The corpus entities can be understood as target entities or associated entities. By using the corpus constraint conditions to control the large language model to combine the corpus entities to generate a sample corpus that matches the semantic expression mode, the sample corpus can perform semantic expression according to the expression mode required by the corpus demand intention, so that the sample corpus can perform natural language generation on target entities and associated entities with a relatively high relevance to the specified business scenario based on the semantic expression mode required by the corpus demand information, thereby improving the matching degree between the sample corpus and the training task requirements or test task requirements of the large model, and improving the flexibility and adaptability of the expression mode of the sample corpus.

[0105] It should be noted that the corpus constraint conditions can constrain the sample corpus to perform natural language expression based on any type of semantic expression mode such as interrogative expression mode, rhetorical question expression mode, negative expression mode, parallelism expression mode, etc. The embodiments of the present disclosure do not limit the specific types of semantic expression modes constrained by the corpus constraint conditions.

[0106] According to an embodiment of the present disclosure, using the large language model to process the corpus constraint conditions in the target entity, associated entity, and the corpus demand intention includes: using the large language model to process the target entity, associated entity, corpus constraint conditions, and a corpus constraint prompt template representing the corpus constraint conditions.

[0107] According to an embodiment of the present disclosure, the corpus constraint prompt template can be used as a prompt to improve the large language model to semantically combine the target entity and the associated entity according to the semantic expression mode represented by the corpus constraint conditions, so as to obtain a sample corpus that can perform natural language expression according to the corpus constraint conditions in the corpus demand intention, thereby improving the accuracy, diversity, and matching degree with the corpus demand of the expression mode of the sample corpus.

[0108] According to an embodiment of the present disclosure, the corpus constraint conditions include at least one of a first constraint condition, a second constraint condition, and a third constraint condition.

[0109] The first constraint condition indicates constraining the corpus entity based on a negative expression mode. For example, the corpus entity can be constrained based on negative expression modes such as "not" and "non" to obtain a sample corpus.

[0110] In one example, the target entity and the associated entity can include "AA brand" and "device". The sample corpus obtained by constraining the corpus entity based on the negative expression mode can be "What are the devices that are not of the AA brand?"

[0111] In one example, the attribute relationships between different corpus entities can be constrained based on negative expressions such as "no" and "not" to obtain the sample corpus. For example, the target entity and the associated entity can include "charging system" and "system type". The sample corpus can be "What are the devices that do not belong to the charging system?"

[0112] The second constraint condition means constraining the corpus entity based on the time requirement information in the corpus requirement information; semantic reasoning can be performed based on the time requirement information in the corpus requirement information to determine the actual required time needed for the corpus requirement intention, and the time attribute words related to the corpus entity in the sample corpus can be constrained based on the actual required time represented by the corpus requirement intention, so as to add the time dimension attribute to the corpus entity in the sample corpus.

[0113] For example, the time requirement information is "today, yesterday, weekend, weekday, last week, last quarter, last month, first half of the year, second half of the year, last year", etc., or the time requirement information can also be a specific time, such as "2023-06-20 00:00:00". Based on the time requirement intention representing "the previous year of the current time" in the corpus requirement intention, combined with the corpus entity and the current time being 2025, the sample corpus "What is the number of devices in 2024?" can be generated based on the time attribute word "2024" and the corpus entity "device". Among them, the corpus constraint condition can be used to represent the time requirement intention in the corpus requirement intention.

[0114] The third constraint condition can mean constraining the expression mode of the corpus entity based on the pre-designed calculation logic requirement in the corpus requirement intention. The pre-designed calculation logic can represent any type of algorithm or model calculation process. The pre-designed calculation logic can, for example, represent calculation processes such as summation, calculating the average value, sorting, grouping, determining the maximum value or the minimum value of the computing device.

[0115] For example, the pre-designed calculation logic in the corpus requirement intention can be expressed as the "summation" algorithm. The corpus entities "A brand" and "device" can be constrained based on the semantic expression mode related to the "summation" algorithm, so that the large language model can be controlled based on the corpus constraint condition to generate the sample corpus "How many devices in total are of brand A?"

[0116] Another example is that the pre-designed calculation logic in the corpus requirement intention can be expressed as the "sorting" algorithm. The corpus entities "A brand", "device type" and "device" can be constrained based on the semantic expression mode related to the "sorting" algorithm, so that the large language model can be controlled based on the corpus constraint condition to generate the sample corpus "What are the top 3 device types of devices of brand A?"

[0117] According to an embodiment of the present disclosure, the corpus constraint condition may further include other constraint conditions characterizing in addition to the above first to third constraint conditions. For example, the corpus constraint condition may further include the content shown in Table 1 below.

[0118] Table 1

[0119]

[0120] According to an embodiment of the present disclosure, the sample corpus includes an upper-context sample corpus and a lower-context sample corpus having a context semantic relationship, and the corpus constraint condition may further include at least one of the following: replacing at least one corpus entity based on a preset pronoun in the lower-context sample corpus; constraining the corpus entity in the lower-context sample corpus based on an interrogative expression.

[0121] According to an embodiment of the present disclosure, the upper-context sample corpus and the lower-context sample corpus may include the question corpus and the answer corpus in the Q&A corpus, where the upper-context sample corpus and the lower-context sample corpus may be the question corpus and the answer corpus in the same Q&A pair, or the upper-context sample corpus and the lower-context sample corpus may also be the question corpus and the answer corpus in different Q&A pairs having a context semantic relationship. Alternatively, the upper-context sample corpus and the lower-context sample corpus may also include two different question corpora having a context relationship. The embodiment of the present disclosure does not limit the specific context semantic relationship between the upper-context sample corpus and the lower-context sample corpus.

[0122] According to an embodiment of the present disclosure, replacing at least one corpus entity based on a preset pronoun in the lower-context sample corpus, the preset pronoun may include pronouns such as "which", "those", "this", etc. The preset pronoun may replace service entities such as devices and events in the upper-context sample corpus to simulate the expression mode of people in a real communication scenario.

[0123] For example, the upper-context sample corpus is: "What are the devices of the charging system?" The lower-context sample corpus may be "Q2: How many of these devices are faulty?"

[0124] Another example is that the upper-context sample corpus is: "Q1: What are the devices in normal operation?", and the lower-context sample corpus may be, for example, "Q2: Which ones have the brand A?" Among them, "which" can be used as a preset pronoun to replace the "devices in normal operation" in "Q1". The sample corpus generated based on the corpus constraint condition of replacing at least one corpus entity based on a preset pronoun in the lower-context sample corpus can make the multiple sample corpora having a context semantic relationship use the preset pronoun in the lower-context sample corpus to approximate the expression mode of people replacing the upper-context keyword with a vague pronoun in a real natural language communication scenario, improving the authenticity and naturalness of the overall sample corpus.

[0125] According to an embodiment of the present disclosure, the corpus constraint condition may further include constraining corpus entities in the following sample corpus based on the follow-up expression manner.

[0126] For example, the above sample corpus is: "Q1: Where are the locations where the littering events occurred today", and the following sample corpus may be, for example: "Q2: What other events occurred besides this event for K100013".

[0127] For another example, the above sample corpus is: "Q1: What are the devices of the charging system?", and the following sample corpus may be, for example: "Q2: How many of these devices are faulty?". Among them, "these devices" can be used as a preset referring pronoun to replace "the devices of the charging system" in "Q1", and "these devices" can also be questioned to implement a follow-up form of the context sample corpus, so that the sample corpus can express the follow-up form existing in natural language communication based on the context semantic relationship, thereby improving the richness of the sample corpus and the communication authenticity close to the real communication scenario, and improving the quality of the sample corpus.

[0128] According to an embodiment of the present disclosure, the corpus constraint condition may also constrain the semantic expression manner of the above corpus or the following corpus through other types of constraint conditions. For example, the corpus constraint condition is based on the content shown in Table 2-1 and Table 2-2 below.

[0129] Table 2-1

[0130]

[0131] Table 2-2

[0132]

[0133] In one example, a large language model can be used to process the target entity, associated entity, and corpus constraint condition represented by the target entity node, and the above sample corpus is obtained: "Q1: At which pile numbers did the congestion events occur today". Based on the corpus entity "pile number" and the attribute relationship "location" in the above sample corpus, query the associated entity node with an edge relationship with the target entity node corresponding to the corpus entity "pile number" in the knowledge graph, and the associated entity represented by this associated entity node is "fault". Use the large language model to process the associated entity "fault" and the corpus constraint prompt template representing the composite constraint condition and the context semantic inheritance constraint condition, and obtain the following sample corpus "Q2: What devices with faults and brand A are there at these locations".

[0134] It should be noted that for the above-sample corpus and the below-sample corpus with context semantic relationships, the sample corpus can also be constrained based on any one or more of the first to third constraint conditions described in the above embodiments, as well as the constraint conditions shown in Table 1 of the above embodiments. The embodiments of the present disclosure will not be elaborated herein.

[0135] Figure 5 Schematically shows a flowchart of a method for generating a sample corpus for a large model according to another embodiment of the present disclosure.

[0136] As Figure 5 shown, the method for generating a sample corpus for a large model in this embodiment includes operations S501 to S507.

[0137] In operation S501, the intent of the corpus requirement information is recognized to obtain the corpus requirement intent. The corpus requirement intent may include scenario requirement attributes, sample corpus requirement types, corpus constraint conditions, etc. The sample corpus requirement type may include a single-round corpus requirement type for generating a corpus for a single-round corpus question, or may also include a multi-round corpus requirement type for a multi-round inquiry and communication scenario.

[0138] In operation S502, based on the sample corpus requirement type, it is determined whether the requirement intent is a single-round sample corpus generation requirement. If the judgment result in operation 502 is yes, operation S503 can be executed to determine the corpus constraint conditions for the single-round sample corpus. The single-round corpus constraint conditions can be used to constrain the semantic expression mode of the single-round sample corpus for the corpus entity. Then operation S504 is executed to determine the corpus entity for generating the single-round sample corpus. For example, the target entity node can be determined from the knowledge graph based on the scenario requirement attributes, and the associated entity nodes having an edge relationship with the target entity node. And the business entities corresponding to the target entity node and the associated entity nodes are used as the corpus entities.

[0139] In operation S505, input prompt information is generated according to the corpus entity, the corpus constraint conditions, and the corpus constraint prompt template. The input prompt information can be used to input into a large language model to prompt the large language model to constrain the corpus entity according to the semantic expression mode characterized by the constraint conditions and output a single-round sample corpus. In operation S507, the single-round sample corpus is pushed to the user.

[0140] If the judgment result in operation 502 is negative, operation S508 can be executed to determine the corpus constraint conditions for the multi-round sample corpus. The corpus constraint conditions for the multi-round sample corpus can include irrelevance constraint conditions, context semantic inheritance constraint conditions, and so on. In operation S509, based on the corpus entities of the above sample corpus, the corpus entity nodes of the below sample corpus are determined in the knowledge graph. Then operations S505 to S506 are executed to generate the multi-round sample corpus. In operation S507, the multi-round sample corpus is pushed to the user.

[0141] Figure 6 Schematically shows a schematic diagram of a method for generating a sample corpus for a large model according to another embodiment of the present disclosure.

[0142] As Figure 6 shown, the knowledge graph 600 may include multiple business entity nodes and the edge relationships between multiple business entity nodes. In the case where the target entity matching the scenario requirement attribute is the business entity node "K001", and the corpus constraint condition can be the context inheritance constraint condition, "K001" is used as the in-degree central node of the edge relationship, and the associated entities are determined to be "pile number", "event", "vehicle speeding", and "spillage" respectively using the business entity "K001". By processing the target entity, the associated entities, and the corpus constraint conditions using a large language model, the multi-round sample corpus "Q1: How many spillage events does K001 have, Q2: How many vehicle speeding events does K001 have." can be obtained.

[0143] In the case where the target entity matching the scenario requirement attribute is the business entity node "K001", and the corpus constraint conditions can be the context inheritance constraint condition and the composite condition, "K001" is used as the in-degree central node of the edge relationship, and the associated entities are determined to be "pile number", "event", "congestion", and "spillage" respectively using the business entity "K001". By processing the target entity, the associated entities, and the corpus constraint conditions using a large language model, the multi-round sample corpus "Q1: How many spillage events does K001 have; Q2: How long has K001 been congested." can be obtained.

[0144] Based on the method for generating a sample corpus for a large model provided by the embodiments of the present disclosure, embodiments of the present disclosure also provide a training method for a large model and a testing method for a large model.

[0145] Figure 7 Schematically shows a flowchart of a training method for a large model according to an embodiment of the present disclosure.

[0146] As Figure 7 shown, the training method for the large model includes operations S710 to S720.

[0147] In operation S710, a sample corpus is obtained.

[0148] According to an embodiment of the present disclosure, the sample corpus is determined based on the method for generating a sample corpus for a large model provided by the embodiment of the present disclosure.

[0149] In operation S720, an initial large model is trained based on the sample corpus to obtain a trained large model.

[0150] According to an embodiment of the present disclosure, diverse and high-quality sample corpora can be generated based on the method for generating a sample corpus provided in any of the above embodiments, so as to fine-tune the parameters of the initial large model based on the sample corpus. The obtained trained large model can more accurately generate response texts for problem information related to a specified business scenario, so as to improve the response accuracy in the specified business scenario.

[0151] Figure 8 A flowchart of a method for testing a large model according to an embodiment of the present disclosure is schematically shown.

[0152] As Figure 8 shown, the method for testing a large model includes operations S810 to S830.

[0153] In operation S810, a sample corpus is obtained.

[0154] According to an embodiment of the present disclosure, the sample corpus is determined based on the method for generating a sample corpus for a large model provided by the embodiment of the present disclosure.

[0155] In operation S820, the sample corpus is processed by the large model to obtain an output text.

[0156] In operation S830, the test result of the large model is determined based on the output text.

[0157] According to an embodiment of the present disclosure, a large model can be used to process diverse and high-quality sample corpora generated based on the method for generating a sample corpus provided in any of the above embodiments to obtain an output text. An output text evaluation result is obtained by evaluating the text quality of the output text. Thus, the output text evaluation result can be used as the test result. So as to more accurately understand the response accuracy and response quality of the large model for the problem text of the specified business scenario based on the test result, and improve the accuracy of performance evaluation of the large model for executing the business scenario.

[0158] Figure 9 A block diagram of a device for generating a sample corpus for a large model according to an embodiment of the present disclosure is schematically shown.

[0159] As Figure 9 shown, the device 900 for generating a sample corpus for a large model includes: an identification module 910, a first determination module 920, and a sample corpus acquisition module 930.

[0160] An identification module 910 for identifying the intention of the corpus requirement information to obtain the corpus requirement intention.

[0161] A first determination module 920 for determining, from the business entity nodes of the knowledge graph, a target entity node that matches the scenario requirement attribute in the corpus requirement intention, where the business entity node represents a business entity extracted from business basic data, the scenario relevance attribute of the business entity represents the matching degree between the business entity and the execution conditions of the specified business scenario, and the scenario requirement attribute represents the requirement intention for the matching degree.

[0162] A sample corpus obtaining module 930 for processing the target entity node by using a large language model to obtain a sample corpus.

[0163] According to an embodiment of the present disclosure, the scenario relevance attribute includes a first relevance attribute, and the first relevance attribute is determined based on the following operations: extracting scenario description entities from the business scenario description information; determining a first detection result based on the semantic similarity between the scenario description entities and the business entity; and determining the scenario relevance attribute of the business entity as the first relevance attribute when the first detection result indicates that the semantic similarity meets a preset similarity condition.

[0164] According to an embodiment of the present disclosure, the scenario relevance attribute includes a second relevance attribute, and the second relevance attribute is determined based on the following operations: processing the scenario description information, the business entity, and the sub-condition prompt template by using a large language model to obtain the second relevance attribute of the business entity, where the sub-condition prompt template is used to control the large language model to determine whether the business entity meets the sub-conditions for executing the business process indicated by the scenario description information, the execution conditions include multiple sub-conditions, and the second relevance attribute represents that the business entity meets a preset number of the multiple sub-conditions.

[0165] According to an embodiment of the present disclosure, the sub-conditions include at least one of the following: the entity attribute of the business entity is used for at least one process task in the business process; the execution status data associated with the business entity is used for the business process.

[0166] According to an embodiment of the present disclosure, the knowledge graph further includes edges between nodes, and the edges represent the attribute relationships between different business entities; the sample corpus obtaining module 930 includes a sample corpus obtaining unit.

[0167] A sample corpus acquisition unit is configured to process a target entity node, associated entity nodes having edge relationships with the target entity node in a knowledge graph, and corpus constraint conditions in a corpus requirement intention using a large language model to obtain a sample corpus, where the corpus constraint conditions are used to constrain the semantic expression mode of corpus entities in the sample corpus, and the corpus entities are the target entity represented by the target entity node or the associated entity represented by the associated entity node.

[0168] According to an embodiment of the present disclosure, the corpus constraint conditions include at least one of the following: constraining the corpus entity based on a negative expression mode; constraining the corpus entity based on time requirement information in the corpus requirement information; constraining the expression mode of the corpus entity based on a pre-designed calculation logic requirement in the corpus requirement intention.

[0169] According to an embodiment of the present disclosure, the sample corpus includes an upper-context sample corpus and a lower-context sample corpus having a context semantic relationship, and the corpus constraint conditions include at least one of the following: replacing at least one corpus entity based on a preset pronoun in the second sample corpus; constraining the corpus entity in the second sample corpus based on a follow-up question expression mode.

[0170] According to an embodiment of the present disclosure, the sample corpus acquisition unit includes a processing subunit.

[0171] The processing subunit is configured to process the target entity represented by the target entity node, the associated entity represented by the associated entity node, the corpus constraint conditions, and a corpus constraint prompt template representing the corpus constraint conditions using a large language model.

[0172] Figure 10 A block diagram of a training device for a large model according to an embodiment of the present disclosure is schematically shown.

[0173] As Figure 10 shown, the training device 1000 for a large model includes: a first acquisition module 1010 and a training module 1020.

[0174] The first acquisition module 1010 is configured to acquire a sample corpus, where the sample corpus is determined based on the sample corpus generation method for a large model provided by an embodiment of the present disclosure.

[0175] The training module 1020 is configured to train an initial large model based on the sample corpus to obtain a trained large model.

[0176] Figure 11 A block diagram of a testing device for a large model according to an embodiment of the present disclosure is schematically shown.

[0177] As Figure 11 shown, the testing device 1100 for a large model includes: a second acquisition module 1110, an output text acquisition module 1120, and a second determination module 1130.

[0178] The second acquisition module 1110 is configured to acquire a sample corpus, where the sample corpus is determined based on the sample corpus generation method for a large model provided in the embodiments of the present disclosure.

[0179] The output text obtaining module 1120 is configured to process the sample corpus by using a large model to obtain an output text.

[0180] The second determination module 1130 is configured to determine the test result of the large model based on the output text.

[0181] Figure 12 A structural block diagram of an intelligent agent of artificial intelligence according to an embodiment of the present disclosure is schematically shown.

[0182] In an embodiment of the present disclosure, as Figure 12 shown, the AI intelligent agent 1200 may include an input module 1210, a processing module 1220, and an output module 1230.

[0183] The input module 1210 is configured to receive input information;

[0184] The processing module 1220 is configured to determine a target task based on the input information received by the input module, determine a large language model based on the target task, execute the sample corpus generation method for a large model provided in the embodiments of the present disclosure by calling the large language model, or obtain output information by calling a large model to execute the training method or the test method of the large model provided in the embodiments of the present disclosure;

[0185] The output module 1230 is configured to output the output information obtained by the processing module.

[0186] According to an embodiment of the present disclosure, the input module 1210 is responsible for receiving or sensing information such as queries, requests, instructions, signals, or data from the outside world (such as a user or an external environment), and converting it into a format that the AI intelligent agent 1200 can understand and process. The input module 1210 is the primary link for the AI intelligent agent 1200 to interact with the outside world, enabling the AI intelligent agent 1200 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.

[0187] In an example, the input module 1210 may input the corpus requirement information, target entity, etc. described above.

[0188] In an example, the processing module 1220 is the core support for the AI intelligent agent 1200 to process complex tasks. The processing module 1220 may execute the sample corpus generation method, the training method, or the test method of the large model described above.

[0189] In the example, the performance of the processing module 1220 can be closely related to the large model on which the AI agent 1200 is based. To fully utilize the capabilities of the large model, the internal structure of the processing module 1220 can be designed to be highly configurable and extensible to handle various different types of tasks and requirements in real-world scenarios.

[0190] In the example, after the AI agent 1200 obtains the demand speech, the processing module 1220 can use the large language model to process the target entity to obtain sample corpus and pass the sample corpus to the output module 1230.

[0191] It can be understood that although the large language model has excellent language understanding and generation capabilities, like humans, without any tools, the tasks it can solve are very limited. When the AI agent 1200 is given the ability to call tools, it can achieve tasks such as performing mathematical operations with the help of a calculator, performing data analysis with the help of Python, and obtaining weather forecasts with the help of a search engine.

[0192] In the example, the output module 1230 can output the sample corpus or test results described above.

[0193] The AI agent 1200 according to the embodiments of the present disclosure can simply and effectively improve the degree of intelligence and enhance flexibility and versatility.

[0194] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0195] According to the embodiments of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.

[0196] According to the embodiments of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.

[0197] According to the embodiments of the present disclosure, a computer program product includes a computer program, and the computer program implements the method as described above when executed by a processor.

[0198] Figure 13A schematic block diagram of an example electronic device for a method of generating sample corpus for a large model, a method of training a large model, and a method of testing a large model, which can be used to implement embodiments of the present disclosure, is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0199] As Figure 13 shown, the device 1300 includes a computing unit 1301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1302 or a computer program loaded from a storage unit 1308 into a random access memory (RAM) 1303. In the RAM 1303, various programs and data required for the operation of the device 1300 can also be stored. The computing unit 1301, the ROM 1302, and the RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.

[0200] A plurality of components in the device 1300 are connected to the I / O interface 1305, including: an input unit 1306, such as a keyboard, a mouse, etc.; an output unit 1307, such as various types of displays, speakers, etc.; a storage unit 1308, such as a magnetic disk, an optical disk, etc.; and a communication unit 1309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1309 allows the device 1300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0201] The computing unit 1301 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1301 executes the various methods and processes described above, such as the sample corpus generation method for large models, the training method for large models, and the testing method for large models. For example, in some embodiments, the sample corpus generation method for large models, the training method for large models, and the testing method for large models can be implemented as a computer software program that is tangibly included in a machine-readable medium, such as the storage unit 1308. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1300 via the ROM 1302 and / or the communication unit 1309. When the computer program is loaded into the RAM 1303 and executed by the computing unit 1301, one or more steps of the sample corpus generation method for large models, the training method for large models, and the testing method for large models described above can be executed. Alternatively, in other embodiments, the computing unit 1301 can be configured to execute the sample corpus generation method for large models, the training method for large models, and the testing method for large models by any other suitable means (e.g., by means of firmware).

[0202] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0203] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0204] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0205] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0206] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0207] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0208] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0209] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A sample corpus generation method for a large model, comprising: Perform intent recognition on the corpus demand information to obtain the corpus demand intent; Determine, from a plurality of business entities, a target entity that matches a scenario requirement attribute in the corpus requirement intent, wherein the business entity is extracted from business basic data, the scenario relevance attribute of the business entity represents a matching degree between the business entity and an execution condition for executing a specified business scenario, and the scenario requirement attribute represents a requirement intent for the matching degree; The target entity is processed using a large language model to obtain a sample corpus.

2. The method according to claim 1, wherein: The scene relevance attribute comprises a first relevance attribute, and the first relevance attribute is determined based on the following operations: Extracting scenario description entities from business scenario description information; Determining a first detection result based on the semantic similarity between the scene description entity and the business entity; When the first detection result indicates that the semantic similarity satisfies a preset similarity condition, the scenario relevance attribute of the business entity is determined to be the first relevance attribute.

3. The method according to claim 1 or 2, wherein: The scene relevance attribute includes a second relevance attribute, and the second relevance attribute is determined based on the following operation: The scenario description information, the business entity and the sub-condition prompt template are processed using a large language model to obtain a second correlation attribute of the business entity, wherein the sub-condition prompt template is used to control the large language model to determine whether the business entity satisfies the sub-condition for executing the business process indicated by the scenario description information, the execution condition includes multiple sub-conditions, and the second correlation attribute represents that the business entity satisfies a preset number of the multiple sub-conditions.

4. The method according to claim 3, wherein: The sub-conditions include at least one of the following: The entity attribute of the business entity is used to execute at least one process task in the business process; The execution status data associated with the business entity is used to execute the business process.

5. The method according to claim 1, wherein: The preset knowledge graph includes a business entity node representing the business entity, and edges between a plurality of the business entity nodes, wherein the edges represent attribute relationships between different business entities; The using of the large language model to process the target entity to obtain the sample corpus comprises: Based on the knowledge graph, the large language model is used to process the target entity, the associated entity and the corpus constraints in the corpus demand intention to obtain the sample corpus, wherein the associated entity node in the knowledge graph represents the associated entity, the associated entity node has an edge relationship with the target entity node representing the target entity, the corpus constraints are used to constrain the semantic expression of the corpus entity in the sample corpus, and the corpus entity is the target entity or the associated entity.

6. The method according to claim 5, wherein: The corpus constraint condition includes at least one of the following: constraining the corpus entity based on a negative expression; Constraining the corpus entity based on time requirement information in the corpus requirement information; The expression mode of the corpus entity is constrained based on the preset computing logic requirements in the corpus requirement intention.

7. The method according to claim 5 or 6, wherein: The sample corpus includes a preceding sample corpus and a following sample corpus having a contextual semantic relationship, and the corpus constraint condition includes at least one of the following: Replacing at least one of the corpus entities in the second sample corpus based on a preset pronoun; The corpus entities in the second sample corpus are constrained based on the question expression.

8. The method according to claim 5, wherein: The use of the large language model to process the target entity, the associated entity, and the corpus constraint conditions in the corpus requirement intent includes: The target entity, the associated entity, the corpus constraint condition, and a corpus constraint prompt template representing the corpus constraint condition are processed using the large language model.

9. A large model training method, comprising: Acquire a sample corpus, wherein the sample corpus is determined based on the method according to any one of claims 1 to 8; and An initial large model is trained based on the sample corpus to obtain a trained large model.

10. A large model testing method comprising: Acquire a sample corpus, wherein the sample corpus is determined based on the method according to any one of claims 1 to 8; Processing the sample corpus using the large model to obtain output text; and A test result of the large model is determined based on the output text.

11. A sample corpus generation device for a large model, comprising: The recognition module is used to identify the intent of the corpus demand information and obtain the corpus demand intent; A first determination module is used to determine, from a plurality of business entities, a target entity that matches a scenario requirement attribute in the corpus requirement intent, wherein the business entity is extracted from business basic data, the scenario relevance attribute of the business entity represents a matching degree between the business entity and an execution condition for executing a specified business scenario, and the scenario requirement attribute represents a requirement intent for the matching degree; The sample corpus acquisition module is used to process the target entity node using a large language model to obtain a sample corpus.

12. A large model training device, comprising: a first acquisition module, configured to acquire a sample corpus, wherein the sample corpus is determined based on the method of any one of claims 1 to 8; and The training module is used to train the initial large model based on the sample corpus to obtain a trained large model.

13. A large-scale model testing device, comprising: A second acquisition module, configured to acquire a sample corpus, wherein the sample corpus is determined based on the method according to any one of claims 1 to 8; An output text obtaining module, used to process the sample corpus using a large model to obtain an output text; and The second determination module is used to determine the test result of the large model based on the output text.

14. An artificial intelligence agent, comprising: An input module, used for receiving input information; A processing module, configured to determine a target task based on the input information received by the input module, determine a large language model based on the target task, and obtain output information by calling the large language model to execute the method of any one of claims 1 to 8, or by calling the large model to execute the method of claim 9 or 10; An output module is used to output the output information obtained by the processing module.

15. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 10.

17. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Interaction method and device based on agent cooperation, agent and storage medium

    CN120654731A

  • Interaction method and device based on agent cooperation, agent, and storage medium

    CN120654731B