Methods, devices, and intelligent agents for generating corpus data based on large models

CN120851212BActive Publication Date: 2026-09-01BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511053125.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2026-09-01
Estimated Expiration
2045-07-29

AI Technical Summary

Benefits of technology

[0009] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods provided in embodiments of this disclosure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120851212B_ABST
    Figure CN120851212B_ABST
Patent Text Reader

Abstract

This disclosure provides a method, apparatus, and intelligent agent for generating corpus data based on large models, relating to the field of artificial intelligence technology, particularly deep learning, large models, and intelligent question answering. The method for generating corpus data based on large models includes: using multiple large role models to conduct dialogues on a preset topic to obtain the speech content of at least one large role model; using a specified large model to plan a dialogue strategy based on the speech content to obtain a dialogue strategy, which constrains the speech patterns of the large role models during the dialogue; and determining target corpus data related to the preset topic based on the target speech content generated by the large role models based on the dialogue strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to the fields of deep learning, large models, and intelligent question answering. Background Technology

[0002] Large Language Model (LLM) is a deep learning-based artificial intelligence model that can be used to understand natural language and generate corresponding corpus content. Summary of the Invention

[0003] This disclosure provides a method, apparatus, intelligent agent, electronic device, and storage medium for generating corpus data based on a large model.

[0004] According to one aspect of this disclosure, a method for generating corpus data based on large-scale models is provided, comprising: using multiple large-scale role models to conduct dialogues on a preset topic to obtain the speech content of at least one large-scale role model; using a specified large-scale model to plan a dialogue strategy based on the speech content to obtain a dialogue strategy, wherein the dialogue strategy is used to constrain the speech patterns of the large-scale role models during the dialogue process; and determining target corpus data related to the preset topic based on the target speech content generated by the dialogue generated by the large-scale role models based on the dialogue strategy.

[0005] According to another aspect of this disclosure, a corpus data generation device based on a large model is provided, comprising: a speech content acquisition module, used to conduct dialogues on a preset topic using multiple role large models to obtain speech content of at least one role large model; a dialogue strategy acquisition module, used to plan a dialogue strategy based on the speech content using a specified large model to obtain a dialogue strategy, wherein the dialogue strategy is used to constrain the speech patterns of the role large model during the dialogue process; and a target corpus data determination module, used to determine target corpus data related to the preset topic based on the target speech content generated by the dialogue based on the dialogue strategy of the role large model.

[0006] According to another aspect of this disclosure, an artificial intelligence agent is provided, comprising: an input module for receiving input information; a processing module for determining a target task based on the input information received by the input module, determining a role model and a specified model based on the target task, and obtaining output information by invoking the role model and the specified model to execute the method provided according to the embodiments of this disclosure; and an output module for outputting the output information obtained by the processing module.

[0007] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the methods provided in embodiments of this disclosure.

[0008] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the methods provided in embodiments of this disclosure.

[0009] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods provided in embodiments of this disclosure.

[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0012] Figure 1 The illustration schematically shows an exemplary system architecture for applying large-model-based corpus data generation methods and apparatus according to embodiments of the present disclosure;

[0013] Figure 2 A flowchart illustrating a method for generating corpus data based on a large model according to an embodiment of the present disclosure is shown schematically.

[0014] Figure 3 The diagram illustrates an application scenario of the corpus data generation method based on a large model according to an embodiment of the present disclosure.

[0015] Figure 4 The diagram illustrates an application scenario of a corpus data generation method based on a large model according to another embodiment of the present disclosure.

[0016] Figure 5 A flowchart illustrating a method for generating corpus data based on a large model according to another embodiment of the present disclosure is shown schematically.

[0017] Figure 6 A block diagram of a large-model-based corpus data generation apparatus according to an embodiment of the present disclosure is shown schematically.

[0018] Figure 7A schematic diagram illustrating the structure of an intelligent agent of artificial intelligence according to embodiments of the present disclosure; and

[0019] Figure 8 A schematic block diagram of an example electronic device is shown that can be used to implement the large-model-based corpus data generation method of the embodiments of this disclosure. Detailed Implementation

[0020] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0021] In the technical solution disclosed herein, the acquisition, storage, and application of user personal information comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and there is no violation of public order and good morals.

[0022] The inventors discovered that users can use large language models to generate various types of documents based on their needs. For example, the large language model can be used to generate press releases, scripts, and other text content by processing the user's requested text. Furthermore, the large language model can also be used to generate information in various formats such as tables and code to meet diverse user needs. However, the semantic quality of the corpus content generated by the large language model may be flawed, making it difficult to accurately meet the user's actual intent.

[0023] This disclosure provides a method, apparatus, intelligent agent, electronic device, and storage medium for generating corpus data based on large-scale models. The method for generating corpus data based on large-scale models includes: using multiple role-based large-scale models to conduct dialogues on a preset topic, obtaining the speech content of at least one role-based large-scale model; using a specified large-scale model to plan a dialogue strategy based on the speech content, obtaining a dialogue strategy, which is used to constrain the speech patterns of the role-based large-scale models during the dialogue; and determining target corpus data related to the preset topic based on the target speech content generated by the role-based large-scale models in the dialogue based on the dialogue strategy.

[0024] According to embodiments of this disclosure, by utilizing a designated large model to constrain the speaking patterns of at least one large role model during dialogues involving multiple large role models on a preset topic, dialogue strategy planning is performed based on the speaking content. This allows the large role model constrained by the dialogue strategy to output target speaking content with a high semantic match to the preset topic, avoiding situations where the semantics of the speaking content of multiple large role models deviates from the preset topic during dialogue, thus preventing the illusion that the content is semantically off-topic. Therefore, target corpus data related to the preset topic can be constructed based on the target speaking content, improving the semantic relevance between the target corpus data and the preset topic, thereby improving the quality of the target corpus data to meet the actual needs of the relevant target objects.

[0025] Figure 1 The illustration schematically shows an exemplary system architecture for applying large-model-based corpus data generation methods and apparatus according to embodiments of the present disclosure.

[0026] It is important to note that Figure 1 The examples shown are merely examples of system architectures applicable to embodiments of this disclosure, intended to help those skilled in the art understand the technical content of this disclosure. However, they do not imply that embodiments of this disclosure cannot be used in other devices, systems, environments, or scenarios. For instance, in another embodiment, an exemplary system architecture applicable to the large-model-based corpus data generation method and apparatus may include a terminal device. However, the terminal device can implement the large-model-based corpus data generation method and apparatus provided by embodiments of this disclosure without interacting with a server.

[0027] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0028] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).

[0029] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0030] Server 105 can be a server that provides various services, such as a backend management server that supports the content browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0031] It should be noted that the large-model-based corpus data generation method provided in this disclosure can generally be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the large-model-based corpus data generation device provided in this disclosure can also be disposed in the first terminal device 101, the second terminal device 102, or the third terminal device 103.

[0032] Alternatively, the large-model-based corpus data generation method provided in this disclosure can generally also be executed by server 105. Correspondingly, the large-model-based corpus data generation apparatus provided in this disclosure can generally be located in server 105. The large-model-based corpus data generation method provided in this disclosure can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the large-model-based corpus data generation apparatus provided in this disclosure can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.

[0033] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0034] Figure 2 A flowchart illustrating a method for generating corpus data based on a large model according to an embodiment of the present disclosure is shown.

[0035] like Figure 2 As shown, the corpus data generation method based on the large model includes operations S210~S230.

[0036] In operation S210, multiple large character models engage in dialogue on a preset topic to obtain the speech content of at least one large character model.

[0037] In operating S220, a dialogue strategy is planned based on the content of the speech using a specified large model to obtain the dialogue strategy.

[0038] In operation S230, the target speech content generated based on the dialogue strategy according to the role model is determined, and the target corpus data related to the preset topic is determined.

[0039] According to embodiments of this disclosure, a large role model can be constructed based on a large language model to generate speech content based on input information. Multiple large role models can be different models. For example, multiple large role models can be obtained by fine-tuning based on different training data, or multiple large role models can also generate speech content matching the prompt information indicated by different prompt words.

[0040] According to embodiments of this disclosure, the speech content may include text expressed in natural language, such as conversational text, intelligent customer service reply text, script dialogue text, etc. However, it is not limited to this; the speech content may also include content in any data format such as tables, code, and icons. Embodiments of this disclosure do not limit the specific data format of the speech content.

[0041] According to embodiments of this disclosure, using multiple large role models to conduct dialogue on a preset topic may include using multiple large role models to conduct interactive dialogue, so as to generate new speech content by using the large role models to perform semantic understanding of the preset topic and the already output speech content.

[0042] According to embodiments of this disclosure, a dialogue strategy is used to constrain the speaking patterns of large role models during a dialogue. The speaking pattern may include the speaking order of multiple large role models, the speaking frequency of a specified large role model during the dialogue, etc., to control the output content of multiple large role models during the dialogue.

[0043] In some embodiments, the speaking mode can also instruct multiple large role models to conduct group dialogues on a preset topic, thereby flexibly controlling the modes of multiple large role models to conduct group dialogues on a preset topic to further enhance the richness and diversity of speaking content, and thus improve the diversity and quality of the target corpus data.

[0044] In some embodiments, the speaking pattern represented by the dialogue strategy may also include content attributes such as thematic semantics, expression method, and expression style used to prompt the speaking content output by the large character model, so as to improve the quality of the speaking content output by the large character model based on the dialogue strategy.

[0045] In some embodiments, the designated large model can be a large language model different from the role large model. The designated large model can perform semantic understanding on one or more speech contents generated during the dialogue process to plan the dialogue strategy for multiple role large models. This enables the dialogue strategy to prompt multiple role large models to conduct high-quality dialogue based on the speech patterns, content attribute requirements, and other strategy content prompted by the dialogue strategy, and output target speech content that meets the quality requirements related to the preset topic.

[0046] According to embodiments of this disclosure, determining target corpus data based on target speech content may include fusing the target speech content and speech content output by multiple large role models to obtain target corpus data. Alternatively, the target corpus data may also include annotation information related to the target speech content and speech content. For example, the target corpus data may include multiple target speech contents and corresponding dialogue strategy content to facilitate convenient annotation of the target speech content and speech content, thereby improving the efficiency and quality of corpus data annotation.

[0047] In some embodiments, the target corpus data can be sample data used to test or train a specified language model. The target corpus data may include the output speech content of multiple role-based large models, target speech content, and policy content in the dialogue strategy. By fine-tuning the specified language model based on the target corpus data, the language model can learn to speak or engage in dialogue based on preset topics according to the dialogue strategy. This allows the trained language model to communicate or respond to questions from real users in specific scenarios such as after-sales service for specified products or press release writing on specified topics, thereby improving the quality of responses in those specific scenarios and enhancing the user experience.

[0048] In some embodiments, the preset topic may include press release topics, financial accounting project topics, discussion topic topics, etc., and the target corpus related to the preset topic may include press release content, financial accounting document content, dialogue content related to discussion topic content, and other corpus content related to specific demand scenarios. By planning dialogue strategies for the speaking content based on a specified large model during the dialogue process of the role large model, and by controlling multiple role large models to conduct dialogues according to the speaking models that match the demand conditions of the preset topic based on the dialogue strategy, the target speaking content can be further matched with the preset topic, and the target corpus data can meet the demand conditions of specific demand scenarios, thereby satisfying the actual demand intentions of the target audience and improving the quality of the generated target corpus data.

[0049] In some embodiments, using a designated large model to plan dialogue strategies based on speech content, the resulting dialogue strategy may include: using the demand information input by the target object as demand prompt information to prompt the designated large model to perform dialogue strategy planning by semantically understanding the speech content and preset topics, so that the generated dialogue strategy can control the role large model to output target speech content that meets the demand conditions represented by the demand information more accurately.

[0050] In some embodiments, using multiple large role models to conduct dialogue on a preset topic may include: based on an initial dialogue strategy, using multiple large role models to conduct dialogue on a preset topic.

[0051] According to embodiments of this disclosure, the initial dialogue strategy is determined by using a specified large model to plan the dialogue strategy based on a preset topic.

[0052] For example, before multiple large role models engage in dialogue, a designated large model can be used to process a preset topic and model capability descriptions of the capabilities of the multiple large role models, resulting in an initial dialogue strategy to control the dialogue between the multiple large role models on the preset topic. This allows the designated large model, with a full understanding of the individual capabilities of the multiple large role models, to control the speaking order, frequency, style, word count range of each output, and structured format requirements of the speaking content during the dialogue process by outputting the initial dialogue strategy. This ensures that the multiple large role models can engage in relatively accurate dialogue around the preset topic from the initial stage, so that the output speaking content meets the requirements of the strategy content representation in the initial dialogue strategy. This improves the accuracy of subsequent dialogue strategy generation and further enhances the quality of the target speaking content and the content quality of the target corpus data.

[0053] In some embodiments, the target object can input requirement configuration information related to each of the multiple role models in the interactive interface before engaging in dialogue, so as to specify the role model to plan the dialogue strategy based on the input requirement configuration information and output the initial dialogue strategy.

[0054] In one embodiment, the requirement configuration information may include a preset topic, role attributes of the role model, model name, model version identifier, and role responsibility description information. Role attributes may include domain expert, user role, commentator role, arbitrator role, etc. The role model corresponding to different role attributes can use the role responsibility description as prompts to process the contextual speech content during the dialogue process according to the role attribute requirements indicated in the role responsibility description, generating speech content that matches the role attribute requirements. For example, a role model with the arbitrator role attribute can arbitrate the viewpoints of multiple other role models' arguments on a preset discussion topic, outputting arbitrated speech content to regulate the dialogue development process among multiple role models.

[0055] In one embodiment, the requirement configuration information may further include a speaking style attribute, which can represent speaking styles such as concise, lively, technical, or dialogue-direction-guiding. The speaking style attribute can be used as style cues in the initial dialogue strategy to prompt the character's output speech content to match the speaking style represented by the corresponding speaking style attribute.

[0056] For example, the initial dialogue strategy can include speaking style prompts corresponding to the main character model A. These prompts can indicate the direction of the dialogue, and the main character model A can output statements such as "Please answer from a data perspective" or "Please give three suggestions." This guides other main character models to speak based on the guiding statements output by the main character model A, thereby improving the quality of their statements.

[0057] It should be noted that the initial dialogue strategy and the dialogue strategy can include the same or similar types of strategy content. For example, the strategy content in the dialogue strategy can also include style cues, so as to continuously control multiple large character models to conduct dialogue during the dialogue process.

[0058] In some embodiments, the dialogue strategy and the initial dialogue strategy can be based on at least one of parallel speaking mode, serial speaking mode and mixed speaking mode to control multiple large role models to conduct dialogue.

[0059] Parallel speaking mode is suitable for scenarios such as brainstorming or clashes of ideas that require active participation from multiple large role models. This mode can simultaneously invoke multiple large role models for parallel output. Furthermore, it can divide these large role models into several "speaking groups," with each group containing multiple large role models discussing and conversing on a specific topic. By controlling multiple large role models to execute the dialogue process in parallel using parallel speaking mode, multiple versions of the target dialogue content can be generated simultaneously. This facilitates the rapid generation of target corpus data based on the target dialogue content, reducing the waiting time for generating high-quality target corpus data.

[0060] In some embodiments, an initial dialogue strategy or dialogue policy including a parallel speaking mode can be triggered at a specific stage by specifying the host attributes of the large model simulation, so as to control each role's large model to output speaking content according to the context and preset topic during the dialogue. The output speaking content can be associated with the group number identifier in the parallel speaking mode process, and the target corpus content can be obtained by semantic fusion based on multiple group number identifiers and the corresponding group speaking content.

[0061] The sequential speaking mode is suitable for specific scenarios such as "deductive discussion" or "gap filling," where multiple large role models need to speak in sequence. The sequential speaking mode in dialogue strategies can instruct multiple large role models on the order and frequency of their speeches during a dialogue on a preset topic.

[0062] For example, a large-scale model with the commentator role attribute can only perform the speech generation task after the large-scale model with the expert role outputs its speech content and passes its self-check. This allows the large-scale model with the expert role to first output its answer, and then the large-scale model with the commentator role to supplement or question the expert's answer. This sequential speech pattern ensures complete contextual transmission between multiple dialogue contents during the conversation, improves semantic coherence and consistency among multiple target speech contents, and enhances the quality of the target corpus data.

[0063] Hybrid speaking mode can represent a dialogue process involving multiple large role models that combines parallel and sequential speaking modes. For example, multiple initial viewpoints can be represented by controlling the output of multiple large role models in parallel speaking mode. Then, by scoring the quality of the output or target output or through human interaction, sequential speaking mode can be triggered to control multiple large role models to conduct in-depth discussions on specific viewpoints related to a preset topic, thereby generating new target output.

[0064] In some embodiments, the dialogue strategy planning based on the speech content using a specified large model may further include: performing semantic understanding on the speech content using the specified large model to obtain a content quality understanding result; adjusting the speech weights of the role large model corresponding to the speech content based on the content quality understanding result to obtain the speech weights for the role large model; and determining the dialogue strategy based on the speech weights.

[0065] According to embodiments of this disclosure, the content quality understanding result can represent a quality evaluation result of the spoken content. For example, the content quality understanding result may include scores for any one or more dimensions of quality indicators, such as semantic logical coherence, semantic consistency, the relevance of the spoken content to the topic, and the quality of the speaking style. The scores for these quality indicators can serve as the quality evaluation result.

[0066] In some embodiments, a trained evaluation big model can be used to perform semantic understanding of one or more statements based on evaluation cues to obtain content quality understanding results. The trained big model can be a different designated big model from the moderator big model used for dialogue strategy planning, thereby improving the accuracy of content quality understanding results and further enhancing the effectiveness and real-time performance of dialogue strategies.

[0067] According to embodiments of this disclosure, the speaking weight characterizes the expected level of participation of a role model in a dialogue. For example, the speaking weight represents the weight of the corresponding role model's speaking frequency, word count limit, speaking order, and other attributes related to the speaking content. A higher speaking weight for a role model indicates that the role model will output more words in its speech or speak more frequently in the dialogue, thereby increasing the role model's participation. Correspondingly, reducing the speaking weight of a specific role model can reduce the word count of its speech or the frequency of its speech in the dialogue.

[0068] In some embodiments, the speech weight can also be used to indicate the level of attention that multiple role models participating in the dialogue should pay to the speech content output by the role model corresponding to the speech weight. For example, the prompt "Pay close attention to the speech content of role model A regarding artificial intelligence technology" can be input to role model B, and the corresponding reply content can be output.

[0069] In some embodiments, the higher the speaking weight of the large role model, the earlier the speaking order of that large role model can be indicated. This is so that the speaking content output by the large role model with higher speaking weight can guide other large role models to output high-quality content based on the context, thereby further improving the corpus content quality of the target speaking content generated by multiple large role models in dialogue according to the dialogue strategy.

[0070] In some embodiments, the content quality understanding result includes topic relevance information, which characterizes the degree of semantic relevance between the spoken content and a preset topic. The degree of semantic relevance can indicate whether the spoken content deviates from the semantic range of the preset topic, or it can also indicate the degree of deviation between the semantics of the spoken content and the semantic range of the preset topic.

[0071] In some embodiments, adjusting the speaking weight of the role model corresponding to the speaking content based on the content quality understanding results may include: adjusting the speaking weight of the role model corresponding to the speaking content based on topic relevance information.

[0072] For example, if the topic relevance information corresponding to the speech content output by the large role model A indicates that the semantic relevance between the speech content and the preset topic is lower than the preset threshold, the current speech weight of the large role model A can be reduced.

[0073] For example, if the topic relevance information corresponding to the speech content output by the large role model B indicates that the semantic relevance between the speech content and the preset topic is higher than the preset threshold, the current speech weight of the large role model B can be increased.

[0074] For example, if the topic relevance information corresponding to the speech content output by the large role model A indicates a semantic relevance value of 0.9 between the speech content and the preset topic, a speech weight of 0.9 associated with the semantic relevance value of 0.9 can be assigned to the large role model A. This allows for dynamic adjustment of the speech weights of large role models whose output speech content is semantically closely related to the preset topic by analyzing the semantic relevance information between the speech content output by multiple large role models. Furthermore, by adjusting the speech weights, the large role models can engage in dialogue based on a speech strategy, thereby improving the quality of the output content during the dialogue and ultimately enhancing the data quality of the target corpus. This avoids excessive human intervention in the dialogue process or excessive modification of the target speech content, improving the efficiency of target corpus data generation and reducing the complexity of interactive operations in generative tasks such as data annotation and document generation.

[0075] In some embodiments, the designated large model can also send prompt words representing the corresponding topic relevance information to the role large model, so that the role large model can regenerate new speech content based on the prompt words, or adjust the content quality of subsequent output speech content.

[0076] For example, the topic relevance information for role model B might be the text message "Highly deviates from the historical research topic." A designated role model can send the prompt "Your speech content deviates highly from the historical research topic; please regenerate your speech content" to role model B. This allows for the automatic adjustment of the content quality output by multiple role models during the dialogue process.

[0077] In some embodiments, the content quality understanding result may further include context relevance information, which characterizes the semantic relevance between the spoken content and the context content generated during the dialogue. The context content generated during the dialogue may be at least one of the already generated spoken content and the target spoken content. The context relevance information in the content understanding result obtained by performing semantic understanding on the spoken content and context content using a specified large model may include context relevance values, context relevance levels, etc.

[0078] In some embodiments, adjusting the speaking weight of the role model corresponding to the speaking content based on the content quality understanding results includes: adjusting the speaking weight of the role model corresponding to the speaking content based on contextual relevance information.

[0079] For example, if the context relevance information corresponding to the speech content output by the large role model A indicates that the semantic relevance between the speech content and the context content is lower than a preset threshold, the current speech weight of the large role model A can be reduced.

[0080] For example, if the context relevance information corresponding to the speech content output by the large character model B indicates that the semantic relevance between the speech content and the context content is higher than a preset threshold, the current speech weight of the large character model B can be increased.

[0081] For example, if the context relevance information corresponding to the speech content output by the large role model A indicates a semantic relevance value of 0.9 between the speech content and the context, a speech weight of 0.9 associated with the semantic relevance value of 0.9 can be assigned to the large role model A. This allows for dynamic adjustment of the speech weights of large role models whose output speech content is semantically closely related to the context by analyzing the semantic relevance information between the speech content output by multiple large role models and the preset topic. Furthermore, by adjusting the speech weights, the large role models can engage in dialogue based on a speech strategy, thereby improving the quality of the output content during the dialogue and ultimately enhancing the quality of the target corpus data. This avoids excessive human intervention in the dialogue process or excessive modification of the target speech content, improving the efficiency of target corpus data generation and reducing the complexity of interactive operations in generative tasks such as data annotation and document generation.

[0082] In some embodiments, the designated large model can also send prompt words representing corresponding contextual relevance information to the role large model, so that the role large model can regenerate new speech content based on the prompt words, or adjust the content quality of subsequent output speech content.

[0083] For example, the contextual relevance information for role model B might be the text message "poor semantic relevance with other speech content." A designated role model can send the prompt "Your speech content has poor semantic relevance with other speech content; please regenerate your speech content" to role model B. This allows role model B to regenerate high-quality speech content with a higher degree of contextual relevance, thus automatically adjusting the quality of content output by multiple role models during the dialogue.

[0084] Since contextual relevance information can represent the degree of semantic relevance between the speech content output by multiple role models in a dialogue and other speech content generated in the dialogue, such as semantic logical coherence and semantic completeness, the speech weights of role models corresponding to the speech content can be adjusted according to contextual relevance information. This allows role models with higher contextual semantic relevance in their output speech content to increase their participation in the dialogue, thereby improving the corpus data quality and generation efficiency of the target corpus data.

[0085] In one embodiment, a context deviation warning message indicating "poor semantic relevance with other speech content" can also be displayed in the interactive interface, so that the target object can pay attention to the content quality of the speech content output by each role's large model in the current dialogue process, and thus control the output of the specified large model to meet the requirements by inputting the requirement information through interactive operation.

[0086] It should be noted that the type of strategy content included in the initial dialogue speaking strategy involved in the embodiments of this disclosure may be the same as or similar to the type of strategy content in the dialogue strategy, and the embodiments of this disclosure will not be described again here.

[0087] In some embodiments, the speech weights in the dialogue strategy or initial dialogue strategy can also be used to remove at least one role-based large model currently participating in the dialogue process, or to designate a target role-based large model from candidate large models that are not participating in the dialogue process based on the speech weights. This allows for flexible scheduling of multiple role-based large models to participate in the dialogue process and generate new, high-quality speech content, thereby improving the quality of the target corpus data. Furthermore, by designating large models, automated and flexible scheduling for the dialogue process can be achieved, improving the efficiency of target corpus data generation.

[0088] In one embodiment, a designated large model can determine a speech token from a token pool by using speech weights, and control the output of speech content by distributing speech tokens to the role large model.

[0089] In some embodiments, during a dialogue, the moderator (or designated moderator) can send a speaking weight token to the designated role moderator based on the "Add Role" or "Remove Role" options displayed on the interactive interface by the target object. The speaking weight token indicates the speaking weight, which can be used to instruct the designated role moderator to stop speaking or participate in speaking. This allows for the rapid insertion of a new role moderator based on the target object's interactive actions to output speaking content according to the context and a preset topic, should an unexpected viewpoint be needed during the dialogue.

[0090] In some embodiments, the dialogue strategy may also include a cooldown period for issuing tokens: It supports setting a "cooldown period" after a speech weight token is assigned to a role-based large model to perform a speech content generation task. The cooldown duration can be configured as 1 or 2 rounds, so that the role-based large model does not speak during the cooldown period after the current speech, avoiding excessive continuous output of speech content by the same role-based large model. The dialogue strategy may also include timeout protection for issuing tokens: If a call times out or fails after a role-based large model is assigned a speech weight token, the current speech weight token is automatically skipped and assigned to the next role-based large model. Speech weight token downgrading and rollback: When multiple attempts to issue a token to a specified role-based large model still fail to meet the triggering conditions, the role-based large model can be downgraded or removed, allowing manual editing to input speech content or calling another role-based large model to intervene, and the rollback operation is recorded.

[0091] In some embodiments, the strategy content in the dialogue strategy may further include topic convergence decisions. These decisions indicate that the current contextual content output by the multiple role-based large models already satisfies the requirements for a preset topic or any subtopic within that preset topic. This allows for summarizing the dialogue, stopping the dialogue, or starting a dialogue for a new preset topic or a new subtopic. This enables control over the termination of the dialogue process and subtopic transitions based on the semantic understanding capabilities of the specified large model, thereby improving the flexibility and efficiency of generating the target corpus data.

[0092] In some embodiments, the topic convergence information indicated by the target object's interactive operations can be used to prompt a designated large model to control multiple role-based large models: to conduct a summary dialogue, stop the dialogue, or start a dialogue on a new preset topic or a new subtopic. This can improve the controllability of the target object in the target corpus data generation process.

[0093] Figure 3 The diagram illustrates an application scenario of the large-model-based corpus data generation method according to embodiments of the present disclosure.

[0094] like Figure 3 As shown, the designated large model in this application scenario is a host large model with host attributes. Multiple role large models can include a first role large model, a second role large model, and a viewpoint summarizer large model. Before the dialogue begins, the host large model processes the preset topic, the model capability descriptions of each role large model, and the target object's configuration requirements for the preset topic to plan the dialogue strategy, thus obtaining an initial dialogue strategy. The host large model uses the initial dialogue strategy to control the speaking order, frequency, and other speaking patterns of the first and second role large models regarding the preset topic during the dialogue. The speaking content of multiple role large models during the dialogue will be displayed in the first interactive interface 300.

[0095] During the dialogue between the first-role big model and the second-role big model, the moderator big model processes the already generated speech content A, speech content B, and speech content C, where speech content A and speech content B serve as contextual content. The output content understanding result can include contextual relevance information related to speech content C, indicating that speech content C has a deficiency in low contextual semantic relevance. This can reduce the speech weight of the second-role big model in subsequent dialogue processes and instruct the second-role big model to re-execute the new speech content generation task to generate new speech content C.

[0096] In the process of controlling multiple role-based large models to conduct dialogue based on a dialogue strategy that includes adjusted speech weights, the moderator large model can also output new dialogue strategies by semantically understanding the speech content, so as to adjust the speech patterns and the quality of the output speech content of the multiple role-based large models in a timely manner. After the moderator large model performs semantic understanding on the multiple speech contents that have been generated, the output content understanding result includes an indication that multiple speech contents related to a preset topic need to be summarized. In this case, the dialogue strategy output by the moderator large model includes a prompt message for the viewpoint summarizer large model to summarize, so as to control the viewpoint summarizer large model to perform semantic summarization on the preset topic and the already generated context speech content, and output the summary speech content.

[0097] Based on speech content A, speech content B, speech content C... up to the summary speech content, as well as multiple dialogue strategies and initial dialogue strategies output by the host's big model, target corpus data can be generated.

[0098] According to embodiments of this disclosure, the target speech content is obtained by engaging in dialogue on a preset topic using multiple role models based on a dialogue strategy.

[0099] In some embodiments, based on a dialogue strategy, using multiple role models to engage in dialogue on a preset topic to obtain target speech content may include: using a specified role model to control multiple role models to engage in dialogue based on a dialogue strategy to obtain intermediate speech content; and determining the target speech content in response to a target object's target interactive operation on the intermediate speech content.

[0100] According to embodiments of this disclosure, the intermediate speech content can be output by a large role model. Controlling multiple large role models to conduct dialogue based on a dialogue strategy using a specified large model can include using strategy content information such as the speaking order and speaking frequency in the speaking mode represented by the dialogue strategy output by the specified large model to schedule the corresponding large role model to perform the speech content generation task based on the context content and preset topic, and output the intermediate speech content.

[0101] In some embodiments, intermediate speech content can be the output of the initial speech content by the large role model through a self-checking task, and the output of intermediate speech content that meets the speech content quality conditions.

[0102] In some embodiments, the target interaction operation of the target object for intermediate speech content may include an adoption operation for adopting intermediate speech content as target speech content. This allows for the rapid generation of intermediate speech content for target corpus data by adopting any intermediate speech content generated by the target object for multiple role models in the process of a preset topic, thereby improving the control efficiency and accuracy of multiple role models in the dialogue process.

[0103] In some embodiments, the target corpus data may include associated target speech content and operation information for target interactive operations. For example, it may carry editing information for performing editing operations on intermediate speech content.

[0104] In one example, the interactive interface can display intermediate speech content output by any character's large model during the dialogue. This intermediate speech content can be displayed in a speech content box, using light-colored characters. If the target object performs an adoption operation on at least one text or paragraph within the intermediate speech content, it can be determined that the text or paragraph of that intermediate speech content is displayed using dark characters, and the text or paragraph corresponding to the dark characters is identified as the target speech content.

[0105] In some embodiments, the target interaction operation may further include an update operation. In response to the target object's target interaction operation on the intermediate speech content, determining the target speech content may include: in response to the update operation on the intermediate speech content, updating the intermediate speech content using the role big model according to the update operation to obtain the updated intermediate speech content; in response to the adoption operation on the intermediate speech content currently displayed in the interaction interface, determining the currently displayed intermediate speech content as the target speech content.

[0106] According to embodiments of this disclosure, intermediate speech content and updated intermediate speech content are displayed in the interactive interface. For example, intermediate speech content can be displayed in the interactive interface, and after updated intermediate speech content is generated, the interactive interface can delete the intermediate speech content and display the updated speech content. Alternatively, for example, intermediate speech content and updated intermediate speech content can be displayed simultaneously in the interactive interface so that the target object can perform an adoption operation by comparison.

[0107] In one example, the update operation can be a cancellation operation. The interactive interface can display the intermediate speech content output by any character's large model during the dialogue. The intermediate speech content can be displayed in a speech content box, using light-colored characters. If the target object cancels at least one text or paragraph in the intermediate speech content, the intermediate speech content can be cleared, and the character's large model can be instructed to re-execute the speech content generation task to generate and display updated intermediate speech content. The target speech content is determined until the target object performs an accept operation on the currently generated intermediate speech content.

[0108] In some embodiments, the update operation may further include an editing operation that modifies the currently displayed intermediate speech content. The editing operation can represent the target object making content edits to the intermediate speech content, such as deleting or adding text or tables.

[0109] In some embodiments, the same character in a dialogue process can have multiple character models perform the speech content generation task separately, resulting in intermediate speech content output by each of the multiple character models. The update operation may also include a fusion operation, which is used to merge multiple intermediate speech contents related to the same character to obtain the target speech content.

[0110] For example, a specified large model can be used to semantically fuse multiple interpretations of philosophical questions selected by the user as intermediate speech content to obtain the target speech content.

[0111] Therefore, by setting up multiple large role models for the same role participating in the dialogue to generate intermediate speech content separately, the diversity and differentiation of intermediate speech content can be improved. Furthermore, by semantically fusing multiple intermediate speech content, high-quality target speech content can be generated, thereby improving the quality and generation efficiency of the target corpus data.

[0112] In some embodiments, the interactive interface may also display content quality scores related to one or more intermediate speech contents. The content quality scores may represent the evaluation results of the intermediate speech contents in terms of content quality indicators such as context relevance, topic relevance, and language fluency.

[0113] For example, the content quality score can be determined based on the intermediate speech content output by each character's large model, as follows.

[0114] For the language fluency index: the language fluency of the intermediate speech content is evaluated based on a specified large model to obtain the language fluency evaluation results.

[0115] For the context relevance metric: the cosine similarity between the intermediate speech content and the context content is calculated using an attention network algorithm as the context relevance evaluation result.

[0116] For coverage and completeness metrics: The coverage and completeness evaluation results of the intermediate speech content are evaluated based on a predefined list of key elements such as representation entities and intent points.

[0117] Regarding the semantic-logical consistency index: Based on the large language model, detect whether there are logical contradictions or logical jumps in the internal semantics of the intermediate speech content, and obtain the logical consistency detection result.

[0118] By integrating and quantifying the results of language fluency assessment, coverage and completeness evaluation, context relevance assessment, and logical consistency detection, a content quality score is obtained for the intermediate speech content. This content quality score can be displayed in the interactive interface to help the target audience select higher-quality intermediate speech content to determine the target speech content.

[0119] For example, a target object can perform a fusion operation on multiple intermediate statements from the same person. A specified large model can automatically select multiple intermediate statements whose content quality scores meet preset scoring criteria for semantic fusion, facilitating the fusion of high-scoring paragraphs or key phrases to generate updated intermediate statements. The target object can then perform an adoption operation on the updated intermediate statements to obtain the target statement.

[0120] Among them, multiple large character models for the same role can be determined based on the interactive operations of the target object. For example, the target object can choose based on information such as the version number and model name of each of the multiple large character models.

[0121] In some embodiments, the target corpus data may also include detailed generation process information such as intermediate speech content, content quality scores corresponding to the intermediate speech content, and generation timestamps, so that relevant personnel can review or learn the target speech content based on the detailed generation process information in the target corpus data. Alternatively, the target corpus data used to train a language model can enable the language model to learn the detailed speech content generation process more clearly and fully, thereby improving the language model's ability and adaptability to perform content generation tasks on preset topics.

[0122] In some embodiments, the target corpus data is determined based on the target speech content and operation information of the target interaction operation related to the target speech content.

[0123] In one embodiment, the target speech content in the target corpus data is related to the adoption operations performed by the target object. The target corpus data may include operation information with mapping relationships, intermediate speech content, and target speech content. Therefore, a language model can be trained based on the target corpus data. This allows the language model to learn the differences between target speech content that meets quality criteria and intermediate speech content that does not, as well as the interaction information of target interaction operations such as fusion, editing, cancellation, and adoption operations for intermediate speech content. This improves the language model's content interaction capabilities for scenarios related to a predefined topic.

[0124] In some embodiments, determining the target corpus data based on operational information and target speech content may further include determining the adoption rate of each role's large model based on the operational information. The adoption rate can represent the statistical proportion of intermediate speech content output by the role's large model that is adopted by the target object. For example, adoption rate = number of adoption operations performed / number of intermediate speech contents output. Therefore, the performance of the language model can be enhanced and the quality of the trained language model's output content improved based on the target corpus data containing the adoption rate.

[0125] In one embodiment, if the adoption rate of the large role model is > 80%, a speech content generation operation can be performed based on the content information edited by the target object during the user's editing operation of intermediate speech content. If the adoption rate of the large role model is < 30%, the large role model performs a speech content generation operation based on the complete content information edited by the target object, or prompts the target object to generate target dialogue content based on the editing operation.

[0126] In some embodiments, controlling multiple role models to engage in dialogue based on a dialogue strategy using a specified large model may further include: in response to the triggering of a specified event related to the dialogue strategy, controlling the target role model to perform semantic understanding based on the specified event using the specified large model based on event prompt information corresponding to the specified event in the dialogue strategy, and obtaining event feedback information; the target corpus data is determined based on the event feedback information and the target speech content, and the specified event is determined by detecting the speech content during the dialogue process.

[0127] According to embodiments of this disclosure, the specified event can be triggered based on preset rules. Specified events may include, for example, events related to risky speech content attributes, severe semantic conflicts, or content quality scores below a threshold. A specified large model can be used to detect information such as the speech content generated by multiple role models during the dialogue, the operation information of the target interaction, and the operating status of the computing device to determine whether the specified event is triggered. Alternatively, a detection tool built based on preset rules can be used to detect information such as the speech content generated during the dialogue, the operation information of the target interaction, and the operating status of the computing device to determine whether the specified event is triggered.

[0128] In some embodiments, event prompts corresponding to a specified event in a dialogue strategy can be used to represent the event response measures triggered by the specified event. By utilizing a specified large model to control the target role's large model to semantically understand the triggered specified event based on the event prompts, the generated event feedback information can be matched with the response intent represented by the event response measures. This allows target corpus data to be generated based on event feedback information, target speech content, and other information such as target interaction operation information. This enables the language model to learn to respond promptly by outputting event feedback information when a relevant specified event is triggered during a dialogue on a preset topic, improving the quality of the language model's output in relevant scenarios and further enhancing the language model's capabilities.

[0129] In one embodiment, the specified event may include any one or more of the following events.

[0130] Risk content trigger event: Based on the detection of preset risk keywords in the context of the dialogue generated by the specified large model, or the content understanding result indicating that the sentiment of the context content shows a specified negative emotion type. The corresponding event prompt information indicates the event response measures: prompting the risk keywords, risk emotion type and other risk detection results, to control at least one role's large model to output speech content that meets the risk requirements.

[0131] Content Quality Rating Event: A content quality rating event is triggered when the content quality rating of any intermediate message or a preset number of intermediate messages falls below a preset rating threshold. The corresponding event notification indicates the event response measures: The event model with the commenter role attribute among multiple role models outputs the type of quality defect in the context content based on the context content and the content quality rating, so that other role models can re-engage the dialogue based on the quality defect type as event feedback information.

[0132] Response Duration Event: A response duration event is triggered when the interval between different dialogue rounds exceeds a preset duration threshold, or when the number of dialogue rounds exceeds a preset number of rounds threshold. The corresponding event prompt message indicates the event response measure: The target character's large model, capable of summarizing viewpoints, is prompted to summarize and generalize the context of the dialogue and output the summarized viewpoint as event feedback information.

[0133] Specified Interaction Event: This refers to the action information detected during the dialogue for a specified type of interaction. For example, the action information could be input from the target object such as "change the direction of the discussion" or "discuss in the next round." The corresponding event prompt information indicates the event response: Multiple role models are prompted to perform a speech content generation task based on the input action information, outputting speech content that meets the requirements as event feedback information.

[0134] Opinion Conflict Event: Opinion conflict is detected based on the context. If the conditions for opinion contradiction or opinion conflict are met, an opinion conflict event is triggered. The corresponding event prompt information indicates the event response measures: by calling the target role model with the opinion arbitrator attribute to respond to the opinion conflict in the context, the event feedback response information is obtained.

[0135] Arbitrator rule trigger: This event is triggered when the content quality score of a question posted by the designated role's main model is lower than the required score threshold, or when the content quality score of a post posted by the expert role's main model on a designated subtopic is lower than the required score threshold. The corresponding event message indicates the event response: the target main model, based on the arbitrator role's attributes, will output a response based on the context.

[0136] Custom events: You can customize event triggering rules based on the interaction of the target object. For example, you can trigger an event after a specified time if the length of the answer output by the expert role model exceeds 5000 words, or if the delay in the output of the commenter role model exceeds 120 seconds.

[0137] In one embodiment, the system log can record log data such as the timestamp of a specified event trigger, the type of a specified event, the content of speech related to the specified event, and event feedback information. The log data and the dialogue strategy can be used together as annotation metadata related to the speech content and the target speech content for annotation, generating target corpus data.

[0138] In some embodiments, the target corpus data may also include the content understanding results corresponding to the speech content output by each role's large model, so that the target object can intuitively understand the content quality and content attributes corresponding to each target speech content in the target corpus data. This allows the language model to be trained to learn based on the content understanding results and the target speech content, and to fine-tune the language model to adapt the trained language model to the interactive scenarios related to the preset topic, thereby improving the performance of the language model.

[0139] For example, the target corpus data can include the speech content output by multiple role-based large models and the target speech content, as well as the strategy content of the specified large model output to each role-based large model, such as speech order, triggering events, adding or deleting roles, inserting convergence nodes, and scheduling commands. The target corpus data can be stored on a structured data file to store the strategy content and target speech content, enabling the large language model to learn dialogue capabilities for dialogue scenarios related to a preset topic more accurately, thereby improving the training efficiency of the language model.

[0140] In some embodiments, the target corpus data is determined based on dialogue strategies, target speech content, and the thought process information of the role big model during the speech content generation task.

[0141] According to embodiments of this disclosure, the thought process information can represent the thought chain, thought tree, etc., of the large role model during the execution of the speech content generation task. The thought process information can include multiple target tasks and the dependencies between them, enabling the large role model to execute multiple target tasks according to the dependencies based on the thought process information and output speech content.

[0142] For example, multiple target tasks may include thinking tasks, which may be based on multiple task representations as shown in the example below.

[0143] Preliminary reasoning task: Quickly generate candidate answers or solutions based on the context of the dialogue.

[0144] The self-check task verifies the consistency, logic, and completeness of the results of the previous step, and outputs a check report or a list of questions.

[0145] Iterative thinking tasks involve supplementing, correcting, or reconstructing previous thinking based on self-check feedback or tool results.

[0146] The task of summarizing and extracting involves extracting key elements, concepts, or facts from complex information to generate concise summaries or lists of key points.

[0147] Format conversion tasks convert text content into different formats, such as question-and-answer pairs, step lists, code comments, or tables.

[0148] The classification and intent recognition task determines the user's intent, sentiment, or text type, providing labels for subsequent processing.

[0149] The task of decision-making and strategy formulation involves evaluating the advantages and disadvantages of various options in a multi-choice scenario and providing the optimal or feasible strategy.

[0150] For example, multiple target tasks may include tool invocation tasks. A tool invocation task represents a large role model executing a tool processing task by invoking a specified tool resource and obtaining the tool resource's output. A tool invocation task may include one or more tool resources for execution. Tool resources may include, for example, document retrieval tools, web search tools, image processing tools, language translation tools, and multimodal data fusion tools.

[0151] Document retrieval tool: Extracts information such as document summaries and question-and-answer functions, performs database queries, and parses the results. Web search tool: Searches web pages and retrieves relevant content. Code execution and debugging tool: Runs program scripts in a specified language and returns the execution results. Image processing tool: Understands image content, provides question-and-answer responses, and performs natural language plotting.

[0152] In some embodiments, the thinking process information can be displayed in the interactive interface in the form of a thinking topology that includes the relationships between nodes and edges, so as to modify the target task and dependencies in the thinking process of the large role model performing the speech content generation task, so that the speech content output by the large role model can meet the needs and intentions of the target object.

[0153] In some embodiments, target corpus data is used to train a language model to be trained. This language model can be constructed based on the principles of large language models, and its parameter count can be smaller than that of a large role model used to perform speech generation tasks. This allows the language model to be trained based on structured thought process information and corpus content, effectively distilling the capabilities of a large model with a large parameter scale and strong performance into a smaller language model. Furthermore, deploying the trained language model on computing devices such as servers enhances the capabilities of these devices for specific scenarios such as intelligent customer service, knowledge-based question answering, and scriptwriting, thereby improving user experience and reducing computational overhead.

[0154] Figure 4 The diagram illustrates an application scenario of a large-model-based corpus data generation method according to another embodiment of the present disclosure.

[0155] like Figure 4 As shown, the second interactive interface 400 displays a thought topology 410 for the role-based large model to perform the speech content generation task. The thought topology 410 can include multiple nodes and the edge relationships between them. Multiple nodes can represent multiple target tasks, and these target tasks are executed according to the dependencies shown by the edge relationships. Specifically, the first node can represent the first target task, which could be a thought task involving task planning and consideration of the question "Predicting changes in Company A's electricity consumption in 2026". The second node represents a tool scheduling task for obtaining Company A's revenue reports for the past three years. The third node represents a tool scheduling task for performing semantic understanding on Company A's revenue reports for the past three years to generate an analysis of Company A's product output changes over the past three years and a projected product output for 2026.

[0156] Node 4 could be a data search task to obtain electricity consumption change data for Company A over the past three years. Node 5 could be a tool scheduling task that uses a role-based big data model to perform semantic analysis on Company A's electricity consumption change data over the past three years, product output change analysis, and 2026 product output forecast, outputting information about Company A's electricity consumption analysis. Node 6 represents a self-examination task regarding the electricity consumption analysis of Company A. Therefore, the output from the role-based big data model could be an electricity consumption analysis of Company A for 2026. This analysis could include diverse data such as multiple text paragraphs, tables, and charts.

[0157] Based on the multiple nodes and edge relationships represented by the thinking topology 410, the large role model can be instructed to perform multiple target tasks and mapping relationships between target tasks in the thinking process of generating speech content. Target objects can perform interactive operations on the thinking topology 410 to add, delete, or replace any node or edge relationship, thereby updating the thinking process information based on interactive operations on the second interactive interface 400. Thus, the large role model can perform the speech content generation task based on the multiple target tasks and dependencies in the updated thinking process information to output electricity consumption analysis content.

[0158] It should be noted that the acquisition of information involved in any embodiment of this disclosure, including but not limited to revenue reports, electricity consumption data, etc., is authorized by relevant personnel or structures before the data is acquired. Before acquiring the data, the actual purpose of acquiring the data is to meet the actual needs of the target object with data access rights. Necessary encryption or desensitization measures are taken for the acquired data to avoid information leakage. This complies with the provisions of relevant laws and regulations and does not violate public order and good morals.

[0159] In some embodiments, the large-model-based corpus data generation method provided in this disclosure may further include setting up a tool invocation subsystem. The target object can configure and manage complex tool resource invocation processes within the same interactive interface as the actual intermediate speech content or speech content, based on the tool invocation subsystem. This obtains reliable tool execution example data and, with the help of simulation, monitoring, and automatic recommendation functions, significantly reduces configuration difficulty and operational risks, providing high-fidelity, multi-scenario tool resources for diverse scenarios such as language model training and research reports with specified requirements.

[0160] The tool registration and description component can be used to manage tool metadata. Tool metadata includes the tool description information required for each tool resource when it is invoked. The tool description information of a tool resource may include: tool name and version, invocation parameter template, result validation rules, retry strategies such as timeout and error retries, and tool invocation trigger conditions.

[0161] The tool invocation subsystem may also include service components for providing intelligent completion and verification services. These service components can automatically recommend suitable tool resources for invocation within the interactive interface based on the context and historical invocation records during the dialogue, and suggest input parameter values ​​such as search keywords and code snippets. Furthermore, when configuring parameters for a tool invocation task, they can check the validity of configuration parameters such as parameter types, required fields, and ranges in real time, and provide correction prompts.

[0162] Figure 5 A flowchart illustrating a method for generating corpus data based on a large model according to another embodiment of the present disclosure is shown.

[0163] like Figure 5 As shown, the corpus data generation method based on the large model includes operations S510~S520.

[0164] When operating S510, preset instruction elements related to preset prompt instructions are displayed.

[0165] During operation S520, in response to the trigger operation for the preset instruction element, at least one character model is prompted to perform the speech content generation task according to the speech prompt information in the preset prompt instruction.

[0166] According to embodiments of this disclosure, a large role model engages in dialogue with other large role models by performing a speech content generation task. For example, a large role model can output speech content by performing semantic understanding of contextual content and the policy content of the dialogue strategy.

[0167] According to embodiments of this disclosure, preset prompts can be used to indicate the task requirements of the target object's current large-scale character model to perform the speech content generation task. For example, preset prompts can instruct the large-scale character model to perform the speech content generation task based on speech prompts corresponding to requirements such as "based on table description content", "simplify answer", "expand details", "convert format", and "list key points", and output speech content that matches the task requirements indicated by the speech prompts.

[0168] In some embodiments, the target object can be edited based on the initial preset instruction element to generate preset prompt instruction elements that match the task requirements. This allows for the configuration of speaking prompt information for the preset prompt instructions, facilitating quick instruction control for at least one large character model during the dialogue process and improving the efficiency of target corpus data generation.

[0169] In some embodiments, isolated file systems and network access can be run in a controlled sandbox environment by invoking tool resources to prevent malicious code or data leakage. After the tool resources are invoked, the execution results of the tool resources, along with the mapped speech content or intermediate speech content, and the context content, can be sent to a designated large model for quality assessment for detection. The detection scope includes, but is not limited to: the completeness of the results (e.g., whether the target task execution result contains the required fields), semantic matching degree (e.g., whether it is related to the expected answer), and reliability assessment (e.g., the credibility score of the search results obtained after the tool invokes the task execution).

[0170] Figure 6 A block diagram of a large-model-based corpus data generation apparatus according to an embodiment of the present disclosure is shown schematically.

[0171] like Figure 6 As shown, the large model-based corpus data generation device 600 includes: a speech content acquisition module 610, a dialogue strategy acquisition module 620, and a target corpus data determination module 630.

[0172] The speech content acquisition module 610 is used to obtain the speech content of at least one character model by having multiple large character models engage in dialogue on a preset topic.

[0173] The dialogue strategy acquisition module 620 is used to plan dialogue strategies based on the speech content of a specified large model to obtain a dialogue strategy. The dialogue strategy is used to constrain the speech patterns of the role large model during the dialogue process.

[0174] The target corpus data determination module 630 is used to determine the target speech content generated by the dialogue based on the dialogue strategy according to the role big model, and to determine the target corpus data related to the preset topic.

[0175] According to embodiments of this disclosure, the dialogue strategy acquisition module includes: a content quality understanding result acquisition unit, a speech weight acquisition unit, and a dialogue strategy acquisition unit.

[0176] The content quality understanding result acquisition unit is used to perform semantic understanding of the speech content using a specified large model to obtain the content quality understanding result.

[0177] The speech weight acquisition unit is used to adjust the speech weight of the role model corresponding to the speech content based on the content quality understanding results, and obtain the speech weight for the role model. The speech weight represents the expected speech participation of the role model in the dialogue process.

[0178] The dialogue strategy acquisition unit is used to determine the dialogue strategy based on the speaking weight.

[0179] According to embodiments of this disclosure, the content quality understanding result includes topic relevance information, which characterizes the semantic relevance between the speech content and a preset topic; wherein, the speech weight acquisition unit includes a first adjustment subunit.

[0180] The first adjustment subunit is used to adjust the speaking weight of the role model corresponding to the speaking content based on topic relevance information.

[0181] According to embodiments of this disclosure, the content quality understanding result further includes context relevance information, which characterizes the semantic relevance between the speech content and the context content generated during the dialogue; wherein, the speech weight acquisition unit includes a second adjustment subunit.

[0182] The second adjustment subunit is used to adjust the speaking weight of the role model corresponding to the speaking content based on contextual relevance information.

[0183] According to embodiments of this disclosure, the corpus data generation apparatus based on a large model further includes: an intermediate speech content acquisition module and a target speech content acquisition module.

[0184] The intermediate speech content acquisition module is used to control multiple role-based large models to conduct dialogue based on a specified large model and a dialogue strategy, and obtain the intermediate speech content.

[0185] The target speech content acquisition module is used to determine the target speech content in response to the target object's target interactive operation on the intermediate speech content.

[0186] According to embodiments of this disclosure, the target speech content acquisition module includes an update unit and a target speech content acquisition unit.

[0187] The update unit is used to respond to the update operation on the intermediate speech content. It uses the large role model to update the intermediate speech content according to the update operation to obtain the updated intermediate speech content. The intermediate speech content and the updated intermediate speech content are displayed in the interactive interface.

[0188] The target speech content acquisition unit is used to respond to the adoption operation of the currently displayed intermediate speech content in the interactive interface and determine the currently displayed intermediate speech content as the target speech content.

[0189] According to embodiments of this disclosure, the target corpus data is determined based on the target speech content and the operation information of the target interactive operation related to the target speech content.

[0190] According to embodiments of this disclosure, the intermediate speech content acquisition module includes an event feedback information acquisition unit.

[0191] The event feedback information acquisition unit is used to respond to the triggering of a specified event related to the dialogue strategy. It uses a specified large model to control the target role large model to perform semantic understanding based on the specified event according to the event prompt information corresponding to the specified event in the dialogue strategy, and obtains event feedback information. The target corpus data is determined based on the event feedback information and the target speech content, and the specified event is determined by detecting the speech content in the dialogue process.

[0192] According to embodiments of this disclosure, the speech content acquisition module includes a dialogue unit.

[0193] The dialogue unit is used to conduct dialogues on a preset topic using multiple role models based on an initial dialogue strategy. The initial dialogue strategy is determined by planning dialogue strategies based on a specified role model according to the preset topic.

[0194] According to embodiments of this disclosure, the corpus data generation apparatus based on a large model further includes: a display module and a speech content generation task execution module.

[0195] The display module is used to display preset instruction elements related to preset prompts.

[0196] The speech content generation task execution module is used to respond to the trigger operation of the preset instruction element. According to the speech prompt information in the preset prompt instruction, it prompts at least one character model to perform the speech content generation task according to the preset prompt instruction. The character model can then engage in dialogue with other character models by performing the speech content generation task.

[0197] According to embodiments of this disclosure, the target corpus data is determined based on dialogue strategies, target speech content, and the thought process information of the role model during the execution of the speech content generation task.

[0198] Figure 7 A schematic block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is shown.

[0199] In embodiments of this disclosure, such as Figure 7 As shown, the AI ​​agent 700 may include an input module 710, a processing module 720, and an output module 730.

[0200] Input module 710 is used to receive input information;

[0201] The processing module 720 is used to determine the target task based on the input information received by the input module, determine the role big model and the specified big model based on the target task, and obtain output information by calling the role big model and the specified big model to execute the corpus data generation method based on the big model provided in the embodiments of this disclosure;

[0202] Output module 730 is used to output the output information obtained by the processing module.

[0203] According to embodiments of this disclosure, the input module 710 is responsible for receiving or sensing information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment), and converting it into a format that the AI ​​agent 700 can understand and process. The input module 710 is the primary link for the AI ​​agent 700 to interact with the outside world, enabling the AI ​​agent 700 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.

[0204] In the example, input module 710 can input the preset topics, dialogue content, etc. described above.

[0205] In the example, processing module 720 is the core support for the AI ​​agent 700's ability to handle complex tasks. Processing module 720 can execute the large-model-based corpus data generation method described above.

[0206] In the example, the performance of the processing module 720 is closely related to the large model on which the AI ​​agent 700 is based. To fully leverage the capabilities of the large model, the internal structure of the processing module 720 can be designed to be highly configurable and scalable to handle various types of tasks and requirements in real-world scenarios.

[0207] In the example, after the AI ​​agent 700 obtains a preset topic, the processing module 720 can use multiple role models to conduct dialogues on the preset topic, obtain dialogue content, use a specified role model to plan a dialogue strategy based on the dialogue content, obtain a dialogue strategy, and then pass the target speech content generated by the role model based on the dialogue strategy to the output module 730.

[0208] Understandably, while large language models possess excellent language understanding and generation capabilities, like humans, their ability to solve tasks is limited without the aid of any tools. Once the AI ​​agent 700 is given the ability to invoke tools, it can perform tasks such as using a calculator to complete mathematical calculations, using Python to perform data analysis, and using a search engine to create weather forecasts.

[0209] In the example, output module 730 can output the target speech content and dialogue strategy described above.

[0210] The AI ​​agent 700 according to the embodiments of this disclosure can simply and effectively improve the level of intelligence, as well as enhance flexibility and versatility.

[0211] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0212] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.

[0213] According to embodiments of the present disclosure, a non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to cause a computer to perform the method described above.

[0214] According to an embodiment of this disclosure, a computer program product includes a computer program that, when executed by a processor, implements the method described above.

[0215] Figure 8 A schematic block diagram of an example electronic device for generating corpus data based on a large model, which can be used to implement embodiments of the present disclosure, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0216] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0217] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0218] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the large-model-based corpus data generation method. For example, in some embodiments, the large-model-based corpus data generation method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the large-model-based corpus data generation method described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured in any other suitable manner (e.g., by means of firmware) to perform a corpus data generation method based on a large model.

[0219] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0220] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0221] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0222] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0223] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0224] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.

[0225] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0226] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for generating corpus data based on a large model, comprising: Multiple large character models engage in dialogue on a preset topic to obtain the speech content of at least one of the large character models, where the preset topic is the text processed by the large character model; The content of the speech is semantically understood using a specified large model to obtain the content quality understanding result; Based on the content quality understanding results, the speaking weights of the role model corresponding to the speaking content are adjusted to obtain the speaking weights for the role model. The speaking weights represent the expected speaking participation level of the role model in the dialogue process. The dialogue strategy is determined based on the speaking weight, and the dialogue strategy is used to constrain the speaking pattern of the large role model during the dialogue process. Based on the target speech content generated by the dialogue strategy according to the character model, target corpus data related to the preset topic is determined.

2. The method of claim 1, wherein, The content quality understanding result includes topic relevance information, which characterizes the degree of semantic relevance between the speech content and the preset topic; The step of adjusting the speech weight of the role model corresponding to the speech content based on the content quality understanding result includes: The speaking weights of the large role models corresponding to the speaking content are adjusted based on the topic relevance information.

3. The method of claim 2, wherein, The content quality understanding result also includes context relevance information, which characterizes the semantic relevance between the spoken content and the context content generated during the dialogue. The step of adjusting the speech weight of the role model corresponding to the speech content based on the content quality understanding result includes: The speech weights of the large role models corresponding to the speech content are adjusted based on the context relevance information.

4. The method of claim 1, wherein, The target speech content is determined based on the following operations: Using the specified large model, multiple role models are controlled to engage in dialogue based on the dialogue strategy to obtain intermediate speech content; In response to a target object's target interactive operation on the intermediate speech content, the target speech content is determined.

5. The method of claim 4, wherein, The step of determining the target speech content in response to a target object's target interaction operation on the intermediate speech content includes: In response to the update operation for the intermediate speech content, the intermediate speech content is updated using the large character model according to the update operation to obtain the updated intermediate speech content, wherein the intermediate speech content and the updated intermediate speech content are displayed in the interactive interface; In response to the adoption operation of the currently displayed intermediate speech content in the interactive interface, the currently displayed intermediate speech content is determined as the target speech content.

6. The method of claim 4 or 5, wherein, The target corpus data is determined based on the target speech content and the operation information of the target interactive operation related to the target speech content.

7. The method of claim 4, wherein, The step of using the designated large model to control multiple character large models to conduct dialogue based on the dialogue strategy includes: In response to the triggering of a specified event related to the dialogue strategy, the specified big model is used to control the target role big model to perform semantic understanding based on the specified event and obtain event feedback information by using the specified big model based on the event prompt information corresponding to the specified event in the dialogue strategy. The target corpus data is determined based on the event feedback information and the target speech content, and the specified event is determined by detecting the speech content during the dialogue process.

8. The method according to claim 1, wherein, The method of using multiple large character models to conduct dialogues on preset topics includes: Based on the initial dialogue strategy, multiple role models are used to conduct dialogues on the preset topic. The initial dialogue strategy is determined by planning dialogue strategies using the specified role models based on the preset topic.

9. The method according to claim 1, wherein, The method further includes: Displays preset instruction elements related to preset prompt instructions; In response to a trigger operation on the preset instruction element, at least one of the character models is prompted to perform a speech content generation task according to the speech prompt information in the preset prompt instruction, wherein the character model engages in dialogue with other character models by performing the speech content generation task.

10. The method according to claim 1 or 9, wherein, The target corpus data is determined based on the dialogue strategy, the target speech content, and the thought process information of the role model during the speech content generation task.

11. A corpus data generation device based on a large model, comprising: The speech content acquisition module is used to conduct dialogues with multiple large character models on a preset topic to obtain the speech content of at least one of the large character models, wherein the preset topic is the text processed by the large character model; The dialogue strategy acquisition module is used to plan a dialogue strategy based on the speech content using a specified large model to obtain a dialogue strategy. The dialogue strategy is used to constrain the speech pattern of the role large model during the dialogue process. The target corpus data determination module is used to determine the target corpus data related to the preset topic based on the target speech content generated by the dialogue strategy based on the role model. The dialogue strategy acquisition module includes: The content quality understanding result acquisition unit is used to perform semantic understanding on the speech content using a specified large model to obtain the content quality understanding result. The speech weight acquisition unit is used to adjust the speech weight of the role model corresponding to the speech content based on the content quality understanding result, so as to obtain the speech weight for the role model. The speech weight represents the expected speech participation degree of the role model in the dialogue process. A dialogue strategy acquisition unit is used to determine the dialogue strategy based on the speaking weight.

12. An artificial intelligence agent system, comprising: The input module is used to receive input information; The processing module is configured to determine a target task based on the input information received by the input module, determine a large role model and a specified large model based on the target task, and execute the method of any one of claims 1 to 10 by calling the large role model and the specified large model to obtain output information; An output module is used to output the output information obtained by the processing module.

13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 10.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 10.

15. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Data generation method, electronic equipment and storage medium

    CN117556026A

  • Task-based dialogue script construction method and device, equipment and storage medium

    CN118550612A