Large model generation control method and device, electronic equipment and storage medium

By generating and replacing text blocks in the thinking chain statement, using the agent to generate the target text, the problem of missing context in the generation process of big model is solved, the generation control of the big model is achieved, and the accuracy of the generation results is improved.

CN120494096APending Publication Date: 2025-08-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510592098.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the process of generating large models, the prior art cannot effectively solve the insufficient generation control effect caused by the missing context information or the untimely database update, and it is difficult to ensure the accuracy of the model generation results.

Method used

By obtaining the pending question text, generating thinking chain statements, and splitting them into each text block, using the agent in the functional feature set to generate target text to replace similar text blocks, and generating reply texts with the preset big model to achieve adjustments to thinking chain statements.

Benefits of technology

Effectively correct logical errors, reduce semantic pollution and knowledge illusions, break through the black box limitations of traditional big model reasoning, and improve the accuracy of generated results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494096A_ABST
    Figure CN120494096A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a large model generation control method and device, electronic equipment and a storage medium, and the method comprises the steps: firstly obtaining a to-be-processed question text; generating a thinking chain statement based on the question text by adopting a preset reasoning model; for each text block in the thinking chain statement, executing the following operations: when a function feature meeting a feature similarity condition with the text feature of one text block exists in a preset function feature set, generating a target text based on one text block by adopting an intelligent agent corresponding to the function feature; replacing one text block in the thinking chain statement with the target text; generating a reply text of the question text based on the processed thinking chain statement by adopting a preset large model; therefore, by adjusting the thinking chain statement, the black box limitation in the traditional large model reasoning process can be broken through, so that the generation control of the large model can be fundamentally realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method, device, electronic device and storage medium for generating and controlling a large model. Background Art

[0002] With the development of large model technology, question-answering tasks can be completed well with the help of large models.

[0003] Currently, to improve the generative controllability of large models, optimization and adjustment are usually performed on the model input content. Specifically, contextual information can be added to the prompt template, or relevant knowledge retrieved based on the input content can be added to the prompt module. This can improve the output of large models to a certain extent.

[0004] However, the above-mentioned optimization and adjustment method can only add auxiliary information related to the input content in addition to the processing of the large model. Therefore, once there is a lack of contextual information or the database is not updated in a timely manner, it is impossible to provide an effective reasoning basis for the large model, thereby failing to fundamentally improve the generation control effect of the large model, and thus it is difficult to ensure the accuracy of the model generation results. Summary of the Invention

[0005] The embodiments of the present application provide a large model generation control method, device, electronic device and storage medium, which are used to control and improve the generation effect of the large model and improve the generation accuracy.

[0006] First, a generation control method for a large model is proposed, including:

[0007] Get the question text to be processed;

[0008] Using a preset reasoning model, based on the question text, a thought chain sentence is generated;

[0009] For each text block obtained by splitting the thought chain sentence, perform the following operations respectively:

[0010] When a function feature that satisfies a feature similarity condition with a text feature of a text block exists in a preset function feature set, an agent corresponding to the function feature is used to generate a target text based on the text block; wherein the function feature set includes the function features of each agent that realizes the content generation function in the same domain;

[0011] Replacing the text block in the thought chain sentence with the target text;

[0012] A preset large model is used to generate a reply text for the question text based on the processed thought chain sentences.

[0013] In a second aspect, a large model generation control device is proposed, comprising:

[0014] An acquisition unit, used to acquire the question text to be processed;

[0015] A first generating unit is configured to generate a thought chain sentence based on the question text by using a preset reasoning model;

[0016] The execution unit is configured to perform the following operations on each text block obtained by splitting the thought chain sentence:

[0017] When a function feature that satisfies a feature similarity condition with a text feature of a text block exists in a preset function feature set, an agent corresponding to the function feature is used to generate a target text based on the text block; wherein the function feature set includes the function features of each agent that realizes the content generation function in the same domain;

[0018] Replacing the text block in the thought chain sentence with the target text;

[0019] The second generating unit is used to generate a reply text of the question text based on the processed thought chain sentence by using a preset large model.

[0020] Optionally, when performing an operation on a text block obtained by splitting the thought chain sentence, before replacing the text block in the thought chain sentence with the target text, the execution unit is further configured to perform any one or combination of the following generation operations:

[0021] Using an agent that satisfies a feature matching condition with the text block among the agents implementing the content recommendation function, and generating a recommended text based on the text block;

[0022] Using the agents that meet the feature screening conditions between each agent and the text block among the agents that realize each knowledge acquisition function, and generating a knowledge text associated with an insertion position based on the text block;

[0023] Among the intelligent agents that realize the multimodal data generation function, the intelligent agent that meets the feature hit condition with the one text block generates associated multimodal data based on the one text block.

[0024] Optionally, the generated content further includes a recommendation text, at least one knowledge text, and multimodal data; when the target text is used to replace the text block in the thought chain statement, the execution unit is configured to:

[0025] Generate a control text based on the multimodal data; wherein the control text is used to instruct the large model to reserve an embedding position for the multimodal data in the output result;

[0026] Inserting the corresponding knowledge text in the thought chain sentence according to the insertion position, and splicing the recommended text and the control text with the target text to obtain a processed target text;

[0027] The processed target text is used to replace the text block in the thought chain sentence.

[0028] Optionally, after generating the corresponding reply text, the execution unit is further configured to:

[0029] Adding corresponding at least one multimodal data to at least one embedding position reserved for multimodal data in the reply text to obtain processed integrated data;

[0030] The comprehensive data is displayed to the target object that sent the question text.

[0031] Optionally, when an agent among the agents implementing the content recommendation function that satisfies a feature matching condition with the text block generates a recommended text based on the text block, the execution unit is configured to:

[0032] The feature similarity between each descriptive feature in the descriptive feature set and the text feature of the text block is calculated respectively, and the agents whose corresponding feature similarity reaches a preset recommendation threshold and whose associated priority information meets a preset matching condition are regarded as agents that meet the feature matching condition; the descriptive feature set includes: the descriptive features of each agent that implements the content recommendation function; and a pre-configured priority information indicating the reference degree of an output result of an agent;

[0033] The filtered intelligent agent is used to generate a recommended text based on the text block.

[0034] Optionally, when the intelligent agents that respectively implement each knowledge acquisition function and the one text block satisfy the feature screening condition and generate the knowledge text associated with the insertion position based on the one text block, the execution unit is configured to:

[0035] For each type of knowledge acquisition function, the following operations are performed on the corresponding agent feature sets:

[0036] The feature similarity between each agent feature in an agent feature set and the text feature of the text block is calculated respectively, and the agents whose corresponding feature similarity reaches a preset functional threshold and whose associated priority information meets the preset screening condition are regarded as agents that meet the feature screening condition; an agent feature set includes: the agent features of each agent that realizes a type of knowledge acquisition function;

[0037] The filtered intelligent agent is used to generate a knowledge text associated with an insertion position based on the text block.

[0038] Optionally, when using an agent among the agents implementing the multimodal data generation function that satisfies a feature hit condition with the text block to generate associated multimodal data based on the text block, the execution unit is configured to:

[0039] The feature similarity between each descriptive feature in the identification feature set and the text feature of the text block is calculated respectively, and the agents whose corresponding feature similarity reaches a preset data threshold and whose associated priority information meets the preset hit condition are regarded as agents that meet the feature hit condition; the identification feature set includes: the descriptive features of each agent that realizes the multimodal data generation function;

[0040] The filtered intelligent agent is used to generate multimodal data based on the one text block.

[0041] Optionally, the device further comprises a construction unit, and the construction unit is further configured to:

[0042] The function feature set is constructed in the following manner: using a trained text feature extraction network, based on the function description texts of the corresponding agents, the function features are extracted respectively to obtain the function feature set;

[0043] When constructing a descriptive feature set for each intelligent agent that realizes the content recommendation function, the descriptive feature set is constructed in the following manner: using the text feature extraction network, based on the function description text of each corresponding intelligent agent, respectively extracting descriptive features to obtain a descriptive feature set;

[0044] When constructing an agent feature set for each agent that realizes each knowledge acquisition function, each agent feature set is constructed in the following manner: using the text feature extraction network, based on the function description text of each corresponding agent, the agent features are extracted respectively to obtain the agent feature set;

[0045] When constructing an identification feature set for each intelligent agent that realizes the multimodal data generation function, the identification feature set is constructed in the following way: using the text feature extraction network, based on the corresponding function description text of each intelligent agent, the identification features are extracted respectively to obtain the identification feature set.

[0046] Optionally, the construction unit is further used to:

[0047] In response to a content removal instruction triggered for a target feature set, deleting at least one target feature in the target feature set that matches the content removal instruction; and

[0048] In response to a new content addition indication triggered for a target feature set, a corresponding function description text is obtained for an intelligent agent to be stored, and the target features extracted by the trained text feature extraction network based on the function description text are added to the target feature set; wherein the target feature set is at least one of the description feature set, the function feature set, the feature sets of each intelligent agent, and the identification features.

[0049] Optionally, the functional feature that satisfies a feature similarity condition with a text feature of a text block is determined by the execution unit in the following manner:

[0050] Calculating feature similarities between each functional feature in the functional feature set and the text feature of the text block;

[0051] When there is a functional feature whose corresponding feature similarity reaches a preset data threshold and the associated priority information meets the preset filtering conditions, it is determined that there is a functional feature that meets the feature similarity conditions with the text feature, and the intelligent agent corresponding to the functional feature is used as the intelligent agent that meets the feature similarity conditions.

[0052] Optionally, each text block is split by the execution unit in any one of the following ways:

[0053] In the thought chain sentence, the sentence content is split based on the preset identifier to obtain each text block separated by the identifier;

[0054] In the thought chain sentence, the sentence content is split based on preset sentence keywords to obtain various text blocks starting with the sentence keywords.

[0055] Optionally, after generating the thought chain statement and before performing operations on each text block obtained by splitting the thought chain statement, the execution unit is further configured to:

[0056] Based on the thought chain sentences, splitting to obtain various text blocks;

[0057] A preset text encoding method is used to encode each text block to obtain corresponding text features.

[0058] In a third aspect, an electronic device is proposed, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above method when executing the computer program.

[0059] In a fourth aspect, a computer-readable storage medium is proposed, on which a computer program is stored, and the computer program implements the above method when executed by a processor.

[0060] In a fifth aspect, a computer program product is proposed, comprising a computer program, which implements the above method when executed by a processor.

[0061] The beneficial effects of this application are as follows:

[0062] The present application proposes a generation control method, device, electronic device and storage medium of a large model, and proposes to first obtain the question text to be processed; then use a preset reasoning model to generate a thought chain sentence based on the question text; then, for each text block obtained by splitting the thought chain sentence, perform the following operations respectively: when there is a functional feature in the preset functional feature set that meets the feature similarity condition with the text feature of a text block, use the intelligent agent corresponding to the functional feature to generate the target text based on a text block; wherein the functional feature set includes: the functional features of each intelligent agent that realizes the content generation function in the same field; use the target text to replace a text block in the thought chain sentence This block; in this way, by generating a thought chain sentence based on the question text, the reasoning process and task analysis results based on the prompt text can be obtained; then, for each text block in the thought chain sentence, by screening for feature similarity, the matching intelligent agent can be flexibly triggered, and, in the case of determining that a text block has a matching intelligent agent, by using the target text generated by the intelligent agent based on the text block to replace the text block in the thought chain sentence, the thought chain sentence can be adjusted, thereby intervening in the model reasoning process, so that the logical errors in the thought chain sentence can be corrected in a timely manner, greatly reducing the reasoning impact caused by problems such as semantic pollution and knowledge illusion;

[0063] Furthermore, a preset big model is used to generate a reply text for the question text based on the processed thought chain sentences; in this way, the reply file finally obtained is generated by the big model based on the adjusted thought chain sentences. This makes it possible to break through the black box limitations of the traditional big model reasoning process by adjusting the thought chain sentences, thereby fundamentally realizing the generation control of the big model, ensuring the accuracy of the model generation results, and improving the generation effect of the big model. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 A schematic diagram of possible application scenarios in the embodiments of the present application;

[0065] Figure 2 This is a schematic diagram of the generation control process of the large model in the embodiment of the present application;

[0066] Figure 3 This is a schematic diagram of an input page for a prompt text in an embodiment of the present application;

[0067] Figure 4 This is a schematic diagram of another input page for prompt text in an embodiment of the present application;

[0068] Figure 5 Schematic diagram of the screening process of matching agents in an embodiment of the present application;

[0069] Figure 6 This is a schematic diagram of an adjustment of a thought chain statement in an embodiment of the present application;

[0070] Figure 7 This is a schematic diagram of the internal composition of the processing device in the embodiment of the present application;

[0071] Figure 8 This is a schematic diagram of the generation control process of the large model in the embodiment of the present application;

[0072] Figure 9 A schematic diagram of the process of determining a matching agent in an embodiment of the present application;

[0073] Figure 10 This is a schematic diagram of the result after adjusting the thought chain sentence in the embodiment of this application;

[0074] Figure 11 Schematic diagram of the logical structure of the generation control device of the large model in the embodiment of the present application;

[0075] Figure 12 A schematic diagram of the hardware structure of an electronic device to which the embodiments of the present application are applied;

[0076] Figure 13 The figure is a schematic diagram of the hardware structure of another electronic device to which the embodiments of the present application are applied. DETAILED DESCRIPTION

[0077] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of the technical solutions of this application, but not all of them. Based on the embodiments described in this application document, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the technical solutions of this application.

[0078] The terms "first," "second," and the like in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate, such that the embodiments of the invention described herein can be practiced in sequences other than those illustrated or described herein.

[0079] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0080] The following explains some of the terms used in the embodiments of the present application to facilitate understanding by those skilled in the art.

[0081] An agent is an entity that can autonomously perceive the environment, make decisions, and perform corresponding actions to achieve a specific task. It can be pure software or a combination of software and hardware. In addition, the agent can realize a closed loop of perception, decision-making, and execution. In the decision-making process, the perceived information can be deeply processed to ultimately obtain the decision result. In the embodiments of this application, the agent involved can make decisions based on a large model or based on preset logic or strategy during the decision-making process. This application does not impose specific restrictions on this.

[0082] Large model: also known as Large Language Model (LLM), is an ultra-large deep learning model pre-trained based on a large amount of data.

[0083] Chain-of-Thought (CoT) is an AI method that enables large models to reason step by step. It breaks down complex tasks into a series of logical steps that ultimately lead to a solution, simulating a human-like reasoning process. This approach reflects the fundamental characteristics of AI and provides a structured problem-solving mechanism. In other words, Chain-of-Thought, based on cognitive strategies, breaks down complex problems into manageable intermediate ideas, which in turn guide the model to the final result.

[0084] Thought chain sentence: refers to the sentence content obtained by the reasoning model based on the question text with the help of thought chain technology. In some feasible embodiments of the present application, the reasoning model and the preset big model can be the same model. In this case, it can be understood that content generation is performed separately under different thinking models of a big model. For example, assuming that thought chain sentences can be generated in the deep thinking mode, then in the generation control method requested for protection in the present application, it can be set to first use the big model in the deep thinking mode to obtain the thought chain sentences generated based on the question text, and set the big model to stop generating after outputting the thought chain sentences, and then after completing the processing of the thought chain sentences, use the big model under the thinking model that does not need to generate thought chain sentences, and generate the final reply text based on the processed thought chain sentences; in other feasible embodiments of the present application, the reasoning model and the preset big model can be different models.

[0085] Context Loss: This refers to the situation where the model forgets or confuses previous information when processing long texts or multi-round conversations, resulting in inconsistent responses.

[0086] Knowledge Hallucination: refers to the situation where a model generates content that is inconsistent with the facts but appears to be reasonable. In essence, it is the result of the model "fictitious" knowledge rather than real retrieval or reasoning.

[0087] Semantic contamination: refers to the phenomenon in natural language processing (NLP) where a language model's semantic understanding of words, phrases, or sentences is distorted or contaminated due to low-quality data, noisy input, or model bias. This is manifested by the model making incorrect associations with certain words or concepts, and outputting content that is inconsistent with expectations or biased.

[0088] The following is a brief introduction to the design concept of the embodiment of this application:

[0089] With the development of artificial intelligence (AI) technology, question-answering tasks can be completed well with the help of large models.

[0090] At present, in order to improve the generation control effect of large models, context information can be added to the prompt template, or relevant knowledge retrieved based on the input content can be added to the prompt template, so that the output effect of the large model can be improved to a certain extent.

[0091] However, the above-mentioned optimization and adjustment method can only insert relevant knowledge or context information before reasoning, so it can only intervene in the input content of the large model. Therefore, once there is a lack of context information or the database is not updated in time, it is impossible to provide an effective reasoning basis for the large model, thereby failing to fundamentally improve the generation control effect of the large model, and thus it is difficult to ensure the accuracy of the model generation results.

[0092] In view of this, the present application proposes a generation control method, device, electronic device and storage medium of a large model, and proposes to first obtain the question text to be processed; then use a preset reasoning model to generate a thought chain sentence based on the question text; thereafter, for each text block obtained by splitting the thought chain sentence, the following operations are performed respectively: when there is a functional feature in the preset functional feature set that meets the feature similarity condition with the text feature of a text block, the intelligent agent corresponding to the functional feature is used to generate a target text based on a text block; wherein the functional feature set includes: the functional features of each intelligent agent that realizes the content generation function in the same field; the target text is used to replace the text in the thought chain sentence a text block; in this way, by generating a thought chain sentence based on the question text, the reasoning process and task analysis results based on the prompt text can be obtained; then, for each text block in the thought chain sentence, by screening for feature similarity, the matching intelligent agent can be flexibly triggered. Moreover, when it is determined that a text block has a matching intelligent agent, the target text generated by the intelligent agent based on the text block is used to replace the text block in the thought chain sentence, thereby adjusting the thought chain sentence and intervening in the model reasoning process, so that the logical errors in the thought chain sentence can be corrected in time, greatly reducing the reasoning impact caused by problems such as semantic pollution and knowledge illusion.

[0093] Furthermore, a preset big model is used to generate a reply text for the question text based on the processed thought chain sentences; in this way, the reply file finally obtained is generated by the big model based on the adjusted thought chain sentences. This makes it possible to break through the black box limitations of the traditional big model reasoning process by adjusting the thought chain sentences, thereby fundamentally realizing the generation control of the big model, ensuring the accuracy of the model generation results, and improving the generation effect of the big model.

[0094] The preferred embodiments of the present application are described below in conjunction with the drawings in the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application and are not used to limit the present application. In addition, the embodiments of the present application and the features in the embodiments may be combined with each other if there is no conflict.

[0095] See Figure 1 , which is a schematic diagram of a possible application scenario in an embodiment of the present application. The schematic diagram of the application scenario includes a client device 110 and a processing device 120.

[0096] In a feasible embodiment of the present application, the processing device 120 can receive the question text sent by the target object on the client device 110, and perform the following processing based on the prompt text: using a preset reasoning model, based on the question text, generate a thought chain sentence; for each text block obtained by splitting the thought chain sentence, perform the following operations respectively: when there is a functional feature in the preset functional feature set that meets the feature similarity condition with the text feature of a text block, use the intelligent agent corresponding to the functional feature to generate the target text based on a text block; wherein the functional feature set includes: the functional features of each intelligent agent that realizes the content generation function in the same field; use the target text to replace a text block in the thought chain sentence; use the preset large model to generate a reply text for the question text based on the processed thought chain sentence.

[0097] Optionally, in some feasible embodiments of the present application, after obtaining the processed thought chain sentence, the processing device can use a preset large model to generate a reply text for the prompt text based on the question text and the processed thought chain sentence.

[0098] Afterwards, the processing device 120 sends the reply text to the client device 110 to be displayed to the target object.

[0099] Optionally, in a feasible embodiment of the present application, after the processing device obtains the reply text, it can add additional determined multimodal data on the basis of the reply text to obtain comprehensive data, and send the comprehensive data to the client device 110 to display it to the target object, wherein the multimodal data can cover any one or combination of the following data types: video, audio, and image.

[0100] The client device 110 includes but is not limited to mobile phones, tablet computers, notebooks, e-book readers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc.

[0101] The processing device 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0102] In the embodiment of the present application, the client device 110 and the processing device 120 can communicate via a wired network or a wireless network. The following description only describes the generation control process of the large model from the perspective of the processing device 120.

[0103] The following is a schematic illustration of the generation and control process of a large model based on possible application scenarios:

[0104] Application scenario 1: Generate and control the intelligent robot in the question-answering scenario of the intelligent robot.

[0105] Among them, intelligent robots can specifically be intelligent customer service robots, as well as any type of question-answering robots in various fields; various fields include but are not limited to education, medical care, catering, travel, etc.; intelligent robots can be manifested as software built based on large models.

[0106] Specifically, after obtaining the question text sent by the target subject through the chat window associated with the intelligent robot, the processing device can first obtain the thought chain sentence generated by the intelligent robot in deep thinking mode, then split the thought chain sentence into each text block, and perform the following operations: by calculating feature similarity, determine that there is a matching intelligent agent among the intelligent agents that implement content generation functions in the same domain, and use the target text generated by the matching intelligent agent based on the text block to replace one text block in the thought chain sentence. Then, in a mode that does not require the output of the thought chain sentence, obtain the reply text output by the intelligent robot based on the processed thought chain sentence.

[0107] Then, the processing device displays the reply text to the target object through the chat window associated with the intelligent robot.

[0108] Application scenario 2: In the content generation scenario of a dedicated large model, the generation of the large model is controlled.

[0109] The processing equipment can develop a large model that implements the specified business processing function according to the processing needs of the developer. Then, after the training of the large model is completed, generation control is performed in the content generation scenario.

[0110] Specifically, after receiving the question text sent by the target subject, the processing device uses a preset reasoning model to generate a thought chain statement based on the question text. It then performs the following operations on each text block of the split thought chain statement: By calculating feature similarity, it determines that a matching agent exists among the agents that implement content generation functions in the same domain. If a matching agent exists, the target text generated by the matching agent based on the text block is used to replace one of the text blocks in the thought chain statement. Then, the preset macro model is used to generate a response text based on the processed thought chain statement.

[0111] Then, the reply text is displayed to the target object; wherein the target object may be the user of the generation function of the large model.

[0112] In addition, it should be understood that in the specific implementation of this application, the generation control process of the large model is involved. When the embodiments recorded in this application are applied to specific products or technologies, the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0113] The following describes the generation and control process of the large model from the perspective of the processing equipment, with reference to the accompanying drawings:

[0114] See Figure 2 As shown, it is a schematic diagram of the generation control process of the large model in the embodiment of the present application. Figure 2 , for specific instructions:

[0115] Step 201: The processing device obtains the question text to be processed.

[0116] In the embodiment of the present application, according to the processing needs, the processing device can obtain the question text based on the large model processing in any scenario where large model processing is required.

[0117] For example, the question text can be entered in a chat window associated with the large model; or, the question text can be entered in a web page associated with the large model; or, the question text can be received by other devices and forwarded to the processing device. This application does not impose any specific restrictions on this.

[0118] For example, see Figure 3 As shown, it is a schematic diagram of an input page for a prompt text in an embodiment of the present application. Figure 3 As can be seen from the illustrated content, the processing device can obtain the question text entered by the target object in the web page.

[0119] For example, see Figure 4 As shown, it is a schematic diagram of another input page for prompt text in the embodiment of the present application. Figure 4As can be seen from the illustrated content, the processing device can obtain the question text entered by the target object in the chat window.

[0120] Step 202: The processing device uses a preset reasoning model to generate a thought chain sentence based on the question text.

[0121] In some feasible embodiments of the present application, the preset reasoning model may be the same as the big model based on which the reply text is finally generated; in this case, the big model has at least two reasoning modes, namely, a reasoning mode that can output thought chain statements, and a reasoning mode that does not require the output of thought chain statements, so that the generation control of the big model can be achieved by adjusting the thought chain statements obtained by reasoning the big model itself.

[0122] It should be understood that, in a reasoning mode capable of outputting thought chain statements, generally speaking, the reasoning results generated include thought chain statements and output text. However, this application controls the large model through pre-programmed settings to stop content generation after outputting the thought chain statement in a reasoning mode capable of outputting thought chain statements.

[0123] In other feasible embodiments of the present application, the preset reasoning model may be different from the big model based on which the reply text is finally generated; in this case, the big model used for generation control in the present application is different from the big model used for generating thought chain sentences, so that the generation control of the big model preset in the present application can be realized based on the internal reasoning process of other big models.

[0124] In the specific process of generating thought chain sentences, the processing device can obtain the thought chain sentences obtained by the reasoning model based on the question text under the action of thought chain technology based on the reasoning ability of the reasoning model.

[0125] Step 203: The processing device performs operations on each text block obtained by splitting the thought chain sentence.

[0126] In an embodiment of the present application, after the processing device obtains the thought chain sentence, it can split the thought chain sentence to obtain various text blocks; wherein, a text block refers to a paragraph in the thought chain sentence; or, refers to a processing step in the thought chain sentence.

[0127] When splitting the thought chain sentence to obtain each text block, in some feasible embodiments, the processing device can split the sentence content in the thought chain sentence based on a preset identifier to obtain each text block separated by the identifier; in other feasible embodiments, the processing device can split the sentence content in the thought chain sentence based on a preset sentence keyword to obtain each text block starting with the sentence keyword.

[0128] It should be noted that in the application embodiment, a text block can be understood as a logical unit split based on a thought chain statement, wherein a logical unit includes an inference step, or a sub-question conclusion, or a content paragraph; in addition, when the processing device splits the thought chain statement to obtain each text block, it can optionally use a preset parsing engine for processing, so that real-time block processing of the thought chain statement can be realized.

[0129] In an embodiment of the present application, the preset identifier may be a symbol indicating a segment.

[0130] For example, paragraph breaks, periods, semicolons, etc.

[0131] In the embodiment of the present application, the preset statement keywords may be words that can distinguish logical processing levels; or, they may be words that can reflect the order of execution.

[0132] For example, first, second, then, next, next…, first, second, third…

[0133] In this way, the thought chain sentences generated by the reasoning model during the reasoning process can be segmented in real time to obtain text blocks that can be processed independently.

[0134] Optionally, in an embodiment of the present application, the processing device can split the text blocks based on the thought chain statement after generating the thought chain statement and before performing operations on each text block separately; and then use a preset text encoding method to encode the corresponding text features for each text block separately.

[0135] Among them, the preset text encoding method can be any encoding algorithm that can realize text encoding, and this application does not impose any specific restrictions on this; according to actual processing needs, the encoding algorithm used to obtain text features can be the same as or different from the algorithm used to extract the functional features of the intelligent agent, but it is necessary to ensure that the text features and the functional features of the intelligent agent are in the same vector space for similarity comparison, and this application does not impose any specific restrictions on this.

[0136] For example, the encoding algorithms that can be used include but are not limited to any of the following: Simple Contrastive Learning of Sentence Embeddings (SimCSE), Bidirectional Encoder Representations from Transformers (BERT), Robust Optimized BERT Method (RoBERTa), Lightweight BERT (ALite BERT, ALBERT), Distilled BERT (Distilled BERT, DistilBERT), Embeddings from Language Models (ELMo), Universal Sentence Encoder (USE), and Document to Vector (Doc2Vec), etc.

[0137] In this way, by performing feature encoding on each text block, a high-dimensional feature representation corresponding to each text block can be obtained, that is, each text feature can be obtained.

[0138] The following describes the processing involved, taking the operation performed on an arbitrary text block as an example:

[0139] Step 2031: When a functional feature that satisfies a feature similarity condition with a text feature of a text block exists in the preset functional feature set, the processing device uses the intelligent agent corresponding to the functional feature to generate a target text based on a text block; wherein the functional feature set includes: the functional features of each intelligent agent that realizes the content generation function in the same field.

[0140] In some feasible embodiments of the present application, when performing an operation on a text block, the processing based on the agent may be the generation of a target text. In this case, the processing device needs to first obtain a preset set of functional features to determine each agent that can generate replaceable content (i.e., the target text); wherein the functional feature set includes: the functional features of each agent that realizes the function of generating content in the same domain, and content in the same domain refers to content in the same domain as the content of the text block.

[0141] Optionally, in order to ensure the diversity and effectiveness of the selectable agents in the functional feature set, the processing device can add functional features of available agents and delete functional features of unavailable agents in response to an update operation on the functional feature set.

[0142] Alternatively, a functional feature set can be constructed by using a trained text feature extraction network to extract functional features from the functional description text of each agent, thereby generating a functional feature set. The training method for the text feature extraction network will be detailed in the following sections. The functional description text is used to describe the processing functions that an agent can perform.

[0143] It should be understood that in the embodiments of the present application, there is no specific restriction on the construction method of the intelligent agent. Moreover, in the embodiments of the present application, intelligent agents that realize the content generation function of various fields can be flexibly obtained, thereby generating target text that can effectively replace text blocks; moreover, an intelligent agent is used to generate content in at least one field, so it has a very good generation effect in the content generation of the specified field, which can guarantee the accuracy of the replacement content to a certain extent, thereby effectively realizing the content correction of the text blocks in the thought chain sentences.

[0144] In some feasible methods of generating intelligent agents, after determining the processing tasks that the intelligent agent can perform, the processing device constructs a training sample set that is adapted to the processing tasks, and selects a basic model from various optional large models; then, the basic model is iteratively trained using the training sample set to obtain a trained target model, and the trained target model is used as an intelligent agent; wherein, a training sample includes: a sample input sentence in a field, and a sample output sentence obtained after content regeneration for the sample input sentence; the sample output sentence in a training sample is generally more accurate than the sample input sentence; this application does not impose specific restrictions on the method of selecting the basic model, for example, any one can be selected as the basic model.

[0145] In other feasible intelligent agent generation schemes, after determining the processing tasks that the intelligent agent can perform, different prompt templates can be constructed based on different processing tasks, so that different task processing can be guided by different prompt templates.

[0146] When specifically determining whether there are functional features in a preset functional feature set that meet the feature similarity condition with the text features of a text block, in some feasible embodiments, the processing device can process based on feature similarity and use the intelligent agent with the highest feature similarity as the intelligent agent that meets the feature similarity condition; in other feasible embodiments, the processing device can process based on feature similarity and associated priority information, and use the intelligent agent whose feature similarity meets the data threshold and has the highest priority as the intelligent agent that meets the feature similarity condition.

[0147] When screening based on feature similarity, the processing device may use an agent with the highest feature similarity between the corresponding functional feature and the text feature as the agent that meets the feature similarity condition.

[0148] When screening based on feature similarity and associated priority information, the processing device can respectively calculate the feature similarity between each functional feature in the functional feature set and the text feature of a text block; then, when there is a functional feature whose corresponding feature similarity reaches a preset data threshold and the associated priority information meets the preset filtering conditions, it is determined that there is a functional feature that meets the feature similarity conditions with the text feature, and the intelligent agent corresponding to the functional feature is regarded as the intelligent agent that meets the feature similarity conditions.

[0149] Among them, there is a relative priority between each intelligent agent of the same type; the priority information is used to indicate the reference degree for the output content of different intelligent agents; the preset filtering condition can be: determining the relatively highest priority based on the associated priority information; the data threshold is set according to actual processing needs, and this application does not impose specific restrictions on this. The relatively highest is relative to the intelligent agent whose feature similarity meets the data threshold, so as to achieve the determination of the intelligent agent with the highest priority among the intelligent agents whose feature similarity meets the data threshold.

[0150] For example, an agent that implements fixed-domain content generation is prioritized over an agent that implements general-domain content generation.

[0151] For example, the priority of an agent that implements "pathological analysis" is higher than that of an agent that implements "general question answering".

[0152] Optionally, when screening intelligent agents whose feature similarity reaches a preset data threshold, the processing device can use various retrieval algorithms for processing, so as to quickly match intelligent agents that can process functions similar to the content of the text block from the perspective of feature similarity. The retrieval algorithms used include but are not limited to any one of the following: Hierarchical Navigable Small World (HNSW), Facebook AI Similarity Search (FAISS) algorithm, and Locality-Sensitive Hashing (LSH) algorithm, etc.

[0153] For example, see Figure 5 As shown, it is a schematic diagram of the screening process of the matching intelligent agent in the embodiment of the present application. Figure 5As can be seen from the schematic content, for each intelligent agent that realizes the content generation function in the same field, a functional feature set of each intelligent agent and its associated priority information are determined respectively; then, by calculating the text features of the text block and the feature similarity between each functional feature, after screening out the intelligent agents that reach the data threshold (assuming they are intelligent agents s1 and intelligent agents s2), by comparing the difference in priority information of the two screened intelligent agents, the intelligent agent with the relatively highest priority is taken as the intelligent agent that meets the feature similarity condition, that is, intelligent agent s1.

[0154] In this way, each functional feature in the functional feature set is equivalent to the feature anchor point of each intelligent agent to be matched, so that by comparing the feature similarity and screening the priority information, the intelligent agent finally screened out can greatly meet the processing needs of the text block; moreover, by defining the priority information, there is a trigger priority between different intelligent agents, which can resolve the competition conflicts between multiple intelligent agents.

[0155] Furthermore, among the various intelligent agents that realize the content generation function in the same field, after determining an intelligent agent that matches a text block, the matching intelligent agent can be used to generate the target text based on a text block.

[0156] In particular, when there is no functional feature in the preset functional feature set that meets the feature similarity condition with the text feature of a text block, it can be determined that there is no intelligent agent that matches the text block among the intelligent agents that implement the content generation function in the same field, and then the processing of the current text block can be stopped, and the operation of step 2031 can be performed for the next text block.

[0157] In other feasible embodiments of the present application, when performing an operation on a text block, the processing based on the intelligent agent can not only generate the target text, but also generate any one or combination of the following: recommended text, various types of knowledge text, and multimodal data.

[0158] Based on this, when performing an operation on a text block, before using the target text to replace a text block in the thought chain statement, you can also perform any one or combination of the following generation operations: use the intelligent agent that meets the feature matching conditions between the intelligent agents that implement the content recommendation function and a text block to generate a recommended text based on a text block; use the intelligent agent that meets the feature screening conditions between the intelligent agents that implement each knowledge acquisition function and a text block to generate associated knowledge text with an insertion position based on a text block; use the intelligent agent that meets the feature hit conditions between the intelligent agents that implement the multimodal data generation function and a text block to generate associated multimodal data based on a text block.

[0159] That is, in the embodiments of the present application, the types of agents used may also include any one or a combination of the following: agents that implement content recommendation functions, various agents that implement knowledge recommendation functions, and agents that implement multimodal data generation functions. Furthermore, each agent is associated with priority information, and a priority relationship exists between agents of the same type. Optionally, based on actual processing needs, for each type of agent that implements knowledge acquisition functions, a corresponding insertion position can be configured for each agent to indicate the insertion position of the knowledge text processed by the agent in the thought chain sentence.

[0160] When generating the recommended text, the processing device can respectively calculate the feature similarity between each descriptive feature in the descriptive feature set and the text feature of a text block, and use the intelligent agent whose corresponding feature similarity reaches a preset recommendation threshold and whose associated priority information meets the preset matching conditions as the intelligent agent that meets the feature matching conditions; the descriptive feature set includes: the descriptive features of each intelligent agent that realizes the content recommendation function; a pre-configured priority information is used to indicate the reference degree of the output result of an intelligent agent; and then the screened intelligent agent is used to generate the recommended text based on a text block.

[0161] Among them, in a feasible embodiment, the content recommendation function can be specifically reflected as advertising content recommendation, so that advertising content that matches the text block can be recommended; the value of the recommendation threshold is set according to actual processing needs, and the preset matching condition can be: the priority is relatively the highest, so as to determine the intelligent agent with the highest priority among the intelligent agents whose feature similarity reaches the recommendation threshold.

[0162] An intelligent agent that implements the content recommendation function specifically implements determining a recommended text that matches a text block in a category of recommended content.

[0163] For example, the various intelligent agents that implement content recommendation functions may include intelligent agents that implement advertising recommendation functions for various types of products, such as intelligent agents that implement advertising recommendation functions for sports shoes, intelligent agents that implement advertising recommendation functions for ski suits, etc., wherein each intelligent agent can determine an advertising sentence that matches a text block among the various advertising sentences for a category of goods.

[0164] In this way, from the perspective of content recommendation, recommended text that is suitable for the text block can be determined.

[0165] When generating each knowledge text, the processing device performs the following operations for the agent feature sets corresponding to each type of knowledge acquisition function: calculates the feature similarity between each agent feature in an agent feature set and the text feature of a text block, and takes the agent whose corresponding feature similarity reaches a preset function threshold and whose associated priority information meets the preset screening conditions as the agent that meets the feature screening conditions; an agent feature set includes: the agent features of each agent that realizes a type of knowledge acquisition function; and then uses the screened agents to generate a knowledge text with an associated insertion position based on a text block.

[0166] In the embodiment of the present application, there may be multiple agents with knowledge acquisition functions, which can provide knowledge texts of different dimensions; the multiple knowledge acquisition functions include but are not limited to the following functions: real-time knowledge search, acquisition of formula calculation knowledge, search and generation of database knowledge, etc. When determining matching agents under different types of knowledge acquisition functions, the functional thresholds may be the same or different, and the specific values are set according to actual processing needs; the preset screening conditions can be: the priority is relatively the highest, so that when there is an agent that meets the feature screening conditions among the agents of a knowledge acquisition function, the agent selected according to the feature screening conditions is specifically the agent with the highest priority among the agents whose feature similarity between the agent feature and the text feature reaches the functional threshold.

[0167] In addition, real-time knowledge can be real-time knowledge obtained by the intelligent agent through the application program interface. The intelligent agents included in the class of intelligent agents involved are, for example, intelligent agents that realize real-time weather information acquisition, intelligent agents that realize implementation location information acquisition, etc.; a class of intelligent agents that obtain formula calculation knowledge can use pre-built calculation program formulas to process and obtain corresponding calculation results. The intelligent agents included in the class of intelligent agents involved are, for example, intelligent agents that realize formula 1 calculation, intelligent agents that realize formula 2 calculation, etc.; intelligent agents that realize database knowledge search generation can search for relevant knowledge of text blocks in the database. The intelligent agents included in the class of intelligent agents involved are, for example, intelligent agents that search in the database of domain 1, intelligent agents that search in the database of domain 2, etc.

[0168] In this way, with the help of various knowledge acquisition functions, intelligent agents can obtain knowledge content that matches the text block from different knowledge dimensions, thereby realizing the expansion of knowledge content based on the text block, thereby providing more reference basis for the analysis and generation of subsequent large models.

[0169] When specifically generating multimodal data, the processing device can respectively calculate the feature similarity between each descriptive feature in the identification feature set and the text feature of a text block, and use the intelligent agent whose corresponding feature similarity reaches a preset data threshold and whose associated priority information meets the preset hit condition as the intelligent agent that meets the feature hit condition; the identification feature set includes: the descriptive features of each intelligent agent that realizes the multimodal data generation function; and then use the screened intelligent agent to generate multimodal data based on a text block.

[0170] Among them, in a feasible embodiment, the multimodal data generation function is specifically used to generate multimodal data related to the text block, such as pictures, videos, audio, etc.; the value of the data threshold is set according to the actual processing needs, and the preset hit condition can be: the priority is relatively the highest, so that when there is an intelligent agent that meets the feature hit condition, an intelligent agent selected according to the feature hit condition is specifically an intelligent agent with the highest priority among the intelligent agents whose feature similarity between the description feature and the text feature reaches the data threshold.

[0171] The specific function implemented by an intelligent agent that realizes the multimodal data generation function may be to generate a multimodal type of data related to a text block.

[0172] For example, the various intelligent agents that implement multimodal data generation may include intelligent agents that generate various types of multimodal data, such as intelligent agents that implement image generation, intelligent agents that implement audio generation, and intelligent agents that implement video generation.

[0173] In this way, relevant multimodal data can be expanded and generated based on the text blocks, which can greatly improve the richness of the content.

[0174] In summary, when performing operations on a text block, the processing device can also selectively generate recommended texts, various knowledge texts, and any one or combination of multimodal data, thereby improving the flexibility of content generation and enabling the content of the text block to be expanded from different dimensions; moreover, it can determine whether there are agents that match the text block from the perspective of different agent types, thereby increasing the range of agents involved in adjusting the logical chain statements.

[0175] In addition, in an embodiment of the present application, when a description feature set is constructed for each intelligent agent that realizes the content recommendation function, the description feature set is constructed in the following manner: a text feature extraction network is used to extract description features based on the function description text of each corresponding intelligent agent to obtain a description feature set; when an intelligent agent feature set is constructed for each intelligent agent that realizes each knowledge acquisition function, each intelligent agent feature set is constructed in the following manner: a text feature extraction network is used to extract intelligent agent features based on the function description text of each corresponding intelligent agent to obtain an intelligent agent feature set; when an identification feature set is constructed for each intelligent agent that realizes the multimodal data generation function, the identification feature set is constructed in the following manner: a text feature extraction network is used to extract identification features based on the function description text of each corresponding intelligent agent to obtain an identification feature set.

[0176] It should be understood that in the embodiments of the present application, the features extracted for each intelligent agent can be extracted using the same feature extraction network; optionally, the processing device can train different feature extraction networks to perform feature extraction for different types of intelligent agents, but it is necessary to ensure that the features extracted for the intelligent agent and the text features of the text block are in the same vector space in order to calculate the feature similarity. This application does not impose specific restrictions on this.

[0177] In this way, the trained text feature extraction network can be used to implement feature extraction for each intelligent agent based on the functional description text of each intelligent agent.

[0178] That is to say, in a feasible embodiment of the present application, when performing an operation on a text block, in addition to using the functional feature set, any one or combination of the following feature sets may also be used: a description feature set, each intelligent agent feature set, and an identification feature set; in order to ensure the validity of the content in each feature set, the processing device can update the feature set used during the operation execution.

[0179] Specifically, the processing device can respond to a content removal instruction triggered for a target feature set, and delete at least one target feature in the target feature set that matches the content removal instruction; and respond to a content addition instruction triggered for a target feature set, obtain the corresponding function description text for an intelligent agent to be stored, and add the target features extracted by the trained text feature extraction network based on the function description text to the target feature set; wherein the target feature set is at least one of a description feature set, a function feature set, a feature set of each intelligent agent, and an identification feature.

[0180] Among them, the intelligent body to be stored in the warehouse can be obtained by the processing equipment, or it can be developed by the processing equipment itself, and this application does not impose specific restrictions on this; the update process of the target feature set can be carried out in real time, or it can be carried out before the operation is executed, and this application does not impose specific restrictions on this.

[0181] In this way, by updating the content of the target feature set that may be used during the execution of the operation, the target feature set used during the execution of the operation is the latest, ensuring the reliability of the content in the target feature set, thereby improving the matching effect of the intelligent agent; moreover, it can realize the hot update of the intelligent agent, thereby adding and removing the available intelligent agents in real time.

[0182] In addition, in a feasible embodiment of the present application, the text feature extraction network used can be trained in an unsupervised or supervised manner. The text feature extraction network, which can also be called a text encoding network, is used to extract features from text content. In other feasible embodiments of the present application, the text feature extraction network used can be directly obtained from other devices.

[0183] For example, assuming that the text feature extraction network is BERT, an unsupervised method can be used for training to obtain the text feature extraction network; this application does not specifically limit the training method of the text feature extraction network.

[0184] Step 2032: The processing device uses the target text to replace a text block in the thought chain sentence.

[0185] In some feasible embodiments of the present application, when performing an operation on a text block, if the only sentence that the agent can generate is the target text, the generated target text can be used to replace the text block in the thought chain.

[0186] In other feasible embodiments of the present application, in the process of performing operations on a text block, when in addition to generating the target text, recommended texts, various knowledge texts, and any one or combination of multimodal data sentences are also generated, it is also necessary to configure other sentences besides the target text in the thinking chain sentence.

[0187] Specifically, for recommended text, the recommended text can be directly inserted into the thought chain statement; for each knowledge text, it can be inserted into the thought chain statement according to the associated insertion position; for multimodal data, a corresponding control text can be generated and inserted into the thought chain statement, where the control text is used to instruct the large model to reserve the embedding position of the multimodal data in the output result.

[0188] Taking the case where, during the process of performing an operation on a text block, the generated content also includes recommended text, at least one knowledge text, and multimodal data, when executing step 2023, the processing device can generate a control text based on the multimodal data; then, in the thinking chain statement, the corresponding knowledge text is inserted according to the insertion position, and the recommended text and the control text are spliced with the generated target text to obtain a processed target text; thereafter, the processed target text is used to replace a text block in the thinking chain statement.

[0189] See Figure 6 As shown, it is a schematic diagram of an adjustment of the thought chain sentence in the embodiment of the present application. Figure 6 As can be seen from the illustrated content, assuming that a text block currently being processed is text block 1, when an operation is performed on text block 1, target text 1, recommended text 1, knowledge text 1 (the associated insertion position is: after the text block), knowledge text 2 (the associated insertion position is: after the thinking chain statement) and control text generated by the corresponding multimodal data are generated accordingly; then, in the process of operating on text block 1, when the thinking chain statement is adjusted once, the splicing result of target text 1, recommended text 1, control text, and knowledge text 1 can replace text block 1, and knowledge text 2 can be added at the end of the thinking chain statement.

[0190] In this way, each time an operation is performed on a text block, if a statement is generated during the operation, an adjustment operation can be performed on the thought chain statement, so that the content of the text block in the thought chain statement can be adjusted based on the regenerated statement.

[0191] Similarly, the processing device can perform operations on each text block in the thought chain sentence respectively until all text blocks in the thought chain sentence are processed to obtain the processed thought chain sentence.

[0192] Step 204: The processing device uses a preset large model to generate a reply text for the question text based on the processed thought chain sentences.

[0193] In some feasible embodiments of the present application, after the processing device obtains the processed thought chain sentence, it can use a preset large model to generate a reply text for the question text based on the processed thought chain sentence.

[0194] In some other feasible embodiments of the present application, after the processing device obtains the processed thought chain sentence, it can use a preset large model to generate a reply text for the question text based on the processed thought chain sentence and the question text.

[0195] The present application does not impose any specific restrictions on the type of large model selected.

[0196] In particular, when the processed thought chain sentence contains control text, it means that an embedding position is reserved for multimodal data in the reply text output by the large model; in this case, after generating the reply text, the processing device adds at least one corresponding multimodal data to at least one embedding position reserved for multimodal data in the reply text to obtain the processed comprehensive data; and then displays the comprehensive data to the target object that sent the question text.

[0197] It should be noted that the reason why there is at least one embedding position in the reply text is that multimodal data may be generated for each text block. Therefore, from an overall perspective, it may be necessary to reserve multiple embedding positions for multimodal data.

[0198] In this way, in the generation scenario of large models, the data presented to the target object is no longer limited to text content, but can also present multimodal data, which greatly improves the diversity of content generated by large models.

[0199] The following describes the core components of the processing device from the perspective of its internal functions.

[0200] See Figure 7 As shown, it is a schematic diagram of the internal composition of the processing equipment in the embodiment of the present application. Figure 7 As shown in the figure, the components inside the processing device include: input content receiving component, thought chain sentence parsing engine, intelligent agent collaboration component, dynamic semantic routing component, hybrid generation control component, and large model result output component; among them,

[0201] Input content receiving component: used to receive the question text input by the target object.

[0202] Thought chain statement parsing engine: internal functions include real-time block processing, as well as state extraction and encoding; real-time block processing refers to the real-time segmentation of thought chain text streams (or thought chain statements) according to logical units (such as reasoning steps, sub-problem conclusions) during the model reasoning process to generate independently processable text blocks; state extraction and encoding refers to the feature encoding of the segmented thought chain fragments (i.e., text blocks), thereby converting them into high-dimensional semantic representations (i.e., text features).

[0203] Dynamic semantic routing component: Its internal functions include feature anchor library maintenance and feature similarity matching. Feature anchor library maintenance refers to extracting features based on the function description text for each matching agent, thereby forming the feature anchor points corresponding to each agent. Feature similarity matching refers to quickly determining the agent that matches the text block based on an efficient retrieval algorithm. Among them, the various thresholds for measuring feature similarity (such as data thresholds, etc.) can be dynamically adjusted according to actual processing needs.

[0204] Agent collaboration component: internal functions include plug-and-play agent and priority scheduling mechanism processing; plug-and-play agent pool refers to support for the access of diverse agents, where the connected agents can include: agents with functions such as knowledge retrieval, data analysis, and domain experts. Each agent can independently process input and return structured results; the priority scheduling mechanism can trigger priorities according to the defined agents to solve the problem of multi-agent competition and conflict.

[0205] Hybrid generation control component: Its internal functions include multi-source result fusion and agent screening process optimization; multi-source result fusion refers to the fusion of sentences generated by various agents to ensure that the generated content is both universal and professional; agent screening process optimization refers to the adjustment of the various thresholds used when evaluating feature similarity, where the various thresholds involved can be any one or a combination of the recommended thresholds, functional thresholds, and data thresholds mentioned in the above description.

[0206] In this way, the technical solution proposed in this application can achieve efficient collaboration between a large language model (or large model) and an external intelligent agent. By means of the characteristics of each intelligent agent that can be dynamically added and the "bus" processing around the thought chain statements, the reasoning process of the large model can be transformed into a programmable and interventional modular system. Moreover, in the case where the reasoning model and the control-generated large model are the same model, by real-time parsing of the thought chain statements generated by the thought chain technology within the large model, and dynamically triggering the pre-configured intelligent agent according to the feature similarity, and integrating the output of the intelligent agents matched in multiple aspects, the accuracy and professionalism of the generated results can be enhanced.

[0207] The following describes the generation and control process of a large model using a specific application scenario as an example with reference to the accompanying drawings:

[0208] It is assumed that the reasoning model and the preset large model are the same model, and that thought chain sentences and reply texts can be generated in the deep thinking mode, and reply texts can be directly generated in the normal thinking mode.

[0209] See Figure 8 As shown, it is a schematic diagram of the generation control process of the large model in the embodiment of the present application. Figure 8 As can be seen from the schematic generation control process, the processing device first obtains the question text input by the user; then, it obtains the logical chain sentence generated based on the question text in the deep thinking mode of the large model; then, it performs real-time block analysis on the logical chain sentence to determine the text blocks included in the logical chain sentence; then, it performs text vectorization processing on each text block to obtain the text features of each text block.

[0210] Furthermore, in the feature anchor library of each type of intelligent agent, the anchor features matched by each text block are determined, as well as the intelligent agent corresponding to the matched anchor features. Taking the processing of a text block as an example, assuming that the matched feature anchors are functional feature 1 and descriptive feature 1, the intelligent agents corresponding to the two feature anchors can be determined in the intelligent agent pool that can be configured and updated in real time. Then, the matched intelligent agents can be used to output content based on a text block respectively, and then the thinking chain sentences can be adjusted based on the output results of each intelligent agent.

[0211] Furthermore, the processing device can adopt a large model to generate a reply text based on the processed thought chain sentences in a normal thinking mode that does not require the output of thought chain sentences; in particular, according to actual processing needs, when generating multimodal data, the reply text and multimodal data can be fused to obtain comprehensive data.

[0212] See Figure 9 As shown, it is a schematic diagram of the process of determining a matching agent in the embodiment of the present application. Figure 9 As can be seen from the illustrated content, the feature anchor library includes features extracted for different types of agents, among which each feature anchor in the feature anchor library includes: descriptive features extracted for each agent that realizes the content recommendation function, agent features extracted for each agent that realizes each type of knowledge generation function, identification features extracted for each agent that realizes the multimodal feature generation function, and functional features extracted for each agent that realizes the content generation function in the same field.

[0213] Continue to combine Figure 9 As can be seen from the explanation, in the feature space, the influence range of each feature anchor can be understood as a circle. When a text feature triggers the influence range of the feature anchor, it is equivalent to matching a corresponding agent based on feature similarity. For example, the feature similarity between text block i and the agent corresponding to descriptive feature q meets the requirements. Then, after the agent screening is finally completed, the configured agent can be triggered for processing, so that the agent can be used to generate sentences based on specific thinking logic and integrate the generated sentences into the thought chain sentence. Among them, each feature anchor in the feature anchor library has a one-to-one correspondence with each agent in the agent pool.

[0214] In this way, with the help of various intelligent agents, a single model can have both domain breadth and domain depth, and the thinking process of the large model can also be adjusted through intelligent agents, avoiding the problem of excessive confusion in the thinking of the large model caused by the black box of thinking.

[0215] In addition, the technical solution proposed in this application takes into account that when a large model is in deep thinking mode, it is often the case that the content obtained through deep thinking cannot be searched for by the existing search method of the large model. With the help of the technical solution proposed in this application, the processing device can asynchronously use an intelligent agent to process the thinking results of the large model, thereby achieving knowledge completion, which can greatly reduce the knowledge hallucinations and corpus pollution of the large model. In areas where the processing effect of the large model was poor before, it can also obtain good generation effects, thereby improving the overall accuracy of the model.

[0216] Furthermore, by adopting the technical solution proposed in this application, the intelligent agent that can be used when modifying the thought chain statement can be adjusted through the real-time updated feature anchor library, thereby indirectly realizing the modification of the associative content of the large model; for example, if there is an advertiser who hopes to buy advertising for a specific game, the traditional large model cannot effectively implant game-related content, and adding game-related promotional copy through a template will also affect the user's original question text due to the problem of template pollution. The method of adjusting the thought chain statement proposed in this application can ensure that the model's ability to understand the question text and advertising push is improved, and the possibility of advertising content replacing user input is reduced; at the same time, this application constructs a feature anchor library and uses the intelligent agent pool as support, so that the processing equipment can dynamically configure the intelligent agents contained in the intelligent agent pool when the large model is running, thereby realizing hot updates of the intelligent agents, which can be enabled and offline in real time, and will not affect the user's use. In particular, for hot issues, there is no need to configure text filters or make compliance adjustments for the big model separately. You only need to configure the corresponding compliance intelligent agent to generate instructions for generating sentences with compliant content. After adjusting the thinking chain sentences based on the generated sentences, the output results of the big model can be made to meet the expectations of the platform and developers without intruding the model, thereby improving the output value of the big model and ensuring the accuracy of the big model output.

[0217] See Figure 10 As shown, it is a schematic diagram of the result after adjusting the thought chain sentence in the embodiment of the present application. Figure 10 From the content shown, it can be seen that if the target object's question text is "What should I pay attention to when traveling to City A at the end of April", then the thinking chain sentences of the large model in the deep thinking mode can be obtained; then, after adjusting the processing method requested by this application, the attached Figure 10As shown in the figure, the adjusted thought chain statement has been modified. Compared to the initially generated thought chain statement, knowledge text has been added to text block 1, and text block 3 has been replaced with the target text to correct errors in text block 3. Furthermore, knowledge text has been added to the end of the thought chain statement to increase restrictions on the output format and content. The labels for knowledge text, recommended text, and other content in the processed thought chain statement are merely schematic descriptions for enhanced readability and may not be used during the actual generation process.

[0218] Moreover, taking scenic spot tourism as an example, by associating the relationship between specific scenic spots and performance activities in the corresponding intelligent body, when the relevant scenic spots are involved, they can be linked to the background configuration link.

[0219] In this way, the thought chain statements can be adjusted without the target object's awareness, thereby greatly improving the accuracy of the large model generation results. Moreover, according to the actual processing needs, the comprehensive data finally displayed to the target object may include pictures, links and other content, so that the recommended content can be delivered in a targeted manner and the delivery effect can be improved. In addition, with the help of the processed thought chain statements, a capability for generation control outside the large model is provided. By dynamically using various intelligent agents to adjust the associative content of the model, the processed thought chain statements are obtained to ensure the final generation quality of the large model. Moreover, it provides the possibility for the large-scale development of AI large models + advertising delivery. Based on the thought chain statements and assisted by the functions of intelligent agents, the traditional large models can achieve more dynamic and rich model capabilities. Moreover, the processing equipment can adjust the direction and results of the large model generation in real time according to the actual processing needs, broadening the business prospects.

[0220] In summary, this application proposes a generation control method based on the thinking chain of a large language model, which determines the intelligent agent that processes different text blocks through text features, and effectively integrates the output results of the matching intelligent agent with the thinking chain sentence of the large model to obtain the adjusted thinking chain sentence. Specifically, it is to configure a specific feature for each intelligent agent of each large language model. As long as the text feature of the text block is close to this feature, the intelligent agent will be triggered; in addition, the thinking chain sentence of the large model during deep thinking is detected in real time, and based on the thinking chain sentence, feature similarity matching is performed for each text block obtained by real-time segmentation to determine whether it can match the intelligent agent, and the sentence output by the matching intelligent agent is integrated with the thinking chain sentence to obtain the processed thinking chain sentence, so that the dynamic programmability of the large model can be realized, the interpretability of the generation result can be improved, and the accuracy and customization of the conclusion of the model result can be greatly improved.

[0221] Moreover, the present application is equivalent to adopting an engineering approach, which additionally implants real-time editable thinking methods for different thinking contents (i.e., different text blocks), wherein real-time editing is reflected in the fact that the agent can be plug-and-play and the available agents can be flexibly added. For a text block, a matching agent is used to execute a specific thinking method for the text block, obtain a generated sentence, and adjust the thinking chain sentence by adding the generated sentence to the thinking chain sentence. For the platform or the user, it is only necessary to set up a specific special task agent to solve a specific task; at the same time, since most of the data injection and data generation can be observed in the thinking chain sentence, the knowledge illusion of the model itself can be avoided to a large extent, and the desired content can be implanted in the thinking chain sentence or a specific result can be obtained without affecting the input.

[0222] Therefore, this application abstracts the internal reasoning process of the large model into a "bus"-style data flow, that is, obtains the thought chain statement, and by adjusting the thought chain statement, it can break through the black box limitations of traditional end-to-end generation, enhance the controllability of the model, and support key scenarios such as manual correction (such as detecting logical errors) and compliance review (such as content filtering) in a feasible implementation method. Furthermore, feature similarity is used instead of hard-coded rules to achieve flexible matching and triggering of intelligent agents, which can adapt to complex tasks in open domains and achieve good generation effects in various generation tasks. For example, the data analysis tool chain is automatically called during report analysis, thereby reducing the cost of domain migration. Only feature anchors need to be updated to quickly adapt to new scenarios, such as migrating from medical consultation to legal consultation, from development assistance to advertising. In addition, by dynamically integrating the general capabilities of large models with the domain expertise of intelligent agents, the breadth and accuracy of generated content can be taken into account, solving problems such as the lag of traditional RAG knowledge and the bloated hybrid expert model. While ensuring real-time performance, the professionalism of the content is improved, realizing the transition from single output to collaborative enhancement. Moreover, thanks to the scalability of intelligent agents, processing equipment can support hot-swap and distributed deployment of intelligent agents, and can dynamically expand the capability boundary according to demand, so that in the processing of complex enterprise-level systems, the ability to collaborate across departments can be met, thereby achieving content generation control with a flexible processing architecture.

[0223] Based on the same inventive concept, see Figure 11 As shown, it is a schematic diagram of the logical structure of the generation control device of the large model in the embodiment of the present application. The generation control device 1100 of the large model includes an acquisition unit 1101, a first generation unit 1102, an execution unit 1103, and a second generation unit 1104, wherein,

[0224] An acquisition unit 1101 is used to acquire a question text to be processed;

[0225] The first generating unit 1102 is configured to generate a thought chain sentence based on the question text using a preset reasoning model;

[0226] The execution unit 1103 is configured to perform the following operations on each text block obtained by splitting the thought chain sentence:

[0227] When a function feature in a preset function feature set satisfies a feature similarity condition with a text feature of a text block, an agent corresponding to the function feature is used to generate a target text based on the text block; wherein the function feature set includes the function features of each agent that realizes the content generation function in the same domain;

[0228] Use the target text to replace a text block in the thought chain statement;

[0229] The second generating unit 1104 is used to generate a reply text to the question text based on the processed thought chain sentences by using a preset large model.

[0230] Optionally, for a text block obtained by splitting a thought chain sentence, when performing an operation, before replacing a text block in the thought chain sentence with the target text, the execution unit 1103 is further configured to perform any one or combination of the following generation operations:

[0231] Among the intelligent agents implementing the content recommendation function, an intelligent agent that meets the feature matching conditions with a text block is used to generate a recommended text based on the text block;

[0232] Among the agents that realize each knowledge acquisition function, the agent that meets the feature screening conditions with a text block is respectively used to generate a knowledge text associated with an insertion position based on the text block;

[0233] Among the intelligent agents that realize the multimodal data generation function, the intelligent agent that meets the feature hit condition with a text block generates associated multimodal data based on the text block.

[0234] Optionally, the generated content further includes a recommendation text, at least one knowledge text, and multimodal data; when the target text is used to replace a text block in a thought chain statement, the execution unit 1103 is configured to:

[0235] Generate control text based on the multimodal data; the control text is used to instruct the large model to reserve the embedding position of the multimodal data in the output result;

[0236] In the thought chain sentence, the corresponding knowledge text is inserted according to the insertion position, and the recommended text and the control text are spliced with the target text to obtain the processed target text;

[0237] Use the processed target text to replace a text block in the thought chain sentence.

[0238] Optionally, after generating the corresponding reply text, the execution unit 1103 is further configured to:

[0239] Adding corresponding at least one multimodal data to at least one embedding position reserved for multimodal data in the reply text to obtain processed integrated data;

[0240] Display comprehensive data to the target object who sent the question text.

[0241] Optionally, when generating a recommended text based on a text block using an agent among the agents implementing the content recommendation function that meets a feature matching condition with the text block, the execution unit is configured to:

[0242] The feature similarity between each descriptive feature in the descriptive feature set and the text feature of a text block is calculated respectively, and the agents whose corresponding feature similarity reaches a preset recommendation threshold and whose associated priority information meets the preset matching condition are regarded as agents that meet the feature matching condition; the descriptive feature set includes: the descriptive features of each agent that implements the content recommendation function; a pre-configured priority information is used to indicate the reference degree of the output result of an agent;

[0243] The selected intelligent agents are used to generate recommended text based on a text block.

[0244] Optionally, when generating a knowledge text associated with an insertion position based on a text block by using an agent that satisfies a feature screening condition between the agents implementing each knowledge acquisition function, the execution unit 1103 is configured to:

[0245] For each type of knowledge acquisition function, the following operations are performed on the corresponding agent feature sets:

[0246] The feature similarity between each agent feature in an agent feature set and the text feature of a text block is calculated respectively, and the agents whose corresponding feature similarity reaches a preset functional threshold and whose associated priority information meets the preset screening conditions are regarded as agents that meet the feature screening conditions; an agent feature set includes: the agent features of each agent that realizes a type of knowledge acquisition function;

[0247] The selected intelligent agents are used to generate knowledge text associated with insertion positions based on a text block.

[0248] Optionally, when generating associated multimodal data based on a text block using an agent among agents that implement the multimodal data generation function and that satisfies a feature hit condition with a text block, the execution unit 1103 is configured to:

[0249] The feature similarity between each descriptive feature in the identification feature set and the text feature of a text block is calculated respectively, and the agents whose corresponding feature similarity reaches a preset data threshold and whose associated priority information meets the preset hit condition are regarded as agents that meet the feature hit condition; the identification feature set includes: the descriptive features of each agent that realizes the multimodal data generation function;

[0250] The selected agents are used to generate multimodal data based on a text block.

[0251] Optionally, the apparatus further includes a construction unit 1105, and the construction unit 1105 is further configured to:

[0252] The function feature set is constructed in the following way: using the trained text feature extraction network, based on the function description text of each agent, the function features are extracted respectively to obtain the function feature set;

[0253] When constructing a descriptive feature set for each intelligent agent that realizes the content recommendation function, the following method is used to construct the descriptive feature set: a text feature extraction network is used to extract descriptive features based on the function description text of each corresponding intelligent agent to obtain a descriptive feature set;

[0254] When constructing an agent feature set for each agent that realizes each knowledge acquisition function, the following method is used to construct each agent feature set: using a text feature extraction network, based on the function description text of each corresponding agent, the agent features are extracted respectively to obtain the agent feature set;

[0255] When constructing an identification feature set for each intelligent agent that realizes the multimodal data generation function, the identification feature set is constructed in the following way: using a text feature extraction network, based on the corresponding function description text of each intelligent agent, the identification features are extracted respectively to obtain the identification feature set.

[0256] Optionally, the construction unit 1105 is further configured to:

[0257] In response to a content removal instruction triggered for a target feature set, deleting at least one target feature in the target feature set that matches the content removal instruction; and

[0258] In response to a new content addition indication triggered for a target feature set, a corresponding function description text is obtained for an intelligent agent to be stored, and the target features extracted by the trained text feature extraction network based on the function description text are added to the target feature set; wherein the target feature set is at least one of a description feature set, a function feature set, a feature set of each intelligent agent, and an identification feature.

[0259] Optionally, the functional feature that satisfies the feature similarity condition with the text feature of a text block is determined by the execution unit 1103 in the following manner:

[0260] Calculate the feature similarity between each functional feature in the functional feature set and the text feature of a text block respectively;

[0261] When there is a functional feature whose corresponding feature similarity reaches a preset data threshold and the associated priority information meets the preset filtering conditions, it is determined that there is a functional feature that meets the feature similarity conditions with the text feature, and the intelligent agent corresponding to the functional feature is regarded as the intelligent agent that meets the feature similarity conditions.

[0262] Optionally, each text block is split by the execution unit 1103 using any of the following methods:

[0263] In the thought chain sentence, the sentence content is split based on the preset identifier to obtain each text block separated by the identifier;

[0264] In thought chain sentences, the sentence content is split based on preset sentence keywords to obtain text blocks starting with the sentence keywords.

[0265] Optionally, after generating the thought chain statement, before performing operations on each text block obtained by splitting the thought chain statement, the execution unit 1103 is further configured to:

[0266] Based on the thought chain sentences, split the text blocks;

[0267] Using the preset text encoding method, each text block is encoded to obtain the corresponding text features.

[0268] For the convenience of description, the above parts are divided into modules (or units) according to their functions and described separately. Of course, when implementing this application, the functions of each module (or unit) can be implemented in the same or multiple software or hardware.

[0269] After introducing the method and apparatus for generating prompt audio according to an exemplary embodiment of the present application, an electronic device according to another exemplary embodiment of the present application will be introduced next.

[0270] Those skilled in the art will appreciate that various aspects of the present application can be implemented as systems, methods, or program products. Therefore, various aspects of the present application can be specifically implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation that combines hardware and software aspects, which may be collectively referred to herein as a "circuit," "module," or "system."

[0271] Based on the same inventive concept as the above method embodiment, an electronic device is also provided in the embodiment of the present application. Figure 12 FIG. 1 is a schematic diagram of the hardware structure of an electronic device using an embodiment of the present application. In one embodiment, the electronic device may be Figure 1 The processing device 120 shown in FIG. In this embodiment, the structure of the electronic device can be as follows: Figure 12 As shown, it includes a memory 1201 , a communication module 1203 and one or more processors 1202 .

[0272] Memory 1201 is used to store computer programs executed by processor 1202. Memory 1201 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and programs required for running instant messaging functions, while the data storage area may store various instant messaging messages and operating instruction sets.

[0273] Memory 1201 may be a volatile memory, such as random-access memory (RAM); a non-volatile memory, such as read-only memory, flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or any other medium capable of carrying or storing a desired computer program in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 1201 may be a combination of the aforementioned memories.

[0274] The processor 1202 may include one or more central processing units (CPUs) or digital processing units, etc. The processor 1202 is configured to implement the aforementioned large model generation control method when calling the computer program stored in the memory 1201 .

[0275] The communication module 1203 is used to communicate with the client device and the server.

[0276] The specific connection medium between the memory 1201, the communication module 1203 and the processor 1202 is not limited in the embodiment of the present application. Figure 12 In the embodiment, the memory 1201 and the processor 1202 are connected via a bus 1204. The bus 1204 is connected to the processor 1202 via a bus 1204. Figure 12 The connections between the other components are shown in bold lines, which are only for illustration and are not intended to be limiting. The bus 1204 can be divided into an address bus, a data bus, a control bus, etc. For ease of description, Figure 12 The diagram shows a single thick line, but this does not indicate that there is only one bus or one type of bus.

[0277] The memory 1201 stores a computer storage medium, which stores computer executable instructions. The computer executable instructions are used to implement the generation control method of the large model of the embodiment of the present application. The processor 1202 is used to execute the generation control method of the large model, such as Figure 2 shown.

[0278] In another embodiment, the electronic device may also be other electronic devices, see Figure 13 FIG. 1 is a schematic diagram of the hardware structure of another electronic device using an embodiment of the present application. Specifically, the electronic device may be Figure 1 The client device 110 shown in FIG. In this embodiment, the structure of the electronic device can be as follows: Figure 13 As shown, it includes: a communication component 1310, a memory 1320, a display unit 1330, a camera 1340, a sensor 1350, an audio circuit 1360, a Bluetooth module 1370, a processor 1380 and other components.

[0279] The communication component 1310 is used to communicate with the server. In some embodiments, it may include a wireless fidelity (WiFi) module. The WiFi module is a short-range wireless transmission technology. Electronic devices can help users send and receive information through the WiFi module.

[0280] Memory 1320 can be used to store software programs and data. Processor 1380 executes the software programs or data stored in memory 1320 to perform various functions and data processing of client device 110. In this application, memory 1320 can store an operating system and various application programs, and can also store and execute computer programs related to the generation and control method of the large model in the embodiment of this application.

[0281] The display unit 1330 may also be used to display information input by the user or information provided to the user, as well as a graphical user interface (GUI) of various menus of the client device 110. Specifically, the display unit 1330 may include a display screen 1332 disposed on the front of the client device 110. The display unit 1330 may be used to display web pages, etc.

[0282] The display unit 1330 can also be used to receive input digital or character information and generate signal input related to user settings and function control of the client device 110. Specifically, the display unit 1330 may include a touch screen 1331 set on the front of the client device 110, which can collect user touch operations on or near it.

[0283] The touch screen 1331 can be covered on the display screen 1332, or the touch screen 1331 and the display screen 1332 can be integrated to realize the input and output functions of the client device 110. The integrated touch screen can be simply called a touch display screen. In this application, the display unit 1330 can display applications and corresponding operation steps.

[0284] Camera 1340 can be used to capture still images, and users can post comments on images captured by camera 1340 through the app. The lens generates an optical image of an object and projects it onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, which is then transmitted to processor 1380 for conversion into a digital image signal.

[0285] The client device may further include at least one sensor 1350, such as an accelerometer 1351, a distance sensor 1352, a fingerprint sensor 1353, and a temperature sensor 1354. The client device may also be equipped with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, a light sensor, and a motion sensor.

[0286] The audio circuit 1360, speaker 1361, and microphone 1362 provide an audio interface between the user and the client device 110. The audio circuit 1360 converts received audio data into electrical signals and transmits them to the speaker 1361, which then converts the signals into sound signals for output. Meanwhile, the microphone 1362 converts collected sound signals into electrical signals, which are then received by the audio circuit 1360 and converted into audio data. The audio data is then output to the communication component 1310 for transmission to, for example, another client device 110, or to the memory 1320 for further processing.

[0287] The Bluetooth module 1370 is used to exchange information with other Bluetooth devices having a Bluetooth module through the Bluetooth protocol.

[0288] The processor 1380 is the control center of the client device. It uses various interfaces and lines to connect various parts of the entire terminal. It executes various functions of the client device and processes data by running or executing software programs stored in the memory 1320 and calling data stored in the memory 1320. In some embodiments, the processor 1380 may include at least one processing unit; the processor 1380 may also integrate an application processor and a baseband processor. In this application, the processor 1380 can run an operating system, application programs, user interface display and touch response, as well as methods related to the generation control of the large model of the embodiment of the present application. In addition, the processor 1380 is coupled to the display unit 1330.

[0289] In some possible implementations, various aspects of the large model generation control method provided by the present application can also be implemented in the form of a program product, which includes a computer program. When the program product is run on an electronic device, the computer program is used to enable the electronic device to execute the steps of the large model generation control method according to various exemplary embodiments of the present application described above in this specification. For example, the electronic device can execute the following steps: Figure 2 Follow the steps shown in .

[0290] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0291] The program product of the embodiment of the present application may be a portable compact disc read-only memory (CD-ROM) and include a computer program, and can be run on an electronic device. However, the program product of the present application is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with a command execution system, apparatus, or device.

[0292] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a readable computer program. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with a command execution system, apparatus, or device.

[0293] The computer program embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0294] The computer program for performing the operations of the present application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The computer program can be executed entirely on the user electronic device, partially on the user electronic device, as a separate software package, partially on the user electronic device and partially on a remote electronic device, or entirely on a remote electronic device or server. In cases involving remote electronic devices, the remote electronic device can be connected to the user electronic device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external electronic device (for example, using an Internet service provider to connect through the Internet).

[0295] It should be noted that although several units or subunits of the device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, depending on the embodiment of the application, the features and functions of two or more units described above can be embodied in a single unit. Conversely, the features and functions of a single unit described above can be further divided and embodied by multiple units.

[0296] Furthermore, although the operations of the method of the present application are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in this particular order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0297] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain a computer-usable computer program.

[0298] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program commands. These computer program commands can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the commands executed by the processor of the computer or other programmable data processing device generate commands for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0299] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0300] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A method for controlling the generation of a large model, characterized in that: include: Get the question text to be processed; Using a preset reasoning model, based on the question text, a thought chain sentence is generated; For each text block obtained by splitting the thought chain sentence, perform the following operations respectively: When a function feature that satisfies a feature similarity condition with a text feature of a text block exists in a preset function feature set, an agent corresponding to the function feature is used to generate a target text based on the text block; wherein the function feature set includes the function features of each agent that realizes the content generation function in the same domain; Replacing the text block in the thought chain sentence with the target text; A preset large model is used to generate a reply text for the question text based on the processed thought chain sentences.

2. The method according to claim 1, wherein For a text block obtained by splitting the thought chain sentence, when performing an operation, before replacing the text block in the thought chain sentence with the target text, the operation further includes generating any one or a combination of the following operations: Using an agent that satisfies a feature matching condition with the text block among the agents implementing the content recommendation function, and generating a recommended text based on the text block; Using the agents that meet the feature screening conditions between each agent and the text block among the agents that realize each knowledge acquisition function, and generating a knowledge text associated with an insertion position based on the text block; Among the intelligent agents that realize the multimodal data generation function, the intelligent agent that meets the feature hit condition with the one text block generates associated multimodal data based on the one text block.

3. The method according to claim 2, wherein When the generated content further includes a recommendation text, at least one knowledge text, and multimodal data, replacing the text block in the thought chain statement with the target text includes: Generate a control text based on the multimodal data; wherein the control text is used to instruct the large model to reserve an embedding position for the multimodal data in the output result; Inserting the corresponding knowledge text in the thought chain sentence according to the insertion position, and splicing the recommended text and the control text with the target text to obtain a processed target text; The processed target text is used to replace the text block in the thought chain sentence.

4. The method according to claim 3, wherein After generating the corresponding reply text, the method further includes: Adding corresponding at least one multimodal data to at least one embedding position reserved for multimodal data in the reply text to obtain processed integrated data; The comprehensive data is displayed to the target object that sent the question text.

5. The method according to claim 2, wherein The intelligent agents among the intelligent agents implementing the content recommendation function that meet the feature matching condition with the text block generate a recommended text based on the text block, including: The feature similarity between each descriptive feature in the descriptive feature set and the text feature of the text block is calculated respectively, and the agents whose corresponding feature similarity reaches a preset recommendation threshold and whose associated priority information meets a preset matching condition are regarded as agents that meet the feature matching condition; the descriptive feature set includes: the descriptive features of each agent that implements the content recommendation function; and a pre-configured priority information indicating the reference degree of an output result of an agent; The filtered intelligent agent is used to generate a recommended text based on the text block.

6. The method according to claim 2, wherein The method of using, among the agents that respectively implement each knowledge acquisition function, an agent that satisfies a feature screening condition with the text block, and generating a knowledge text associated with an insertion position based on the text block, includes: For each type of knowledge acquisition function, the following operations are performed on the corresponding agent feature sets: The feature similarity between each agent feature in an agent feature set and the text feature of the text block is calculated respectively, and the agents whose corresponding feature similarity reaches a preset functional threshold and whose associated priority information meets the preset screening condition are regarded as agents that meet the feature screening condition; an agent feature set includes: the agent features of each agent that realizes a type of knowledge acquisition function; The filtered intelligent agent is used to generate a knowledge text associated with an insertion position based on the text block.

7. The method according to claim 2, wherein Among the agents that realize the multimodal data generation function, the agents that meet the feature hit condition with the text block generate associated multimodal data based on the text block, including: The feature similarity between each descriptive feature in the identification feature set and the text feature of the text block is calculated respectively, and the agents whose corresponding feature similarity reaches a preset data threshold and whose associated priority information meets the preset hit condition are regarded as agents that meet the feature hit condition; the identification feature set includes: the descriptive features of each agent that realizes the multimodal data generation function; The filtered intelligent agent is used to generate multimodal data based on the one text block.

8. The method according to any one of claims 1 to 7, wherein: The method further comprises: The functional feature set is constructed in the following manner: using a trained text feature extraction network, based on the functional description text of each corresponding agent, functional features are extracted respectively to obtain a functional feature set; When constructing a description feature set for each intelligent agent that realizes the content recommendation function, the description feature set is constructed in the following manner: using the text feature extraction network, based on the function description text of each corresponding intelligent agent, respectively extracting description features to obtain a description feature set; When constructing an agent feature set for each agent that realizes each knowledge acquisition function, each agent feature set is constructed in the following manner: using the text feature extraction network, based on the function description text of each corresponding agent, the agent features are extracted respectively to obtain the agent feature set; When constructing an identification feature set for each intelligent agent that realizes the multimodal data generation function, the identification feature set is constructed in the following way: using the text feature extraction network, based on the corresponding function description text of each intelligent agent, the identification features are extracted respectively to obtain the identification feature set.

9. The method according to claim 8, wherein The method further comprises: In response to a content removal instruction triggered for a target feature set, deleting at least one target feature in the target feature set that matches the content removal instruction; and In response to a new content addition indication triggered for a target feature set, a corresponding function description text is obtained for an intelligent agent to be stored, and the target features extracted by the trained text feature extraction network based on the function description text are added to the target feature set; wherein the target feature set is at least one of the description feature set, the function feature set, the feature sets of each intelligent agent, and the identification features.

10. The method according to any one of claims 1 to 7, wherein: The functional features that satisfy the feature similarity condition with the text features of a text block are determined in the following manner: Calculating feature similarities between each functional feature in the functional feature set and the text feature of the text block; When there is a functional feature whose corresponding feature similarity reaches a preset data threshold and the associated priority information meets the preset filtering conditions, it is determined that there is a functional feature that meets the feature similarity conditions with the text feature, and the intelligent agent corresponding to the functional feature is used as the intelligent agent that meets the feature similarity conditions.

11. The method according to any one of claims 1 to 7, wherein: Each text block is split using any of the following methods: In the thought chain sentence, the sentence content is split based on the preset identifier to obtain each text block separated by the identifier; In the thought chain sentence, the sentence content is split based on preset sentence keywords to obtain various text blocks starting with the sentence keywords.

12. The method according to any one of claims 1 to 7, wherein: After generating the thought chain sentence and before performing operations on each text block obtained by splitting the thought chain sentence, the method further includes: Based on the thought chain sentences, splitting to obtain various text blocks; A preset text feature extraction method is adopted to extract corresponding text features from each text block.

13. A large model generation control device, characterized in that: include: An acquisition unit, used to acquire the question text to be processed; A first generating unit is configured to generate a thought chain sentence based on the question text by using a preset reasoning model; The execution unit is configured to perform the following operations on each text block obtained by splitting the thought chain sentence: When a function feature that satisfies a feature similarity condition with a text feature of a text block exists in a preset function feature set, an agent corresponding to the function feature is used to generate a target text based on the text block; wherein the function feature set includes the function features of each agent that realizes the content generation function in the same domain; Replacing the text block in the thought chain sentence with the target text; The second generating unit is used to generate a reply text of the question text based on the processed thought chain sentence by using a preset large model.

14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 12 is implemented.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.

16. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.

Citation Information

Cited By

  • Large language model generation process real-time intervention method based on thinking chain verification

    CN121615798A

  • A large language model generation process real-time intervention method based on thought chain verification

    CN121615798B