Agent-based training sample construction method and device, and agent

CN122817877APending Publication Date: 2026-09-25BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611116165.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-24
Publication Date
2026-09-25

Smart Images

  • Figure CN122817877A_ABST
    Figure CN122817877A_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for constructing a training sample based on an agent, a method for training a text review model, a text review method, an apparatus, an agent, an electronic device and a storage medium, relating to the technical field of computers, and particularly to the fields of artificial intelligence, deep learning, natural language processing, etc. The method comprises: obtaining a review track, the review track indicating a record generated by a review agent in the process of performing a text review task; and deleting a prohibited field from the review track to obtain training input data, the prohibited field being related to an output result of a model called by the review agent in the process of performing the text review task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to the fields of artificial intelligence, deep learning, and natural language processing. More specifically, this disclosure provides a method for constructing training samples based on an intelligent agent, a method for training a text review model, a text review method, an apparatus, an intelligent agent, an electronic device, and a storage medium. Background Technology

[0002] Text review is widely used in legal compliance, content moderation, and quality inspection, aiming to assess the compliance, completeness, and potential risks of text content. For example, in risk reviews before signing commercial contracts, key clauses such as payment terms, liability for breach of contract, and ownership of intellectual property rights need to be systematically checked. With the application of machine learning technology, artificial intelligence-based review models are increasingly being used for text review tasks. The performance of such models is highly dependent on the quality and richness of the training dataset. Summary of the Invention

[0003] This disclosure provides a method for constructing training samples based on an intelligent agent, a method for training a text review model, a text review method, an apparatus, an intelligent agent, an electronic device, and a storage medium.

[0004] According to one aspect of this disclosure, a method for constructing training samples based on an agent is provided, applied to a governance agent, comprising: acquiring a review trajectory, the review trajectory indicating records generated by the review agent during the execution of a text review task; deleting prohibited fields from the review trajectory to obtain training input data, wherein the prohibited fields are related to the output results of the model invoked by the review agent during the execution of the text review task.

[0005] According to another aspect of this disclosure, a method for training a text censorship model is provided, the method comprising: acquiring training input data, the training input data being obtained according to the method described above; and training an initial text censorship model based on the training input data to obtain a target text censorship model.

[0006] According to another aspect of this disclosure, a text review method is provided, the method comprising: acquiring a text to be reviewed; inputting the text to be reviewed into a target text review model for processing, and obtaining a review result, wherein the target text review model is trained by the method described above.

[0007] According to another aspect of this disclosure, an apparatus for constructing training samples based on an agent is provided. The apparatus includes an acquisition module applied to a governance agent, comprising: an acquisition module for acquiring review trajectories, the review trajectories indicating records generated by the review agent during the execution of a text review task; and a generation module for deleting prohibited fields from the review trajectories to obtain training input data, wherein the prohibited fields are related to the output results of a model invoked by the review agent during the execution of the text review task.

[0008] According to another aspect of this disclosure, a training apparatus for a text censorship model is provided. The apparatus includes: an acquisition module for acquiring training input data, the training input data being obtained according to the method described above; and a training module for training an initial text censorship model based on the training input data to obtain a target text censorship model.

[0009] According to another aspect of this disclosure, a text review apparatus is provided, the apparatus comprising: an acquisition module for acquiring text to be reviewed; and a review module for inputting the text to be reviewed into a target text review model for processing to obtain a review result, wherein the target text review model is trained by the method described above.

[0010] According to another aspect of this disclosure, an artificial intelligence agent is provided, comprising: an input module for receiving input information; a processing module for executing the method described above based on the input information received by the input module to obtain output information; and an output module for outputting the output information obtained by the processing module.

[0011] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.

[0012] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the methods described above.

[0013] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method described above.

[0014] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0015] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0016] Figure 1 The illustration schematically depicts an exemplary system architecture applicable to methods for constructing training samples, training text review models, and text review methods and apparatus according to embodiments of the present disclosure;

[0017] Figure 2 A flowchart illustrating a method for constructing training samples according to an embodiment of the present disclosure is shown schematically;

[0018] Figure 3 A system block diagram according to an embodiment of the present disclosure is illustrated schematically;

[0019] Figure 4 The diagram illustrates a process for generating the input payload of a training task according to an embodiment of the present disclosure.

[0020] Figure 5 A flowchart illustrating a method for training a text review model according to an embodiment of the present disclosure is shown schematically.

[0021] Figure 6 A flowchart illustrating a text review method according to an embodiment of this disclosure is shown schematically;

[0022] Figure 7 A block diagram of an apparatus for constructing training samples according to an embodiment of the present disclosure is shown schematically;

[0023] Figure 8 A block diagram illustrating a training apparatus for a text censorship model according to an embodiment of the present disclosure is shown schematically.

[0024] Figure 9 A block diagram of a text review apparatus according to an embodiment of the present disclosure is shown schematically;

[0025] Figure 10 A schematic diagram illustrating the structure of an intelligent agent of artificial intelligence according to embodiments of the present disclosure; and

[0026] Figure 11 The diagram illustrates a block diagram of an electronic device suitable for implementing a method for constructing training samples, a method for training a text review model, and a text review method, according to embodiments of the present disclosure. Detailed Implementation

[0027] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0028] As used in the specification and appended claims of this disclosure, the singular expressions “a,” “an,” and “the” are intended to include expressions such as “one or more,” unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of this disclosure, “at least one,” “at least one,” and “one or more” refer to one, two, or more than two. “First,” “second,” and various numerical designations are merely distinctions for ease of description and are not intended to limit the scope of the embodiments of this disclosure. “And / or” describes the correspondence between corresponding objects, indicating that three relationships can exist. For example, “A and / or B” can represent: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character “ / ” generally indicates that the preceding and following related objects are in an “or” relationship. The sequence numbers of the processes below do not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this disclosure. For example, in the embodiments of this disclosure, the words “301,” “401,” and “501” are merely identifiers for ease of description and do not limit the order of execution steps.

[0029] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this disclosure include a specific feature, structure, or characteristic described in connection with that embodiment. In this disclosure, words such as "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" or "for example" in this disclosure should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized. In embodiments of this disclosure, descriptions such as "when," "in the case of," "if," and "if" indicate that the device will perform a corresponding action under certain objective circumstances, not a limitation on time, and do not require the device to necessarily perform a judgment action during implementation, nor do they imply any other limitations. In this disclosure, "for indicating" can include both direct and indirect indication. When describing an indication message as indicating A, it can include whether the indication message directly indicates A or indirectly indicates A, but does not necessarily mean that the indication message carries A.

[0030] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure, and application of any type of information, such as user personal information, comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and they do not violate public order and good morals.

[0031] In the technical solution disclosed herein, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.

[0032] Text review is widely used in legal compliance, content moderation, and quality inspection, aiming to assess the compliance, completeness, and potential risks of text content. For example, in risk reviews before signing commercial contracts, key clauses such as payment terms, liability for breach of contract, and ownership of intellectual property rights need to be systematically checked. With the application of machine learning technology, artificial intelligence-based review models are increasingly being used for text review tasks. The performance of such models is highly dependent on the quality and richness of the training dataset.

[0033] In agent-based contract review scenarios, the agent generates corresponding review trajectory data during the review process. This trajectory data typically contains a large amount of information, the trajectories are quite long, and the content items and data structures of different trajectories often differ. If conventional template-based methods are used to construct training inputs, it is difficult to extract all the effective information from the trajectories, easily leading to the omission of key review steps or decision-making criteria, thus affecting the information completeness of the training samples.

[0034] In view of this, embodiments of the present disclosure provide a scheme to remove fields related to model output from the review trajectory and construct training input data based on this. This helps to ensure the completeness of training sample information and avoid the leakage of supervision information, thereby helping to ensure the accuracy of performance indicators during the training phase and thus ensuring the actual performance of the model after deployment.

[0035] Figure 1 This is a schematic diagram of an exemplary system architecture for methods and apparatus applicable to embodiments of this disclosure, based on one embodiment of the present disclosure. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.

[0036] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0037] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0038] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using terminal devices 101, 102, and 103 (for example only). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0039] A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system. A server can also be a server for a distributed system, or a server that incorporates blockchain technology.

[0040] It should be noted that the agent-based training sample construction method, text review model training method, and text review method provided in this embodiment can be executed by the first terminal device 101, the second terminal device 102, and the third terminal device 103. Correspondingly, the agent-based training sample construction device, text review model training device, and text review device provided in this embodiment can also be disposed in the first terminal device 101, the second terminal device 102, and the third terminal device 103.

[0041] Alternatively, the method for constructing training samples based on intelligent agents, the method for training text review models, and the text review method provided in this disclosure embodiment can be executed by server 105. Accordingly, the apparatus for constructing training samples based on intelligent agents, the apparatus for training text review models, and the text review apparatus provided in this disclosure embodiment can generally be located in server 105.

[0042] Alternatively, the agent-based training sample construction method, text review model training method, and text review method provided in this disclosure can also be executed by a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103, and / or server 105. Correspondingly, the agent-based training sample construction apparatus, text review model training apparatus, and text review apparatus provided in this disclosure can also be located in a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103, and / or server 105.

[0043] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0044] Figure 2 This is a flowchart of a method for constructing training samples according to an embodiment of the present disclosure. Figure 2 The method shown can be executed by a server, by a terminal device, or by both a server and a terminal device. For example, Figure 2 The method 200 shown can be executed by the system architecture described above.

[0045] For example, this text review method can be applied to agents (such as governance agents) or agent systems. Method 200 can also be referred to as a method for constructing training samples based on agents.

[0046] like Figure 2 As shown, method 200 may include the following steps.

[0047] Step S210: Obtain the review trajectory, which indicates the records generated by the review agent during the execution of the text review task.

[0048] Step S220: Remove the prohibited fields from the review trajectory to obtain training input data. The prohibited fields are related to the output results of the model called by the review agent during the text review task.

[0049] A review agent, also known as a text review agent, is capable of performing text review tasks. As an example, this review agent can be used in risk review scenarios involving various types of text, such as contract review, compliance review, and rule-driven document review. For instance, the text can be a contract text, or other types of documents; this disclosure does not limit the scope of the document.

[0050] The records generated by the review agent during the execution of a text review task constitute the review trajectory for that task. Alternatively, the review trajectory can also be understood as the records generated by a review instance during its execution. A review instance is a single text review agent operation targeting a single review unit, which can generate one review trajectory. A text typically corresponds to multiple review instances.

[0051] The model invoked by the review agent during the text review process can also be called a text review model, which is used to output predicted review results for the review unit. For example, this model can be a large language model (LLM), or it can be other specialized models. The embodiments of this disclosure do not limit the structure of the model.

[0052] Forbidden fields can also be understood as fields in a blacklist. In method 200, a blacklist can be set, and certain fields can be removed from the review trajectory based on this blacklist to obtain training data.

[0053] The prohibited fields are related to the output of the text review model. Alternatively, the prohibited fields can be related to the supervision information (supervision conclusions). That is, in step S220, fields related to supervision information are removed from the review trajectory to obtain training input data.

[0054] If supervisory information remains in the training input data, the model can obtain this information during the input phase, which is called supervisory information leakage. This will lead to an overestimation of the model's performance metrics during the training phase, which will not reflect the actual performance of the model after deployment, and will also affect the training effect of the model.

[0055] According to the embodiments of this disclosure, removing fields related to the model's output from the review trajectory and constructing training input data based on this helps ensure the completeness of the training input data, thereby improving the training effect of the model. Simultaneously, removing fields related to the model's output also helps prevent supervisory information from appearing in the model's visible training input, thus ensuring the accuracy of performance metrics during the training phase and ultimately guaranteeing the actual performance of the model after deployment.

[0056] In some embodiments, the review trajectory indicates the records generated by the review agent during the process of performing a text review task based on the risk points corresponding to the review trajectory.

[0057] In other words, this review trajectory is a risk-point level review trajectory. The granularity of the text review task is at the risk-point level. Based on the records generated by the review agent during the review of text targeting a risk point, the review trajectory corresponding to that risk point can be constructed.

[0058] Accordingly, the "review instance" mentioned above can be a risk point review instance, with the review unit being the risk point. A risk point review instance is a single text review agent operation targeting a single risk point, which can generate a risk point-level review trajectory. A text typically corresponds to multiple risk point review instances.

[0059] It should be understood that using the review unit as a risk point is only one possible implementation. In other implementations, the granularity of the review unit can be set as needed; for example, the granularity of the review unit can be a clause in the contract. For ease of understanding, the embodiments of this disclosure are mainly described using the review unit as a risk point, and do not constitute a limitation on the implementation of the embodiments of this disclosure.

[0060] According to the embodiments of this disclosure, the review trajectory and risk point are corresponding, and each review trajectory can be used to construct training input data. Compared with review trajectories of other review granularities (which may contain a large amount of review information of risk points), it avoids the data processing overhead and information omission caused by extracting information corresponding to each risk point from a large amount of data, improves data construction efficiency, and further ensures information completeness.

[0061] In some embodiments, the review trajectory includes at least one of the following: model input data, tool call records, execution environment feedback records, model output results, supervision target values, reward values, quality evaluation, correctness judgment results, manual review information, sensitive content, or governance control content.

[0062] Furthermore, the review process can also include identity verification.

[0063] The following is an explanation of each of the above fields.

[0064] The input data of the model is the input data read by the review agent when performing the text review task (i.e., review input), or the set of input references, which may include at least one of the following: contract fragment references, risk point rule references, and candidate evidence references.

[0065] The tool call log is used to indicate the query tools, reading and writing tools, search tools, calculation tools, and their parameters called by the review agent during the execution of text review tasks.

[0066] The execution environment feedback log is used to indicate what the tool returns.

[0067] For example, tool call logs and execution environment feedback logs can be used for auditing and review.

[0068] The model's output, namely the model prediction result or model prediction label, is used to indicate whether the model predicts whether a risk point has been hit (i.e. whether the risk point exists in the text).

[0069] The monitoring target value is the ground truth of whether the risk point is hit. For example, the monitoring target value can be manually labeled.

[0070] Reward values, which can also be replaced by other descriptions such as evaluation signals, are used to indicate a score and / or label for the quality of the review process or model output. For example, reward values ​​can be calculated based on deterministic rules; for instance, a lower reward value is given if the model outputs results directly without calling a tool, while a higher reward value is given if the model outputs results after calling a tool. Alternatively, reward values ​​can also be assigned by a larger model to score the model's output.

[0071] Quality assessment refers to the evaluation of the quality of the review process; for example, quality assessment could be a score of the quality of the review process.

[0072] The correctness determination result is used to indicate whether the model's output is correct.

[0073] The manual review information is used to instruct manual annotation of information.

[0074] Sensitive content is used to indicate sensitive information related to the review process. For example, in a contract review scenario, sensitive content may include excerpts from the original contract and information about the parties involved.

[0075] For example, sensitive content fields need to be de-identified.

[0076] Governance control content is used to indicate governance metadata, and may include at least one of the following: data isolation subject, data source, anonymization status, authorization information for training purposes, legal retention status, privileged confidentiality status, etc.

[0077] For example, governance control content can come from various sources, such as governance ledgers, manual approval systems, data asset ledgers, automated verifiers that meet preset verification conditions, or external compliance systems.

[0078] For example, governance control content can be used for governance counting.

[0079] A data isolation entity is used to indicate the data ownership unit of the audit trail. Different data ownership units are configured to be isolated from each other. For example, the data ownership unit can be determined according to at least one of the following isolation requirements: data ownership, usage authorization, or access control requirements. For instance, the data ownership unit can include any one or more of the following: tenant, customer, organization, department, project, or data provider, etc. In a multi-tenant application scenario, the identifier of the data isolation entity can be the tenant's identifier.

[0080] The data source can also be replaced with other descriptive methods such as data partitioning, used to indicate the set to which the review trajectory belongs, or in other words, the purpose of the review trajectory. For example, the data source may include: training set, development set, and evaluation reserve set; correspondingly, the purpose of the review trajectory may include: training, development, and evaluation, etc.

[0081] The desensitization status is used to indicate whether the content in the original review track has been desensitized.

[0082] The training use authorization information is used to indicate whether authorization has been obtained for the original review trajectory to be used for training.

[0083] Legal reservation status is used to indicate whether there are non-trainable restrictions.

[0084] The privileged confidentiality status is used to indicate whether privileged confidentiality is involved.

[0085] For example, the identity identifier may include at least one of the following: the identifier of the review trajectory, the corresponding review unit (such as a risk point), the data isolation subject, or the data source.

[0086] Below is an example of the fields in the review trajectory for a risk point.

[0087] For example, review risk point R-PAY-014: "Does the payment term rely excessively on the client's acceptance?" The fields in the review track generated from a single run are shown below.

[0088] Contract excerpt: Payment shall be made within 30 days after acceptance and receipt of the invoice… Acceptance shall be subject to written confirmation by Party A.

[0089] Rules for risk points: verification and acceptance standards, time limits, and handling of objections.

[0090] Tool call history: Search for "acceptance", read payment terms, submit review items.

[0091] Execution environment feedback log: The search found 3 instances of the word "acceptance".

[0092] Model output: No risk found.

[0093] Monitoring target value: Risk exists.

[0094] Correctness determination result: False.

[0095] Reward value: 0.

[0096] Quality rating: 0.42.

[0097] Data source: Training set.

[0098] Data isolation subject: Tenant A.

[0099] Desensitization status: Correct (true).

[0100] Training purpose authorization: true.

[0101] Taking the review unit as a risk point as an example, the review trajectory can be generated in the following way.

[0102] For example, for each risk point review instance, at least one of the following is read: input references, tool call records, execution environment feedback records, model output results, reward values, and governance control content. The above information for each risk point review instance is then structured to generate a review trajectory corresponding to each risk point.

[0103] The types of fields in different review tracks may be the same or different.

[0104] It should be understood that the above is merely an example of a specific field in the review track and does not constitute a limitation on the fields in the review track. For example, the review track may also include other fields. For example, the review track may also include the error reason, i.e., explanatory text. Taking the example above, the error reason could be the omission of the client's unilateral acceptance and the lack of objection handling.

[0105] In some embodiments, the data sources for reviewing trajectories include training sets and / or development sets.

[0106] The review tracks to be exported (candidate review tracks) may come from different sets, such as the training set, the development set, and the evaluation reserve set. In step S210, the obtained review tracks come from the training set and / or the development set.

[0107] For example, before exporting the candidate review trajectory, the data source field of the candidate review trajectory is read. If the data source of the candidate review trajectory indicates that the trajectory comes from the training set or the development set, the trajectory is exported as the review trajectory in step S210. If the data source of the candidate review trajectory indicates that the trajectory comes from the evaluation retention set, the trajectory is intercepted.

[0108] Data sources can be represented in various ways, such as the name of the source file, the source directory, or the source list.

[0109] This avoids data used for evaluation from entering the training phase, prevents pollution of the evaluation retainer, and helps avoid the problem of supervision leakage.

[0110] The review tracks are not entirely consistent in data structure; different review tracks may include different fields, and may also include complex fields such as nested levels. Using conventional field filtering methods can easily miss supervisory information, thus causing supervisory leakage.

[0111] In some embodiments, prohibited fields include first type fields, which include at least one of the following: model output, supervision target value, reward value, quality evaluation, or correctness determination result.

[0112] The first type of field can be a field in the review track that directly indicates the supervision information. This type of field can be a field that is usually present in the review track and is related to the supervision information, or it can be replaced with other descriptions such as the direct supervision conclusion field.

[0113] The above items are merely examples of Type I fields and do not constitute a limitation on the specific content of Type I fields. Type I fields can be set according to the needs of the training task.

[0114] As mentioned earlier, different review tracks may include different fields. In addition to fields that directly indicate supervisory information, review tracks may also include other fields that imply supervisory information.

[0115] In some embodiments, the prohibited field may also include a second type field, which is derived from the first type field.

[0116] In other words, a hierarchical field blacklist can be set up to divide prohibited fields into multiple levels, with the first level being the first type of field and the second level being the second type of field. Fields derived from directly supervised fields may also include supervisory information. In step S220, in addition to deleting the directly supervised fields, fields derived from the directly supervised fields (derived supervisory fields or indirect supervisory fields) must also be deleted.

[0117] For example, derived supervision fields can be statistics, rankings, confidence levels, rule hit indicators, and explanatory text that references supervision conclusions, obtained from direct supervision fields through processing methods such as calculation, aggregation, or normalization.

[0118] The rule hit indicator is used to indicate whether the rule corresponding to the risk point has been hit.

[0119] Explanatory text is an explanation of the model's output. For example, if the model's output is incorrect, the explanatory text can include the reason for the error.

[0120] Second-type fields can be determined in several ways. For example, second-type fields can be manually set. Or, they can be generated from a large model.

[0121] According to the embodiments of this disclosure, in addition to deleting the direct supervision fields from the review trajectory, fields derived from the direct supervision fields are also deleted to prevent supervision information from being provided to the model to be trained through the derived fields. This helps to further avoid supervision leakage and thus further ensures the training effect.

[0122] In some embodiments, step S220 may include: removing the prohibited field from the top-level field in the review track.

[0123] In some embodiments, step S220 may further include: performing a recursive traversal of the nested structure in the review trajectory and deleting prohibited fields from the fields of the nested structure.

[0124] For example, the nested structure in the review trajectory may include one or more of the following: dictionary, list, tool call log, execution environment feedback log, or evaluation object.

[0125] For example, objects in the review trajectory can be recursively traversed according to nested structures such as dictionaries, lists, tool call records, execution environment feedback records, and evaluation objects. The key name of each leaf node and intermediate node in each nested structure can be matched with the blacklist, and prohibited fields can be deleted.

[0126] Taking the tool call record in the previous example as an example, the field "tool_calls" is a list of tools called. The return result of the third tool call (submit_review_item: tool_calls[2]) contains a nested key "gold_hint":

[0127] tool_calls[2] = {

[0128] "name": "submit_review_item",

[0129] "result": {

[0130] "status": "ok",

[0131] "gold_hint": "Risk exists"}}

[0132] The "gold_hint" field indicates a risk and is considered supervisory information. If only the top-level field is filtered, this field will be ignored. However, by recursively traversing the data, this field can be matched and the data removed.

[0133] According to the embodiments of this disclosure, the nested structure is recursively traversed to remove prohibited fields. This method can remove supervision information nested inside objects (e.g., tool call records, execution environment feedback records, or quality evaluations) that cannot be reached by filtering based on the top-level field. This helps to ensure the comprehensiveness of the filtering results, further avoid supervision leakage, and thus further ensure the training effect.

[0134] In some embodiments, step S220 may include: deleting at least one of the following from the review track: a field that matches the name of a prohibited field, or a field that matches the value of a prohibited field.

[0135] For example, a field matching the name of a prohibited field can include a field with the same name as the prohibited field. For instance, if a key name in the review track is in the blacklist, the field corresponding to that key name is removed from the review track. Alternatively, a field matching the name of a prohibited field can also include a field with the same part of the prohibited field's name. For example, if the blacklist includes a monitoring target value with the key name "gold_answer", and in the previous example there is a key named "gold_hint" that has the same part of the name as "gold_answer", then the "gold_hint" field is removed from the review track.

[0136] As one possible implementation, fields matching the values ​​of prohibited fields could include fields whose values ​​have a deterministic derivation relationship with the values ​​of direct supervision conclusion fields. For such fields, even if their key names are not on the blacklist, their values ​​are related to the values ​​of direct supervision conclusion fields and should also be removed from the review process.

[0137] For example, a field that matches the value of a prohibited field may include one or more of the following: the value is the same as the value of a prohibited field, the value is a hash or truncation of the value of a prohibited field, the value is a difference or ratio calculated based on the value of a prohibited field, or the value has a mapping relationship with the value of a prohibited field.

[0138] Taking the supervisory target value in the prohibited field as an example, the field that matches the value of the supervisory target value can include one or more of the following: the value is the same as the value of the supervisory target value, the value is a hash or truncation of the value of the supervisory target value, the value is a difference or ratio calculated based on the value of the supervisory target value, and the value has a mapping relationship with the value of the supervisory target value.

[0139] For example, the review track includes a risk level field with a value of "high," and the key name of this field is not in the blacklist. The value of the monitoring target value indicates "risk exists," which, after risk level mapping, equals "high." This establishes a deterministic derivation relationship between the risk level and the monitoring target value, allowing the field to be removed from the review track.

[0140] For example, the above values ​​can be literals, such as text literals. Text exists in some fields, for example, tool call logs and execution environment feedback logs may contain text. The system checks if there are literals in the text that are the same as or similar to literals in prohibited fields (e.g., monitoring target values, correctness judgment results, or reward values). If so, the corresponding field is deleted.

[0141] For example, the execution environment feedback record (env_feedback) field in the review trajectory includes the following search hit details:

[0142] env_feedback = {

[0143] "hits": [

[0144] {"text": "Acceptance shall be subject to written confirmation by Party A"},

[0145] {"text": "Risk exists (gold)"} ]}

[0146] The example above includes two hit texts: the first is "Acceptance is subject to written confirmation from Party A," and the second is "There is a risk (gold)." The second text contains the literal phrase "There is a risk (gold)," which is identical to the literal phrase of the monitoring target value. Remove this field from the review trajectory.

[0147] According to the scheme of this disclosure embodiment, fields that match the values ​​of prohibited fields can be deleted, avoiding the fields with residual supervision information in the values ​​from being provided as training inputs to the model to be trained. This helps to ensure the comprehensiveness of the filtering results, further avoid supervision leakage, and thus further ensure the training effect.

[0148] The data object obtained after removing prohibited fields from the review trajectory can be called the training input view.

[0149] Furthermore, method 200 may also include performing a residual scan on the data object obtained after removing the prohibited field from the review trajectory.

[0150] In this scenario, the data object obtained after removing prohibited fields from the review trajectory can also be understood as a candidate training input view. This candidate training input view serves as the object of the residual scan; if the residual scan passes, this candidate training input view can be used as the training input view. For example, upon successful scanning, a pass marker and a scan fingerprint can be output. The scan fingerprint is used to record relevant information about this scan, such as the scan time and the rules used in the scan.

[0151] For example, in step S220, fields (first type fields and second type fields) can be recursively deleted and prohibited from the review trajectory to obtain a candidate training input view.

[0152] Residual scanning may include at least one of the following:

[0153] Perform a recursive traversal of the objects in the candidate training input view to check for the existence of prohibited fields; or...

[0154] The values ​​of fields in the candidate training input view are scanned to detect whether there are any fields whose values ​​match those of prohibited fields. Scanning the values ​​of fields in the candidate training input view can include at least the following effects: derivation detection or fragment detection.

[0155] Derivation detection refers to detecting whether there are fields whose values ​​have a deterministic derivation relationship with the values ​​of directly supervised conclusion fields.

[0156] Fragment detection refers to detecting whether there are fields whose literals are exactly the same or partially the same as the literals of the directly supervised conclusion fields.

[0157] Forbidden fields in residual scans can be handled using a hierarchical blacklist.

[0158] For example, if residual supervision information is detected, the candidate training input view can be marked as residual blocking. This marking indicates that the candidate training input view cannot be used as training input data, that is, the candidate training input view is blocked.

[0159] Furthermore, a scan report can be output, which may include at least one of the following: the path of the field to which the residual supervision information belongs (i.e., the residual field path, such as tool_calls[2].result.gold_hint), the matching level (e.g., the first level of the blacklist, the second level of the blacklist, the derived detection or the fragment detection), or the evidence of a hit (e.g., a key name hit or a literal hit).

[0160] For example, the scan report can also serve as an input to the governance count, which will be discussed later.

[0161] Furthermore, the above-mentioned detection methods can be configured according to the training task, such as specific fields in the blacklist, derived detection rules, and scanning range.

[0162] According to the embodiments of this disclosure, the introduction of a residual scanning mechanism helps to further reduce the probability of supervisory information entering the visible training input of the model.

[0163] In some training scenarios, such as supervised fine-tuning, the supervised target value can be stored separately as the target output of the training input data and associated with the training input data through sample identifiers. In this way, during the training phase, the model can be trained based on the training input data and the corresponding target supervised value.

[0164] In some training scenarios, such as reinforcement learning, training requires training with preference pairs. In the scheme disclosed herein, training preference pairs can be constructed based on comparable review trajectories.

[0165] In some embodiments, step S210 may include: acquiring raw data, which includes multiple review trajectories.

[0166] Method 200 may further include: constructing one or more training preference pairs based on multiple training input data obtained from multiple review trajectories, wherein the training input data in the same training preference pair belong to the same data governance isolation domain, and the data governance isolation domain is determined based on at least one of the following: data isolation subject or data source.

[0167] For example, multiple training input data can be grouped according to data governance isolation domains, and training preference pairs can be constructed in the data belonging to the same group.

[0168] The training input data in the same training preference pair corresponds to the same review unit, for example, the same risk point.

[0169] The following example illustrates the process of constructing training preference pairs.

[0170] For example, the original data includes four original review trajectories, and the training input data obtained based on these four original review trajectories are samples A, B, C, and D, respectively. The relevant fields of the four samples are shown in Table 1.

[0171] Table 1

[0172] A Tenant A training set (Tenant A, Training Set) R-PAY-014 better B Tenant A training set (Tenant A, Training Set) R-PAY-014 Poor C Tenant B training set (Tenant B, Training Set) R-PAY-014 Poor D Tenant A Evaluation Reservation Set (Tenant A, evaluation retention set) R-PAY-014 Poor

[0173] Of the four samples mentioned above, only samples A and B belong to the unified data governance isolation domain and can be used to construct training preference pairs.

[0174] It should be understood that the above scheme is only one implementation of the data governance isolation domain. In other implementations, the data governance isolation domain can also be divided based on other information. For example, the data governance isolation domain can also be divided based on one or more of the following: authorized scope for training purposes, privileged confidentiality status, and data source category.

[0175] According to the embodiments of this disclosure, the preference relationship is constructed by using training input data belonging to the same data isolation subject, which avoids the mixing of data belonging to different data isolation subjects. This helps to ensure that the use of data follows the compliance boundary. At the same time, the consistent distribution of data belonging to the same data isolation subject helps to avoid the introduction of training bias by the differences in data distribution across subjects, thereby helping to ensure the training effect of the model.

[0176] Furthermore, when constructing training preference pairs across data governance isolation domains is required, the training input data involved in the pairing can first undergo a de-identification operation to obtain authorization for training purposes from the corresponding data isolation entity, and cross-domain authorization evidence can be separately marked in the governance list. Pairing is not allowed without cross-domain authorization. This facilitates the construction of training preference pairs across data governance isolation domains while meeting compliance requirements.

[0177] Furthermore, paired training input data within the same data isolation domain can serve as candidate training preference pairs. Only when the approval criteria are met can they be used as training preference pairs; otherwise, they remain in the candidate state.

[0178] In some embodiments, the training preference meets approval criteria, which are related to at least one of the following:

[0179] (1) The confidence that the preferred training input data in the training preference pair is better than the non-preferred training input data in the training preference pair;

[0180] (2) The consistency between the structure of the preference training input data and the structure of the non-preference training input data in the preference pair; or

[0181] (3) The consistency between the evidence in the preference training input data and the review rules corresponding to the preference training input data in the training preference pair, and the consistency between the evidence in the non-preference training input data and the review rules corresponding to the non-preference training input data in the training preference pair.

[0182] The above item (1) can be simply referred to as the confidence condition, which is used to ensure that the better sample in the training preference pair is actually a better sample. For example, if the confidence that the preferred sample (i.e., the preferred training input data) in the candidate training preference pair is better than the non-preferred sample (i.e., the non-preferred training input data) is greater than or equal to a preset threshold, the candidate training preference pair satisfies the confidence condition.

[0183] For example, this confidence level can be output by a large model. For instance, taking samples A and B mentioned earlier as examples, sample A is a preferred sample and sample B is a non-preferred sample. The large model is instructed by prompt words to output a confidence level that sample A is better than sample B.

[0184] The above term (2) can be called the structural consistency condition, which is used to ensure that the two training input data in the training preference pair have the same structure.

[0185] For example, the training input data may include at least one of the following: review instructions, text fragments, or review rules (e.g., rules for risk points). For instance, if the two training input data in a candidate training preference pair have the same review instructions, the same text fragments, and the same review rules, the candidate training preference pair satisfies the structural consistency condition.

[0186] The above item (3) can be called the evidence consistency condition, which is used to ensure that the evidence in the preference samples of the training preference pair is more consistent with the corresponding review rules.

[0187] For example, if the evidence in the preferred sample hits the rule, but the non-preferred sample does not, then the candidate training preference pair satisfies the evidence consistency condition.

[0188] For example, the validator can automatically verify whether the candidate training preference pair meets the above approval conditions. If it does, the training preference pair is output; if it does not, the candidate state of the candidate training preference pair is retained.

[0189] Alternatively, the review can be conducted manually, but this disclosure does not limit the scope of the embodiments.

[0190] According to the embodiments of this disclosure, by setting approval conditions to screen training preference pairs, it is beneficial to ensure the quality of training preference pairs, thereby further improving the training effect of the model.

[0191] In some embodiments, method 200 may further include: outputting a training input data set, wherein the training input data set includes training input data, and the output condition is related to at least one of the following of the training input data: the anonymization status of the training input data, the authorization information for the training purpose of the training input data, the data source of the training input data, the legal reservation status of the training input data, the privileged confidentiality status of the training input data, the data isolation subject of the training input data, the data structure of the training input data, or the data size of the training input data set.

[0192] In other words, in method 200, governance counting can also be performed, and the output condition can also be called the governance condition. If the governance condition is not met, the training input data set is not allowed to be output. The input data set can also be called the training task input payload.

[0193] If training preference pairs are constructed, the training input dataset can include training preference pairs.

[0194] For example, the output conditions may include at least one of the following:

[0195] The desensitization status is true (desensitized);

[0196] The training purpose authorization is set to true (training use is allowed);

[0197] The data source is either the training set or the development set;

[0198] There are no restrictions on the non-trainability of the legal reservation status instruction;

[0199] The privileged confidentiality status indication does not involve privileged confidentiality;

[0200] The data isolation entities of the training input data in the training preference pair are the same;

[0201] The training input data for each training preference pair includes censorship instructions, text fragments, and censorship rules, and all censorship instructions, text fragments, and censorship rules are identical; or

[0202] The number of training input data in the training input dataset exceeds the size threshold.

[0203] Different training input data formats can be generated for different training tasks and different training methods, and this disclosure does not limit the specific training input data.

[0204] Furthermore, a governance checklist can be output, which includes at least one of the following: the identifier of the review trajectory, the identifier of the training input data, the data source, the scan report, the governance status, access control credentials, and payload integrity information. Access control credentials are used to indicate that the training input data meets the output conditions.

[0205] This is beneficial for subsequent auditing, backtracking, and pre-training verification.

[0206] According to the embodiments of this disclosure, output conditions can be set for the training input dataset. Output is only allowed when these conditions are met, which helps to further ensure the reliability of the data entering the training process. For example, output conditions may include anonymization status. If the anonymization status is unknown or does not meet the conditions, output is not allowed, which helps to reduce the probability of unanonymized data entering the training process. As another example, output conditions may include authorization information for training purposes. If this information is unknown or does not meet the conditions, output is not allowed, which helps to reduce the probability of unauthorized data entering the training process and ensure data compliance. Furthermore, output conditions may include the size of the training input dataset. The payload is only derived when the number of candidate training input data corresponding to the current training task's input payload (i.e., the size of the training input dataset) meets a size threshold, which helps to ensure the size of the training input data and thus ensure the stability of the training.

[0207] The methods of the embodiments of this disclosure can be applied to a variety of scenarios. For example, the solutions of the embodiments of this disclosure can be applied to intelligent contract review products, legal document retrieval enhancement and generation systems, intelligent agent tool execution frameworks, model training data governance systems, rule engines, and business rule management systems.

[0208] Figure 3 A block diagram of a system architecture according to an embodiment of the present disclosure is shown as an example. Figure 3 The system 300 shown can be used to perform the method for constructing training samples according to the embodiments of this disclosure, such as method 200 described above.

[0209] like Figure 3 As shown, system 300 may include trajectory source 310, trajectory review layer 320, training input view layer 330 and governance access control layer 340.

[0210] Trajectory source 310 is the source of the raw data. For example, such as... Figure 3 As shown, trajectory source 310 may include baseline data (i.e., review input), supervision target value, reward value, quality evaluation, model output results, tool call records, execution environment feedback records, sensitive content, and governance control content.

[0211] The review trajectory layer 320 includes an isolation module 321 and an export module 322.

[0212] The isolation module 321 is used to physically isolate the training set, development set, and evaluation retention set into different source files, source directories, or source lists before exporting data from the trajectory source, so as to avoid the data of the evaluation retention set being exported by module 322.

[0213] Export module 322 is used to combine review inputs, model outputs, tool call records, execution environment feedback records, supervision target values, reward values, and quality evaluations into a review trajectory. Alternatively, as described above, it can combine review inputs, model outputs, tool call records, execution environment feedback records, supervision target values, reward values, quality evaluations, sensitive content, and governance control content into a review trajectory.

[0214] The training input view layer 330 includes a generation module 331, a scanning module 332, and a construction module 333.

[0215] The generation module 331 is used to recursively delete prohibited fields.

[0216] The scanning module 332 is used to perform residual scanning on the results output by the generation module 331, and output a scan report and residual blocking markers.

[0217] Construction module 333 is used to group data according to data governance isolation domains, and to construct training preference pairs only within the same group. Furthermore, if the approval criteria are not met, the training preference pair is marked as a candidate state and is not allowed to be output.

[0218] The access control layer 340 includes a management counting module 341 and an access control module 342.

[0219] The governance counting module 341 is used to count whether the anonymization status, training purpose authorization, data source, data isolation subject, legal retention status, privileged confidentiality status, and audit confirmation status meet the requirements.

[0220] Access control module 342 is used to comprehensively assess the results of governance counts and determine whether training task input payloads can be generated. Specifically, it executes a failure-closed export access control when any condition is not met, governance control content fields are missing, the status is unknown, or evidence citations are incomplete, preventing the generation of training task input payloads. Furthermore, it can also output a failure-closed list and audit artifacts. For example, training task input payloads can be prevented from being generated by default, with the access control defaulting to a failure-closed state. This means that samples remain in the audit or candidate stage by default, unable to generate training task input payloads. Only when governance conditions are met can training task input payloads be generated, further reducing the probability of training with non-compliant data.

[0221] like Figure 3 As shown in the embodiments of this disclosure, the review trajectory, training input view, and training task input payload can be hierarchically isolated to avoid direct training based on the review trajectory. Simultaneously, the states in the training task input payload, review trajectory, training input view, scan report, and governance count are correlated, which is beneficial for subsequent auditing, backtracking, location, and pre-training verification.

[0222] To better illustrate the method for constructing training samples in this disclosure, the following example demonstrates the method for constructing training samples in this disclosure using a specific processing flow.

[0223] Figure 4 The illustration shows a schematic diagram of the execution process of a method flow for constructing training samples according to an embodiment of the present disclosure. Figure 4 The execution process described herein can be considered a specific example of method 200 above, and does not constitute a limitation on the scheme disclosed herein. To avoid repetition, in the description... Figure 4 The description of method 400 shown is omitted appropriately. For the specific process, please refer to method 200 above.

[0224] like Figure 4 As shown, method 400 may include the following steps.

[0225] S1, Read the review data.

[0226] Read relevant data from each risk point review instance.

[0227] For example, for each risk point review instance, at least one of the following is read: input references, tool call records, execution environment feedback records, model output results, reward values, and governance control content.

[0228] S2, data partitioning.

[0229] The read data is divided according to the source file.

[0230] Export data for training or development, and intercept data retained for evaluation.

[0231] S3 generates the review trajectory.

[0232] The above information for each risk point review instance is structured to generate the review trajectory corresponding to each risk point.

[0233] For example, the risk point-level review trajectory includes identity identifiers, model input data, tool call records, execution environment feedback records, model output results, governance control content, and sensitive content.

[0234] S4, Delete prohibited fields.

[0235] S5, perform residual scan.

[0236] If the scan passes, proceed to step S6.

[0237] If the scan fails, proceed to step S14.

[0238] S6, determine whether they belong to the same governance isolation domain.

[0239] Determine whether two training input data sets to be paired belong to the same governance isolation domain.

[0240] If so, proceed to step S7.

[0241] If not, proceed to step S15.

[0242] S7, construct candidate training preference pairs.

[0243] S8, determine whether the approval conditions are met.

[0244] If so, proceed to step S9.

[0245] If not, proceed to step S10.

[0246] S9 generates training preference pairs.

[0247] Candidate preference pairs that meet the approval criteria can be used as training preference pairs. For example, a state can be set to indicate that the approval criteria are met.

[0248] Continue with step S11.

[0249] S10 is marked as a candidate state.

[0250] Candidate preference pairs that do not meet the approval criteria are not allowed to be used as training preference pairs, and their states are marked as candidate states. Proceed to step S11.

[0251] S11, determine whether the output condition is met.

[0252] If so, proceed to step S12.

[0253] If not, proceed to step S13.

[0254] S12 outputs the training task input payload and governance list.

[0255] S13, maintain the access control system in a failed closed state and output an audit list (i.e., a closed failure list and audit artifacts). Cases where it cannot be determined whether the output conditions are met (e.g., missing information) can be treated as cases where the output conditions are not met.

[0256] S14, Output scan report.

[0257] S15, refuse to construct candidate training preference pairs.

[0258] Figure 5 A flowchart of a method for training a text review model according to an embodiment of the present disclosure is shown as an example. Figure 5 The method shown can be executed by a server, by a terminal device, or by both a server and a terminal device. For example, Figure 5 The method 500 shown can be executed by the system architecture described above. Figure 5 The training input data in the method shown is generated using the training sample construction method in the embodiments of this disclosure. For a detailed description, please refer to method 200 above. To avoid repetition, some descriptions of method 500 are omitted appropriately.

[0259] Figure 5 The method 500 shown can also be understood as a training method for reviewing intelligent agents.

[0260] like Figure 5 As shown, method 500 may include the following steps.

[0261] Step S510: Obtain training input data, which is obtained according to the training sample construction method in the embodiments of this disclosure.

[0262] Step S520: Train the initial text review model based on the training input data to obtain the target text review model.

[0263] The initial text review model and the target text review model are only used to distinguish between the text review models before and after training, and have no other limiting function.

[0264] Step S520 can also be replaced by training the review agent based on the training input data to obtain a trained review agent.

[0265] For a detailed description of the review agent and text review model, please refer to Method 200 above, which will not be repeated here.

[0266] In some embodiments, the training input data is obtained by removing prohibited fields from the review trajectories. The prohibited fields are related to the output of the model invoked by the review agent during the text review task, and the review trajectories indicate the records generated by the review agent during the text review task.

[0267] In some embodiments, the training input data is obtained by removing prohibited fields from the top-level fields of the review trajectory and performing a recursive traversal of the nested structure in the review trajectory to remove prohibited fields from the fields of the nested structure.

[0268] In some embodiments, the review trajectory includes at least one of the following: model input data, tool call records, execution environment feedback records, model output results, supervision target values, reward values, quality evaluation, correctness judgment results, manual review information, sensitive content, or governance control content.

[0269] In some embodiments, prohibited fields include first type fields, which include at least one of the following: model output, supervision target value, reward value, quality evaluation, or correctness determination result.

[0270] In some embodiments, the prohibited field also includes a second type field, which is derived from the first type field.

[0271] In some embodiments, the training input data is obtained by removing at least one of the following from the review trajectory: fields that match the name of a prohibited field, or fields that match the value of a prohibited field.

[0272] In some embodiments, training an initial text review model based on training input data includes: training the initial text review model based on one or more training preference pairs, wherein the one or more training preference pairs are constructed from multiple training input data, and the training input data in the same training preference pair belong to the same data governance isolation domain, wherein the data governance isolation domain is determined based on at least one of the following: data isolation subject or data source.

[0273] In some embodiments, the training preference pair meets approval criteria, which are related to at least one of the following: the confidence that the preferred training input data in the training preference pair is superior to the non-preferred training input data in the training preference pair; the consistency between the structure of the preferred training input data and the structure of the non-preferred training input data in the training preference pair; or the consistency between the evidence in the preferred training input data and the review rules corresponding to the preferred training input data, and the consistency between the evidence in the non-preferred training input data and the review rules corresponding to the non-preferred training input data in the training preference pair.

[0274] In some embodiments, the training input data belongs to a training input data set, which is output under certain output conditions. The output conditions are related to at least one of the following: the anonymization status of the training input data, the authorization information for the training purpose of the training input data, the data source of the training input data, the legal reservation status of the training input data, the privileged confidentiality status of the training input data, the data isolation subject of the training input data, the data structure of the training input data, or the data size of the training input data set.

[0275] In some embodiments, the review trajectory indicates the records generated by the review agent during the process of performing a text review task based on the risk points corresponding to the review trajectory.

[0276] Figure 6 A flowchart of a text review method according to an embodiment of the present disclosure is shown as an example. Figure 6The method shown can be executed by a server, by a terminal device, or by both a server and a terminal device. For example, Figure 6 The method 600 shown can be executed by the system architecture described above. Figure 6 The target text review model in the method shown is trained using the text review model training method in the embodiments of this disclosure, and the training input data is generated using the training sample construction method in the embodiments of this disclosure. For detailed descriptions, please refer to methods 200 and 500 above. To avoid repetition, some descriptions are omitted when describing method 600.

[0277] like Figure 6 As shown, method 600 may include the following steps.

[0278] Step S610: Obtain the text to be reviewed.

[0279] Step S620: Input the text to be reviewed into the target text review model for processing to obtain the review result. The target text review model is trained by the training method of the text review model in the embodiments of this disclosure.

[0280] Step S620 can also be replaced by inputting the text to be reviewed into the review agent for processing to obtain the review result. The review agent is trained using the training method for the review agent in the embodiments of this disclosure.

[0281] For a detailed description of the review agent and text review model, please refer to Method 200 above, which will not be repeated here.

[0282] In some embodiments, the target text review model is obtained by training an initial text review model based on training input data, which is obtained according to the method for constructing training samples in the embodiments of this disclosure.

[0283] In some embodiments, the training input data is obtained by removing prohibited fields from the review trajectories. The prohibited fields are related to the output of the model invoked by the review agent during the text review task, and the review trajectories indicate the records generated by the review agent during the text review task.

[0284] In some embodiments, the training input data is obtained by removing prohibited fields from the top-level fields of the review trajectory and performing a recursive traversal of the nested structure in the review trajectory to remove prohibited fields from the fields of the nested structure.

[0285] In some embodiments, the review trajectory includes at least one of the following: model input data, tool call records, execution environment feedback records, model output results, supervision target values, reward values, quality evaluation, correctness judgment results, manual review information, sensitive content, or governance control content.

[0286] In some embodiments, prohibited fields include first type fields, which include at least one of the following: model output, supervision target value, reward value, quality evaluation, or correctness determination result.

[0287] In some embodiments, the prohibited field also includes a second type field, which is derived from the first type field.

[0288] In some embodiments, the training input data is obtained by removing at least one of the following from the review trajectory: fields that match the name of a prohibited field, or fields that match the value of a prohibited field.

[0289] In some embodiments, the target text review model is obtained by training an initial text review model based on one or more training preference pairs. The one or more training preference pairs are constructed from multiple training input data. The training input data in the same training preference pair belong to the same data governance isolation domain. The data governance isolation domain is determined based on at least one of the following: data isolation subject or data source.

[0290] In some embodiments, the training preference pair meets approval criteria, which are related to at least one of the following: the confidence that the preferred training input data in the training preference pair is superior to the non-preferred training input data in the training preference pair; the consistency between the structure of the preferred training input data and the structure of the non-preferred training input data in the training preference pair; or the consistency between the evidence in the preferred training input data and the review rules corresponding to the preferred training input data, and the consistency between the evidence in the non-preferred training input data and the review rules corresponding to the non-preferred training input data in the training preference pair.

[0291] In some embodiments, the training input data belongs to a training input data set, which is output under certain output conditions. The output conditions are related to at least one of the following: the anonymization status of the training input data, the authorization information for the training purpose of the training input data, the data source of the training input data, the legal reservation status of the training input data, the privileged confidentiality status of the training input data, the data isolation subject of the training input data, the data structure of the training input data, or the data size of the training input data set.

[0292] In some embodiments, the review trajectory indicates the records generated by the review agent during the process of performing a text review task based on the risk points corresponding to the review trajectory.

[0293] Figure 7 A block diagram of an apparatus for constructing agent-based training samples according to an embodiment of the present disclosure is shown schematically.

[0294] like Figure 7 As shown, the device 700 may include an acquisition module 710 and a generation module 720.

[0295] The operations performed by modules 710 to 720 in device 700 and the effects they can achieve are similar to steps S210 to S220 in the above-mentioned training sample construction method. To avoid repetition, appropriate parts of the description are provided when describing device 700.

[0296] The acquisition module 710 is used to acquire the review trajectory, which indicates the records generated by the review agent during the execution of the text review task.

[0297] The generation module 720 is used to remove prohibited fields from the review trajectory to obtain training input data. The prohibited fields are related to the output results of the model called by the review agent during the text review task.

[0298] In some embodiments, the generation module 720 includes a first deletion module and a second deletion module.

[0299] The first deletion module is used to remove prohibited fields from the top-level fields in the review trajectory;

[0300] The second deletion module is used to perform recursive traversal of nested structures in the review trajectory and delete prohibited fields from the fields of the nested structures.

[0301] In some embodiments, the review trajectory includes at least one of the following: model input data, tool call records, execution environment feedback records, model output results, supervision target values, reward values, quality evaluation, correctness judgment results, manual review information, sensitive content, or governance control content.

[0302] In some embodiments, prohibited fields include first type fields, which include at least one of the following: model output, supervision target value, reward value, quality evaluation, or correctness determination result.

[0303] In some embodiments, the prohibited field also includes a second type field, which is derived from the first type field.

[0304] In some embodiments, the generation module 720 is configured to: delete from the review track at least one of the following: a field that matches the name of a prohibited field, or a field that matches the value of a prohibited field.

[0305] In some embodiments, the acquisition module 710 is used to acquire raw data, which includes multiple review trajectories; the generation module 720 further includes a construction module for constructing one or more training preference pairs based on multiple training input data obtained from the multiple review trajectories, wherein the training input data in the same training preference pair belongs to the same data governance isolation domain, and the data governance isolation domain is determined based on at least one of the following: data isolation subject or data source.

[0306] In some embodiments, the training preference pair meets approval criteria, which are related to at least one of the following: the confidence that the preferred training input data in the training preference pair is superior to the non-preferred training input data in the training preference pair; the consistency between the structure of the preferred training input data and the structure of the non-preferred training input data in the training preference pair; or the consistency between the evidence in the preferred training input data and the review rules corresponding to the preferred training input data, and the consistency between the evidence in the non-preferred training input data and the review rules corresponding to the non-preferred training input data in the training preference pair.

[0307] In some embodiments, the apparatus 700 may further include an output module for outputting a training input data set, which includes training input data, when output conditions are met. The output conditions are related to at least one of the following: the anonymization status of the training input data, the authorization information for the training purpose of the training input data, the data source of the training input data, the legal reservation status of the training input data, the privileged confidentiality status of the training input data, the data isolation subject of the training input data, the data structure of the training input data, or the data size of the training input data set.

[0308] In some embodiments, the review trajectory indicates the records generated by the review agent during the process of performing a text review task based on the risk points corresponding to the review trajectory.

[0309] Figure 8 A block diagram of a training apparatus for a text review model according to an embodiment of the present disclosure is shown schematically.

[0310] like Figure 8 As shown, the device 800 may include an acquisition module 810 and a training module 820.

[0311] The operations performed by modules 810 to 820 in device 800 and the effects they can achieve are similar to steps S510 to S520 in the training method of the above-mentioned text review model. To avoid repetition, appropriate parts of the description are provided when describing device 800.

[0312] The acquisition module 810 is used to acquire training input data, which is obtained by the method of constructing training samples in the embodiments of this disclosure.

[0313] The training module 820 is used to train the initial text review model based on the training input data to obtain the target text review model.

[0314] In some embodiments, the training input data is obtained by removing prohibited fields from the review trajectories. The prohibited fields are related to the output of the model invoked by the review agent during the text review task, and the review trajectories indicate the records generated by the review agent during the text review task.

[0315] In some embodiments, the training input data is obtained by removing prohibited fields from the top-level fields of the review trajectory and performing a recursive traversal of the nested structure in the review trajectory to remove prohibited fields from the fields of the nested structure.

[0316] In some embodiments, the review trajectory includes at least one of the following: model input data, tool call records, execution environment feedback records, model output results, supervision target values, reward values, quality evaluation, correctness judgment results, manual review information, sensitive content, or governance control content.

[0317] In some embodiments, prohibited fields include first type fields, which include at least one of the following: model output, supervision target value, reward value, quality evaluation, or correctness determination result.

[0318] In some embodiments, the prohibited field also includes a second type field, which is derived from the first type field.

[0319] In some embodiments, the training input data is obtained by removing at least one of the following from the review trajectory: fields that match the name of a prohibited field, or fields that match the value of a prohibited field.

[0320] In some embodiments, the training module 820 is used to: train an initial text review model based on one or more training preference pairs, wherein the one or more training preference pairs are constructed from multiple training input data, and the training input data in the same training preference pair belong to the same data governance isolation domain, wherein the data governance isolation domain is determined based on at least one of the following: data isolation subject or data source.

[0321] In some embodiments, the training preference pair meets approval criteria, which are related to at least one of the following: the confidence that the preferred training input data in the training preference pair is superior to the non-preferred training input data in the training preference pair; the consistency between the structure of the preferred training input data and the structure of the non-preferred training input data in the training preference pair; or the consistency between the evidence in the preferred training input data and the review rules corresponding to the preferred training input data, and the consistency between the evidence in the non-preferred training input data and the review rules corresponding to the non-preferred training input data in the training preference pair.

[0322] In some embodiments, the training input data belongs to a training input data set, which is output under certain output conditions. The output conditions are related to at least one of the following: the anonymization status of the training input data, the authorization information for the training purpose of the training input data, the data source of the training input data, the legal reservation status of the training input data, the privileged confidentiality status of the training input data, the data isolation subject of the training input data, the data structure of the training input data, or the data size of the training input data set.

[0323] In some embodiments, the review trajectory indicates the records generated by the review agent during the process of performing a text review task based on the risk points corresponding to the review trajectory.

[0324] Figure 9 A block diagram of a text review apparatus according to an embodiment of the present disclosure is shown schematically.

[0325] like Figure 9 As shown, the device 900 may include an acquisition module 910 and an examination module 920.

[0326] The operations performed by modules 910 to 920 in device 900 and the effects they can achieve are similar to steps S610 to S620 in the above-mentioned text review method. To avoid repetition, appropriate parts of the description are provided when describing device 900.

[0327] Module 910 is used to acquire the text to be reviewed.

[0328] The review module 920 is used to input the text to be reviewed into the target text review model for processing and to obtain the review result. The target text review model is trained by the training method of the text review model of the embodiment of this disclosure.

[0329] In some embodiments, the target text review model is obtained by training an initial text review model based on training input data, which is obtained according to the method for constructing training samples in the embodiments of this disclosure.

[0330] In some embodiments, the training input data is obtained by removing prohibited fields from the review trajectories. The prohibited fields are related to the output of the model invoked by the review agent during the text review task, and the review trajectories indicate the records generated by the review agent during the text review task.

[0331] Figure 10 A schematic block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is shown.

[0332] In embodiments of this disclosure, such as Figure 10 As shown, the AI ​​agent 1000 may include an input module 1010, a processing module 1020, and an output module 1030.

[0333] Input module 1010 is used to receive input information;

[0334] The processing module 1020 is used to execute the agent-based text review method, text review model training method, or text review method provided in the embodiments of this disclosure to obtain output information based on the input information received by the input module;

[0335] Output module 1030 is used to output the output information obtained by the processing module.

[0336] According to embodiments of this disclosure, the input module 1010 is responsible for receiving or sensing information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment), and converting it into a format that the AI ​​agent 1000 can understand and process. The input module 1010 is the primary link for the AI ​​agent 1000 to interact with the outside world, enabling the AI ​​agent 1000 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.

[0337] In the example, input module 1010 can input the review trajectory described above, training input data, or text to be reviewed.

[0338] In the example, the processing module 1020 is the core support for the AI ​​agent 1000's ability to handle complex tasks. The processing module 1020 can execute the agent-based training sample construction method, the text review model training method, or the text review method described above.

[0339] In the example, the performance of the processing module 1020 is closely related to the large model on which the AI ​​agent 1000 is based. To fully leverage the capabilities of the large model, the internal structure of the processing module 1020 can be designed to be highly configurable and scalable to handle various types of tasks and requirements in real-world scenarios.

[0340] In the example, after the AI ​​agent 1000 acquires the review trajectory, the processing module 1020 can remove prohibited fields from the review trajectory to obtain training input data. The prohibited fields are related to the output results of the model called by the review agent during the text review task. The training input data can be passed to the output module 1030.

[0341] Understandably, while large language models possess excellent language understanding and generation capabilities, like humans, their ability to solve tasks is limited without the aid of any tools. Once the AI ​​agent 1000 is given the ability to invoke tools, it can perform tasks such as using a calculator to complete mathematical calculations, using Python to perform data analysis, and using a search engine to create weather forecasts.

[0342] In the example, output module 1030 can output the training input data described above.

[0343] The AI ​​agent 1000 according to the embodiments of this disclosure can simply and effectively improve the level of intelligence, as well as enhance flexibility and versatility.

[0344] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0345] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the methods described above.

[0346] According to embodiments of the present disclosure, a non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to cause a computer to perform the methods described above.

[0347] According to an embodiment of this disclosure, a computer program product includes a computer program that, when executed by a processor, implements the method described above.

[0348] Figure 11A schematic block diagram of an example electronic device 1100 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0349] like Figure 11 As shown, device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 1102 or a computer program loaded from storage unit 1108 into random access memory (RAM) 1103. The RAM 1103 may also store various programs and data required for the operation of device 1100. The computing unit 1101, ROM 1102, and RAM 1103 are interconnected via bus 1104. An input / output (I / O) interface 1105 is also connected to bus 1104.

[0350] Multiple components in device 1100 are connected to I / O interface 1105, including: input unit 1106, such as keyboard, mouse, etc.; output unit 1107, such as various types of monitors, speakers, etc.; storage unit 1108, such as disk, optical disk, etc.; and communication unit 1109, such as network card, modem, wireless transceiver, etc. Communication unit 1109 allows device 1100 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0351] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above, such as methods for constructing training samples based on intelligent agents, methods for training text review models, or text review methods. For example, in some embodiments, methods for constructing training samples based on intelligent agents, methods for training text review models, or text review methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1100 via ROM 1102 and / or communication unit 1109. When the computer program is loaded into RAM 1103 and executed by computing unit 1101, one or more steps of the agent-based training sample construction method, text review model training method, or text review method described above can be performed. Alternatively, in other embodiments, computing unit 1101 can be configured by any other suitable means (e.g., by means of firmware) to execute the agent-based training sample construction method, text review model training method, or text review method.

[0352] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0353] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0354] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0355] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0356] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0357] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.

[0358] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0359] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for constructing training samples based on intelligent agents, applied to governance intelligent agents, comprising: Acquire the review trajectory, which indicates the records generated by the review agent during the execution of the text review task; The prohibited fields are removed from the review trajectory to obtain training input data. The prohibited fields are related to the output of the model called by the review agent during the execution of the text review task.

2. The method according to claim 1, wherein, The removal of prohibited fields from the review track includes: Remove the prohibited field from the top-level field in the review trajectory; A recursive traversal is performed on the nested structures in the review trajectory, and the prohibited fields are deleted from the fields of the nested structures.

3. The method according to claim 1 or 2, wherein, The review trajectory includes at least one of the following: the model's input data, tool call records, execution environment feedback records, the model's output results, supervision target values, reward values, quality evaluation, correctness judgment results, manual review information, sensitive content, or governance control content.

4. The method according to claim 3, wherein, The prohibited fields include a first type of field, which includes at least one of the following: the output of the model, the supervision target value, the reward value, the quality evaluation, or the correctness determination result.

5. The method according to claim 4, wherein, The prohibited field also includes a second type field, which is derived from the first type field.

6. The method according to any one of claims 1 to 5, wherein, The removal of prohibited fields from the review track includes: Delete at least one of the following from the review track: A field that matches the name of the prohibited field, or a field that matches the value of the prohibited field.

7. The method according to any one of claims 1 to 6, wherein, The acquisition of the review trajectory includes: The method further includes: acquiring raw data, which includes multiple review trajectories; and .... One or more training preference pairs are constructed based on multiple training input data obtained from the multiple review trajectories, wherein the training input data in the same training preference pair belong to the same data governance isolation domain, and the data governance isolation domain is determined based on at least one of the following: data isolation subject or data source.

8. The method according to claim 7, wherein, The training preferences meet approval criteria, which are related to at least one of the following: The confidence that the preferred training input data in the training preference pair is superior to the non-preferred training input data in the training preference pair; The consistency between the structure of the preference training input data and the structure of the non-preference training input data in the training preference pair; or The consistency between the evidence in the preference training input data of the training preference pair and the review rules corresponding to the preference training input data, and the consistency between the evidence in the non-preference training input data of the training preference pair and the review rules corresponding to the non-preference training input data.

9. The method according to any one of claims 1 to 8, further comprising: Under the condition that the output conditions are met, a training input data set is output, the training input data set including the training input data, and the output conditions are related to at least one of the following: the anonymization status of the training input data, the authorization information for the training purpose of the training input data, the data source of the training input data, the legal retention status of the training input data, the privileged confidentiality status of the training input data, the data isolation subject of the training input data, the data structure of the training input data, or the data size of the training input data set.

10. The method according to any one of claims 1 to 9, wherein, The review trajectory indicates the records generated by the review agent during the process of performing the text review task based on the risk points corresponding to the review trajectory.

11. A training method for a text censorship model, comprising: Acquire training input data, wherein the training input data is obtained according to the method of any one of claims 1 to 10; The initial text review model is trained based on the training input data to obtain the target text review model.

12. A text review method, comprising: Obtain the text to be reviewed; The text to be reviewed is input into the target text review model for processing to obtain the review result. The target text review model is trained by the method described in claim 11.

13. An apparatus for constructing training samples based on intelligent agents, applied to governing intelligent agents, comprising: The acquisition module is used to acquire the review trajectory, which indicates the records generated by the review agent during the execution of the text review task; The generation module is used to remove prohibited fields from the review trajectory to obtain training input data, wherein the prohibited fields are related to the output results of the model called by the review agent during the execution of the text review task.

14. A training device for a text censorship model, comprising: An acquisition module is configured to acquire training input data, wherein the training input data is obtained according to the method described in any one of claims 1 to 10; The training module is used to train the initial text review model based on the training input data to obtain the target text review model.

15. A text review device, comprising: The acquisition module is used to acquire the text to be reviewed; The review module is used to input the text to be reviewed into the target text review model for processing and to obtain the review result. The target text review model is trained by the method described in claim 11.

16. An intelligent agent of artificial intelligence, comprising: The input module is used to receive input information; The processing module is configured to perform the method as described in any one of claims 1 to 10, 11 or 12 based on the input information received by the input module to obtain output information; An output module is used to output the output information obtained by the processing module.

17. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 10, 11 or 12.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 10, 11 or 12.

19. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 10, 11 or 12.