Training method of auditing attribution model and related device
By integrating multiple first audit categories corresponding to multiple first sample data into a small number of coarse-grained audit categories, sorting out cumbersome audit rules into streamlined rules, and training audit attribution models, the problems of low audit efficiency and accuracy in the existing technology are solved, and efficient and accurate content data audits are achieved.
Patent Information
- Application Number
- CN202510152887.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, the audit model cannot understand the cumbersome audit rules corresponding to a large number of fine-grained audit categories, resulting in low audit efficiency and accuracy.
By integrating multiple first audit categories corresponding to multiple first sample data into a small number of coarse-grained audit categories, and sorting out a large number of cumbersome audit rules into a small number of streamlined audit rules, labeling attribution data whose sample data belongs to the coarse-grained audit category, continuing to train the pre-trained model as an audit attribution model, so that it fully understands the streamlined audit rules corresponding to a small number of coarse-grained audit categories.
It improves the efficiency and accuracy of content data review, so that the audit attribution model can quickly and accurately analyze whether the content data belongs to a small number of coarse-grained audit categories.
Smart Images

Figure CN120067688A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method for training an audit attribution model and related devices. Background Art
[0002] In the content audit scenario, content data of the target audit category, such as harmful data, is the key data that needs to be focused on in content audit. Through the audit model, it can be audited whether the content data belongs to the target audit category, which can, to a certain extent, reduce the audit burden of auditors and improve the audit efficiency of auditors.
[0003] In the related art, the training method of the audit model refers to: obtaining a large amount of content data of the target audit category and the fine-grained audit category of each content data as a large amount of sample data and the audit category of each sample data, and performing fine-tuning training on the large language model to obtain the audit model.
[0004] However, a large amount of content data of the target audit category involves a large number of fine-grained audit categories, and a large number of fine-grained audit categories correspond to a large number of cumbersome audit rules. In the above method, the audit model cannot understand a large number of cumbersome audit rules, and relies on the audit model to audit whether the content data belongs to the target audit category, resulting in low audit efficiency and audit accuracy. Summary of the Invention
[0005] To solve the above technical problems, this application provides a method for training an audit attribution model and related devices, which fuse a large number of fine-grained audit categories corresponding to a large amount of sample data into a small number of coarse-grained audit categories, sort out a large number of cumbersome audit rules into a small number of streamlined audit rules for each coarse-grained audit category, and label the attribution data of a large amount of sample data belonging to the coarse-grained audit category based on the small number of streamlined audit rules corresponding to the small number of coarse-grained audit categories. By using the sample data, the coarse-grained audit category of the sample data, and the attribution data belonging to the coarse-grained audit category, continue to train the pre-trained model into an audit attribution model, so that the audit attribution model can fully understand the small number of streamlined audit rules corresponding to the small number of coarse-grained audit categories, and analyze the attribution data belonging to the coarse-grained audit category when auditing the content data, thereby improving the audit efficiency and audit accuracy.
[0006] The embodiments of this application disclose the following technical solutions:
[0007] On the one hand, the embodiments of this application provide a method for training an audit attribution model, and the method includes:
[0008] Performing category fusion on multiple first audit categories corresponding to multiple first sample data to obtain multiple second audit categories; the total number of the multiple second audit categories is less than the total number of the multiple first audit categories;
[0009] Sort the multiple first review rules corresponding to each second review category to obtain multiple second review rules corresponding to each second review category; the total number of the multiple second review rules is less than the total number of the multiple first review rules;
[0010] Perform attribution annotation on the first sample data according to the multiple second review rules corresponding to the second review category of the first sample data to obtain the attribution annotation data of the first sample data;
[0011] Train a pre-trained model according to the first sample data, the second review category of the first sample data, and the attribution annotation data to obtain an audit attribution model.
[0012] On the other hand, an embodiment of the present application provides a training device for an audit attribution model, and the device includes: a category fusion unit, a rule sorting unit, an attribution annotation unit, and a model training unit;
[0013] The category fusion unit is configured to perform category fusion on multiple first review categories corresponding to multiple first sample data to obtain multiple second review categories; the total number of the multiple second review categories is less than the total number of the multiple first review categories;
[0014] The rule sorting unit is configured to sort the multiple first review rules corresponding to each second review category to obtain multiple second review rules corresponding to each second review category; the total number of the multiple second review rules is less than the total number of the multiple first review rules;
[0015] The attribution annotation unit is configured to perform attribution annotation on the first sample data according to the multiple second review rules corresponding to the second review category of the first sample data to obtain the attribution annotation data of the first sample data;
[0016] The model training unit is configured to train a pre-trained model according to the first sample data, the second review category of the first sample data, and the attribution annotation data to obtain an audit attribution model.
[0017] On the other hand, an embodiment of the present application provides a computer device, and the computer device includes a processor and a memory:
[0018] The memory is used to store a computer program and transmit the computer program to the processor;
[0019] The processor is configured to execute the method described in any of the foregoing aspects according to the instructions in the computer program.
[0020] On the other hand, an embodiment of the present application provides a computer-readable storage medium for storing a computer program. When the computer program runs on a computer device, the computer device is caused to execute the method described in any of the foregoing aspects.
[0021] On the other hand, an embodiment of the present application provides a computer program product including a computer program. When the computer program runs on a computer device, the computer device is caused to execute the method described in any of the foregoing aspects.
[0022] As can be seen from the above technical solutions, multiple first review categories corresponding to multiple first sample data are fused into multiple second review categories, so that the total number of the multiple second review categories is less than the total number of the multiple first review categories, in order to reduce a large number of first review categories corresponding to a large number of first sample data to a small number of second review categories; multiple first review rules corresponding to each second review category are sorted out into multiple second review rules, so that the total number of the multiple second review rules is less than the total number of the multiple first review rules, in order to simplify a large number of first review rules corresponding to each second review category to a small number of second review rules; attribution annotation data indicating that the first sample data belongs to the second review category is marked through the multiple second review rules corresponding to the second review category of the first sample data, so as to clarify attribution data indicating that a large number of first sample data belong to the second review category during content review; a review attribution model is obtained by training a pre-trained model through the first sample data, the second review category of the first sample data, and the attribution annotation data, so that the pre-trained model fully learns the association relationship between the first sample data and the second review category and the attribution annotation data of the first sample data, so that the review attribution model can fully understand a small number of second review rules corresponding to each second review category, and analyze the attribution data belonging to the second review category during content review, and can quickly and accurately review whether the content data belongs to a small number of second review categories.
[0023] Based on this, the method fuses a large number of fine-grained review categories corresponding to a large number of sample data into a small number of coarse-grained review categories, sorts out a large number of cumbersome review rules into a small number of streamlined review rules for each coarse-grained review category, and marks attribution data indicating that the large number of sample data belong to the coarse-grained review category based on the small number of streamlined review rules corresponding to the small number of coarse-grained review categories. The pre-trained model is continuously trained into a review attribution model through the sample data, the coarse-grained review category of the sample data, and the attribution data belonging to the coarse-grained review category, so that the review attribution model can fully understand the small number of streamlined review rules corresponding to the small number of coarse-grained review categories, and analyze the attribution data belonging to the coarse-grained review category when reviewing content data through the review attribution model, thereby improving the review efficiency and review accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0025] Figure 1 A schematic diagram of the verification accuracy rate of whether content data belongs to the target verification category by a plurality of verification models provided by the related art;
[0026] Figure 2 A system diagram of a training method for a verification attribution model provided by an embodiment of the present application;
[0027] Figure 3 A flowchart of a training method for a verification attribution model provided by an embodiment of the present application;
[0028] Figure 4 A schematic diagram of a verification model and a verification attribution model provided by an embodiment of the present application respectively verifying whether content data belongs to the target verification category;
[0029] Figure 5 A schematic flowchart of constructing attribution annotation data for the first sample data provided by an embodiment of the present application;
[0030] Figure 6 A schematic diagram of the structure of a large amount of sample data required to train a verification attribution model based on a pre-trained model provided by an embodiment of the present application;
[0031] Figure 7 A schematic diagram of the verification accuracy rate of different proportions of the first sample data in a large amount of sample data under different verification tasks provided by an embodiment of the present application;
[0032] Figure 8 A schematic diagram of the verification accuracy rate of a verification attribution model and an optimized verification attribution model under different verification tasks provided by an embodiment of the present application;
[0033] Figure 9 A flowchart of training an optimized verification attribution model based on a pre-trained model provided by an embodiment of the present application;
[0034] Figure 10 A schematic diagram of the verification accuracy rate of a verification model and a verification attribution model under multiple second verification categories provided by an embodiment of the present application;
[0035] Figure 11A schematic diagram of the audit accuracy of an audit model and an audit attribution model provided by an embodiment of the present application under different audit categories based on different data magnitudes of a sample data set;
[0036] Figure 12 A schematic diagram of the coverage rate of the audit traffic of an audit attribution model provided by an embodiment of the present application under different sample data volumes;
[0037] Figure 13 A flowchart of secondary audit when the model audit result of the content data in the audit attribution model provided by an embodiment of the present application is different from the manual audit result of the auditor;
[0038] Figure 14 A structural diagram of a training device for an audit attribution model provided by an embodiment of the present application;
[0039] Figure 15 A structural diagram of a server provided by an embodiment of the present application;
[0040] Figure 16 A structural diagram of a terminal provided by an embodiment of the present application. Detailed implementation manners
[0041] The embodiments of the present application will be described below with reference to the accompanying drawings.
[0042] At present, in the content audit scenario, to determine whether the content data belongs to the target audit category, generally, a large amount of content data of the target audit category and the fine-grained audit category of each content data are obtained as a large amount of sample data and the audit category of each sample data, and the large language model is fine-tuned and trained to obtain an audit model. Through the audit model, it can be determined whether the content data belongs to the target audit category. However, through research, it is found that a large amount of content data of the target audit category involves a large number of fine-grained audit categories, and a large number of fine-grained audit categories correspond to a large number of cumbersome audit rules. The audit model cannot understand a large number of cumbersome audit rules, and relying on the audit model to determine whether the content data belongs to the target audit category results in low audit efficiency and audit accuracy.
[0043] See Figure 1 This Figure 1 is a schematic diagram of the audit accuracy of a plurality of audit models provided by the related art for determining whether the content data belongs to the target audit category. Among them, the plurality of audit models include Audit Model 1, Audit Model 2, Audit Model 3, Audit Model 4, Audit Model 5, Audit Model 6, and Audit Model 7. The audit accuracy of the plurality of audit models for determining whether the content data belongs to the target audit category cannot reach the audit accuracy required by the business, let alone the audit accuracy of the auditor.
[0044] The embodiments of the present application provide a training method for an audit attribution model, which integrates a large number of fine-grained audit categories corresponding to a large number of sample data into a small number of coarse-grained audit categories. For each coarse-grained audit category, a large number of cumbersome audit rules are sorted out into a small number of concise audit rules, and based on the small number of concise audit rules corresponding to the small number of coarse-grained audit categories, the attribution data of a large number of sample data belonging to the coarse-grained audit categories is labeled. Through the sample data, the coarse-grained audit categories of the sample data, and the attribution data belonging to the coarse-grained audit categories, the pre-trained model is continuously trained into an audit attribution model, so that the audit attribution model fully understands the small number of concise audit rules corresponding to the small number of coarse-grained audit categories, and analyzes the attribution data belonging to the coarse-grained audit categories when auditing content data through the audit attribution model, thereby improving the audit efficiency and audit accuracy.
[0045] Next, the system architecture of the training method for the audit attribution model will be introduced. See Figure 2 ,which Figure 2 is a system diagram of a training method for an audit attribution model provided by an embodiment of the present application. The system includes a computer device 200, and the computer device 200 is used to execute the training method for the audit attribution model.
[0046] The computer device 200 performs category fusion on multiple first audit categories corresponding to multiple first sample data to obtain multiple second audit categories; the total number of the multiple second audit categories is less than the total number of the multiple first audit categories.
[0047] As an example, the first sample data is the first harmful data, the multiple first audit categories are m first harmful categories, the multiple second audit categories are n second harmful categories, both m and n are positive integers, and n < m; the computer device 200 fuses the m first harmful categories corresponding to the multiple first harmful data into n second harmful categories.
[0048] The computer device 200 sorts out multiple first audit rules corresponding to each second audit category to obtain multiple second audit rules corresponding to each second audit category; the total number of the multiple second audit rules is less than the total number of the multiple first audit rules.
[0049] As an example, on the basis of the above example, the multiple first audit rules are multiple first harmful rules, and the multiple second audit rules are multiple second harmful rules; the computer device 200 sorts out the multiple first harmful rules corresponding to each second harmful category into multiple second harmful rules, so that the total number of the multiple second harmful rules is less than the total number of the multiple first harmful rules.
[0050] The computer device 200 performs attribution annotation on the first sample data according to the multiple second audit rules corresponding to the second audit category of the first sample data to obtain the attribution annotation data of the first sample data.
[0051] As an example, based on the above example, the computer device 200 labels the attribution annotation data of the first harmful data through multiple second harmful rules corresponding to the second harmful category of the first harmful data.
[0052] The computer device 200 performs model training on the pre-trained model according to the first sample data, the second review category of the first sample data, and the attribution annotation data, and obtains a review attribution model.
[0053] As an example, based on the above example, the pre-trained model is a large language model (LLM); the computer device 200 trains the LLM through the first harmful data, the second harmful category of the first harmful data, and the attribution annotation data to obtain a review attribution model.
[0054] That is to say, the multiple first review categories corresponding to the multiple first sample data are fused into multiple second review categories, so that the total number of the multiple second review categories is less than the total number of the multiple first review categories, in order to reduce the large number of first review categories corresponding to the large number of first sample data to a small number of second review categories; the multiple first review rules corresponding to each second review category are sorted into multiple second review rules, so that the total number of the multiple second review rules is less than the total number of the multiple first review rules, in order to simplify the large number of first review rules corresponding to each second review category to a small number of second review rules; the attribution annotation data indicating that the first sample data belongs to the second review category is labeled through the multiple second review rules corresponding to the second review category of the first sample data, in order to clarify the attribution data indicating that the large number of first sample data belongs to the second review category during content review; the pre-trained model is trained through the first sample data, the second review category of the first sample data, and the attribution annotation data to obtain a review attribution model, so that the pre-trained model fully learns the correlation between the first sample data and the second review category and the attribution annotation data of the first sample data, so that the review attribution model can fully understand the small number of second review rules corresponding to each second review category, and analyze the attribution data belonging to the second review category during content review, and can quickly and accurately review whether the content data belongs to the small number of second review categories.
[0055] It should be noted that in the embodiments of the present application, the computer device may be a server or a terminal. The method provided in the embodiments of the present application may be executed independently by the terminal or the server, or may be executed in cooperation by the terminal and the server. Among them, when the method provided in the embodiments of the present application is executed independently by the terminal or the server, its execution method is the same as Figure 2The corresponding embodiments are similar, mainly by replacing the computer device with a terminal or a server. In addition, when the method provided in the embodiments of the present application is executed in cooperation by a terminal and a server, the steps that need to be reflected on the front-end interface can be executed by the terminal, while some steps that require background calculation and do not need to be reflected on the front-end interface can be executed by the server.
[0056] Among them, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart voice interaction device, a vehicle-mounted terminal, an extended reality device, an aircraft, etc., but is not limited thereto. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services, but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions here. For example, the terminal and the server can be connected through a network, and the network can be a wired or wireless network.
[0057] Next, taking the computer device executing the method provided in the embodiments of the present application as an example, the training method of the audit attribution model provided in the embodiments of the present application will be introduced in detail in combination with the accompanying drawings.
[0058] See Figure 3 , Figure 3 which is a flowchart of a training method of an audit attribution model provided in the embodiments of the present application. The method includes:
[0059] S301: Perform category fusion on multiple first audit categories corresponding to multiple first sample data to obtain multiple second audit categories; the total number of the multiple second audit categories is less than the total number of the multiple first audit categories.
[0060] S302: Organize the multiple first audit rules corresponding to each second audit category to obtain multiple second audit rules corresponding to each second audit category; the total number of the multiple second audit rules is less than the total number of the multiple first audit rules.
[0061] In the related art, when auditing whether content data belongs to a target audit category in a content audit scenario, it is usually to obtain a large amount of content data of the target audit category and the fine-grained audit category of each content data as a large amount of sample data and the audit category of each sample data, and perform fine-tuning training on a large language model to obtain an audit model. Through the audit model, it can be audited whether the content data belongs to the target audit category. However, a large amount of content data of the target audit category involves a large number of fine-grained audit categories, and a large number of fine-grained audit categories correspond to a large number of cumbersome audit rules. The audit model cannot understand a large number of cumbersome audit rules, and relies on the audit model to audit whether the content data belongs to the target audit category, resulting in low audit efficiency and audit accuracy.
[0062] In the embodiments of the present application, in order to solve the problems in the related art and prevent the model from being unable to understand a large number of fine-grained review categories corresponding to a large number of cumbersome review rules, first, it is necessary to reduce a large number of fine-grained review categories corresponding to a large number of sample data to a small number of coarse-grained review categories, and for each coarse-grained review category, simplify a large number of cumbersome review rules into a small number of concise review rules. Therefore, multiple first review categories corresponding to multiple first sample data are fused into multiple second review categories, so that the total number of multiple second review categories is less than the total number of multiple first review categories, and multiple first review rules corresponding to each second review category are sorted into multiple second review rules, so that the total number of multiple second review rules is less than the total number of multiple first review rules.
[0063] Among them, the first sample data is the content data of the target review category. For example, the content data of the harmful category, that is, the first harmful data; the first review category is the fine-grained review category under the target review category. For example, the fine-grained review category under the harmful category, that is, the first harmful category; the second review category is the coarse-grained review category under the target review category. For example, the coarse-grained review category under the harmful category, that is, the second harmful category; then the granularity of the second review category is greater than the granularity of the first review category. The first review rule is the cumbersome review rule corresponding to the coarse-grained review category under the target review category. For example, the cumbersome review rule corresponding to the coarse-grained review category under the harmful category, that is, the first harmful rule; the second review rule is the concise review rule corresponding to the coarse-grained review category under the target review category. For example, the concise review rule corresponding to the coarse-grained review category under the harmful category, that is, the second harmful rule; then the complexity of the second review rule is less than the complexity of the first review rule.
[0064] In practical applications, performing category fusion on multiple first review categories corresponding to multiple first sample data to obtain multiple second review categories can be: clustering multiple first review categories corresponding to multiple first sample data to obtain multiple review category clusters; generating multiple second review categories according to the multiple review category clusters, and the granularity of the second review category is greater than the granularity of the first review category.
[0065] Sorting multiple first review rules corresponding to each second review category to obtain multiple second review rules corresponding to each second review category can be: clustering multiple first review rules corresponding to each second review category to obtain multiple review rule clusters, and generating multiple second review rules according to the multiple review rule clusters, and the complexity of the second review rule is less than the complexity of the first review rule.
[0066] In the above S301 - S302, multiple first review categories corresponding to multiple first sample data are first fused into multiple second review categories, such that the total number of the multiple second review categories is less than the total number of the multiple first review categories, so as to reduce a large number of first review categories corresponding to a large amount of first sample data to a small number of second review categories; then, multiple first review rules corresponding to each second review category are sorted into multiple second review rules, such that the total number of the multiple second review rules is less than the total number of the multiple first review rules, so as to simplify a large number of first review rules corresponding to each second review category to a small number of second review rules; this way facilitates the subsequent model to fully understand the small number of second review rules corresponding to each second review category, so as to quickly and accurately review whether the content data belongs to the small number of second review categories.
[0067] As an example of S301 - S302, the first sample data is the first harmful data, the multiple first review categories are m first harmful categories, the multiple second review categories are n second harmful categories, both m and n are positive integers, n < m, the multiple first review rules are multiple first harmful rules, and the multiple second review rules are multiple second harmful rules; the computer device fuses the m first harmful categories corresponding to the multiple first harmful data into n second harmful categories, and the computer device sorts the multiple first harmful rules corresponding to each second harmful category into multiple second harmful rules, such that the total number of the multiple second harmful rules is less than the total number of the multiple first harmful rules.
[0068] In practical applications, the m first harmful categories include First Harmful Category 1, First Harmful Category 2,..., First Harmful Category m, for example, suspected black production, online fraud, etc.; the n second harmful categories include Second Harmful Category 1, Second Harmful Category 2,..., Second Harmful Category n, for example, black production, fraud, etc. Among them, m can be 100 +, and n can be 8.
[0069] S303: Perform attribution annotation on the first sample data according to the multiple second review rules corresponding to the second review category of the first sample data to obtain the attribution annotation data of the first sample data.
[0070] In the embodiment of the present application, after performing the above S301 - S302 to reduce a large number of first review categories corresponding to a large amount of first sample data to a small number of second review categories and simplify a large number of first review rules corresponding to each second review category to a small number of second review rules, in order to solve the problems in the related art and avoid the model being unable to understand a large number of fine - grained review categories corresponding to a large number of cumbersome review rules, it is also necessary to further label the attribution data of a large amount of sample data belonging to the coarse - grained review categories based on the small number of streamlined review rules corresponding to the small number of coarse - grained review categories. Therefore, the attribution annotation data indicating that the first sample data belongs to the second review category is labeled through the multiple second review rules corresponding to the second review category of the first sample data.
[0071] Among them, the attribution annotation data of the first sample data is the attribution data determined based on multiple second review rules corresponding to the second review category of the first sample data, indicating that the first sample data belongs to the second review category.
[0072] In practical applications, when performing attribution annotation on the first sample data according to multiple second review rules corresponding to the second review category of the first sample data to obtain the attribution annotation data of the first sample data, it can be: generating, according to multiple second review rules corresponding to the second review category of the first sample data, the attribution data indicating that the first sample data belongs to the second review category as the attribution annotation data of the first sample data.
[0073] In the above S303, the attribution annotation data indicating that the first sample data belongs to the second review category is marked according to multiple second review rules corresponding to the second review category of the first sample data, so as to clarify the attribution data of a large number of first sample data belonging to the second review category during content review; this method facilitates the subsequent model to fully learn the correlation between the first sample data, the second review category of the first sample data, and the attribution annotation data, and analyze the attribution data belonging to the second review category during content review, so as to quickly and accurately review whether the content data belongs to a small number of second review categories.
[0074] As an example of S303, based on the example of the above 302, the computer device marks the attribution annotation data of the first harmful data according to multiple second harmful rules corresponding to the second harmful category of the first harmful data.
[0075] S304: Train the pre-trained model according to the first sample data, the second review category of the first sample data, and the attribution annotation data to obtain an audit attribution model.
[0076] In the embodiment of the present application, after performing the above S303 to mark the attribution data of a large number of sample data belonging to a small number of coarse-grained review categories based on a small number of refined review rules corresponding to a small number of coarse-grained review categories, in order to solve the problems in the related art and avoid the model's inability to understand a large number of detailed review categories corresponding to a large number of cumbersome review rules, the pre-trained model is continuously trained as an audit attribution model through the sample data, the coarse-grained review category of the sample data, and the attribution data belonging to the coarse-grained review category, so that the audit attribution model fully understands the small number of refined review rules corresponding to the small number of coarse-grained review categories, and analyzes the attribution data belonging to the coarse-grained review category when reviewing content data through the audit attribution model. Therefore, the pre-trained model is trained with the first sample data, the second review category of the first sample data, and the attribution annotation data to obtain an audit attribution model.
[0077] Among them, the audit attribution model is a model that audits whether the content data belongs to the second audit category and determines the attribution data belonging to the second audit category.
[0078] In practical applications, the pre-trained model is trained according to the first sample data, the second audit category of the first sample data, and the attribution annotation data to obtain the audit attribution model, which can be: inputting the first sample data into the pre-trained model for content auditing, and outputting the predicted audit category and predicted attribution data of the first sample data; training the pre-trained model according to the difference between the predicted audit category of the first sample data and the second audit category, and the difference between the predicted attribution data and the attribution annotation data to obtain the audit attribution model.
[0079] In the above S304, the pre-trained model is trained with the first sample data, the second audit category of the first sample data, and the attribution annotation data to obtain the audit attribution model, so that the pre-trained model fully learns the correlation between the first sample data and the second audit category and attribution annotation data of the first sample data, so that the audit attribution model can fully understand a small number of second audit rules corresponding to each second audit category, and analyze the attribution data belonging to the second audit category during content auditing, and can quickly and accurately audit whether the content data belongs to a small number of second audit categories.
[0080] As an example of S304, on the basis of the example in the above 303, the pre-trained model is an LLM; the computer device trains the LLM with the first harmful data, the second harmful category of the first harmful data, and the attribution annotation data to obtain the audit attribution model.
[0081] See Figure 4 , this Figure 4 is a schematic diagram of an audit model and an audit attribution model provided by an embodiment of the present application for respectively auditing whether content data belongs to a target audit category. Among them, the model input of the audit model is content data, and the model output of the audit model is whether the content data belongs to the target audit category; the model input of the audit attribution model is content data, and the model output of the audit attribution model is whether the content data belongs to the target audit category and the attribution data of the content data belonging to the target audit category. This attribution data enhances the model perception of the auditor for the audit attribution model, serves as a reference basis for the auditor to implement content auditing, further reduces the audit burden of the auditor and improves the audit efficiency of the auditor, significantly reduces the training cost of subsequent downstream tasks of the audit attribution model, and effectively shortens the development cycle of subsequent audit services of the audit attribution model.
[0082] As can be seen from the above technical solution, a large number of fine-grained review categories corresponding to a large amount of sample data are fused into a small number of coarse-grained review categories. For each coarse-grained review category, a large number of cumbersome review rules are sorted out into a small number of streamlined review rules, and based on the small number of streamlined review rules corresponding to the small number of coarse-grained review categories, attribution data indicating that a large amount of sample data belongs to the coarse-grained review category is marked. Through the sample data, the coarse-grained review category of the sample data, and the attribution data belonging to the coarse-grained review category, the pre-trained model is continuously trained into a review attribution model, so that the review attribution model fully understands the small number of streamlined review rules corresponding to the small number of coarse-grained review categories. When the review attribution model reviews content data, it analyzes the attribution data belonging to the coarse-grained review category, thereby improving the review efficiency and review accuracy.
[0083] In the embodiment of the present application, when specifically implementing the fusion of multiple first review categories corresponding to multiple first sample data into multiple second review categories in S301 above, considering the superiority of the large language model in learning ability, understanding ability, and generation ability, multiple first review categories corresponding to multiple first sample data can be input into the first language model based on the large language model for category fusion, and multiple second review categories are output; further considering that the first language model needs to fuse and reduce multiple first review categories corresponding to multiple first sample data into multiple main review categories as multiple second review categories under the category fusion prompt; therefore, multiple first review categories corresponding to multiple first sample data are input into the first language model, and the multiple first review categories corresponding to multiple first sample data are fused based on the category fusion prompt and reduced to multiple main review categories to output multiple second review categories. Based on this, the present application provides a possible implementation manner. For example, S301 above may include S301a (not shown in the figure): The first language model fuses multiple first review categories corresponding to multiple first sample data according to the category fusion prompt to obtain multiple second review categories; the category fusion prompt is used to guide the first language model to fuse and reduce multiple first review categories into multiple main review categories.
[0084] Among them, the first language model is a model that fuses and reduces multiple coarse-grained review categories into multiple fine-grained review categories; the category fusion prompt is category fusion text provided to the first language model to guide the first language model to fuse and reduce multiple coarse-grained review categories into multiple fine-grained review categories.
[0085] The above S301a uses category fusion prompts to guide the first language model to fuse and reduce multiple first review categories into multiple major review categories as multiple second review categories. Given the superior learning, understanding, and generation capabilities of the first language model based on the large language model, combined with category fusion prompts, it is possible to more quickly and accurately reduce a large number of first review categories corresponding to a large amount of first sample data to a small number of second review categories.
[0086] As an example of S301a, based on the example of S301 - S302 above, the first language model is Language Model 1, and the category fusion prompt is Prompt1; the computer device inputs m first harmful categories corresponding to multiple first harmful data into Language Model 1, and performs category fusion on the m first harmful categories corresponding to multiple first harmful data based on Prompt1, and outputs n second harmful categories.
[0087] In the embodiments of the present application, when specifically implementing the above S302 to organize multiple first review rules corresponding to each second review category into multiple second review rules, considering the superiority of the large language model in terms of learning, understanding, and generation capabilities, multiple first review rules corresponding to each second review category can be input into the second language model based on the large language model for rule organization, and multiple second review rules are output; further considering that for multiple first review rules corresponding to each second review category, the first language model needs to organize and simplify multiple first review rules into multiple major review rules under the rule organization prompt as multiple second review rules corresponding to each second review category; therefore, multiple first review rules corresponding to each second review category are input into the first language model, and based on the rule organization prompt, multiple first review rules corresponding to each second review category are organized and simplified into multiple major review rules to output multiple second review rules corresponding to each second review category. Based on this, the present application provides a possible implementation method. The above S302 may include, for example, S302a (not shown in the figure): The second language model organizes multiple first review rules corresponding to each second review rule according to the rule organization prompt to obtain multiple second review rules corresponding to each second review rule; the rule organization prompt is used to guide the second language model to organize and simplify multiple first review rules corresponding to each second review rule into multiple major review rules.
[0088] Among them, the second language model is a model that organizes and simplifies multiple cumbersome review rules into multiple concise review rules; the rule organization prompt is a rule organization text provided to the second language model to guide the second language model to organize and simplify multiple cumbersome review rules into multiple concise review rules.
[0089] In the above S302a, through rule sorting prompts, the second language model is guided to sort and simplify the multiple first review rules corresponding to each second review rule into multiple main review rules as multiple second review rules. Given the superior learning, understanding, and generation capabilities of the second language model based on the large language model, combined with rule sorting prompts, it is possible to more quickly and accurately simplify the large number of first review rules corresponding to each second review category into a small number of second review rules.
[0090] As an example of S302a, based on the examples of the above S301 - S302, the second language model is Language Model 2, and the rule sorting prompt is Prompt2; the computer device inputs the multiple first harmful rules corresponding to each second harmful category into Language Model 2, and performs rule sorting on the multiple first harmful rules corresponding to each second harmful category based on Prompt2, and outputs the multiple second harmful rules corresponding to each second harmful category.
[0091] In the embodiments of the present application, when specifically implementing the attribution annotation data indicating that the first sample data belongs to the second review category through the multiple second review rules corresponding to the second review category of the first sample data in the above S303, considering the superiority of the large language model in learning, understanding, and generation capabilities, the first sample data can be input into the third language model based on the large language model, and attribution annotation is performed based on the multiple second review rules corresponding to the second review category of the first sample data, and the attribution annotation data indicating that the first sample data belongs to the second review category is output; further considering that the third language model needs to label the attribution data of the first sample data from the multiple second review rules corresponding to the second review category of the first sample data under the attribution annotation prompt as the attribution annotation data indicating that the first sample data belongs to the second review category; therefore, the first sample data is input into the third language model, and based on the multiple second review rules corresponding to the second review category of the first sample data and the attribution annotation prompt, attribution annotation is performed on the first sample data, and the attribution annotation data indicating that the first sample data belongs to the second review category is output. Based on this, the present application provides a possible implementation manner. The above S303 may, for example, include S303a (not shown in the figure): the third language model performs attribution annotation on the first sample data according to the multiple second review rules corresponding to the second review category of the first sample data and the attribution annotation prompt, and obtains the attribution annotation data of the first sample data; the attribution annotation prompt is used to guide the third language model to label the attribution data of the first sample data from the multiple second review rules corresponding to the second review category of the first sample data.
[0092] Among them, the third language model is a model that annotates the attribution data of the first sample data from multiple refined review rules corresponding to the coarse-grained review categories of the first sample data; the attribution annotation prompt is the attribution annotation text provided to the third language model to guide the third language model to annotate the attribution data of the first sample data from multiple refined review rules corresponding to the coarse-grained review categories of the first sample data.
[0093] The above S303a guides the third language model to annotate the attribution data of the first sample data from multiple second review rules corresponding to the second review category of the first sample data through the attribution annotation prompt, and uses it as the attribution annotation data indicating that the first sample data belongs to the second review category. When the third language model based on the large language model has superior learning ability, understanding ability and generation ability, combined with the attribution annotation prompt, it can more quickly and accurately identify the attribution data indicating that a large number of first sample data belong to the second review category during content review.
[0094] As an example of S303a, based on the example of the above S303, the third language model is the language model 3, and the attribution annotation prompt is Prompt3; the computer device inputs the first harmful data into the language model 3, and based on multiple second harmful rules corresponding to the second harmful category of the first harmful data and Prompt3, performs attribution annotation on the multiple first harmful data, and outputs the attribution annotation data indicating that the first harmful data belongs to the second harmful category.
[0095] See Figure 5 this Figure 5A schematic flowchart of constructing attribution annotation data for the first sample data provided by an embodiment of the present application. In the data cleaning stage, multiple original sample data are repeatedly filtered and low-quality filtered into multiple first sample data; multiple first review categories corresponding to the multiple first sample data are combined with category fusion prompts through a first language model to perform category fusion and category balancing into multiple second review categories, obtaining a sample data pool, which includes the first sample data, the first review categories of the first sample data, and the second review categories; in the rule sorting stage, multiple first review rules corresponding to each second review category are combined with rule sorting prompts through a second language model to perform rule merging and rule optimization into multiple second review rules corresponding to each second review category; in the attribution annotation stage, through a third language model, the first sample data and attribution annotation prompts corresponding to the second review categories of the first sample data are used to annotate the attribution annotation data indicating that the first sample data belongs to the second review category, obtaining an attribution data pool, which includes the first sample data, the second review categories of the first sample data, and the attribution annotation data. In addition, the sample data pool can also derive a large category data pool, which includes the first sample data and the large category review categories to which the first review categories of the first sample data belong, and the large category review categories are target review categories or non-target review categories.
[0096] In addition, in the embodiment of the present application, when specifically implementing the above S303a of inputting the first sample data into the third language model, performing attribution annotation on the first sample data based on multiple second review rules corresponding to the second review categories of the first sample data and attribution annotation prompts, and outputting attribution annotation data indicating that the first sample data belongs to the second review category, considering that adding sample data and attribution annotation data that meet the quality requirements of the sample data in the attribution annotation prompts can improve the quality of the attribution annotation data output by the third language model indicating that the first sample data belongs to the second review category, it is necessary to first implement attribution annotation for some of the first sample data to obtain attribution annotation data, select the attribution annotation data that meets the quality requirements of the first sample data and the first sample data, and update the attribution annotation prompts, so as to implement attribution annotation for each first sample data to obtain attribution annotation data that meets the quality requirements.
[0097] Therefore, select some first sample data from multiple first sample data to form a sample data subset. First, for each first sample data in the sample data subset, input the first sample data into the third language model. Based on multiple second review rules and attribution annotation prompts corresponding to the second review category of the first sample data, perform attribution annotation on the first sample data, and output attribution annotation data indicating that the first sample data belongs to the second review category. From the multiple first sample data in the sample data subset and the multiple attribution annotation data corresponding to the multiple first sample data, select the attribution annotation data of the first sample data that meets the quality requirements, and combine it with the first sample data as the second sample data and the attribution annotation data of the second sample data to update the attribution annotation prompts to obtain updated attribution annotation prompts. Then, for each first sample data in the multiple first sample data, input the first sample data into the third language model. Based on multiple second review rules and the updated attribution annotation prompts corresponding to the second review category of the first sample data, perform attribution annotation on the first sample data, and output attribution annotation data indicating that the first sample data belongs to the second review category. Based on this, the present application provides a possible implementation method. For example, the method may further include S1 (not shown in the figure): perform data selection on multiple first sample data to obtain a sample data subset. Correspondingly, the above S303a may include the following S303aa - S303ad (not shown in the figure).
[0098] S303aa: For each first sample data in the sample data subset, use the third language model to perform attribution annotation on the first sample data according to multiple second review rules and attribution annotation prompts corresponding to the second review category of the first sample data, to obtain the attribution annotation data of the first sample data.
[0099] S303ab: Perform data selection on the multiple first sample data in the sample data subset and the multiple attribution annotation data corresponding to the multiple first sample data, to obtain the second sample data and the attribution annotation data of the second sample data; the attribution annotation data of the second sample data meets the quality requirements.
[0100] S303ac: Update the attribution annotation prompts according to the second sample data and the attribution annotation data of the second sample data, to obtain updated attribution annotation prompts.
[0101] S303ad: For each first sample data in the multiple first sample data, use the third language model to perform attribution annotation on the first sample data according to multiple second review rules and the updated attribution annotation prompts corresponding to the second review category of the first sample data, to obtain the attribution annotation data of the first sample data.
[0102] Among them, the total number of multiple first sample data is greater than the total number of multiple first sample data in the sample data subset; the second sample data is the first sample data for which the attribution annotation data meets the quality requirements; the updated attribution annotation prompt is the attribution annotation prompt that adds the second sample data and the attribution annotation data of the second sample data.
[0103] In the above S1 and S303aa - S303ad, a part of the first sample data is selected from the multiple first sample data to form a sample data subset. First, for each first sample data in the sample data subset, the third language model is guided by the attribution annotation prompt to annotate the attribution data of the first sample data from multiple second review rules corresponding to the second review category of the first sample data, as the attribution annotation data for the first sample data belonging to the second review category, so as to more quickly and accurately determine the attribution data of a large number of first sample data belonging to the second review category during content review; from the multiple first sample data in the sample data subset and the multiple attribution annotation data corresponding to the multiple first sample data, the second sample data and the attribution annotation data of the second sample data that meet the quality requirements are selected to achieve the selection of the second sample data and the high-quality attribution annotation data of the second sample data; through the second sample data and the attribution annotation data of the second sample data, the attribution annotation prompt is updated to obtain the updated attribution annotation prompt, so that the updated attribution annotation prompt can guide the third language model to annotate the attribution data that meets the quality requirements; then, for each first sample data in the multiple first sample data, the third language model is guided by the attribution annotation prompt to annotate the attribution data of the first sample data from multiple second review rules corresponding to the second review category of the first sample data, as the attribution annotation data for the first sample data belonging to the second review category, so as to more quickly and accurately determine the attribution data of a large number of first sample data belonging to the second review category and meeting the quality requirements during content review.
[0104] As an example of S1 and S303aa - S303ad, based on the example of S301a above, the sample data subset is a harmful data subset, and the second sample data is the second harmful data; the computer device selects some of the first harmful data from multiple first harmful data to form a harmful data subset. First, for each first harmful data in the harmful data subset, the first harmful data is input into the language model 3. Based on multiple second harmful rules and Prompt3 corresponding to the second harmful category of the first harmful data, attribution annotation is performed on the multiple first harmful data, and attribution annotation data indicating that the first harmful data belongs to the second harmful category is output; the computer device selects the attribution annotation data that meets the quality requirements of the first harmful data from the multiple first harmful data in the harmful data subset and the multiple attribution annotation data corresponding to the multiple first harmful data, combines it with the first harmful data, and uses it as the second harmful data and the attribution annotation data of the second harmful data to update Prompt3 to obtain the updated Prompt3; then, for each first harmful data in the multiple first harmful data, the first harmful data is input into the language model 3. Based on multiple second harmful rules and the updated Prompt3 corresponding to the second harmful category of the first harmful data, attribution annotation is performed on the first harmful data, and attribution annotation data indicating that the first harmful data belongs to the second harmful category is output.
[0105] In addition, in the embodiments of the present application, in order to further determine a finer - grained review category of the first sample data and use the first sample data and the finer - grained review category of the first sample data for model training, considering that the multiple second review rules corresponding to the second review category are actually multiple refined review rules corresponding to multiple first review rules under the second review category, that is, the multiple second review rules corresponding to the second review category include more fine - grained key information in the multiple first review rules corresponding to the second review category, the third review category of the first sample data can also be labeled based on the multiple second review rules corresponding to the second review category of the first sample data, so that the granularity of the third review category is smaller than the granularity of the first review category. Based on this, the present application provides a possible implementation manner. The method may further include S2 (not shown in the figure): performing category annotation on the first sample data according to the multiple second review rules corresponding to the second review category of the first sample data to obtain the third review category of the first sample data; the granularity of the third review category is smaller than the granularity of the first review category, and the first sample data and the third review category of the first sample data are used for model training.
[0106] Among them, the third review category of the first sample data is to determine the review category of the first sample data based on the multiple second review rules corresponding to the second review category of the first sample data.
[0107] In practical applications, the first sample data is classified according to multiple second review rules corresponding to the second review category of the first sample data to obtain the third review category of the first sample data. It can be: the fourth language model classifies the first sample data according to multiple second review rules corresponding to the second review category of the first sample data and the category annotation prompt to obtain the third review category of the first sample data; the category annotation prompt is used to guide the fourth language model to label the review category of the first sample data from multiple second review rules corresponding to the second review category of the first sample data.
[0108] In the above S2, the third review category of the first sample data is labeled according to multiple second review rules corresponding to the second review category of the first sample data, so that the granularity of the third review category is smaller than that of the first review category, in order to clarify the finer-grained review category of a large number of first sample data during content review; this method facilitates the subsequent model to fully learn the correlation between the first sample data and the third review category of the first sample data, and identify finer-grained review categories during content review, so as to quickly and accurately review whether the content data belongs to the finer-grained review category.
[0109] As an example of S2, based on the example of S301 - S302 above, the computer device labels the third harmful category of the first harmful data according to multiple second harmful rules corresponding to the second harmful category of the first harmful data, so that the granularity of the third harmful category is smaller than that of the first harmful category.
[0110] In summary, the embodiments of the present application can be divided into five different types of sample data sets, specifically including: single review category data set, full review category data set, discrimination task data set, category rule change data set, and review knowledge base data set.
[0111] Single review category data set: The review task corresponding to this single review category data set is to give a single second review category, and let the model determine whether the content data belongs to this single second review category, and analyze the attribution data belonging to the single second review category. This single review category data set includes the first sample data belonging to this single second review category and the first sample data belonging to other second review categories.
[0112] Full review category data set: The review task corresponding to this full review category data set is to give all second review categories, and let the model determine which second review category the content data belongs to among all second review categories, and analyze the attribution data belonging to the second review category.
[0113] Discrimination task data set: The review task corresponding to this discrimination task data set is to give the target review category corresponding to the second review category, and let the model determine whether the content data belongs to the target review category.
[0114] Category rule change dataset: The second review category and the second review rules are not fixed. This category rule change dataset includes the updated second review category and the updated second review rules, enhancing the generalization ability, flexibility, and adaptability of the model.
[0115] Review knowledge base dataset: This review knowledge base dataset includes questions and answers about the second review category and the second review rules, helping the model better understand the second review category and the second review rules, and at the same time providing richer sample data for the model.
[0116] See Figure 6 , the Figure 6 This is a schematic structural diagram of a large number of sample data required for training a review attribution model obtained based on a pre-trained model provided by an embodiment of this application. Among them, the large number of sample data of the pre-trained model includes multiple first sample data, review knowledge base data, plain text data, mathematical / chart data, story / image description data, medical / character recognition / 3D / user interface data.
[0117] See Figure 7 , the Figure 7 This is a schematic diagram of the review accuracy rates of different proportions of the first sample data in a large number of sample data under different review tasks provided by an embodiment of this application. Among them, the different review tasks are review task 1, review task 2, review task 3, review task 4, and review task 5, and the different proportions of the first sample data in the large number of sample data are 30%, 50%, 75%, and 100%. By using 30%, 50%, 75%, and 100% of the first sample data in the large number of sample data, the review accuracy rates under review task 1, review task 2, review task 3, review task 4, and review task 5 are determined, and the best proportion of the first sample data in the large number of sample data is between 20% and 40%.
[0118] In the embodiment of this application, when specifically implementing the training of the pre-trained model to obtain the review attribution model through the first sample data and the second review category and attribution annotation data of the first sample data in S304 above, considering that the training method from easy to difficult makes the model training more efficient and stable; therefore, first, for the first sample data of a single image among multiple first sample data, the pre-trained model is trained through the first sample data and the second review category and attribution annotation data of the first sample data to obtain the first attribution model, and then for the first sample data of multiple images among multiple first sample data, the first attribution model is trained through the first sample data and the second review category and attribution annotation data of the first sample data to obtain the review attribution model. Based on this, this application provides a possible implementation manner. S304 above may, for example, include the following S304a-S304b (not shown in the figure).
[0119] S304a: For the first sample data of a single image among multiple first sample data, model training is performed on the pre-trained model according to the first sample data, the second review category of the first sample data, and the attribution annotation data to obtain a first attribution model.
[0120] S304b: For the first sample data of multiple images among multiple first sample data, model training is performed on the first attribution model according to the first sample data, the second review category of the first sample data, and the attribution annotation data to obtain a review attribution model.
[0121] Among them, the first attribution model has higher review efficiency and review accuracy for whether the content data of a single image belongs to the second review category and for determining the attribution data that belongs to the second review category.
[0122] In the above S304a - S304b, first, through the first sample data of a single image among multiple first sample data, the second review category of the first sample data, and the attribution annotation data, the pre-trained model is trained to obtain a first attribution model, enabling the pre-trained model to fully learn the correlation between the first sample data of a single image and the second review category and attribution annotation data of the first sample data, so that the first attribution model can fully understand a small number of second review rules corresponding to the second review category of the first sample data of a single image. When analyzing the attribution data belonging to the second review category during content review, it can quickly and accurately review whether the content data of a single image belongs to a small number of second review categories. Then, through the first sample data of multiple images among multiple first sample data, the second review category of the first sample data, and the attribution annotation data, the first attribution model is trained to obtain a review attribution model, enabling the first attribution model to fully learn the correlation between the first sample data of multiple images and the second review category and attribution annotation data of the first sample data, so that the review attribution model can further fully understand a small number of second review rules corresponding to the second review category of the first sample data of multiple images. When analyzing the attribution data belonging to the second review category during content review, it can further quickly and accurately review whether the content data of multiple images belongs to a small number of second review categories. This method expands the model training from a simple single-image content review scenario to a more difficult multiple-image content review scenario, making the model training more efficient and stable.
[0123] As an example of S304a - S304b, based on the above example of S304, the first attribution model is the attribution model 1; the computer device first trains the LLM to obtain the attribution model 1 for the first harmful data of a single image among multiple first harmful data through the first harmful data, the second harmful category of the first harmful data, and the attribution annotation data, and then the computer device trains the attribution model 1 to obtain the audit attribution model for the first harmful data of multiple images among multiple first harmful data through the first harmful data, the second harmful category of the first harmful data, and the attribution annotation data.
[0124] In the embodiments of the present application, when specifically implementing the above S304 to train the first attribution model to obtain the audit attribution model for the first sample data of multiple images among multiple first sample data through the first sample data, the second audit category of the first sample data, and the attribution annotation data, further considering that adding a training method for pure text makes the model training more comprehensive; therefore, first train the first attribution model to obtain the second attribution model for the first sample data of multiple images among multiple first sample data through the first sample data, the second audit category of the first sample data, and the attribution annotation data; then train the second attribution model to obtain the audit attribution model for the first sample data of pure text among multiple first sample data through the first sample data, the second audit category of the first sample data, and the attribution annotation data. Based on this, the present application provides a possible implementation manner, and the above S304b may include, for example, the following S304bb - S304bc (not shown in the figure).
[0125] S304bb: For the first sample data of multiple images among multiple first sample data, perform model training on the first attribution model according to the first sample data, the second audit category of the first sample data, and the attribution annotation data to obtain the second attribution model.
[0126] S304bc: For the first sample data of pure text among multiple first sample data, perform model training on the second attribution model according to the first sample data, the second audit category of the first sample data, and the attribution annotation data to obtain the audit attribution model.
[0127] Among them, the second attribution model has higher audit efficiency and audit accuracy for whether the content data of a single image and multiple images belongs to the second audit category and for determining the attribution data belonging to the second audit category.
[0128] In the above S304bb - S304bc, first, the first attribution model is trained with the first sample data of multiple images among multiple first sample data, the second review category of the first sample data, and the attribution annotation data to obtain the second attribution model, enabling the first attribution model to fully learn the correlation between the first sample data of multiple images and the second review category and attribution annotation data of the first sample data. In this way, the second attribution model can further fully understand a small number of second review rules corresponding to the second review category of the first sample data of multiple images. When analyzing attribution data belonging to the second review category during content review, it can further quickly and accurately review whether the content data of multiple images belongs to the small number of second review categories. Then, the second attribution model is trained with the first sample data of plain text among multiple first sample data, the second review category of the first sample data, and the attribution annotation data to obtain the review attribution model, enabling the second attribution model to fully learn the correlation between the first sample data of plain text and the second review category and attribution annotation data of the first sample data. In this way, the review attribution model can further fully understand a small number of second review rules corresponding to the second review category of the first sample data of plain text. When analyzing attribution data belonging to the second review category during content review, it can further quickly and accurately review whether the content data of plain text belongs to the small number of second review categories. This method further expands model training to include content review scenarios for plain text, avoiding the review attribution model forgetting the review attribution ability for the content data of plain text, thus making model training more comprehensive.
[0129] As an example of S304bb - S304bc, based on the example of the above S304a - S304b, the second attribution model is the attribution model 2; the computer device first trains the attribution model 1 with the first harmful data of multiple images among multiple first harmful data, the second harmful category of the first harmful data, and the attribution annotation data to obtain the attribution model 2. Then, the computer device trains the attribution model 2 with the first harmful data of plain text among multiple first harmful data, the second harmful category of the first harmful data, and the attribution annotation data to obtain the review attribution model.
[0130] In addition, in the embodiments of the present application, after training the pre-trained model with the first sample data, the second review category of the first sample data, and the attribution annotation data to obtain the review attribution model in S304 above, considering optimizing the review attribution model with sample data that meets the quality requirements and the attribution annotation data of the sample data to ensure the stability, generality, and generalization of the review attribution model, it is necessary to select the first sample data that meets the quality requirements and the attribution annotation data of the first sample data to optimize the review attribution model. Therefore, from multiple first sample data and multiple attribution annotation data corresponding to the multiple first sample data, the first sample data that meets the quality requirements and the attribution annotation data of the first sample data are selected as the third sample data and the attribution annotation data of the third sample data; the review attribution model is optimized with the third sample data, the second review category of the third sample data, and the attribution annotation data to obtain the optimized review attribution model. Based on this, the present application provides a possible implementation manner, and the method may further include the following S3-S4 (not shown in the figure).
[0131] S3: Perform data selection on multiple first sample data and multiple attribution annotation data corresponding to the multiple first sample data to obtain the third sample data and the attribution annotation data of the third sample data; the third sample data and the attribution annotation data of the third sample data meet the quality requirements.
[0132] S4: Optimize the review attribution model according to the third sample data and the attribution annotation data of the third sample data to obtain the optimized review attribution model.
[0133] Among them, the third sample data is the first sample data for which both the sample data and the attribution annotation data meet the quality requirements; the quality of the attribution data of the review content data belonging to the second review category of the optimized review attribution model is higher than the quality of the attribution data of the review content data belonging to the second review category of the review attribution model.
[0134] In S3 - S4 above, from multiple first sample data and the multiple attribution annotation data corresponding to the multiple first sample data, the third sample data that meets the quality requirements and the attribution annotation data of the third sample data are selected, so as to realize the selection of high - quality third sample data and the attribution annotation data of the third sample data; through the third sample data, the second review category of the third sample data, and the attribution annotation data, the review attribution model is trained to obtain an optimized review attribution model, so that the review attribution model can further fully learn the correlation between high - quality third sample data and the second review category and attribution annotation data of the third sample data, so that when the optimized review attribution model analyzes high - quality attribution data belonging to the second review category during content review, it can quickly and accurately review whether the content data belongs to a small number of second review categories and determine the high - quality attribution data of the content data belonging to the second review category. This method realizes the optimized training of the review attribution model through high - quality sample data and the attribution annotation data of the sample data, so as to ensure the stability, generality and generalization of the review attribution model.
[0135] As an example of S3 - S4, on the basis of the example of S304 above, the third sample data is the third harmful data; the computer device selects the first harmful data that meets the quality requirements and the attribution annotation data of the first harmful data from multiple first harmful data and the multiple attribution annotation data corresponding to the multiple first harmful data as the third harmful data and the attribution annotation data of the third harmful data; the computer device optimizes the review attribution model through the third harmful data, the second harmful category of the third harmful data, and the attribution annotation data to obtain an optimized review attribution model.
[0136] See Figure 8 , the Figure 8 is a schematic diagram of the review accuracy of a review attribution model and an optimized review attribution model provided by an embodiment of this application under different review tasks. Among them, the different review tasks are review task 1, review task 2, review task 3, and review task 4. By the review accuracy of the review attribution model and the optimized review attribution model under review task 1, review task 2, review task 3, and review task 4, it is determined that the review accuracy of the optimized review attribution model is higher than that of the review attribution model.
[0137] In addition, in the embodiments of the present application, in order to make the attribution data of the review attribution model for the review content data belonging to the second review category conform to the review preferences of the reviewers, the attribution preference data of the first sample data can be further obtained; through the first sample data, the second review category of the first sample data, the attribution annotation data, and the attribution preference data, the review attribution model is optimized to obtain an optimized review attribution model, so that the attribution data determined by the optimized review attribution model for the content data belonging to the second review category aligns with the review preferences of the reviewers. Based on this, the present application provides a possible implementation manner. For example, the method may further include the following S5-S6 (not shown in the figure).
[0138] S5: Obtain the attribution preference data of the first sample data.
[0139] S6: Optimize the review attribution model according to the first sample data, the second review category of the first sample data, the attribution annotation data, and the attribution preference data to obtain an optimized review attribution model.
[0140] Among them, the attribution preference data of the first sample data is the attribution data determined by the reviewer for the first sample data belonging to the second review category.
[0141] In practical applications, the first sample data, the second review category of the first sample data, the attribution annotation data, and the attribution preference data form the sample data set of the review attribution model Specifically as follows:
[0142]
[0143] Among them, x represents the first sample data, y w represents the second review category and attribution preference data of the first sample data, y l represents the second review category and attribution annotation data of the first sample data.
[0144] Based on the loss function of optimizing the review attribution model into an optimized review attribution model Specifically as follows:
[0145]
[0146] Among them, π θ and π ref both represent the review attribution model, σ represents the logistic function, and β represents a controllable hyperparameter, initialized to 0.1.
[0147] Based on the attribution annotation data of the first sample data, S5 - S6 further obtains the attribution preference data of the first sample data to clarify the review preferences of the reviewers; through the first sample data, the second review category of the first sample data, the attribution annotation data, and the attribution preference data, the review attribution model is optimized to obtain an optimized review attribution model, enabling the review attribution model to further fully learn the attribution preference data of the first sample data, so that the optimized review attribution model determines that the attribution data of the content data belonging to the second review category aligns with the review preferences of the reviewers.
[0148] As an example of S5 - S6, based on the example of S304 above, the computer device obtains the attribution preference data of the first harmful data; the computer device optimizes the review attribution model through the first harmful data, the second harmful category of the first harmful data, the attribution annotation data, and the attribution preference data to obtain an optimized review attribution model.
[0149] See Figure 9 , which Figure 9 is a flowchart of an optimized review attribution model trained based on a pre - trained model provided by an embodiment of the present application. The pre - trained model is subjected to single - image training, multi - image training, annealing training, and alignment training to obtain an optimized review attribution model. Among them, single - image training means: for the first sample data of a single image among multiple first sample data, through the first sample data, the second review category of the first sample data, and the attribution annotation data, the pre - trained model is trained to obtain a first attribution model. Multi - image training means: for the first sample data of multiple images among multiple first sample data, through the first sample data, the second review category of the first sample data, and the attribution annotation data, the first attribution model is trained to obtain a second attribution model; for the first sample data of pure text among multiple first sample data, through the first sample data, the second review category of the first sample data, and the attribution annotation data, the second attribution model is trained to obtain a review attribution model. Annealing training means: from multiple first sample data and the corresponding multiple attribution annotation data of multiple first sample data, the first sample data and the attribution annotation data of the first sample data that meet the quality requirements are selected as the third sample data and the attribution annotation data of the third sample data; through the third sample data, the second review category of the third sample data, and the attribution annotation data, the review attribution model is optimized. Alignment training means: obtaining the attribution preference data of the first sample data; through the first sample data, the second review category of the first sample data, the attribution annotation data, and the attribution preference data, the review attribution model is further optimized to obtain an optimized review attribution model.
[0150] In addition, in the embodiments of the present application, when specifically implementing the training of the pre-trained model with the first sample data, the second review category of the first sample data, and the attribution annotation data to obtain the review attribution model in S304 above, considering that when the first sample data includes images, the larger the preset number of pixels of the image, the more the number of image blocks when the first sample data is input into the pre-trained model. To improve the training efficiency of the pre-trained model, the maximum number of image blocks and the number of pixels of the image block can also be pre-configured, calculate the maximum number of pixels of each image in the first sample data, and adjust the pixel number of the image when the preset number of pixels of the image is too large to adjust the first sample data, so as to reduce the number of image blocks when input into the pre-trained model. Therefore, when the first sample data includes images, first determine the maximum number of pixels of each image in the first sample data through the maximum number of image blocks and the number of pixels of the image block; then for each image in the first sample data, determine whether the preset number of pixels of the image is greater than the maximum number of pixels. If so, adjust the pixel number of the image through the preset number of pixels and the maximum number of pixels to adjust the first sample data, and obtain the adjusted first sample data, so that the adjusted pixel number of each image in the adjusted first sample data is less than or equal to the maximum number of pixels; finally, train the pre-trained model with the adjusted first sample data, the second review category of the adjusted first sample data, and the attribution annotation data to obtain the review attribution model. Based on this, the present application provides a possible implementation manner. The above S304 may further include the following S304c-S304e (not shown in the figure).
[0151] S304c: If the first sample data includes images, determine the maximum number of pixels of each image in the first sample data according to the maximum number of image blocks and the number of pixels of the image block.
[0152] S304d: For each image in the first sample data, if the preset number of pixels of the image is greater than the maximum number of pixels, adjust the pixel number of the image according to the preset number of pixels and the maximum number of pixels to obtain the adjusted first sample data; the adjusted pixel number of each image in the adjusted first sample data is less than or equal to the maximum number of pixels.
[0153] S304e: Perform model training on the pre-trained model according to the adjusted first sample data, the second review category of the adjusted first sample data, and the attribution annotation data to obtain the review attribution model.
[0154] Among them, the maximum number of image blocks is the maximum number of image blocks input into the model pre-configured; the number of pixels of the image block is the number of pixels of the image block pre-configured; the maximum number of pixels of each image in the first sample data is the quotient between the product of the maximum number of image blocks and the number of pixels of the image block and the total number of multiple images in the first sample data; the preset number of pixels of the image is the original number of pixels of the image.
[0155] In practical applications, for each image in the first sample data, if the preset number of pixels of the image is greater than the maximum number of pixels, the number of pixels of the image is adjusted according to the preset number of pixels and the maximum number of pixels to obtain the adjusted first sample data, which may be: for each image in the first sample data, if the preset number of pixels of the image is greater than the maximum number of pixels, the pixel scaling ratio is determined according to the quotient of the preset number of pixels and the maximum number of pixels; according to the square root of the pixel scaling ratio, the target scaling ratios of the width and height in the preset number of pixels are determined; according to the quotient of the width in the preset number of pixels and the target scaling ratio, and the quotient of the height in the preset number of pixels and the target scaling ratio, the number of pixels of the image is adjusted to obtain the adjusted first sample data.
[0156] In the case where the first sample data includes images, the above S304c - S304e determine the maximum number of pixels of each image in the first sample data through the maximum number of image blocks and the number of pixels of the image blocks, so as to clarify the upper limit of the number of pixels of each image in the first sample data under the maximum number of image blocks; for each image in the first sample data, it is judged whether the preset number of pixels of the image is greater than the maximum number of pixels. If so, the number of pixels of the image is adjusted through the preset number of pixels and the maximum number of pixels to adjust the first sample data, and the adjusted first sample data is obtained, so that the adjusted number of pixels of each image in the adjusted first sample data is less than or equal to the maximum number of pixels, so as to avoid the preset number of pixels of the images in the first sample data exceeding the upper limit of the number of pixels; the pre - trained model is trained with the adjusted first sample data, the second review category of the adjusted first sample data, and the attribution annotation data to obtain the review attribution model, which can avoid the number of image blocks input into the pre - trained model being greater than the maximum number of image blocks, realize reducing the number of image blocks input into the pre - trained model, and reduce the training time of the pre - trained model.
[0157] As an example of S304c - S304e, based on the above example of S304, the maximum number of image blocks is max_image_tokens, the number of pixels per image block is pixels_per_token, the maximum number of pixels per image is per_img_max_pix, and the preset number of pixels is w×h; when the first harmful data includes images, the computer device determines per_img_max_pix in the first harmful data through max_image_tokens and pixels_per_token; for each image in the first harmful data, the computer device determines whether the w×h of the image is greater than per_img_max_pix. If so, it adjusts the pixel count of the image by w×h and per_img_max_pix to adjust the first harmful data, obtaining the adjusted first harmful data such that the adjusted pixel count of each image in the adjusted first harmful data is less than or equal to the maximum number of pixels; the computer device trains the LLM with the adjusted first harmful data, the second harmful category of the adjusted first harmful data, and the attribution annotation data to obtain an audit attribution model.
[0158] Referring to Table 1, this Table 1 shows the number of image blocks input to the pre - trained model and the training time of the pre - trained model. It can be seen that reducing the number of image blocks input to the pre - trained model significantly reduces the training time of the pre - trained model.
[0159] Table 1 The number of image blocks input to the pre - trained model and the training time of the pre - trained model
[0160] # Number of image blocks Training time consumption 1 1500 Approximately 26 hours 2 900 Approximately 16 hours
[0161] It should be noted that based on the implementation methods provided in the above aspects of this application, further combinations can be made to provide more implementation methods.
[0162] The application scenarios of the audit attribution model in the embodiments of this application are as follows:
[0163] 1. Personalized content audit scenario: This personalized content audit scenario has personalized audit categories and audit rules. The audit attribution model has fully learned the correlation between the first sample data and the second audit category and attribution annotation data of the first sample data. It can support one - key fine - tuning training and single - or multiple - time offline evaluation tests, and perform the next application operation according to the accuracy rate index of the audit category, which is superior to open - source large - language models.
[0164] 2. General content review scenario: There are no standardized review rules in the cold start phase for this general content review scenario. However, the general content review scenario needs to introduce content review capabilities to avoid content security risks. Content data is input into the review attribution model for content review, and it outputs whether the content data belongs to the second review category. The covered general review categories meet most content review requirements and are suitable for quick and low-cost access.
[0165] 3. Review attribution scenario: Content data is input into the review attribution model for content review, and it outputs whether the content data belongs to the second review category, as well as the attribution data for the content data belonging to the second review category. This attribution data enhances the reviewer's model awareness of the review attribution model and serves as a reference basis for the reviewer to conduct content review, further reducing the reviewer's review burden and improving the reviewer's review efficiency.
[0166] In summary, referring to Figure 10 , this Figure 10 is a schematic diagram of the review accuracy of a review model and a review attribution model under multiple second review categories provided by an embodiment of the present application. Among them, the multiple second review categories include the second harmful category 1, the second harmful category 2, the second harmful category 3, the second harmful category 4, the second harmful category 5, the second harmful category 6, the second harmful category 7, and the second harmful category 8. The review accuracy of the review attribution model under the second harmful category 1 is higher than that of the review model under the second harmful category 1,..., and the review accuracy of the review attribution model under the second harmful category 8 is higher than that of the review model under the second harmful category 8.
[0167] Referring to Figure 11 , this Figure 11 is a schematic diagram of the review accuracy of a review model and a review attribution model under different review categories based on different data magnitudes of a sample data set provided by an embodiment of the present application. The different data magnitudes based on the sample data set include 1k data magnitude, 2k data magnitude, 5k data magnitude, 1w data magnitude, 10w data magnitude, and 30w data magnitude, and the different review categories include review category 1, review category 2, review category 3, and review category 4. The review accuracy of the review attribution model based on the 30w sample data set under review category 1, review category 2, review category 3, and review category 4 is 2 to 5 percentage points higher than that of the review model based on the 30w data magnitude sample data set under review category 1, review category 2, review category 3, and review category 4. The review accuracy of the review attribution model based on a small data magnitude sample data set under review category 1, review category 2, review category 3, and review category 4 can still reach the available standard.
[0168] Referring to Figure 12 , this Figure 12A schematic diagram of the coverage rate of the audit traffic of an audit attribution model provided by an embodiment of the present application under different sample data volumes. As the sample data volume of the audit attribution model increases, its audit traffic also increases. Currently, the coverage rate of the audit traffic under the global traffic has reached 10.17%.
[0169] See Figure 13 , the Figure 13 A flowchart of a secondary audit when the model audit result of the content data in the audit attribution model provided by an embodiment of the present application is different from the manual audit result of the auditor. The content data is input into the audit attribution model for content audit to output the model audit result. The content data undergoes content audit by the auditor to obtain the manual audit result. It is determined whether the model audit result is the same as the manual audit result. If not, the content data is subjected to a secondary audit.
[0170] Based on Figure 2 For the training method of the audit attribution model provided by the corresponding embodiment, an embodiment of the present application also provides a training device for the audit attribution model. See Figure 14 , Figure 14 A structural diagram of a training device for an audit attribution model provided by an embodiment of the present application. The training device 1400 of the audit attribution model includes: a category fusion unit 1401, a rule sorting unit 1402, an attribution annotation unit 1403, and a model training unit 1404;
[0171] The category fusion unit 1401 is used to perform category fusion on multiple first audit categories corresponding to multiple first sample data to obtain multiple second audit categories; the total number of the multiple second audit categories is less than the total number of the multiple first audit categories;
[0172] The rule sorting unit 1402 is used to sort the multiple first audit rules corresponding to each second audit category to obtain multiple second audit rules corresponding to each second audit category; the total number of the multiple second audit rules is less than the total number of the multiple first audit rules;
[0173] The attribution annotation unit 1403 is used to perform attribution annotation on the first sample data according to the multiple second audit rules corresponding to the second audit category of the first sample data to obtain the attribution annotation data of the first sample data;
[0174] The model training unit 1404 is used to perform model training on the pre-trained model according to the first sample data, the second audit category of the first sample data, and the attribution annotation data to obtain the audit attribution model.
[0175] In a possible implementation manner, the category fusion unit 1401 is used for:
[0176] The first language model performs category fusion on multiple first review categories corresponding to multiple first sample data according to the category fusion prompt, to obtain multiple second review categories; the category fusion prompt is used to guide the first language model to fuse and reduce the multiple first review categories into multiple main review categories.
[0177] In a possible implementation manner, the rule sorting unit 1402 is configured to:
[0178] The second language model performs rule sorting on multiple first review rules corresponding to each second review category according to the rule sorting prompt, to obtain multiple second review rules corresponding to each second review category; the rule sorting prompt is used to guide the second language model to sort and simplify the multiple first review rules corresponding to each second review category into multiple main review rules.
[0179] In a possible implementation manner, the attribution annotation unit 1403 is configured to:
[0180] The third language model performs attribution annotation on the first sample data according to multiple second review rules corresponding to the second review category of the first sample data and the attribution annotation prompt, to obtain the attribution annotation data of the first sample data; the attribution annotation prompt is used to guide the third language model to annotate the attribution data of the first sample data from multiple second review rules corresponding to the second review category of the first sample data.
[0181] In a possible implementation manner, the training device 1400 of the review attribution model further includes: a first selection unit;
[0182] The first selection unit is configured to perform data selection on multiple first sample data to obtain a subset of sample data;
[0183] The attribution annotation unit 1403 is configured to:
[0184] For each first sample data in the subset of sample data, the third language model performs attribution annotation on the first sample data according to multiple second review rules corresponding to the second review category of the first sample data and the attribution annotation prompt, to obtain the attribution annotation data of the first sample data;
[0185] Perform data selection on multiple first sample data in the subset of sample data and multiple attribution annotation data corresponding to the multiple first sample data, to obtain second sample data and the attribution annotation data of the second sample data; the attribution annotation data of the second sample data meets the quality requirements;
[0186] Update the attribution annotation prompt according to the second sample data and the attribution annotation data of the second sample data, to obtain the updated attribution annotation prompt;
[0187] For each of the multiple first sample data, the third language model performs attribution annotation on the first sample data according to multiple second review rules corresponding to the second review category of the first sample data and the updated attribution annotation prompt, to obtain the attribution annotation data of the first sample data.
[0188] In a possible implementation, the training device 1400 of the review attribution model further includes: a category annotation unit;
[0189] The category annotation unit is configured to perform category annotation on the first sample data according to multiple second review rules corresponding to the second review category of the first sample data, to obtain the third review category of the first sample data; the granularity of the third review category is smaller than that of the first review category, and the first sample data and the third review category of the first sample data are used for model training.
[0190] In a possible implementation, the model training unit 1404 is configured to:
[0191] For the first sample data of a single image among the multiple first sample data, perform model training on the pre-trained model according to the first sample data, the second review category of the first sample data, and the attribution annotation data, to obtain the first attribution model;
[0192] For the first sample data of multiple images among the multiple first sample data, perform model training on the first attribution model according to the first sample data, the second review category of the first sample data, and the attribution annotation data, to obtain the review attribution model.
[0193] In a possible implementation, the model training unit 1404 is configured to:
[0194] For the first sample data of multiple images among the multiple first sample data, perform model training on the first attribution model according to the first sample data, the second review category of the first sample data, and the attribution annotation data, to obtain the second attribution model;
[0195] For the first sample data of plain text among the multiple first sample data, perform model training on the second attribution model according to the first sample data, the second review category of the first sample data, and the attribution annotation data, to obtain the review attribution model.
[0196] In a possible implementation, the training device 1400 of the review attribution model further includes: a first selection unit and a first optimization unit;
[0197] The first selection unit is configured to perform data selection on the multiple first sample data and the multiple attribution annotation data corresponding to the multiple first sample data, to obtain the third sample data and the attribution annotation data of the third sample data; the third sample data and the attribution annotation data of the third sample data meet the quality requirements;
[0198] A first optimization unit, configured to optimize an audit attribution model according to third sample data and attribution annotation data of the third sample data, so as to obtain an optimized audit attribution model.
[0199] In a possible implementation manner, the training device 1400 of the audit attribution model further includes: a data acquisition unit and a second optimization unit;
[0200] The data acquisition unit is configured to acquire attribution preference data of first sample data;
[0201] The second optimization unit is configured to optimize the audit attribution model according to the first sample data, the second audit category of the first sample data, the attribution annotation data, and the attribution preference data, so as to obtain an optimized audit attribution model.
[0202] In a possible implementation manner, the model training unit 1404 is configured to:
[0203] If the first sample data includes images, determine the maximum pixel number of each image in the first sample data according to the maximum number of image blocks and the pixel number of each image block;
[0204] For each image in the first sample data, if the preset pixel number of the image is greater than the maximum pixel number, adjust the pixel number of the image according to the preset pixel number and the maximum pixel number to obtain adjusted first sample data; the adjusted pixel number of each image in the adjusted first sample data is less than or equal to the maximum pixel number;
[0205] Train a pre-trained model according to the adjusted first sample data, the second audit category of the adjusted first sample data, and the attribution annotation data, so as to obtain an audit attribution model.
[0206] As can be seen from the above technical solutions, the training device for the review attribution model includes a category fusion unit, a rule sorting unit, an attribution annotation unit, and a model training unit; the category fusion unit fuses multiple first review categories corresponding to multiple first sample data into multiple second review categories, so that the total number of the multiple second review categories is less than the total number of the multiple first review categories, in order to reduce a large number of first review categories corresponding to a large amount of first sample data to a small number of second review categories; the rule sorting unit sorts multiple first review rules corresponding to each second review category into multiple second review rules, so that the total number of the multiple second review rules is less than the total number of the multiple first review rules, in order to simplify a large number of first review rules corresponding to each second review category to a small number of second review rules; the attribution annotation unit annotates the attribution annotation data indicating that the first sample data belongs to the second review category through multiple second review rules corresponding to the second review category of the first sample data, in order to clarify the attribution data indicating that a large amount of first sample data belongs to the second review category during content review; the model training unit trains a pre-trained model to obtain a review attribution model through the first sample data, the second review category of the first sample data, and the attribution annotation data, so that the pre-trained model fully learns the correlation between the first sample data and the second review category and the attribution annotation data of the first sample data, so that the review attribution model can fully understand a small number of second review rules corresponding to each second review category, and analyze the attribution data belonging to the second review category during content review, and can quickly and accurately review whether the content data belongs to a small number of second review categories.
[0207] Based on this, the device fuses a large number of fine-grained review categories corresponding to a large amount of sample data into a small number of coarse-grained review categories, sorts a large number of cumbersome review rules into a small number of concise review rules for each coarse-grained review category, and based on the small number of concise review rules corresponding to the small number of coarse-grained review categories, annotates the attribution data indicating that a large amount of sample data belongs to the coarse-grained review category, and continues to train the pre-trained model into a review attribution model through the sample data, the coarse-grained review category of the sample data, and the attribution data belonging to the coarse-grained review category, so that the review attribution model can fully understand a small number of concise review rules corresponding to a small number of coarse-grained review categories, and analyze the attribution data belonging to the coarse-grained review category when reviewing content data through the review attribution model, thereby improving the review efficiency and review accuracy.
[0208] The embodiment of the present application further provides a computer device, which may be a server, see Figure 15 , Figure 15The accompanying figure is a structural diagram of a server provided by an embodiment of the present application. The server 1500 may vary significantly due to different configurations or performances. It may include one or more processors, such as a central processing unit (CPU) 1522, and a memory 1532, as well as one or more storage media 1530 (such as one or more mass storage devices) for storing application programs 1542 or data 1544. Among them, the memory 1532 and the storage media 1530 may be transient storage or persistent storage. The programs stored in the storage media 1530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations for the server. Further, the central processing unit 1522 may be configured to communicate with the storage media 1530 and execute a series of instruction operations in the storage media 1530 on the server 1500.
[0209] The server 1500 may further include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input / output interfaces 1558, and / or one or more operating systems 1541, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.
[0210] In this embodiment, the method provided in various alternative implementations of the above embodiment may be executed by the central processing unit 1522 in the server 1500.
[0211] The computer device provided by the embodiment of the present application may also be a terminal. Refer to Figure 16 , Figure 16 The accompanying figure is a structural diagram of a terminal provided by an embodiment of the present application. Taking a smart phone as an example of the terminal, the smart phone includes components such as a radio frequency (RF) circuit 1610, a memory 1620, an input unit 1630, a display unit 1640, sensors 1650, an audio circuit 1660, a wireless fidelity (WiFi) module 1670, a processor 1680, and a power supply 1690. The input unit 1630 may include a touch panel 1631 and other input devices 1632. The display unit 1640 may include a display panel 1641. The audio circuit 1660 may include a speaker 1661 and a microphone 1662. Those skilled in the art can understand that Figure 16 the structure of the smart phone shown in
[0212] The memory 1620 can be used to store software programs and modules. The processor 1680 executes various functional applications and data processing of the smart phone by running the software programs and modules stored in the memory 1620. The memory 1620 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the smart phone (such as audio data, phone book, etc.). In addition, the memory 1620 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0213] The processor 1680 is the control center of the smart phone, connecting various parts of the entire smart phone through various interfaces and lines. By running or executing the software programs and / or modules stored in the memory 1620, and by calling the data stored in the memory 1620, it executes various functions of the smart phone and processes data. Optionally, the processor 1680 may include one or more processing units; preferably, the processor 1680 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1680.
[0214] In this embodiment, the processor 1680 in the smart phone can execute the methods provided in various optional implementation manners of the above embodiments.
[0215] According to one aspect of the present application, there is provided a computer-readable storage medium for storing a computer program. When the computer program runs on a computer device, the computer device is caused to execute the methods provided in various optional implementation manners of the above embodiments.
[0216] According to one aspect of the present application, there is provided a computer program product. The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the methods provided in various optional implementation manners of the above embodiments.
[0217] The descriptions of the processes or structures corresponding to the above respective drawings each have their own focuses. For parts not detailed in a certain process or structure, reference may be made to the relevant descriptions of other processes or structures.
[0218] In the description of this application and the above-mentioned accompanying drawings, terms such as "first" and "second" are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0219] In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in electrical, mechanical, or other forms.
[0220] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0221] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0222] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store computer programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), RAM, magnetic disks, or optical discs.
[0223] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of that module or unit.
[0224] As described above, the above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of this application.
Claims
1. A training method for an audit attribution model, characterized in that: The method comprises: Performing category fusion on a plurality of first audit categories corresponding to a plurality of first sample data to obtain a plurality of second audit categories; the total number of the plurality of second audit categories is less than the total number of the plurality of first audit categories; Arrange the multiple first audit rules corresponding to each second audit category to obtain multiple second audit rules corresponding to each second audit category; the total number of the multiple second audit rules is less than the total number of the multiple first audit rules; Performing attribution labeling on the first sample data according to a plurality of second review rules corresponding to the second review category of the first sample data to obtain attribution labeling data of the first sample data; The pre-trained model is trained according to the first sample data and the second audit category and attribution annotation data of the first sample data to obtain an audit attribution model.
2. The method according to claim 1, characterized in that The step of performing category fusion on the multiple first audit categories corresponding to the multiple first sample data to obtain the multiple second audit categories includes: The multiple first audit categories corresponding to the multiple first sample data are categorized and fused by the first language model according to category fusion prompts to obtain the multiple second audit categories; the category fusion prompts are used to guide the first language model to fuse and reduce the multiple first audit categories into multiple main audit categories.
3. The method according to claim 1, characterized in that The step of arranging the plurality of first audit rules corresponding to each second audit category to obtain the plurality of second audit rules corresponding to each second audit category includes: The multiple first audit rules corresponding to each second audit category are sorted according to rule sorting prompts by the second language model to obtain the multiple second audit rules corresponding to each second audit category; the rule sorting prompts are used to guide the second language model to sort and simplify the multiple first audit rules corresponding to each second audit category into multiple main audit rules.
4. The method according to claim 1, characterized in that: The attribution labeling of the first sample data according to the plurality of second audit rules corresponding to the second audit category of the first sample data to obtain attribution labeling data of the first sample data includes: The first sample data is attributed and labeled by a third language model according to a plurality of second audit rules and attribution annotation prompts corresponding to the second audit category of the first sample data, so as to obtain attribution annotation data of the first sample data; the attribution annotation prompt is used to guide the third language model to label the attribution data of the first sample data from the plurality of second audit rules corresponding to the second audit category of the first sample data.
5. The method according to claim 4, characterized in that The method further comprises: Performing data selection on the plurality of first sample data to obtain a sample data subset; The step of performing attribution annotation on the first sample data according to a plurality of second audit rules and attribution annotation prompts corresponding to the second audit category of the first sample data by using a third language model to obtain attribution annotation data of the first sample data includes: For each first sample data in the sample data subset, attribution annotation is performed on the first sample data by the third language model according to a plurality of second review rules corresponding to the second review category of the first sample data and the attribution annotation prompt to obtain attribution annotation data of the first sample data; Selecting a plurality of first sample data in the sample data subset and a plurality of attribution annotation data corresponding to the plurality of first sample data to obtain second sample data and attribution annotation data of the second sample data; the attribution annotation data of the second sample data meets the quality requirements; updating the attribution annotation prompt according to the second sample data and the attribution annotation data of the second sample data to obtain an updated attribution annotation prompt; For each first sample data among the multiple first sample data, the first sample data is attribution-labeled by the third language model according to the multiple second review rules corresponding to the second review category of the first sample data and the updated attribution labeling prompt to obtain the attribution labeling data of the first sample data.
6. The method according to claim 1, characterized in that The method further comprises: The first sample data is categorized according to multiple second audit rules corresponding to the second audit category of the first sample data to obtain a third audit category of the first sample data; the granularity of the third audit category is smaller than that of the first audit category, and the first sample data and the third audit category of the first sample data are used for model training.
7. The method according to claim 1, characterized in that The performing model training on the pre-trained model according to the first sample data and the second audit category and attribution annotation data of the first sample data to obtain the audit attribution model includes: For first sample data of a single image among the plurality of first sample data, the pre-trained model is trained according to the first sample data and the second audit category and attribution annotation data of the first sample data to obtain a first attribution model; For the first sample data of multiple images in the multiple first sample data, the first attribution model is trained according to the first sample data and the second audit category and attribution annotation data of the first sample data to obtain the audit attribution model.
8. The method according to claim 7, characterized in that The first sample data of multiple images among the multiple first sample data, training the first attribution model according to the first sample data and the second audit category and attribution annotation data of the first sample data to obtain the audit attribution model, includes: For first sample data of multiple images among the multiple first sample data, training the first attribution model according to the first sample data and the second audit category and attribution annotation data of the first sample data to obtain a second attribution model; For the first sample data in plain text among the multiple first sample data, the second attribution model is trained according to the first sample data and the second audit category and attribution annotation data of the first sample data to obtain the audit attribution model.
9. The method according to any one of claims 1 to 8, characterized in that: The method further comprises: Performing data selection on the plurality of first sample data and the plurality of attribution annotation data corresponding to the plurality of first sample data to obtain third sample data and the attribution annotation data of the third sample data; the third sample data and the attribution annotation data of the third sample data both meet the quality requirements; The audit attribution model is optimized according to the third sample data and the attribution annotation data of the third sample data to obtain an optimized audit attribution model.
10. The method according to any one of claims 1 to 8, characterized in that: The method further comprises: Obtaining attribution preference data of the first sample data; The audit attribution model is optimized according to the first sample data, the second audit category of the first sample data, the attribution annotation data, and the attribution preference data to obtain an optimized audit attribution model.
11. The method according to claim 1, characterized in that: The performing model training on the pre-trained model according to the first sample data and the second audit category and attribution annotation data of the first sample data to obtain the audit attribution model includes: If the first sample data includes images, determining a maximum number of pixels of each image in the first sample data according to a maximum number of image blocks and a number of pixels of the image blocks; For each image in the first sample data, if the preset number of pixels of the image is greater than the maximum number of pixels, adjusting the number of pixels of the image according to the preset number of pixels and the maximum number of pixels to obtain adjusted first sample data; the adjusted number of pixels of each image in the adjusted first sample data is less than or equal to the maximum number of pixels; The pre-trained model is trained according to the adjusted first sample data and the second audit category and attribution annotation data of the adjusted first sample data to obtain the audit attribution model.
12. A training device for an audit attribution model, characterized in that: The device comprises: a category fusion unit, a rule arrangement unit, an attribution labeling unit and a model training unit; The category fusion unit is used to perform category fusion on a plurality of first audit categories corresponding to a plurality of first sample data to obtain a plurality of second audit categories; the total number of the plurality of second audit categories is less than the total number of the plurality of first audit categories; The rule arrangement unit is used to arrange the multiple first audit rules corresponding to each second audit category to obtain multiple second audit rules corresponding to each second audit category; the total number of the multiple second audit rules is less than the total number of the multiple first audit rules; The attribution labeling unit is used to perform attribution labeling on the first sample data according to a plurality of second audit rules corresponding to the second audit category of the first sample data, so as to obtain attribution labeling data of the first sample data; The model training unit is used to perform model training on a pre-trained model according to the first sample data and the second audit category and attribution annotation data of the first sample data to obtain an audit attribution model.
13. A computer device, characterized in that: The computer device comprises a processor and a memory: The memory is used to store a computer program and transmit the computer program to the processor; The processor is configured to execute the method according to any one of claims 1 to 11 according to the instructions in the computer program.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and when the computer program is executed on a computer device, the computer device executes the method according to any one of claims 1 to 11.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 11.
Citation Information
Cited By
Content auditing model processing method and device, electronic equipment and medium
CN121390195A