Text data processing method and device and related equipment

By building a quality inspection model and using text definitions and tag information to train the initial business model, the problem of low efficiency of manual quality inspection is solved, and efficient and accurate text data quality inspection is achieved.

CN120743890APending Publication Date: 2025-10-03TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410385844.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-30
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing technologies make it difficult to ensure the accuracy and efficiency of quality inspection results during manual quality inspection before training large models. This is especially true when screening training data from massive text data. Manual quality inspection is time-consuming, resulting in low quality inspection efficiency.

Method used

By building a quality inspection model based on text definition information and tag information, using sampled text data to train the initial business model, generating a tag prompt template, and quickly training the target quality inspection model, quality inspection of the input text data can be achieved.

Benefits of technology

It improves the efficiency and accuracy of quality inspection, can quickly train models for performing quality inspection tasks, and improves the quality inspection effect of input text data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743890A_ABST
    Figure CN120743890A_ABST
Patent Text Reader

Abstract

The invention discloses a text data processing method and device and related equipment. The method comprises the steps of determining text definition information of to-be-processed text data based on a quality inspection task; obtaining sampled text data from the to-be-processed text data, and determining text marking information of the sampled text data based on the text definition information; obtaining a text prompt template associated with the text definition information, determining a mark prompt template corresponding to the text prompt template based on the text mark information, performing model training on the initial service model based on the text mark information, the sampling text data and the mark prompt template, and training to obtain a target quality inspection model for executing a quality inspection task; and obtaining input text data used for inputting the target quality inspection model from the to-be-processed text data, and executing a quality inspection task by the target quality inspection model for the input text data to obtain a text quality inspection result of the input text data. According to the invention, the quality inspection efficiency and the quality inspection accuracy can be effectively improved through the quality inspection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a text data processing method, apparatus, and related equipment. Background Art

[0002] Currently, before training a large model for text processing, in order to ensure the accuracy of the large model in text processing, it is necessary to conduct manual quality inspection on the massive training data (for example, massive text data) used to input the large model in advance to ensure the data quality of the training data (for example, massive text data) finally input into the large model.

[0003] However, in practice, the inventors discovered that during manual quality inspection, given that different personnel have different standards for training data quality, it is difficult to accurately filter all training data for input into large models from massive amounts of text data. This makes it difficult to ensure the accuracy of the quality inspection results obtained through manual quality inspection. Furthermore, since training large models often requires a large amount of training data, manually inspecting each piece of training data would consume a long time, thereby reducing the efficiency of quality inspection of these training data. Summary of the Invention

[0004] The embodiments of the present application provide a text data processing method, apparatus, and related equipment, which can improve quality inspection efficiency and accuracy through a quality inspection model obtained through rapid training.

[0005] On the one hand, an embodiment of the present application provides a text data processing method, the method comprising:

[0006] Determining text definition information of the text data to be processed based on a quality inspection task associated with the text data to be processed; the text definition information is determined based on text key fields and text attributes of the text key fields in the text data to be processed;

[0007] Acquire sampled text data from the text data to be processed, and determine text tag information of the sampled text data based on the text definition information; the text tag information is determined based on a sampling key field in the sampled text data and text attributes of the sampling key field; the sampling key field is a field in the text key field;

[0008] Acquire a text prompt template associated with the text definition information, determine a tag prompt template corresponding to the text prompt template based on the text tag information, and use the text tag information and the sampled text data as a training sample data pair for training an initial business model;

[0009] The initial business model is trained by training sample data pairs and labeling prompt templates to obtain the target quality inspection model for performing quality inspection tasks;

[0010] The input text data for inputting into the target quality inspection model is obtained from the text data to be processed, and the target quality inspection model performs a quality inspection task on the input text data to obtain a text quality inspection result of the input text data.

[0011] An embodiment of the present application provides a text data processing device, comprising:

[0012] A definition information determination module is used to determine text definition information of the text data to be processed based on the quality inspection task associated with the text data to be processed; the text definition information is determined based on text key fields in the text data to be processed and text attributes of the text key fields;

[0013] A sampling data acquisition module is used to obtain sampling text data from the text data to be processed;

[0014] The tag information determination module is used to determine the text tag information of the sampled text data based on the text definition information; the text tag information is determined based on the sampling key field in the sampled text data and the text attribute of the sampling key field; the sampling key field is a field in the text key field;

[0015] a sample data pair determination module, configured to obtain a text prompt template associated with the text definition information, determine a tag prompt template corresponding to the text prompt template based on the text tag information, and use the text tag information and the sampled text data as a training sample data pair for training an initial business model;

[0016] The initial business model training module is used to train the initial business model through training sample data pairs and labeling prompt templates, and obtain the target quality inspection model for performing quality inspection tasks;

[0017] The quality inspection task execution module is used to obtain input text data for inputting the target quality inspection model from the text data to be processed, and the target quality inspection model executes the quality inspection task on the input text data to obtain the text quality inspection result of the input text data.

[0018] The definition information determination module includes:

[0019] A data set acquisition unit is used to acquire a training text data set for training a target business model, and use each training text data in the training text data set as text data to be processed; the target business model is a business model associated with the target quality inspection model;

[0020] A quality inspection task acquisition unit, used to acquire quality inspection tasks associated with the training text dataset;

[0021] A quality inspection definition unit is used to perform quality inspection definition on the text data to be processed based on the quality inspection task and the data type of the text data to be processed, and obtain text key fields and text attributes of the text key fields of the text data to be processed;

[0022] The definition information determining unit is used to use the text key field and the text attribute of the text key field as the text definition information of the text data to be processed.

[0023] Among them, the quality inspection task at least includes the watermark quality inspection task;

[0024] The quality inspection definition unit is specifically used to obtain a watermark quality inspection strategy associated with the watermark quality inspection task, and obtain a first target data type to be watermark defined from the data type of the text data to be processed based on the watermark quality inspection strategy;

[0025] The quality inspection definition unit is further specifically configured to use, in the text data to be processed, the text data to be processed that matches the first target data type as first target text data, perform watermark definition on the first target text data, obtain a watermark definition field of the first target text data, and use the watermark definition field as a text key field of the text data to be processed;

[0026] The quality inspection definition unit is further specifically configured to classify the watermark definition field into a watermark type indicated by the watermark quality inspection strategy, and determine the watermark type as a text attribute of the text key field.

[0027] Optionally, the quality inspection task includes at least a noise quality inspection task;

[0028] The quality inspection definition unit is further specifically configured to obtain a noise quality inspection strategy associated with the noise quality inspection task, and obtain a second target data type to be subjected to noise definition from the data type of the to-be-processed text data based on the noise quality inspection strategy;

[0029] The quality inspection definition unit is further specifically configured to use, in the text data to be processed, the text data to be processed that matches the second target data type as the second target text data, perform noise definition on the second target text data to obtain a noise definition field of the second target text data, and use the noise definition field as a text key field of the text data to be processed;

[0030] The quality inspection definition unit is further specifically configured to classify the noise definition field into the noise type indicated by the noise quality inspection strategy, and determine the noise type as a text attribute of the text key field.

[0031] Among them, the sampling data acquisition module includes:

[0032] A sampling processing unit, configured to perform sampling processing on the text data to be processed based on the data type of the text data to be processed, to obtain a sampling data set that matches the data type of the text data to be processed;

[0033] The sampled text determining unit is configured to determine the sampled text data based on a sampled data set that matches the data type of the text data to be processed.

[0034] The data types of the text data to be processed include at least a first data type and a second data type;

[0035] The sampling processing unit is specifically configured to determine a first data type from the data types of the text data to be processed, extract N text data matching the first data type from the text data to be processed, and use the extracted N text data matching the first data type as a first sampling data set; N is a positive integer;

[0036] The sampling processing unit is further specifically configured to determine a second data type from the data types of the text data to be processed, extract N text data matching the second data type from the text data to be processed, and use the extracted N text data matching the second data type as a second sampling data set;

[0037] The sampling processing unit is further specifically configured to determine a sampling data set that matches the data type of the text data to be processed based on the first sampling data set and the second sampling data set.

[0038] The marking information determination module includes:

[0039] A sampled text distribution unit is used to distribute the sampled text data to a plurality of quality inspection marking terminals associated with the quality inspection task, wherein one quality inspection marking terminal is used to perform text annotation on the sampled text data based on the text definition information, and obtain a text annotation information for the sampled text data by the quality inspection marking terminal;

[0040] The annotation information receiving unit is used to receive the text annotation information returned by each quality inspection mark terminal for the sampled text data, cross-check the received text annotation information, and obtain a cross-check result for the sampling key fields in the sampled text data and the text attributes of the sampling key fields;

[0041] The tag information determining unit is configured to determine the text tag information of the sampled text data based on the sampling key fields and the text attributes of the sampling key fields if the cross-check result indicates that the received text annotation information is consistent.

[0042] Optionally, the quality inspection task includes at least a watermark quality inspection task, the text key field includes a watermark definition field, and the text attribute of the text key field includes the watermark type to which the watermark definition field belongs;

[0043] The tag information determination module includes:

[0044] The prompt template construction unit is used to obtain watermark definition information including watermark definition fields and watermark types from the text definition information based on the watermark quality inspection task, and determine the prompt template constructed based on the watermark definition information as the watermark prompt template;

[0045] a reply information output unit, configured to input the sampled text data into a watermark prompt template, and output the marked reply information of the sampled text data based on the watermark definition information by the watermark prompt template;

[0046] The sampled text sending unit is further configured to send the sampled text data to a watermark quality inspection terminal associated with the watermark quality inspection task, so that the watermark quality inspection terminal performs text marking on the sampled text data based on the watermark definition information to obtain watermark marking information of the sampled text data;

[0047] The marking information determining unit is further configured to receive the watermark marking information returned by the watermark quality inspection terminal, and determine the text marking information of the sampled text data based on the watermark marking information and the marking reply information.

[0048] The marking information determining unit is specifically configured to use the watermark text segment of the sampled text data carried in the watermark marking information as the first watermark segment, and use the watermark text segment of the sampled text data carried in the marking reply information as the second watermark segment;

[0049] The marking information determining unit is further specifically configured to, if the first watermark segment contains a field different from the second watermark segment, correct the second watermark segment in the marking reply information to be the first watermark segment, use the watermark text field indicated by the first watermark segment as a sampling key field in the sampled text data, and use the watermark type to which the watermark text field indicated by the first watermark segment belongs as a text attribute of the sampling key field;

[0050] The tag information determining unit is further specifically configured to determine the text tag information of the sampled text data based on the sampling key field and the text attribute of the sampling key field.

[0051] The initial business model training module includes:

[0052] A prediction tag output unit, configured to input the sampled text data in the training sample data pair into the initial business model, have the initial business model predict and output text tag information of the sampled text data, and use the predicted output text tag information as the prediction tag information of the sampled text data;

[0053] A loss value determining unit, configured to determine a model loss value of an initial business model based on text tag information and prediction tag information in a training sample data pair;

[0054] The parameter optimization unit is used to optimize the model parameters of the initial business model if the model loss value of the initial business model does not meet the model optimization conditions indicated by the quality inspection task, until the model loss value of the business model after parameter optimization meets the model optimization conditions, and use the business model whose model loss value meets the model optimization conditions as the target quality inspection model for executing the quality inspection task; the model optimization conditions are used to instruct the business model after parameter optimization to output predicted tag information that matches the text tag information according to the tag prompt template prediction.

[0055] Among them, the quality inspection task execution module includes:

[0056] a target data type determination unit, configured to obtain a target data type from the data type of the text data to be processed, and obtain K data sets deployed under the target data type; one data set corresponds to one data identifier, K is a positive integer, and each of the K data sets includes at least k1 training data;

[0057] A training data extraction unit is configured to obtain k2 training data associated with the target data type based on the k1 training data extracted from each data set, and use the k2 training data as input text data for inputting into the target quality inspection model; k2 = k1*K;

[0058] a quality inspection task execution unit, configured to input the input text data into a target quality inspection model, and when the target quality inspection model performs a quality inspection task on the input text data based on the tag prompt template, predict and output predicted tag information that matches the text output style of the tag prompt template, and use the predicted tag information output as target tag information of the input text data;

[0059] The tag information parsing unit is used to parse the target tag information to obtain text parsing information corresponding to the target tag information. Based on the data identifier of each data set, the text parsing information associated with the training data in each data set is searched in the text parsing information. The text parsing information associated with the training data in each data set is determined as the data set quality inspection result of each data set, and the data set quality inspection result of each data set is used as the text quality inspection result of the input text data.

[0060] Among them, the quality inspection task at least includes the watermark quality inspection task;

[0061] The device also includes:

[0062] The target dataset acquisition module is used to obtain the target dataset from the K datasets, and obtain the dataset quality inspection results that match the data identifier of the target dataset from the text quality inspection results; the total number of quality inspections corresponding to the dataset quality inspection results of the target dataset is k1;

[0063] The watermark detection quantity determination module is used to search for data quality inspection results associated with the watermark quality inspection task in the data set quality inspection results of the target data set, use the data quality inspection results associated with the watermark quality inspection task found as the watermark detection results, and use the cumulative number of watermark detection results as the number of watermark detections corresponding to the target data set;

[0064] The watermark ratio determination module is used to determine the watermark ratio corresponding to the target dataset based on the number of watermark detections and the total number of quality inspections. When each of the K datasets is used as the target dataset, the watermark ratio corresponding to each dataset is obtained;

[0065] The watermark proportion curve display module is used to fit the watermark proportion curve corresponding to each data set based on the watermark proportion corresponding to each data set, and display the watermark proportion curve on the quality inspection display platform associated with the watermark quality inspection task.

[0066] Optionally, the device further comprises:

[0067] A cleaning task acquisition module is used to acquire the cleaning task associated with the quality inspection task and obtain the text quality inspection segment associated with the cleaning task from the text quality inspection result;

[0068] The module for determining samples to be cleaned is used to obtain the initial recall model and the initial cleaning model associated with the cleaning task, and use the training data carrying the text quality inspection fragments obtained from the input text data as the sample data to be cleaned associated with the initial recall model and the initial cleaning model;

[0069] The recall model training module is used to train the initial recall model using the sample data to be cleaned and the text quality inspection fragments in the sample data to be cleaned, so as to obtain the target recall model for data recall;

[0070] The cleaning model training module is used to train the initial cleaning model using the sample data to be cleaned and the text quality inspection fragments in the sample data to be cleaned, so as to obtain the target cleaning model for data cleaning.

[0071] Optionally, the device further comprises:

[0072] A data cleaning module is configured to perform a model cascade process on the target recall model and the target cleaning model to obtain a data cleaning model associated with the cleaning task, and obtain text data to be cleaned associated with a target dataset from a first business database associated with the cleaning task; the target dataset is any dataset deployed under the data type of the text data to be processed;

[0073] The data cleaning module is further configured to input the text data to be cleaned into a target recall model in the data cleaning model, and the target recall model screens target text data to be cleaned that carries text data segments that are identical to the text quality inspection segments from the text data to be cleaned, uses the text data segments in the target text data to be cleaned that match the text quality inspection segments as cleaning mark segments of the text data to be cleaned, and outputs the target text data to be cleaned that carries the cleaning mark segments from the target recall model;

[0074] The data cleaning module is also used to input the target text data to be cleaned into the target cleaning model in the data cleaning model, the target cleaning model identifies cleaning mark segments from the target text data to be cleaned, performs text cleaning on the cleaning mark segments in the target text data to be cleaned, and uses the target text data to be cleaned after text cleaning as the data to be quality inspected for input into the target quality inspection model.

[0075] Optionally, the quality inspection task execution module is further configured to input the data to be quality inspected into the target quality inspection model, and the target quality inspection model executes the quality inspection task on the data to be quality inspected to obtain the data quality inspection result of the data to be quality inspected;

[0076] The quality inspection task execution module is also used to determine the quality inspection ratio corresponding to the target data set based on the data quality inspection results;

[0077] The quality inspection task execution module is also used to use the data to be quality inspected associated with the target data set as model training sample data for training the target business model if the quality inspection ratio meets the quality inspection data strategy indicated by the quality inspection task, and add the model training sample data to the second business database corresponding to the target business model.

[0078] In one aspect, an embodiment of the present application provides a computer device, including a memory and a processor, wherein the memory is connected to the processor, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method provided in the above aspect of the embodiment of the present application.

[0079] On one hand, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. The computer program is suitable for being loaded and executed by a processor, so that a computer device with a processor executes the method provided in the above aspect of the embodiment of the present application.

[0080] According to one aspect of the present application, a computer program product is provided, which includes a computer program. When the computer program is executed by a processor, it implements the method provided in any of the above aspects of the embodiments of the present application.

[0081] In an embodiment of the present application, a computer device may determine text definition information of the text data to be processed based on a quality inspection task associated with the text data to be processed; the text definition information here is determined based on text key fields and text attributes of the text key fields in the text data to be processed; further, the computer device may obtain sampled text data from the text data to be processed, and may determine text tag information of the sampled text data based on the text definition information; the text tag information here is determined based on sampling key fields and text attributes of the sampling key fields in the sampled text data; it should be understood that the sampling key fields here are fields in the text key fields; further, the computer device may obtain text tag information related to the text definition information. The associated text prompt template can be determined based on the text tag information to determine the tag prompt template corresponding to the text prompt template, and then the text tag information and the sampled text data can be used as a training sample data pair for training the initial business model; further, the computer device can use the training sample data pair and the tag prompt template to perform model training on the initial business model (the initial business model here can be any business model that has been trained currently), and the training obtains a target quality inspection model for performing quality inspection tasks, so that when the input text data for inputting the target quality inspection model is obtained from the text data to be processed, the target quality inspection model can be used to perform quality inspection tasks on the input text data to obtain text quality inspection results for the input text data. Thus, the embodiment of the present application can accurately construct a mark prompt template for the text output style currently used to standardize the initial business model through the text definition information of the to-be-processed text data and the text tag information of the limited sampled text data sampled from the to-be-processed text data in the process of training the quality inspection model (i.e., the target quality inspection model), and then can realize the expansion of the business capabilities of the initial business model through the mark prompt model, the limited sampled text data sampled from the to-be-processed text data and the text tag information of the limited sampled text data, so as to quickly train and obtain a quality inspection model for performing data quality inspection. This means that the embodiment of the present application can improve the efficiency of obtaining the target quality inspection model for performing quality inspection tasks by fine-tuning the model when the initial business model (i.e., any business model) is trained, and then can improve the quality inspection efficiency and quality inspection accuracy of data quality inspection of input text data when the quality inspection model obtained by rapid training performs quality inspection tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0083] Figure 1 This is a schematic diagram of the network structure of a text data processing system provided in an embodiment of the present application;

[0084] Figure 2 This is a schematic diagram of a scenario in which a quality inspection model is trained using limited labeled training samples, as provided in an embodiment of the present application;

[0085] Figure 3 This is a flowchart of a text data processing method provided by an embodiment of the present application;

[0086] Figure 4 This is a schematic diagram of a scenario for quality inspection definition provided in an embodiment of the present application;

[0087] Figure 5 This is a schematic diagram of a scenario in which text tag information of sampled text data is determined based on watermark definition information, as provided in an embodiment of the present application;

[0088] Figure 6 This is a schematic diagram of a scenario for obtaining input text data provided by an embodiment of the present application;

[0089] Figure 7 This is a flowchart of a text data processing method provided by an embodiment of the present application;

[0090] Figure 8 This is a schematic diagram of a scenario for training a target recall model and a target cleaning model provided in an embodiment of the present application;

[0091] Figure 9 This is a schematic diagram of a scenario of performing watermark recognition on quality inspection data based on a watermark quality inspection task provided by an embodiment of the present application;

[0092] Figure 10 This is a schematic diagram of a scenario of keyword recall in a recall layer provided by an embodiment of the present application;

[0093] Figure 11 This is a schematic diagram of a scenario in which data quality inspection is performed under different quality inspection dimensions, as proposed in an embodiment of the present application;

[0094] Figure 12 This is a schematic diagram of a quality inspection and analysis report provided in an embodiment of the present application;

[0095] Figure 13 This is a structural diagram of a text data processing device provided in an embodiment of the present application;

[0096] Figure 14 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0097] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0098] See Figure 1 , Figure 1 This is a network structure diagram of a text data processing system provided by an embodiment of the present application. Figure 1 As shown, the text data processing system may include terminal devices (such as device 11a, device 12a, device 13a) and server 100a. It is understood that, Figure 1 The number of terminal devices and servers in the example is merely illustrative; any number of terminal devices and servers may be used depending on implementation requirements. Terminal devices (e.g., device 11a, device 12a, and device 13a) may communicate with servers via a network (i.e., a medium providing a communication link via a wired or wireless communication link or fiber optic cable, etc.) to transmit data.

[0099] It is understood that a client may be running on a terminal device (such as device 12a), and the client may be a program that provides local services to a user (also called a business object or operation object). Server 100a may be a server corresponding to the client, and the server 100a may run a program for providing resources, service data, and other services to the client.

[0100] It is understandable that the client running on the terminal device can also be called an application client, a business client, and so on. For example, the client running on the terminal device can be a client for providing data quality inspection services, where the data quality inspection services are determined by the quality inspection model integrated on the client (i.e., the target quality inspection model), and the quality inspection model (i.e., the target quality inspection model) here is obtained by performing model training (for example, specifically, model fine-tuning) on ​​any of the currently acquired trained business models (i.e., the initial business model) in accordance with the corresponding quality inspection tasks. The initial business model can be a large model for text processing (for example, a language business model), or a mixed-yuan assistant model (also called a mixed-yuan large model) for intelligent question and answering.

[0101] It should be understood that in the embodiment of the present application, model fine-tuning processing refers to optimizing the model parameters of the currently acquired initial business model (any trained business model) until the business model after parameter optimization has the ability to perform the corresponding quality inspection task. The business model that can currently perform the corresponding quality inspection task can be used as the target quality inspection model.

[0102] In this way, when the target quality inspection model obtained by current training is integrated on the client, the target quality inspection model can be used to perform quality inspection tasks (for example, one or more of watermark quality inspection tasks, noise quality inspection tasks, illegal data quality inspection tasks, semantic quality inspection tasks, and format quality inspection tasks) on the currently input text data (i.e., input text data) to obtain text quality inspection results of the input text data, and then the input text data can be distilled using the text quality inspection results to filter out distilled text data (also referred to as distilled sample data) for training the recall model and the cleaning model from the input text data, and then the filtered distilled text data can be used as the sample data to be cleaned, and then the recall model and the cleaning model can be fine-tuned through the sample data to be cleaned and the text quality inspection fragments carried in the sample data to be cleaned (i.e., the text fragments corresponding to the sensitive data carried in the sample data to be cleaned).

[0103] It should be understood that when the quality inspection task specifically includes one or more of a watermark quality inspection task, a noise quality inspection task, an illegal data quality inspection task, a semantic quality inspection task, and a format quality inspection task, the target quality inspection model integrated in the client may specifically include one or more of the following quality inspection models: a watermark quality inspection model for performing a watermark quality inspection task, a noise quality inspection model for performing a noise quality inspection task, an illegal data quality inspection model for performing an illegal data quality inspection task, a semantic quality inspection model for performing a semantic quality inspection task, and a format quality inspection model for performing a format quality inspection task.

[0104] It should be understood that the embodiments of the present application can, based on the different quality inspection tasks, perform quality inspection definitions of different quality inspection dimensions on the text data to be processed obtained from the business database to obtain text definition information associated with different quality inspection tasks, and then train different text definition information to obtain quality inspection models for performing different quality inspection tasks. Furthermore, the embodiments of the present application can collectively refer to the quality inspection models for performing different quality inspection tasks as target quality inspection models. Among them, the text definition information obtained by the embodiments of the present application may include but is not limited to: one or more of: watermark definition information associated with watermark quality inspection tasks, noise definition information associated with noise quality inspection tasks, illegal data definition information associated with illegal data quality inspection tasks, semantic definition information associated with semantic quality inspection tasks, and format definition information associated with format quality inspection tasks.

[0105] The watermark quality inspection task refers to the process of identifying, when obtaining the data to be inspected from the business database, the watermark quality inspection model in the currently trained target quality inspection model, and identifying the watermark segments corresponding to the text segments of the sensitive data (e.g., watermark data) from the data to be inspected. It should be understood that in the embodiments of the present application, the business database used to obtain the data to be inspected can be the same business database or a different business database than the business database used to obtain the input text data, and this is not limited here.

[0106] Among them, the specific process of training the watermark quality inspection model can be summarized as follows: when the text data to be processed is obtained from the business database, the sampled text data obtained from the text data to be processed can be watermarked based on the watermark definition information of the text data to be processed to obtain the watermark marking information of the sampled text data, and then the initial business model can be trained through the marking prompt template, sampled text data and watermark marking information of the sampled text data constructed based on the watermark definition information to obtain a watermark quality inspection model for performing watermark quality inspection tasks.

[0107] Similarly, the noise quality inspection task refers to the process of identifying the text segments corresponding to the various sensitive data (e.g., noise data) from the data to be inspected as noise segments by using the noise quality inspection model in the target quality inspection model currently trained when the data to be inspected is obtained from the business database. The specific process of training the noise quality inspection model can be summarized as follows: when the text data to be processed is obtained from the business database, the sampled text data obtained from the text data to be processed can be noise-labeled based on the noise definition information of the text data to be processed to obtain the noise labeling information of the sampled text data, and then the initial business model can be trained using the labeling prompt template currently constructed based on the noise definition information, the sampled text data, and the noise labeling information of the sampled text data to obtain a noise quality inspection model for performing the noise quality inspection task.

[0108] Similarly, the illegal data quality inspection task refers to the process of identifying the illegal data segments corresponding to the sensitive data (e.g., illegal data) from the data to be inspected by the illegal data quality inspection model in the target quality inspection model currently trained when the data to be inspected is obtained from the business database. The specific process of training the illegal data quality inspection model can be summarized as follows: when the text data to be processed is obtained from the business database, the sampled text data obtained from the text data to be processed can be labeled as illegal data based on the illegal data definition information of the text data to be processed to obtain the illegal data labeling information of the sampled text data, and then the initial business model can be trained using the labeling prompt template currently constructed based on the illegal data definition information, the sampled text data, and the illegal data labeling information of the sampled text data to obtain a noise quality inspection model for performing the illegal data quality inspection task.

[0109] By analogy, the semantic quality inspection task refers to the process of identifying the text segments corresponding to various sensitive data (e.g., semantic data) from the data to be inspected as semantic segments by using the semantic quality inspection model in the target quality inspection model currently trained when the data to be inspected is obtained from the business database. The specific process of training the semantic quality inspection model can be summarized as follows: when the text data to be processed is obtained from the business database, the sampled text data obtained from the text data to be processed can be semantically labeled based on the semantic definition information of the text data to be processed to obtain the semantic labeling information of the sampled text data, and then the initial business model can be trained using the labeling prompt template currently constructed based on the semantic definition information, the sampled text data, and the semantic labeling information of the sampled text data to obtain a semantic quality inspection model for performing the semantic quality inspection task.

[0110] It should be understood that the format quality inspection task refers to the process of identifying the text fragments corresponding to the various sensitive data (e.g., format data) from the data to be inspected as format fragments by using the format quality inspection model in the target quality inspection model currently trained when the data to be inspected is obtained from the business database. The specific process of training the format quality inspection model can be summarized as follows: when the text data to be processed is obtained from the business database, the sampled text data obtained from the text data to be processed can be format-marked based on the format definition information of the text data to be processed to obtain the format marking information of the sampled text data, and then the initial business model can be trained using the marking prompt template currently constructed based on the format definition information, the sampled text data, and the format marking information of the sampled text data to obtain a format quality inspection model for performing the format quality inspection task.

[0111] For ease of understanding, in the embodiments of the present application, one or more of the watermark recognition service for identifying watermark data, the noise recognition service for identifying noise data, the illegal data recognition service for identifying illegal data, the semantic recognition service for identifying semantic data, and the format recognition service for identifying format data provided by the target quality inspection model may be collectively referred to as a data quality inspection service. The text data to be processed for training the watermark quality inspection model, the text data to be processed for training the noise quality inspection model, the text data to be processed for training the illegal data quality inspection model, the text data to be processed for training the semantic quality inspection model, and the text data to be processed for training the format quality inspection model may be the same or different, and are not limited thereto.

[0112] Optionally, when a target quality inspection model is integrated on a server (such as server 100a) associated with a client, an embodiment of the present application may perform a quality inspection task (for example, one or more of the watermark quality inspection task, noise quality inspection task, illegal data quality inspection task, semantic quality inspection task, and format quality inspection task) on the input text data (for example, the above-mentioned input text data or the data to be inspected) through the target quality inspection model integrated on the server (such as server 100a) to obtain a text quality inspection result of the input text data or a data quality inspection result of the data to be inspected. The embodiment of the present application will not limit the computer device integrated with the target quality inspection model. The computer device here can be a server or a terminal device running a client.

[0113] Furthermore, it can be understood that in order to ensure the security and reliability of the model training sample data subsequently added to the business database for training the large model (i.e., the target business model, where the target business model can be any business model in the language business model), the embodiment of the present application can also obtain the sample data to be cleaned for training the first business model (i.e., the initial recall model) and the second business model (i.e., the initial cleaning model) from the input text data by means of data distillation after the target quality inspection model obtained through the current training predicts the text quality inspection result of the output input text data. In this way, after the first business model (i.e., the initial recall model) and the second business model (i.e., the initial cleaning model) are respectively trained by the sample data to be cleaned, the target recall model for data recall and the target cleaning model for data cleaning can be obtained.

[0114] Furthermore, it can be understood that the target recall model here can also be used to obtain text data to be cleaned from the first business database of the business database (for example, a newly added sample database), and then the obtained text data to be cleaned can be screened to filter out most of the text data that does not carry text quality inspection segments from the obtained text data to be cleaned, and then the remaining text data carrying text quality inspection segments in the text data to be cleaned can be used as target text data to be cleaned, and the text quality inspection segments in the target text data to be cleaned can be used as cleaning mark segments of the target text data to be cleaned, so as to ensure that the target recall model can efficiently output the target text data to be cleaned carrying the cleaning mark segments.

[0115] Furthermore, it can be understood that the target cleaning model can perform data cleaning (also known as text cleaning) on ​​the target text data to be cleaned output by the target recall model, and then the target text data to be cleaned after text cleaning (i.e., simply referred to as cleaned data) can be used as the data to be quality inspected for input into the target quality inspection model to obtain the data quality inspection results of the data to be quality inspected, and then data statistics can be performed on these data quality inspection results in units of data sets (i.e., unit granularity) to obtain the quality inspection ratio corresponding to each data set, and then the data sets whose quality inspection ratios meet the quality inspection data strategy indicated by the quality inspection task can be selected from the quality inspection ratios corresponding to these data sets, and then The training text data in the selected data set is used as model training sample data for training the target business model (i.e. the above-mentioned language business model, where the language business model can also be a large model with a large number of model parameters), so that the model training sample data here is added to the business database corresponding to the target business model (i.e. the above-mentioned language business model) (specifically, it can be a second business database in the business database), so that the target business model (i.e. the above-mentioned language business model) can be trained through the model training sample data after data cleaning and data quality inspection to ensure that the trained target business model can subsequently output business data that meets the quality inspection quality in the corresponding business scenario.

[0116] It should be understood that after the embodiment of the present application predicts and outputs the data quality inspection results of the data to be quality inspected through the target quality inspection model, a quality inspection analysis report for the data to be quality inspected can be generated in units of data sets, and then the quality inspection analysis report can be displayed on the quality inspection display platform of the client (for example, the watermark ratio curve of each data set obtained by fitting based on the watermark ratio of each data set can be displayed), so that the quality inspection display platform can help relevant business personnel (for example, quality inspection participants or quality inspection managers, etc.) to monitor and manage the data quality of the model training sample data input into the second business database in real time.

[0117] It is understandable that the terminal devices involved in the embodiments of the present application (such as device 11a) may include but are not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, smart speakers, smart home appliances, etc., and are not limited here. In addition, the server involved in the embodiments of the present application (such as server 100a) may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms, and are not limited here.

[0118] It is understandable that the text data processing method involved in the embodiment of the present application can be applied to a computer device, which can be the above-mentioned Figure 1 If a terminal device is used, the business object can perform data quality inspection on the currently acquired data to be quality inspected (for example, the newly added text data acquired from the above-mentioned first business database in real time, or the model training sample data acquired from the second business database periodically) in the terminal device, and display the quality inspection ratio obtained by performing quality inspection analysis on the data quality inspection results of these data to be quality inspected in the terminal device.

[0119] Alternatively, the computer device may be the above-mentioned Figure 1 In the example, if the server or other data processing device is used for data processing, the business object (for example, user A mentioned above) can flexibly select the data to be inspected that needs to be inspected from the business database through the terminal device, and then notify the server (or other data processing device) to perform data quality inspection on the data to be inspected selected by user A. In this way, the terminal device can receive the data quality inspection results returned by the server (or other data processing device) and display the quality inspection percentage obtained by quality inspection analysis of these data quality inspection results on the quality inspection display platform of the terminal device. It can be understood that the other data processing device can be a server or a terminal device, and there is no limitation here.

[0120] For further understanding, please see Figure 2 , Figure 2 It is a schematic diagram of a scenario in which a quality inspection model is trained through limited labeled training samples provided in an embodiment of the present application. It should be understood that in an embodiment of the present application, in order to quickly and conveniently obtain the data from the business database (for example, Figure 2 In the database Q1 shown in FIG, a large number of safe and reliable desensitized training samples are obtained, and the target business model under the corresponding business scenario (for example, Figure 2 The embodiment of the present application proposes that any text data that needs to be added to the database Q1 can be subjected to data desensitization processing (e.g., watermark removal processing, noise removal processing, etc.) by means of data cleaning and data quality inspection, and then the desensitized text data can be added to the database Q1 as model training sample data. In this way, the model training sample data added to the database Q1 can be extracted as training samples for training the business model A during the training process of the business model A.

[0121] It can be understood that the target business model here (for example, business model A) can be a large model in which the parameter quantity of the model parameters reaches a parameter quantity threshold (for example, tens of billions of parameters). For example, the business model A here may include but is not limited to a large language model (Large Language Model, abbreviated as LLM) for understanding and generating natural language. For ease of understanding, the embodiments of the present application may collectively refer to the LLM model as the above-mentioned language business model.

[0122] Among them, it can be understood that the above-mentioned data desensitization processing includes using a trained data cleaning model to perform data cleaning on each text data (i.e., data to be cleaned, where the data to be cleaned can be the above-mentioned sample data to be cleaned) obtained from the business database (i.e., the first business database, for example, the first business database can be an incremental database in the business database), and then each text data after data cleaning (i.e., the above-mentioned cleaned data) can be used as data to be inspected for input into the trained target quality inspection model. In this way, after data quality inspection is performed on these data to be inspected through the trained target quality inspection model, the data quality inspection results of each data to be inspected can be obtained, and then the data to be inspected that meet the quality inspection data strategy indicated by the quality inspection task and are screened out from these data to be inspected based on these data quality inspection results can be used as the above-mentioned model training text data.

[0123] It should be understood that the data cleaning model is a business model for implementing two-stage data cleaning, determined based on the trained target recall model and target cleaning model. The target recall model and target cleaning model are obtained by training the initial recall model and initial cleaning model, respectively, based on the sample data to be cleaned. It should be understood that the sample data to be cleaned is obtained by performing data distillation on the input text data based on the text quality inspection results output by the target quality inspection model.

[0124] It is understood that the target quality inspection model involved in the embodiment of the present application can be Figure 2 The quality inspection model 2b shown. The quality inspection model 2b can be Figure 2 The steps S11 to S16 shown in FIG. 1 are obtained by training the initial business model. For example, the initial business model here can be Figure 2 The business model 2a shown may be any trained large model obtained by a computer device from a model database.

[0125] For example, the embodiment of the present application can adopt any one of the open source / internal 7B large model bases such as LLaMA, Qwen, and Hunyuan as the initial business model, and use the manually calibrated training data on the basis of the 7B large model base to fine-tune the model of the 7B large model base, so that the 7B large model has the ability to extract text quality inspection fragments (for example, one or more of the above-mentioned watermark fragments, noise fragments, illegal data fragments, semantic fragments, and format fragments) from a piece of text data. Among them, the Hunyuan model is an internally developed Hunyuan large model, which can realize intelligent question and answer in the question and answer scenario (for example, the Hunyuan large model can be used to extract and summarize the content summary of each resource segment in the document resource), and can also realize intelligent recommendation of advertising data in the advertising business scenario.

[0126] Specifically, such as Figure 2 As shown, the computer device (for example, the server mentioned above) can execute step S11 to obtain Figure 2 The database Q1 (i.e., the second business database) shown here acquires a training text dataset for training business model A (i.e., the aforementioned language business model). The acquired training text dataset can be collectively referred to as the text data to be processed. For ease of understanding, the example of 90,000 pieces of text data to be processed is used here. In this case, the number of text data to be processed is 90,000.

[0127] It should be understood that the training text data set here can be all or part of the training text data under different data types obtained from the database Q1. The different data types here can include but are not limited to test question data types, web page data types and encyclopedia data types. It should be understood that in the embodiment of the present application, a data type can include multiple data sets, a data set can correspond to a data identifier, and a data set corresponding to a data identifier can include multiple text data. For example, when the data type includes a test question data type, the test question data set can include test question data sets of different subjects, and the subject identifier corresponding to the test question data set of a subject is the aforementioned data identifier. For example, the test question data sets of multiple subjects can specifically include a Chinese test question data set, a mathematics test question data set, an English test question data set, etc., and the test question data sets of each subject under the test question data type will not be listed one by one here.

[0128] Furthermore, a computer device (eg, a server) may Figure 2 The quality inspection definition component 21a shown executes step S12 to perform quality inspection definition on the text data to be processed to obtain text definition information of the text data to be processed. It should be understood that in the embodiment of the present application, the text definition information here can be further input into Figure 2The quality inspection and calibration component 22a shown can distribute the pre-defined text definition information to different quality inspection and marking terminals, so that different quality inspection and marking terminals can subsequently perform text annotation on subsequently acquired text data (for example, sampled text data) based on the pre-defined text definition information.

[0129] Among them, Figure 2 As shown, the computer device (for example, a server) can execute step S13 through the sampling component 23a, and then can sample the text data to be processed (for example, 90,000 text data) in units (large unit granularity) of data types (for example, the above-mentioned test data types, web page data types, and encyclopedia data types). For example, each data type of the text data to be processed can be uniformly sampled to extract text data of each data type with a sampling quantity of a preset threshold value from the text data to be processed as sampled text data. For example, 3,000 text data of the test data type, 3,000 text data of the web page data type, and 3,000 text data of the encyclopedia data type can be extracted as Figure 2 The sampled text data shown, that is, the number of the sampled text data here, may be 9,000 (ie, 9,000).

[0130] Furthermore, in order to improve the efficiency and quality of data annotation, the computer device (e.g., a server) can execute step S14 through the quality inspection and calibration component 22a to receive multiple text annotation information returned by each quality inspection terminal for the same sampled text data through manual quality inspection, and then the received multiple text annotation information can be cross-checked to ensure that the different text annotation information obtained by each quality inspection terminal after data annotation of the same sampled text data based on pre-defined text definition information can be consistent. For example, if the cross-check result obtained by the cross-check indicates that these text annotation information are consistent, the computer device can determine the text tag information currently obtained based on the text definition information as the text tag information of the sampled text data.

[0131] Further, such as Figure 2 As shown, a computer device (e.g., a server) can execute step S15 through the prompt template construction component 24a to construct a text prompt template associated with the text definition information through the text definition information (the text prompt template here carries the text key fields in the text definition information and the text attributes of the text key fields), and then can fill in the text prompt template based on the text tag information of the currently acquired sampled text data to obtain a tag prompt template corresponding to the text prompt template.

[0132] Further, such as Figure 2 As shown, the computer device (for example, the server) can execute step S16 through the quality inspection model training component 25a, so that after obtaining the current business model 2a to be trained (that is, the above-mentioned initial business model) from the above-mentioned model database, the business model 2a can be trained (that is, model fine-tuning processing is performed on the business model 2a) through limited annotated sample data (that is, sampled text data carrying text tag information, and the sampled text data here specifically refers to text data extracted from massive text data to be processed and annotated with data, for example, the model fine-tuning processing here specifically refers to parameter optimization of the model parameters of the business model 2a), and then the business model 2a obtained after the model training can be used as the target quality inspection model (for example, Figure 2 The quality inspection model shown in 2b).

[0133] Furthermore, it is understood that the computer device (eg, server) verifies the target quality inspection model (eg, Figure 2 The quality inspection model 2b) shown in the figure shows the quality inspection accuracy when performing the quality inspection task, and it is proposed that it can be further improved by Figure 2 The sampling component 23a shown executes step S17, and can then perform more fine-grained sampling processing on the text data to be processed in units of data sets (i.e., small unit dimensions). For example, for ease of understanding, here, the number of multiple data sets under the same data type is 4 as an example. At this time, the computer device (e.g., a server) can uniformly sample different data sets under the same data type to extract a fixed number (e.g., 100) of text data from each of the 4 different data sets under the same data type as input text data.

[0134] Further, such as Figure 2 As shown, a computer device (e.g., a server) can input the input text data into the quality inspection model 2b, and then execute step S18 through the quality inspection model 2b. In this way, when the computer device performs the quality inspection task on the input text data through the quality inspection model 2b, the quality inspection model 2b can perform data quality inspection on the input text data to identify whether there is a text segment matching the current quality inspection task in the current input text data (i.e., determine whether the input text data carries a text quality inspection segment). If so, the quality inspection model 2b can output the text quality inspection result of the input text data according to the text output style indicated by the mark prompt template.

[0135] It should be understood that in an embodiment of the present application, if the above-mentioned quality inspection task includes a multidimensional quality inspection task of different quality inspection dimensions, for example, the multidimensional quality inspection task here can specifically include the above-mentioned watermark quality inspection task, noise quality inspection task, illegal data quality inspection task, semantic quality inspection task and format quality inspection task, then the quality inspection model 2b obtained by the current training can be a multidimensional quality inspection model corresponding to the multidimensional quality inspection task, and the multidimensional quality inspection model can include multiple quality inspection sub-models. For example, the multiple quality inspection sub-models here can specifically include a watermark quality inspection model for performing a watermark quality inspection task, a noise quality inspection model for performing a noise quality inspection task, an illegal data quality inspection model for performing an illegal data quality inspection task, a semantic quality inspection model for performing a semantic quality inspection task and a format quality inspection model for performing a format quality inspection task.

[0136] In this way, for the input text data in the input quality inspection model 2b, the multidimensional quality inspection model (for example, the watermark quality inspection model, the noise quality inspection model, the illegal data quality inspection model, the semantic quality inspection model and the format quality inspection model) can be used to perform data quality inspection on the input text data in different quality inspection dimensions to obtain the data quality inspection results of the same input text data in different quality inspection dimensions. Furthermore, the data quality inspection results of the same input text data in different quality inspection dimensions (for example, the watermark quality inspection result predicted and output by the watermark quality inspection model, the noise quality inspection result predicted and output by the noise quality inspection model, the illegal data quality inspection result predicted and output by the illegal data quality inspection model, the semantic quality inspection result predicted and output by the semantic quality inspection model and the format quality inspection result predicted and output by the format quality inspection model) can be collectively referred to as the text quality inspection result of the input text data.

[0137] Among them, it should be understood that the data quality inspection involved in the embodiment of the present application refers to the process of pre-quality inspection of the training samples for training business model A (i.e., the above-mentioned language business model). This is because the business performance and output quality of business model A (i.e., the above-mentioned language business model) in the corresponding business scenario depend to a large extent on the data quality of the input sample data. Therefore, before obtaining the sample data for training business model A (i.e., the above-mentioned language business model), the embodiment of the present application needs to use the text quality inspection result output by the target quality inspection model obtained by the current training to perform data distillation on the input text data, so as to obtain text data carrying text quality inspection fragments from the input text data as sample data to be cleaned, and then the sample data to be cleaned can be used to train other business models (e.g., the above-mentioned initial recall model and the initial cleaning model) whose parameter amount of model parameters is smaller than that of the target quality inspection model, so as to train the target recall model for data recall and the target cleaning model for data cleaning.

[0138] For example, the embodiment of the present application can perform two-stage data cleaning on any text data currently acquired (for example, the text data to be cleaned under a certain data set acquired from the above-mentioned first business database) through a data cleaning model composed of a target recall model and a target cleaning model, so as to accurately filter out the target text data to be cleaned that needs to be cleaned from the text data to be cleaned through the data cleaning model, and perform data desensitization processing on the text data to be cleaned.

[0139] It can be understood that the embodiment of the present application can perform binary classification processing on the text data to be cleaned through the target recall model in the data cleaning model to determine whether the text data to be cleaned carries text quality inspection fragments that need to be cleaned. If the judgment is yes, it is necessary to screen out the text data carrying the text quality inspection fragments from the text data to be cleaned as the target text data to be cleaned, and then the target cleaning model in the data cleaning model can be used to accurately locate the text quality inspection fragments in the target text data to be cleaned, so as to perform data cleaning on the accurately located text quality inspection fragments (for example, the text quality inspection fragments in the target text data to be cleaned can be removed), so that the target text data to be cleaned after data cleaning can be used as the data to be inspected after data desensitization, and then input into the target quality inspection model again for data quality inspection, so that the target quality inspection model can be used to screen out the data to be inspected that meets the data quality inspection strategy from the data to be inspected as the above-mentioned training samples, so as to improve the sample data quality of the training samples obtained for business model A (that is, the above-mentioned language business model). In this way, when business model A (that is, the above-mentioned language business model) is trained through the training samples, the business performance and output quality of the trained business model can be effectively improved.

[0140] It should be understood that data cleaning of the text data to be processed can be understood as optimizing the quality of the problem text data (i.e., cleaning mark segments) in the target text data to be cleaned through the target cleaning model, or deleting the problem text data (i.e., cleaning mark segments) in the target text data to be cleaned, which is not limited here. The problem text data (i.e., cleaning mark segments) refers to the text quality inspection segments that are identified and marked from the target text data to be cleaned through the target recall model and match the current quality inspection task.

[0141] Thus, in the embodiment of the present application, any one or more of the multiple quality inspection sub-models included in the multi-dimensional quality inspection model can be used to implement data quality inspection on the data to be inspected obtained in units of data sets, thereby achieving diversity in data quality inspection during the process of obtaining training samples. In addition, when performing data quality inspection, data quality inspection can be performed separately on multiple quality inspection dimensions to improve the accuracy of data quality inspection on the data to be inspected, thereby ensuring the sample quality of the training samples determined from the data to be inspected.

[0142] Among them, the key points of the embodiments of the present application may specifically include: for different quality inspection dimensions, a quality inspection sub-model for performing corresponding quality inspection tasks can be trained and fine-tuned based on the open source large model, and the text quality inspection results associated with different quality inspection tasks can be accurately output through the corresponding quality inspection sub-model, and then the input text data can be distilled based on the text quality inspection results to obtain sample data to be cleaned for training the recall model and the cleaning model. In this way, after the recall model and the cleaning model are respectively trained with the sample data to be cleaned, a data cleaning model for realizing two-stage data cleaning can be obtained through the trained recall model (i.e., the target recall model) and the trained cleaning model (i.e., the target cleaning model). In addition, the embodiment of the present application can also automatically realize the automated data cleaning and data quality inspection of the text data to be cleaned (for example, the data to be cleaned sampled from the training sample data sets corresponding to different data sets) through the data cleaning model and the target quality inspection model, and then after the quality inspection analysis of the data quality inspection results of the data to be inspected in the text data to be cleaned, the quality inspection ratio of each data set in the same data to be inspected under a certain quality inspection dimension can be obtained in units of data sets, so that the text data in the data sets that meet the quality inspection data strategy can be found from these data sets through the quality inspection ratio of each data set under a certain quality inspection dimension, as the model training sample data for training the above-mentioned business model A. Compared with the existing manual quality inspection method, the embodiment of the present application can fine-tune the initial business model through effectively labeled sample data. In this way, in the training process of the quality inspection model, not only can the training cost of the target quality inspection model be minimized, but the efficiency of the target quality inspection model can also be improved. Then, when the quality inspection task is performed by the target quality inspection model, the quality inspection efficiency and quality of the data quality inspection of the quality inspection data can be improved.

[0143] It can be seen from this that the embodiment of the present application can utilize the algorithm model (i.e., the target quality inspection model) to realize automated data quality inspection, which can not only effectively reduce the time cost and manpower cost of data quality inspection for input text data, but also ensure the efficiency of data quality inspection, such as each A10 card (i.e., one of the graphics processors) on the computer equipment can process 10+ text data per second. In addition, in the process of data quality inspection by the target quality inspection model, the accuracy and recall rate of the model quality inspection are much higher than those of manual and rule-based (accuracy and recall rate are 80%+), and it can be well applied to different data sources. In addition, the embodiment of the present application can ensure that the training samples finally obtained for training business model A have higher data quality through data cleaning and data cleaning, thereby improving the model convergence speed when the business model A is trained, thereby reducing training time, which is particularly important for large models, because by reducing training time, not only can training costs be reduced, but resource utilization can also be improved. In addition, when business model A is trained using high-quality training samples, it can not only help improve the accuracy, generalization ability and stability of the trained business model A, but also generate more accurate and valuable outputs through the trained business model A, thereby improving user experience. This can ensure that the data results output by the trained business model A can be as compliant as possible, thereby avoiding legal and ethical risks caused by data quality issues.

[0144] It can be understood that the trained business model A involved in the embodiment of the present application (i.e., the trained language business model, which can also be referred to as the trained target business model) can be applied to the field of artificial intelligence technology. For example, when the text data currently obtained after desensitization processing is a document resource, the embodiment of the present application can input the document resource into the trained business model A (for example, business model A'), and then the trained business model A (for example, business model A') can be used to extract the content summaries of the multiple resource fragments of the document resource, and the trained business model A (for example, business model A') can summarize the content summaries of the multiple resource fragments to predict and output the summary summary of the document resource.

[0145] Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive field within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. AI technology is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI domains. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0146] It is understandable that the initial business model in the embodiment of the present application can be a pre-training model (Pre-training model), also known as a cornerstone model (or base model). Among them, it is understandable that as the initial business model of the large model, it refers to a deep neural network (DNN) with massive model parameters (i.e., large parameters). The present application can use limited labeled sample data and combine fine tuning (fine tune), parameter efficient fine tuning (PEFT) and prompt-tuning technologies to achieve model training of the large-parameter DNN network model, so that the large-parameter DNN network model can have the business capability to perform the above-mentioned quality inspection tasks.

[0147] Among them, it can be understood that in an embodiment of the present application, the business model deployed in the model database (for example, a 7B large model base) may include at least: a language model (LLaMA, Qwen, ELMO, BERT, FastTest, GPT), a visual model (swin-transformer, ViT, V-MOE), a speech model (VALL-E), a multimodal model (ViBERT, CLIP, Flamingo, Gato) and a mixed large model, etc., wherein a multimodal model refers to a model that establishes two or more data modal feature representations. The pre-trained model is an important tool for outputting artificial intelligence generated content (AIGC), and can also be used as a general interface for connecting multiple specific task models. For example, in one or more implementable methods, the target business model involved in this application (for example, the above-mentioned model A) can be a mixed large model, and the initial business model involved in the embodiment of the present application (for example, the above-mentioned business model 2a) can be an LLaMA business model (i.e., a speech business model, for example, specifically an LLaMA13B large model). It should be understood that in the above-mentioned intelligent question and answer scenario, if the text data entered by the business object (for example, user U1) is a document resource, if the content data volume of each slice content of the document resource is large, that is, the content of each slice content is long, then a large language model of 6B or more is required to extract the summary of the slice content. For example, based on the text processing method involved in the embodiment of the present application, training samples for training the large model can be obtained, and then the LLaMA 13B large model for summary extraction and summary summary can be specially trained with the obtained training samples, so that in the subsequent intelligent question and answer scenario, the summary summary of the document resource that does not carry sensitive data can be safely and reliably output to the user in the form of a conversation message.

[0148] It should be noted that before collecting relevant user data and during the process of collecting relevant user data (such as obtaining training text data related to the user, for example, text data under the above-mentioned test question data type, text data under the web page data type, and text data under the encyclopedia data type), this application can display a prompt interface, pop-up window, or output a voice prompt message. The prompt interface, pop-up window, or voice prompt message is used to remind the user that its relevant data is currently being collected, so that this application only starts to execute the relevant steps of obtaining user-related data after obtaining the user's confirmation operation on the prompt interface or pop-up window. Otherwise (that is, when the user's confirmation operation on the prompt interface or pop-up window is not obtained), the relevant steps of obtaining user-related data are terminated, that is, the user's relevant data is not obtained. In other words, all user data collected by this application are collected with the user's consent and authorization, and the collection, use, and processing of relevant user data must comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0149] It is understood that the above scenarios are merely examples and do not limit the application scenarios of the technical solutions provided in the embodiments of this application. The technical solutions of this application can also be applied to other scenarios. For example, those skilled in the art will appreciate that with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application will also be applicable to similar technical problems.

[0150] Further, see Figure 3 , Figure 3 This is a flow chart of a text data processing method provided by an embodiment of the present application. The method can be executed by a computer device, where the computer device can be a server, for example, the server can be the above-mentioned Figure 1 Server 100a in. Figure 3 As shown, the method may at least include the following steps S101 to S105.

[0151] Step S101, determining text definition information of the text data to be processed based on the quality inspection task associated with the text data to be processed;

[0152] The text definition information is determined based on the text key fields and text attributes of the text key fields in the text data to be processed;

[0153] Specifically, the server can obtain a training text data set for training the target business model, and use each training text data in the training text data set as the text data to be processed; the target business model is a business model associated with the target quality inspection model; further, the server can obtain the quality inspection task associated with the training text data set; further, the server can perform quality inspection definition on the text data to be processed based on the quality inspection task and the data type of the text data to be processed, and obtain the text key fields of the text data to be processed and the text attributes of the text key fields; further, the server can use the text key fields and the text attributes of the text key fields as text definition information of the text data to be processed.

[0154] It can be understood that the embodiment of the present application can perform quality inspection definitions of different quality inspection dimensions on the text data to be processed in the currently acquired training text data set based on different quality inspection tasks to obtain text definition information under different quality inspection dimensions.

[0155] For further understanding, please see Figure 4 , Figure 4 This is a schematic diagram of a scenario for quality inspection definition provided by an embodiment of the present application. Figure 4The user terminal 40a shown may be a sample data management terminal. The display page (eg, quality inspection management page) corresponding to the sample data management terminal may display the data generated by the business object (eg, Figure 4 The quality inspection tasks entered or selected by the user U1) shown require sensitive data identification. For example, the quality inspection tasks here may include the above-mentioned watermark quality inspection tasks, noise quality inspection tasks, illegal data quality inspection tasks, semantic quality inspection tasks, and format resource tasks.

[0156] It should be understood that in one implementable method, user U1 can select the quality inspection task that needs to be quality inspection defined as the target quality inspection task from the quality inspection tasks displayed on the quality inspection management page, and then enter the manually mined data obtained by data mining text data of different data types in the text definition area corresponding to the target quality inspection task. For example, user U1 can mine email, social account information and other data related to the quality inspection definition details from the text data under the web page data type.

[0157] It should be understood that, in one practicable manner, the user terminal 40a may use the data expressed in natural language entered by the user U1 as watermark definition information.

[0158] Alternatively, as Figure 4 As shown, in another possible implementation, when user U1 triggers a confirmation operation for a quality inspection task, the user terminal 40a can respond to the confirmation operation and send the quality inspection task to the server 40b, so that the server 40b can perform a quality inspection definition of the corresponding quality inspection dimension based on the quality inspection task. It should be understood that at this time, the task data of the quality inspection task can specifically include the quality inspection task (for example, a watermark quality inspection task) specified by the user U1 that needs to be defined for quality inspection, and can also include the artificial mining data obtained by the user U1 performing data mining on text data of different data types. At this time, when the server 40b performs a quality inspection definition based on the quality inspection task, it can further generate and output definition information with the same or similar semantics as these artificial mining data based on the current quality inspection task through the large model (for example, the mixed-element large model), and then the generated definition information carrying the artificial mining data can be used as text definition information expressed in natural language.

[0159] It can be understood that the embodiments of the present application can define quality inspections of different dimensions for the training samples commonly used for pre-training the target business model (for example, the mixed-element large model obtained from the above-mentioned model database) according to the data types of the training samples commonly used for pre-training the large model (for example, web page data types, encyclopedia data types, test question data types, etc.).

[0160] Specifically, such as Figure 4 As shown, the server 40b can obtain the data from the business database (for example, Figure 2 The database Q1) shown in FIG1 obtains training samples corresponding to different data types for training the target business model, which are collectively referred to as training text data sets associated with the quality inspection task. It should be understood that a training sample is a piece of training text data.

[0161] Among them, the training text dataset here can include Figure 4 The training text dataset 4a, training text dataset 4b, and training text dataset 4c are shown. The training text dataset 4a is a training sample corresponding to data type D1 (e.g., test question data type) obtained from the business database; the training text dataset 4b is a training sample corresponding to data type D2 (e.g., web page data type) obtained from the business database; and the training text dataset 4c is a training sample corresponding to data type D3 (e.g., encyclopedia data type) obtained from the business database.

[0162] Furthermore, it is understood that the server 40b can refer to the obtained batch of training text data sets as the to-be-processed text data during the quality inspection definition process, and perform quality inspection definitions under different quality inspection dimensions on the batch of to-be-processed text data based on the quality inspection task, so as to pre-define the obtained Figure 4 The text definition information of the same batch of text data to be processed under different quality inspection dimensions is shown. The text definition information here can include Figure 4 Shown are watermark definition information 41a, noise definition information 42a, illegal data definition information 43a, semantic definition information 44a, and format definition information 45a.

[0163] Among them, Figure 4 As shown, in the embodiment of the present application, when the quality inspection task includes a watermark quality inspection task, watermark definition can be performed on the currently acquired text data to be processed of different data types to obtain watermark definition information of the text data to be processed under the watermark quality inspection dimension (for example, Figure 4 The watermark definition information 41a shown in FIG. 4 is further described. Figure 4 As shown, the user terminal 40a can receive the text definition information returned by the server 40b. In this way, in the process of data management of the training samples, the user terminal 40a can further send the text definition information to multiple quality inspection and annotation terminals used for manual quality inspection to ensure the data reliability of the training samples obtained subsequently, so that multiple quality inspection standard terminals can obtain text annotation information of the sampled text data based on the text definition information.

[0164] Specifically, it can be understood that the server 40b can obtain the watermark quality inspection strategy associated with the watermark quality inspection task, and can obtain the first target data type to be watermarked from the data type of the text data to be processed based on the watermark quality inspection strategy. For example, the first target data type here can refer to a data type traversed and obtained from the above-mentioned multiple data types (for example, test question data type, web page data type and encyclopedia data type). For ease of understanding, the first target data type is taken as the test question data type as an example; further, the server 40b can use the text data to be processed that matches the first target data type (for example, test question data type) as the first target text data in the text data to be processed, and then define the watermark for the first target text data to obtain the watermark definition field of the first target text data, and can use the watermark definition field as the text key field of the text data to be processed; further, the server 40b can classify the watermark definition field into the watermark type indicated by the watermark quality inspection strategy to determine the watermark type as the text attribute of the text key field.

[0165] It should be understood that the specific implementation method of the server 40b defining watermarks for text data under other data types can refer to the description of the specific process of defining watermarks for text data under the test question data type, which will not be further described here.

[0166] In the embodiment of the application itself, the watermark definition refers to the watermark data (i.e., watermark fragments) with watermark problems in this batch of text data to be processed, which can be roughly divided into two major watermark types. One is the first watermark type associated with personal privacy (i.e., personal sensitive data) (for example, personal privacy data class), and the other refers to the second watermark type associated with source information (i.e., source data) (for example, source data class).

[0167] Among them, personal privacy (i.e., personal sensitive data) refers to the data identifier (i.e., data identity document, referred to as data ID) or descriptive data content that can be obtained from text data and used to directly or indirectly locate and identify a business object (e.g., a person, an object, or a virtual object). For example, the data ID here may include but is not limited to the social account information of a person or a virtual object obtained from the text data (e.g., Weibo ID, WeChat account, mobile phone number, QQ number, email address, etc.). The descriptive data content here may include but is not limited to the account information of a person or a virtual object obtained from the text data (e.g., online nickname, private address, blog name, contact website, etc.).

[0168] The watermark definition fields defined in the first watermark type specifically include the following personal sensitive data:

[0169] 1. Need to label real names in text data

[0170] ● Label the names of minors in text data, including text data in self-media reports, text data in school publications, signatures in primary and secondary school essays, and other text data.

[0171] ●For medical / legal consultation questions and answers, the name of the doctor / lawyer that appears.

[0172] ●All names in legal documents must be marked.

[0173] No need to mark:

[0174] ●x is the surname, such as Mr. x / Ms. x / Uncle x, etc.

[0175] ●The author / translator / editor of a publication / the signature of a particular article;

[0176] ● Names of people in the main text of a publication;

[0177] ●The names of the protagonists in novels / games are fictitious names and do not need to be marked;

[0178] ● Names come from the introduction of the subject matter (e.g., the name of the legal representative when introducing a company / institution / organization, or the full name of the teacher when introducing a course), heads of public institutions, public figures, and protagonists of regular news reports (not self-media).

[0179] 2. Contact information in text data needs to be annotated: QQ (group) number, WeChat ID, various social media IDs, phone numbers (except public utility numbers with less than 6 digits), email address, etc.;

[0180] 3. Personal IDs in text data need to be labeled: police ID, doctor ID, practice ID, bank card number, license plate number, real estate ID, ID card number, external IP address, computer MAC address, etc.

[0181] 4. Private addresses and dates of birth in text data need to be annotated (except for searchable public figures). Addresses at the city level need to be annotated.

[0182] Among them, source information (i.e. source data) refers to the data content obtained from the text data that can identify the text source of the current text data, such as website URL, website name, public account name, reporter name, etc. (excluding text subject).

[0183] The watermark definition fields defined in the second watermark type specifically include the following personal sensitive data:

[0184] 1. Identifiable source information such as web links, website names, organization names, and public account names.

[0185] 2. The source of the news (newspaper, media, etc.), including the name of the editor and reporter.

[0186] Exceptions:

[0187] 1. The institution / unit / company of the publication is not considered as source information (such as novels / books / literary works / classical Chinese texts / songs / ancient poems / modern poems / research reports / papers / patents / journals, etc.)

[0188] 2. If the information comes from the questions or answers in the test paper / exam questions / textbook, it does not need to be marked.

[0189] 3. Organizations / groups / institutions / schools publish their own news without labeling

[0190] 4. References, such as mentioning the original author / where the original article was published / where it was selected from, are not considered watermarks as long as the site or URL is not disclosed.

[0191] For another example, when the quality inspection task includes a noise quality inspection task, the embodiment of the present application can perform noise definition on the currently acquired text data of different data types to be processed to obtain noise definition information of the text data to be processed under the noise quality inspection dimension (for example, Figure 4 Noise definition information 42a) is shown.

[0192] Specifically, the server 40b can obtain the noise quality inspection strategy associated with the noise quality inspection task, and obtain the second target data type to be noise-defined from the data type of the text data to be processed based on the noise quality inspection strategy; for example, the second target data type here can refer to a data type traversed and obtained from the above-mentioned multiple data types (for example, test question data type, web page data type and encyclopedia data type). For ease of understanding, the second target data type is also taken as the test question data type as an example; at this time, the server 40b can use the text data to be processed that matches the second target data type (for example, test question data type) as the second target text data in the text data to be processed, and then perform noise definition on the second target text data to obtain the noise definition field of the second target text data, and can use the noise definition field as the text key field of the text data to be processed; further, the server 40b can classify the noise definition field into the noise type indicated by the noise quality inspection strategy, and determine the noise type as the text attribute of the text key field.

[0193] It should be understood that the specific implementation method of the server 40b to define noise for text data under other data types can refer to the description of the specific process of defining noise for text data under the test question data type, which will not be further described here.

[0194] Among them, the noise types can be roughly divided into three categories (i.e., the first noise type, the second noise type, and the third noise type). It should be understood that the noise definition field under the first noise type includes other text data that is not related to the text data body (for example, page numbers and page turning buttons that are not related to the text); the noise definition field under the second noise type includes text annotation data related to the display abnormality of multimedia data (for example, picture data) (for example, missing pictures); the noise definition field under the third noise type contains advertising data that is not related to the text theme of the multimedia data.

[0195] By analogy, in the embodiment of the present application, when the quality inspection task includes an illegal data quality inspection task, illegal data definition can be performed on the currently acquired text data to be processed of different data types to obtain illegal data definition information of the text data to be processed under the illegal data quality inspection dimension (for example, Figure 4 Similarly, when the quality inspection task includes a semantic quality inspection task, the embodiment of the present application can perform semantic definition on the currently acquired text data of different data types to be processed, so as to obtain semantic definition information of the text data to be processed under the illegal data quality inspection dimension (for example, Figure 4 Similarly, when the quality inspection task includes a format quality inspection task, the embodiment of the present application can perform format definition on the currently acquired text data of different data types to be processed, so as to obtain format definition information of the text data to be processed under the illegal data quality inspection dimension (for example, Figure 4 Format definition information 45a) shown.

[0196] Among them, the watermark definition information includes a watermark definition field and a watermark type to which the watermark definition field belongs. It can be understood that if the text data to be processed is text data extracted from a web page containing a watermark (such as a watermark for indicating the provider information of the web page), then the text data to be processed obtained from the business database is also likely to include the web page watermark in the web page; for example, if the text data to be processed contains text data extracted from image data containing a watermark (such as a watermark for indicating the provider information of the image), then the text data to be processed may include the image watermark in the image data. Based on this, the embodiment of the present application proposes that all watermark definition fields in the text data to be processed can be pre-defined and detected by means of watermark definition, and the detected watermark definition fields can be divided into the watermark types indicated by the watermark quality inspection strategy, so that the watermark quality inspection model for intelligent watermark recognition can be constructed in the future by assisting with the pre-defined watermark definition information.

[0197] Among them, the noise definition information includes the noise definition field and the noise type to which the noise definition field belongs. For example, when the text data to be processed is text data extracted from a web page, when extracting text data from the web page, information in the web page that is not related to the content of the web page may also be extracted as text data, such as descriptive information in the web page similar to "click to return to the previous page" and "exit the web page". Based on this, the embodiment of the present application proposes that all noise definition fields in the text data to be processed can be pre-defined and detected by noise definition, and the detected noise definition fields can be divided into the noise types indicated by the noise quality inspection strategy, so that the noise definition information obtained in advance can be used to assist in the construction of a noise quality inspection model for intelligent noise recognition.

[0198] The illegal data definition fields in the illegal data definition information 43a may be sensitive data that may negatively impact an individual, organization, or institution when leaked, used without authorization, or abused. For example, the sensitive data may be illegal data of an individual, organization, or institution, and the like, which is not limited here. Similarly, the embodiments of the present application propose that all illegal data definition fields in the text data to be processed can be pre-defined and detected through illegal data definition, and the detected illegal data definition fields can be classified into illegal data types indicated by the illegal data quality inspection strategy, so that an illegal data quality inspection model for intelligent illegal data identification can be constructed subsequently using the pre-defined illegal data definition information.

[0199] The semantic definition fields in the semantic definition information 44a may be stored in Figure 4The meaning expressed in the content of the training data in the business database shown. Therefore, the semantic quality inspection dimension, which is the dimension that needs to detect semantic errors in the training data, indicates that if semantic errors are detected in the training data, such as incoherent or incomplete sentences in the training data, the training data fails the semantic quality inspection dimension, that is, the training data is problematic training data in this semantic quality inspection dimension.

[0200] The format definition field in the format definition information 45a may be stored in Figure 4 The format of the data content in the training data in the business database shown (such as the punctuation and typesetting format of the training data). For example, when the training data is text data, the format can be used to indicate the punctuation and paragraph typesetting in the text data. For another example, when the training data is image data, the format can be the punctuation and paragraph typesetting involved in the text in the image data. For another example, if the training data is a piece of code data, the format can be a code format.

[0201] Step S102, obtaining sampled text data from the text data to be processed, and determining text tag information of the sampled text data based on the text definition information;

[0202] The text tag information is determined based on the sampling key field in the sampled text data and the text attributes of the sampling key field; the sampling key field is a field in the text key field;

[0203] It is understood that when executing step S102, the server may sample the text data to be processed based on the data type of the text data to be processed (e.g., the test question data type, web page data type, or encyclopedia data type) to obtain a sample data set that matches the data type of the text data to be processed. Further, the server may determine the sampled text data based on the sample data set that matches the data type of the text data to be processed.

[0204] For ease of understanding, the embodiment of the present application takes the data type of the text data to be processed including the first data type and the second data type as an example to illustrate the specific process of sampling the text data to be processed. Specifically, the server can determine the first data type from the data type of the text data to be processed, and extract N text data matching the first data type from the text data to be processed, and then use the extracted N text data matching the first data type as the first sampling data set; N is a positive integer; further, the server can determine the second data type from the data type of the text data to be processed, extract N text data matching the second data type from the text data to be processed, and use the extracted N text data matching the second data type as the second sampling data set; further, the server can determine the sampling data set that matches the data type of the text data to be processed based on the first sampling data set and the second sampling data set.

[0205] For example, in the embodiment of the present application, the test question data type (ie, the first data type) can be used as a sampling unit in the text data to be processed, and the training text data set corresponding to the test question data type (eg, the above Figure 4 3000 text data are extracted from the training text data set 4a) shown in the figure and added to the first sampling data set. At the same time, the embodiment of the present application can also use the web page data type (i.e., the second data type) as the sampling unit in the text data to be processed, and extract 3000 text data from the training text data set corresponding to the web page data type (e.g., the above Figure 4 3000 text data are extracted from the training text data set 4b) shown and added to the second sampling data set.

[0206] It should be understood that in the embodiment of the present application, optionally, the second data type may also be the above-mentioned encyclopedia data type. Similarly, the server may use the encyclopedia data type (i.e., another second data type) as a sampling unit in the text data to be processed, and select the training text data set corresponding to the encyclopedia data type (e.g., the above-mentioned encyclopedia data type). Figure 4 3000 text data are extracted from the training text data set 4c) and added to another second sampling data set. It should be understood that the embodiment of the present application does not limit the specific data types of the first data type and the second data type.

[0207] It is understandable that the embodiment of the present application can be based on the first sampling data set and the second sampling data set, which are sampling data sets that match the data type of the text data to be processed, and the text data in these sampling data sets can be collectively referred to as sampled text data.

[0208] Furthermore, the server can distribute the sampled text data to multiple quality inspection marking terminals associated with the quality inspection task, and one quality inspection marking terminal is used to perform text annotation on the sampled text data based on the text definition information, and obtain a text annotation information for the sampled text data from the quality inspection marking terminal; further, the server can receive the text annotation information returned by each quality inspection marking terminal for the sampled text data, cross-check the received text annotation information, and obtain a cross-check result for the sampling key fields in the sampled text data and the text attributes of the sampling key fields; further, if the cross-check result indicates that the received text annotation information is consistent, the server can determine the text marking information of the sampled text data based on the sampling key fields and the text attributes of the sampling key fields.

[0209] For ease of understanding, the present application embodiment takes the text definition information including watermark definition information as an example to illustrate the specific process of obtaining text mark information of sampled text data. Figure 5 , Figure 5 This is a schematic diagram of a scenario in which text mark information of sampled text data is determined based on watermark definition information provided by an embodiment of the present application. Figure 5 The text data to be processed shown can be the above Figure 4 The text data to be processed is shown. It is understood that the sampled text data here can be text data of a specified number threshold (for example, 3000) uniformly extracted from multiple training text data sets included in the text data to be processed, with data type as the unit (i.e., sampling unit).

[0210] The specific process of the server performing watermark definition on the text data to be processed to obtain the watermark definition information of the text data to be processed can refer to the specific process of obtaining the watermark definition information 41a mentioned above, which will not be further described here.

[0211] It should be understood that in the embodiment of the present application, in order to improve the efficiency of data tagging of sampled text data, the embodiment of the present application proposes that before performing intelligent tagging, a watermark prompt template (for example, here it can be a Prompt template) for intelligent tagging can be constructed based on the pre-defined watermark definition information. Figure 5 As shown, the server can obtain watermark definition information from the text definition information based on the watermark quality inspection task (for example, the above Figure 4 The watermark definition information 41a shown includes a watermark definition field and a watermark type, and the prompt template constructed based on the watermark definition information can be determined as the watermark prompt template.

[0212] Further, such as Figure 5As shown, the server can pre-mark the sampled text data with the help of the annotation model (e.g., ChatGPT) model for intelligent annotation obtained from the above-mentioned model database, so that the annotation model (e.g., ChatGPT) can predict the marking reply information of the output sampled text data according to the text output style indicated by the watermark prompt template (e.g., Prompt template here).

[0213] It should be understood that the text output style of the mark reply information here is consistent with the text output style indicated by the watermark prompt template (for example, it can be the Prompt template here), and the mark reply information includes the text segment detected in the text data to be processed that matches the watermark definition field, and the watermark type to which the text segment matching the watermark definition field belongs.

[0214] Among them, it can be understood that the template content of the Prompt template is as follows:

[0215] Please determine whether the following text contains a text watermark. Text watermarks are divided into two categories: personal privacy and source.

[0216] Personal privacy: The ability to directly or indirectly locate a person's ID or descriptive content, such as real name, social media ID, WeChat account, mobile phone number, QQ number, email address, website nickname, etc.

[0217] Source: Content that can identify the source of the text, such as website links, website names, public account names, reporter names, etc. (excluding the text subject).

[0218] Please input the category name + ":" first, then output the watermark text. Multiple categories should be output in multiple lines.

[0219] For each category, if the watermark of that category is included, all texts containing the watermark are output in the original order, and multiple watermark texts are separated by the "@@" symbol. If the watermark is not included, the output does not contain it.

[0220] Text: [fill in the text to be recognized here];

[0221] Whether to include text watermark:

[0222] Based on this, in the embodiment of the present application, after the sampled text data is input into the above-mentioned annotation model (for example, ChatGPT model), the text output style of the marked reply information output by the annotation model can be:

[0223] "Personal privacy: does not contain source attribution: [a fragment of the text that contains the source]".

[0224] It should be understood that the embodiments of the present application can parse the watermark fragment (i.e., the text fragment corresponding to the watermark definition field belonging to the watermark type) in the text (i.e., the sampled text data) based on the marked reply information replied by the ChatGPT model, and can map the parsed watermark fragment back to the original text (i.e., the sampled text data) to obtain text parsing information similar to the text from the Xth word to the Yth word in the original text belonging to the corresponding watermark type.

[0225] Further, such as Figure 5 As shown in the figure, in order to maximize the quality of the sample data of the training samples finally obtained, the server proposes that when the quality inspection task includes the watermark quality inspection task, the marked reply information replied by the ChatGPT model can be further verified by manual annotation. Figure 5 As shown, the server can send the sampled text data input into the ChatGPT model to the watermark quality inspection terminal associated with the watermark quality inspection task (for example, Figure 5 The user terminal 51a and the user terminal 52a shown in the figure are used to make the watermark quality inspection terminal (for example, Figure 5 The user terminal 51a and the user terminal 52a shown in FIG. 5 perform text marking on the sampled text data based on the watermark definition information to ultimately obtain the watermark information of the sampled text data. Figure 5 As shown, the watermark information here can be obtained by cross-checking the watermark information received by the server based on the watermark information of each watermark quality inspection terminal (for example, watermark information 51b corresponding to user terminal 51a and watermark information 52b corresponding to user terminal 52a). In this way, when the server determines that the watermark information received from each watermark quality inspection terminal is consistent during the cross-check, it is determined that the watermark information obtained by each watermark quality inspection terminal for the same sampled text data is legitimate and reliable.

[0226] At this time, the server can receive the watermark marking information with legality and reliability returned by the watermark quality inspection terminal, and then perform data calibration on the marking reply information based on the watermark marking information, so that the marking reply information or watermark marking information after data calibration can be determined as the text marking information of the sampled text data, so that the following steps S103-step S105 can be further executed.

[0227] Among them, the specific process of the server calibrating the data of the mark reply information through the watermark marking information can be described as: the server can use the watermark text segment of the sampled text data carried in the watermark marking information as the first watermark segment, and the watermark text segment of the sampled text data carried in the mark reply information as the second watermark segment; further, if there is a field different from the second watermark segment in the first watermark segment, the server can correct the second watermark segment in the mark reply information to the first watermark segment, and use the watermark text field indicated by the corrected first watermark segment in the mark reply information as the sampling key field in the sampled text data, and use the watermark type to which the watermark text field indicated by the first watermark segment belongs as the text attribute of the sampling key field; further, the server can determine the text marking information of the sampled text data based on the sampling key field and the text attribute of the sampling key field.

[0228] Optionally, it is understandable that if all fields in the first watermark segment are identical to all fields in the second watermark segment, the server may directly use the watermark information as text marking information of the sampled text data.

[0229] It is understood that, in the embodiment of the present application, for different quality inspection tasks under different quality inspection dimensions, different annotation models can be used when performing intelligent annotation, or the Figure 5 The steps of obtaining the marked reply information by outputting the annotation model are shown. In this way, when the server obtains the watermarked information through manual quality inspection, it can skip the step of calibrating the marked reply information through the watermarked information. In other words, at this time, the server can directly cross-validate (also called cross-check) the text annotation information obtained by multiple quality inspection and annotation terminals in the process of obtaining the watermarked information, so as to ensure that the text annotation information obtained by manual quality inspection (for example, Figure 5 The legitimacy and reliability of the watermark information 51b and watermark information 52b shown.

[0230] Specifically, the specific process of cross-validation can be described as follows: the server can distribute the sampled text data to multiple quality inspection marking terminals associated with the quality inspection task (for example, Figure 5As shown in the user terminal 51a and user terminal 52a), a quality inspection mark terminal is used to perform text annotation on the sampled text data based on the text definition information, and obtain a text annotation information of the sampled text data by the quality inspection mark terminal; further, the server can receive the text annotation information returned by each quality inspection mark terminal for the sampled text data, and cross-check the received text annotation information, and then verify and obtain the cross-check results for the sampling key fields in the sampled text data and the text attributes of the sampling key fields; it should be understood that if the cross-check result indicates that the received text annotation information is consistent, the server can quickly determine the text marking information of the sampled text data based on the sampling key fields and the text attributes of the sampling key fields.

[0231] Among them, it can be understood that in order to maximize the data quality of training samples, the embodiment of the present application can also use manual quality inspection methods for different quality inspection dimensions to perform secondary calibration on the ChatGPT labeling results, and in the process of secondary calibration, formulate a number of measures to ensure the labeling quality and efficiency.

[0232] To ensure the quality of annotation, the following annotation scheme is formulated in the embodiment of the present application:

[0233] 1) For a piece of text, all watermark segments must be accurately marked, no more and no less, and the segment boundaries must be clear to improve the text annotation standards.

[0234] 2) Develop and record training videos and design test questions to ensure that quality inspectors involved in manual quality inspection can learn and simulate annotation before taking up their posts, so as to screen out quality inspectors with a high learning pass rate (for example, quality inspectors who can achieve a score of 90 or above in the simulation test) for annotation, so as to fundamentally ensure the accuracy of manual annotation.

[0235] 3) Different quality inspectors are required to cross-label the same text data. For example, when the text data is a question in the test data, different quality inspectors (for example, 2-3 quality inspectors) are required to cross-label the text data corresponding to the same question, and then the consistency of the cross-validation results obtained by cross-labeling can be verified. The quality inspectors who perform cross-labeling can include but are not limited to Figure 5 The user terminal 51a corresponds to a quality inspector, and the user terminal 52a corresponds to a quality inspector.

[0236] To ensure labeling efficiency, the following labeling scheme is formulated in the embodiment of the present application:

[0237] 1) Task decomposition: For example, the task of labeling watermarked segments can be divided into three steps. The first step is to determine whether there is a watermark, the second step is to fine-tune the segments with watermarks, and the third step is to perform a second watermark screening for segments without watermarks.

[0238] 2) Batch delivery: For sampled text data, small batches of about 1K are delivered to ensure that watermark information for training quality inspection models can be obtained quickly, thereby improving the quality of the data. Figure 5 The model acquisition efficiency of the watermark quality inspection model shown.

[0239] It should be understood that in the example of this application, the manually calibrated samples are complementary to manual labeling and ChatGPT intelligent labeling, which can maximize the accuracy of the model output of the trained target quality inspection model, with an acceptance accuracy of 90%.

[0240] Step S103, obtaining a text prompt template associated with the text definition information, determining a tag prompt template corresponding to the text prompt template based on the text tag information, and using the text tag information and the sampled text data as a training sample data pair for training an initial business model;

[0241] It is understood that in the embodiment of the present application, in the process of training the quality inspection model that matches the corresponding quality inspection task, the server ensures the output reliability of the trained quality inspection model. The embodiment of the present application proposes that before training the corresponding quality inspection model, a prompt template consistent with the above-mentioned annotation model can be constructed and requested, and the correct watermark segment after manual calibration can be filled in it to obtain the correct watermark segment. Figure 5 It should be understood that the watermark marking template here can be a prompt template in the marking prompt template, and the marking prompt template here can also include other marking templates constructed based on text definition information under other quality inspection dimensions (for example, noise marking template, illegal data marking template, semantic marking template and format marking template, etc.).

[0242] Step S104: Perform model training on the initial business model using training sample data pairs and labeling prompt templates to obtain a target quality inspection model for performing quality inspection tasks;

[0243] Specifically, the server can input the sampled text data in the training sample data pair into the initial business model, and the initial business model predicts and outputs the text tag information of the sampled text data, and uses the predicted output text tag information as the predicted tag information of the sampled text data; further, the server can determine the model loss value of the initial business model based on the text tag information and the predicted tag information in the training sample data pair; further, if the model loss value of the initial business model does not meet the model optimization conditions indicated by the quality inspection task, the server can optimize the model parameters of the initial business model until the model loss value of the business model after parameter optimization meets the model optimization conditions, and use the business model whose model loss value meets the model optimization conditions as the target quality inspection model for performing the quality inspection task; the model optimization condition is used to indicate that the business model after parameter optimization predicts and outputs predicted tag information that matches the text tag information according to the tag prompt template.

[0244] For example, when the target quality inspection model includes a watermark quality inspection model, as mentioned above Figure 5 As shown, in the process of training the initial business model through text tag information, watermark tag template, and sampled text data to obtain a watermark quality inspection model, the initial business model will calculate the model loss value of the initial business model (i.e., loss, for example, the model loss value here is calculated based on the loss function for the text tag information and the prediction tag information used in the supervised training process) after predicting the corresponding prediction tag information (for example, the text after "Whether it contains text watermark:"), and can further optimize the model parameters of the initial business model through the optimizer when the model loss value does not meet the model optimization conditions to reduce the loss until the model loss value calculated during the quality inspection model training meets the model optimization conditions. The model loss value that meets the model optimization conditions can then be used as the target quality inspection model, and the target quality inspection model (for example, the watermark quality inspection model) finally trained can be enabled to have the ability to output watermark fragments. For example, the watermark quality inspection model obtained by training can subsequently have the ability to output watermark fragments according to the text output style indicated by the watermark tag template.

[0245] It is understood that, in the process of training the initial business model, examples of template data in the watermark template used in fine-tuning the model are as follows:

[0246] Please determine whether the following text contains a text watermark. Text watermarks are divided into two categories: personal privacy and source.

[0247] Personal privacy: IDs or descriptive content that can directly or indirectly locate a person, such as real name, social media ID, WeChat ID, mobile phone number, QQ number, email address, website nickname, etc.

[0248] Source: Content that can identify the source of the text, such as website links, website names, public account names, reporter names, etc. (excluding the text subject);

[0249] Please input the category name + ":" first, then output the watermark text. Multiple categories should be output in multiple lines.

[0250] For each category, if the watermark of that category is included, all texts containing the watermark are output in the original order, and multiple watermark texts are separated by the "@@" symbol. If the watermark is not included, the output does not contain it.

[0251] Text (i.e. sampled text data): Below, the editor of Jiuyou.com brings you a guide to unlocking the essential features of "Identity V".

[0252] The text output style of the predicted mark information of whether the prediction output contains a text watermark is: "Personal Privacy: Not Contained Source: Jiuyou.com".

[0253] Step S105 , obtaining input text data for inputting into a target quality inspection model from the text data to be processed, and having the target quality inspection model perform a quality inspection task on the input text data to obtain a text quality inspection result of the input text data.

[0254] Specifically, the server can obtain the target data type from the data type of the text data to be processed, and obtain K data sets deployed under the target data type; one data set corresponds to one data identifier, K is a positive integer, and each of the K data sets includes at least k1 training data; further, the server can obtain k2 training data associated with the target data type based on the k1 training data extracted from each data set, and use the k2 training data as input text data for inputting the target quality inspection model; k2=k1*K; further, the server can input the input text data into the target quality inspection model, and the target quality inspection model will perform quality inspection on the input data based on the label prompt template. When the quality inspection task is performed on the input text data, the predicted tag information that matches the text output style of the tag prompt template is predicted and output, and the predicted tag information output is used as the target tag information of the input text data; further, the server can perform information analysis on the target tag information to obtain text analysis information corresponding to the target tag information, and based on the data identifier of each data set, search the text analysis information for the training data in each data set in the text analysis information, and determine the text analysis information associated with the training data in each data set as the data set quality inspection result of each data set, and use the data set quality inspection result of each data set as the text quality inspection result of the input text data.

[0255] For further understanding, please see Figure 6 , Figure 6 This is a schematic diagram of a scenario for obtaining input text data provided by an embodiment of the present application. For ease of understanding, the text data to be processed is the above Figure 4 In this case, the data type of the text data to be processed may include Figure 6 The data type D1 shown is (for example, the test question data type), the data type D2 (for example, the web page data type), and the data type D3 (for example, the encyclopedia data type).

[0256] like Figure 6 As shown, in order to verify the data quality inspection quality of the target quality inspection model obtained by training, the embodiment of the present application can select a data type from the above multiple data types as the target data type (for example, Figure 6 The data type D2 shown in FIG2 can be used to determine the data type D2, and then multiple data sets deployed under the data type D2 (for example, K=3, for example, Figure 6 The datasets 6a, 6b and 6c shown in FIG are uniformly sampled in units of datasets (ie, sampling units) to select from a plurality of datasets (eg, Figure 6k1 training data (e.g., 100 pieces of text data) are extracted from the datasets 6a, 6b, and 6c shown respectively to obtain k2 training data (e.g., 100*3=300 pieces of text data) associated with the data type D2.

[0257] like Figure 6 As shown, the server can use k2 training data associated with the data type D2 (for example, 100*3=300 text data) as input text data for the target quality inspection model (for example, watermark quality inspection model), and then the target quality inspection model (for example, watermark quality inspection model) can predict and output predicted tag information that matches the text output style of the watermark tag template when performing data quality inspection (for example, watermark recognition) on the input text data, and then the predicted tag information of the predicted output can be used as the target tag information of the input text data.

[0258] It should be understood that in the embodiment of the present application, the server can perform information analysis on the target tag information output by the prediction to obtain text analysis information corresponding to the target tag information. Furthermore, the server can search for text analysis information associated with the training data in each data set in the text analysis information based on the data identifier of each data set, and then determine the text analysis information associated with the training data in each data set as the data set quality inspection result of each data set (for example, Figure 6 The quality inspection results of dataset 6a are 60a, Figure 6 The quality inspection results of dataset 6b and dataset 60b Figure 6 The quality inspection result 60c of the data set 6c is obtained. At this time, if Figure 6 As shown, the server may collectively refer to the dataset quality inspection results of each dataset as text quality inspection results of the input text data.

[0259] Optionally, it is understandable that after executing step S105, the server may further execute the following steps: For example, the server may select K data sets (for example, Figure 6The target data set (for example, data set 6a) is obtained from the data sets 6a, 6b and 6c shown in the figure, so as to obtain the data set quality inspection results that match the data identifier of the target data set from the text quality inspection results; the total number of quality inspections corresponding to the data set quality inspection results of the target data set is k1 (for example, k1=100); further, the server can search for the data quality inspection results associated with the watermark quality inspection task in the data set quality inspection results of the target data set, and use the data quality inspection results associated with the watermark quality inspection task found as the watermark detection results, and then use the number of watermark detection results accumulated as the number of watermark detections corresponding to the target data set (for example, 2); further, the server can search for the data quality inspection results associated with the watermark quality inspection task in the data set quality inspection results of the target data set, and use the watermark detection results associated with the watermark quality inspection task found as the watermark detection results, and then use the number of watermark detection results accumulated as the number of watermark detections corresponding to the target data set (for example, 2); further, the server can search for the data quality inspection results associated with the watermark quality inspection task based on the watermark detection number. , the total number of quality inspections, determine the watermark ratio corresponding to the target data set (for example, 2 / 100=2%), and when each of the K data sets is used as the target data set, obtain the watermark ratio corresponding to each data set; further, the server can fit the watermark ratio curve corresponding to each data set based on the watermark ratio corresponding to each data set, and then can display the watermark ratio curves corresponding to each data set under the same data type on the quality inspection display platform associated with the watermark quality inspection task with the data set as the unit granularity, so that the training data management personnel can subsequently realize real-time monitoring and management of the quality inspection quality of the target quality inspection model in the data quality inspection process through the watermark ratio curves corresponding to different data sets displayed on the quality inspection display platform.

[0260] In the process of training a quality inspection model (i.e., a target quality inspection model), the embodiment of the present application can obtain all key fields that need to undergo data quality inspection in the text data to be processed as text key fields based on the quality inspection tasks that need to be carried out at present, and use the quality inspection classification attributes to which these text key fields belong as the text attributes of these text key fields, and then pre-define the text definition information of the text data to be processed based on these text key fields and the text attributes of these text key fields. Furthermore, in order to ensure the accuracy of the pre-defined text definition information, the embodiment of the present application proposes that sampling sample data can be obtained from the text data to be processed, and then the text tag information of the sampling text data determined based on the text definition information can be obtained, and when a text prompt template associated with the text definition information is obtained, the obtained text tag information can be filled into the text prompt template to obtain a tag prompt template for standardizing the text tag style output by the model. Furthermore, the embodiment of the present application can use the sampled text data and text marking information as training sample data pairs for training the initial business model (the initial business model here can be any currently trained business model), so as to realize model training of the initial business model through the marking prompt template and the training sample data pair (the model training here refers to fine-tuning training of the model parameters of the initial business model), and obtain a target quality inspection model with data quality inspection capabilities (that is, the target quality inspection model here can refer to a quality inspection model that can perform quality inspection tasks to perform data quality inspection), so that the target quality inspection model can be used to perform quality inspection tasks on the input text data obtained from the text data to be processed, so as to quickly and accurately obtain the text quality inspection results of the input text data. In other words, the embodiment of the present application can construct a marking prompt template for standardizing the text output style of any business model through the text definition information of the text data to be processed and the text marking information of the limited sampled text data sampled from the text data to be processed, and then, a quality inspection model for data quality inspection can be quickly trained through the marking prompt model, the limited sampled text data sampled from the text data to be processed, and the text marking information of the limited sampled text data. Then, when the quality inspection model obtained by the rapid training performs the quality inspection task, the quality inspection efficiency and quality inspection accuracy of the input text data can be improved.

[0261] Further, see Figure 7 , Figure 7 This is a flow chart of another text data processing method provided by an embodiment of the present application. The method can be executed by a computer device, where the computer device can be a server, for example, the server can be the above-mentioned Figure 1 Server 100a in. Figure 7 As shown, the method may at least include the following steps S201 to S212.

[0262] Step S201, determining text definition information of the text data to be processed based on the quality inspection task associated with the text data to be processed;

[0263] The text definition information is determined based on the text key fields and text attributes of the text key fields in the text data to be processed;

[0264] Step S202, obtaining sampled text data from the text data to be processed, and determining text tag information of the sampled text data based on the text definition information;

[0265] The text tag information is determined based on the sampling key field in the sampled text data and the text attributes of the sampling key field; the sampling key field is a field in the text key field;

[0266] Step S203, obtaining a text prompt template associated with the text definition information, determining a tag prompt template corresponding to the text prompt template based on the text tag information, and using the text tag information and the sampled text data as a training sample data pair for training an initial business model;

[0267] Step S204: Perform model training on the initial business model using training sample data pairs and labeling prompt templates to obtain a target quality inspection model for performing quality inspection tasks;

[0268] Step S205 , obtaining input text data for inputting into a target quality inspection model from the text data to be processed, and having the target quality inspection model perform a quality inspection task on the input text data to obtain a text quality inspection result of the input text data.

[0269] The specific process of the server executing steps S201 to S205 can be found in the above Figure 3 The description of steps S101 to S105 in the corresponding embodiment will not be repeated here.

[0270] It should be understood that in the embodiment of the present application, after the server completes the model fine-tuning of the initial business model, it can use the target quality inspection model (for example, watermark quality inspection model) data set obtained by the model fine-tuning as a unit to perform watermark recognition on the text data in each data set, so as to obtain the watermark ratio of each data set by performing quality inspection analysis on the target mark information output by the watermark quality inspection model.

[0271] For example, optionally, in an embodiment of the present application, the server can also randomly sample a fixed number of samples from each data set contained in the text data to be processed to ensure that the amount of input text data finally extracted for inputting the target quality inspection model remains at k2 (for example, k2 = 20,000 items) to ensure that the statistical data obtained from the final quality inspection analysis of the data quality inspection results has a good confidence level.

[0272] It should be understood that the text data in these input text data will be filled into the Prompt template constructed above (i.e., the watermark marking template mentioned above), and the fine-tuned large model (e.g., the watermark quality inspection model) will be requested. After parsing the predicted output of the large model, it can be determined whether the text data contains watermarks, and ultimately the watermark proportion of each data set can be obtained. In actual testing, when using the watermark quality inspection model for watermark recognition, the watermark quality inspection accuracy rate currently analyzed by the watermark quality inspection model can reach 91%, and the recall rate can reach 82%.

[0273] Step S206: obtaining a cleaning task associated with the quality inspection task, and obtaining a text quality inspection segment associated with the cleaning task from the text quality inspection result;

[0274] Step S207: Obtain an initial recall model and an initial cleaning model associated with the cleaning task, and use the training data carrying the text quality inspection fragment obtained from the input text data as the sample data to be cleaned associated with the initial recall model and the initial cleaning model;

[0275] Step S208: training the initial recall model using the sample data to be cleaned and the text quality inspection fragments in the sample data to be cleaned, to obtain a target recall model for data recall;

[0276] Step S209 : training the initial cleaning model using the sample data to be cleaned and the text quality inspection segments in the sample data to be cleaned, to obtain a target cleaning model for data cleaning.

[0277] For further understanding, please see Figure 8 , Figure 8 This is a schematic diagram of a scenario for training a target recall model and a target cleaning model provided by an embodiment of the present application. Figure 8 As shown, after the server inputs the input text data into the target quality inspection model, it can perform data distillation on the input text data through the target quality inspection model to obtain Figure 8 The distillation sample data shown here should be understood that the specific process of obtaining the distillation sample data here can refer to the above description of the process of obtaining the sample data to be cleaned.

[0278] like Figure 8 As shown, the server can input the distilled sample data into the initial recall model and the initial cleaning model respectively to achieve fine-tuning training of the initial recall model and the initial cleaning model, and finally obtain Figure 8 target recall model and target cleaning model.

[0279] Regarding data distillation, it is understandable that in data cleaning tasks, due to performance limitations, it is impossible to use large models to identify watermarks in data sets of tens or even hundreds of TB. Otherwise, the cost of the machine will be too high to bear. Therefore, it is necessary to use a model with about 100 million parameters such as BERT as the initial cleaning model, or even a smaller model to achieve data cleaning.

[0280] Although this type of initial cleaning model is also a business model obtained after pre-training on a large corpus, the natural language knowledge contained in it is significantly weaker than the above-mentioned 7B large model (i.e., the above-mentioned initial business model). This will result in the inability to achieve satisfactory results by directly using a small amount of manually labeled data. The model's extrapolation ability is limited, and it can only identify watermarks related to a small amount of manually labeled data.

[0281] Based on this, the embodiment of the present application proposes that the text data extracted from the data set under each data type can be used as input text data through the data distillation method, so that the hundreds of thousands or even millions of input text data sampled can be predicted using the 7B large model (i.e., the above-mentioned target quality inspection model) trained for data quality inspection to obtain text data of millions to tens of millions of data volumes and the corresponding watermark fragment data. This means that these text data carrying watermark fragments can be subsequently used to train the BERT model (i.e., the initial cleaning model) to allow the teacher model (7B large model) to impart more knowledge to the student (BERT model) as much as possible.

[0282] Step S210: performing model cascade processing on the target recall model and the target cleaning model to obtain a data cleaning model associated with the cleaning task, and obtaining text data to be cleaned associated with the target data set from the first business database associated with the cleaning task;

[0283] The target dataset refers to any dataset deployed under the data type of the text data to be processed;

[0284] like Figure 8 As shown, after the server has trained the target recall model and the target cleaning model, it can obtain the target recall model from the business database (for example, Figure 8 The data to be cleaned is obtained from the database Q2 shown, that is, the first business database mentioned above. Figure 8 The process of obtaining the data to be cleaned can be referred to above. Figure 2The description of the specific process of obtaining the text data to be cleaned in the corresponding embodiment will not be repeated here.

[0285] Step S211: Input the text data to be cleaned into the target recall model in the data cleaning model. The target recall model selects target text data to be cleaned that carries text quality inspection segments from the text data to be cleaned, marks the text quality inspection segments in the target text data to be cleaned, uses the marked text quality inspection segments as cleaning mark segments of the target text data to be cleaned, and outputs the target text data to be cleaned that carries the cleaning mark segments from the target recall model.

[0286] In step S212, the target text data to be cleaned is input into the target cleaning model in the data cleaning model. The target cleaning model identifies cleaning mark segments from the target text data to be cleaned, performs text cleaning on the cleaning mark segments in the target text data to be cleaned, and uses the target text data to be cleaned after text cleaning as the data to be quality inspected for input into the target quality inspection model.

[0287] Further, such as Figure 8 As shown, the server can input the data to be cleaned into the target recall model so that the target recall model can determine which text data in the data to be cleaned carries text quality inspection segments (e.g., watermark segments). This means that at this time, the target recall model can be used to remove text data that does not carry text quality inspection segments (e.g., watermark segments) from the data to be cleaned, and then the text data that carries text quality inspection segments (e.g., watermark segments) remaining in the text data to be cleaned can be regarded as text data that carries text quality inspection segments (e.g., watermark segments) filtered out by the target recall model, and then the filtered text data that carries text quality inspection segments (e.g., watermark segments) can be used as target text data to be cleaned for input into the target cleaning model. Further, as Figure 8 As shown, the server can perform text cleaning on the target text data to be cleaned through the target cleaning model, so that the target text data to be cleaned after the text emotion is collectively referred to as Figure 8 Data after cleaning are shown.

[0288] like Figure 8 As shown, the embodiment of the present application can use the cleaned data as the data to be quality inspected, or can use the text data that does not carry the text quality inspection fragments filtered out by the target recall model as the data to be quality inspected.

[0289] In this way, the server can input these data to be inspected into the target quality inspection model, so that the target quality inspection model can perform data quality inspection on these data to be inspected, and obtain Figure 8 The data quality inspection results are shown.

[0290] Optionally, after executing step S212, the server may further perform the following steps: the server may input the data to be quality inspected into the target quality inspection model, and the target quality inspection model may perform the quality inspection task on the data to be quality inspected to obtain the data quality inspection results of the data to be quality inspected; further, the server may determine the quality inspection ratio corresponding to the target data set based on the data quality inspection results; further, if the quality inspection ratio satisfies the quality inspection data strategy indicated by the quality inspection task, the server may use the data to be quality inspected associated with the target data set as model training sample data for training the target business model, and then may add the model training sample data to the second business database corresponding to the target business model.

[0291] For further understanding, please see Figure 9 , Figure 9 This is a schematic diagram of a scenario in which watermark recognition is performed on quality inspection data based on a watermark quality inspection task provided by an embodiment of the present application. Figure 9 The data to be inspected is determined by the cleaned text data obtained by the above two-stage data cleaning method. Figure 8 The cleaned data in the corresponding embodiment.

[0292] like Figure 9 As shown, when the quality inspection task includes a watermark quality inspection task, after the server obtains the text data to be processed from the database Q1 (i.e., the second business database mentioned above), it can execute step S31, and then define the watermark for the text data to be processed to obtain the watermark definition information of the text data to be processed. Furthermore, before executing step S33 based on the watermark definition information, the server can execute step S32 to obtain the sampled text data for intelligent standardization from the text data to be processed. Among them, the specific implementation method of the server obtaining the marking reply information of the sampled text data based on the watermark definition information can be found in the above Figure 3 A description of the specific process of obtaining the marked reply information in the corresponding embodiment.

[0293] Further, such as Figure 9 As shown, the server can execute step S34 to perform data calibration on the marked reply by manual calibration to obtain Figure 9 The specific implementation method of obtaining the text mark information can be found in the above Figure 3 The specific process of manual calibration in the corresponding embodiment will not be described here. Then, the server can execute step S35 to train the quality inspection model of the initial business model by sampling text data and text tag information to obtain Figure 9 The watermark quality inspection model shown in the figure. For the specific implementation of obtaining the watermark quality inspection model, please refer to the above Figure 3 The description of the specific process of model fine-tuning in the corresponding embodiment will not be repeated here.

[0294] Furthermore, the server may execute step S36 to obtain the text data to be cleaned for training the recall model and the cleaning model from the input text data by means of data distillation (the text data to be cleaned here may be the above-mentioned Figure 8 The distilled sample data generated by the target quality inspection model as shown in the figure can then be executed Figure 9 Steps S37 and S38 are shown to obtain a target recall model and a target cleaning model through training of the text data to be cleaned.

[0295] It should be understood that in the data cleaning process, the embodiment of the present application proposes a two-layer cleaning framework. In this cleaning framework, the first layer is the recall layer, which is used to perform preliminary watermark screening through the target recall model to preliminarily screen out the text data that needs to be cleaned of watermarks. That is, the goal of the recall layer is to eliminate most of the watermark-free data in the text data to be cleaned sampled from the database Q2 to save the computational cost of the second layer; the second layer is the cleaning layer, which is used to accurately identify the watermark fragments in the target text data to be cleaned output by the target recall model through the target cleaning model, and accurately remove the watermark fragments.

[0296] In an embodiment of the present application, the business model used to train the recall layer (i.e., the initial recall model) can adopt the FastText model obtained from the model database, so that the FastText model can be used to perform binary classification with or without watermarks. The model parameters of the FastText model have a parameter volume of about 1 million, which can be run on the CPU and has extremely high performance. It should be noted that the FastText model here does not require a very high accuracy rate (for example, the current accuracy rate of the watermark initial screening is 25%), but it is required to have a relatively high recall rate (for example, the current recall rate is 86%). It should be understood that the training data used to train the FastText model all come from the sample distillation data obtained by data distillation. In the process of training the FastText model, samples containing watermark segments in the sample distillation data can be used as positive samples, and samples that do not contain watermark segments can be used as negative samples to train the FastText model. Figure 9 The target recall model shown.

[0297] In the embodiment of the present application, the business model used to train the cleaning layer (i.e., the initial cleaning model) is sampled from the BERT model. The BERT model has approximately 100 million parameters and has a higher computational cost than the FastText model. Therefore, the BERT model can be used to accurately identify watermark segments in text data and delete the sentences or paragraphs containing the watermark segments. Similarly, the training data used to train the BERT model also comes from sample distilled data obtained through data distillation. In this sample distilled data, the position information of the watermark segments obtained by parsing in the original text will be used to train the BERT model.

[0298] After experimental testing, the target recall model in the recall layer can filter out about 60% of the watermark-free samples from the data to be cleaned, which can help the cleaning layer save 60% of the computing power and only sacrifice 14% of the recall rate, which is acceptable.

[0299] Optionally, it is understood that the embodiment of the present application can also supplement the samples not recalled by the target recall model (i.e. the FastText model after fine-tuning training) by keyword recall during the watermark cleaning process to improve the recall rate of the overall watermark cleaning. For ease of understanding, please refer to Figure 10 , Figure 10 This is a schematic diagram of a scenario for keyword recall in the recall layer provided by an embodiment of the present application. Figure 10 As shown in the figure, for the text data to be cleaned with a data volume of 50TB, in order to ensure the accuracy of the recall rate determined by the target recall model, it is proposed that before the watermark is initially screened by the target recall model, the regular expression constructed based on the watermark definition information in the text definition information can be obtained by keyword recall, so as to pre-screen (i.e., coarse-grained screening) the transition text data that may contain watermark fragments from the massive data to be cleaned through the regular expression, and then these transition text data can be input into the target recall model, and the target recall model can perform a binary classification of these transition text data that may contain watermark fragments into two categories: with or without watermarks, for example, Figure 9 As shown, by combining this keyword recall method with the discriminant model recall method, 20TB of text data without watermark segments can be eliminated from the 50TB of text data to be cleaned. Furthermore, the target recall model can use the discriminant model recall method to select the 30TB of text data with watermark segments from the 50TB of text data to be cleaned as the target text data to be cleaned. In this way, after the server cleans the target text data to be cleaned using the target cleaning model, the resulting cleaned data is 5TB in size.

[0300] Further, such as Figure 9As shown, the server can use the cleaned text data obtained through two-stage cleaning as the data to be quality inspected to further execute step S39. That is, at this time, the server can perform data quality inspection (for example, watermark recognition) on the data to be quality inspected through the watermark quality inspection model to obtain the watermark quality inspection result of the data to be quality inspected.

[0301] For further understanding, please see Figure 11 , Figure 11 This is a schematic diagram of a scenario in which data quality inspection is performed under different quality inspection dimensions as proposed in the embodiment of this application. Figure 11 As shown, when the server obtains the target quality inspection model containing different quality inspection sub-models through training, it can call the quality inspection model used to perform the corresponding quality inspection task through the quality inspection routing.

[0302] For example, Figure 11 As shown, the server can perform periodic routine quality inspection on part of the sampled text data (the first training text data set 1201a) sampled from the full database, and can also perform immediate quality inspection on all the training text data (i.e., the second training text data set 1202a) obtained from the incremental database. At this time, the server can remove the training text data sets (for example, the first training text data set 1201a or the second training text data set 1202a) obtained from different business databases from the warehouse to the area to be processed 1203a. For example, the embodiment of the present application can use the text data in the training text data set (for example, the first training text data set 1201a) obtained from the area to be processed 1203a as the data to be quality inspected, through Figure 11 The target quality inspection model shown is used to perform data quality inspection on the data to be quality inspected in different quality inspection dimensions to obtain data quality inspection results of the same data to be quality inspected in different quality inspection dimensions.

[0303] For example, a watermark quality inspection model can be used to perform watermark identification on the data to be inspected to obtain the watermark quality inspection result of the data to be inspected under the watermark quality inspection dimension. For another example, a noise quality inspection model can be used to perform noise identification on the data to be inspected to obtain the noise quality inspection result of the data to be inspected under the noise quality inspection dimension. For another example, an illegal data quality inspection model can be used to perform illegal data identification on the data to be inspected to obtain the illegal data quality inspection result of the data to be inspected under the illegal data quality inspection dimension. Similarly, a semantic quality inspection model can be used to perform semantic identification on the data to be inspected to obtain the semantic quality inspection result of the data to be inspected under the semantic quality inspection dimension, and a format quality inspection model can be used to perform format identification on the data to be inspected to obtain the format quality inspection result of the data to be inspected under the format quality inspection dimension.

[0304] Further, such as Figure 11As shown, the server can perform quality inspection analysis and statistics on the data quality inspection results under different quality inspection dimensions, and then calculate the quality inspection ratio of the same data under different quality inspection dimensions. For example, the quality inspection ratio here can include watermark ratio, noise ratio, illegal data ratio, semantic ratio, and format ratio. The statistical methods for calculating the noise ratio, illegal data ratio, semantic ratio, and format ratio can be found in the description of the specific process for deriving the watermark ratio above, and will not be further described here.

[0305] like Figure 11 As shown, the server can also generate a quality inspection and analysis report. The quality inspection and analysis report can include the information displayed in the above-mentioned visual result viewing page, and can also include the overall watermark ratio corresponding to the corresponding data type, the ratio of problem data subsets in each data set under the corresponding data type, the ratio ranking list of declining ratios, and other information. The ratio of problem data subsets can be the ratio of the number of subsets consisting of problem training data counted in each data set under a certain data type to the total number of training data subsets. The ratio ranking list of declining ratios can be the ranking list obtained by sorting each problem training data subset based on the ratio of the declining ratio of problem training data in the problem data subset after data cleaning. It can be understood that after determining the quality inspection and analysis report, the quality inspection and analysis report can also be sent to the business object through email, session, etc., so that the business terminal corresponding to the business object can display the quality inspection and analysis report on the quality inspection display platform provided by the server.

[0306] For ease of understanding, please refer to Figure 12 , Figure 12 This is a scene diagram of a quality inspection and analysis report provided by an embodiment of the present application. Figure 12 The quality inspection analysis report shown includes the overall watermark ratio of the data type of data type V1, and the overall watermark ratio of the data type of data type V1.

[0307] The overall watermark ratio is obtained by weighting the number of data items to be inspected under the corresponding data type. The data to be inspected here refers to all text data obtained during the sampling period from the sampling start date (T1) to the sampling end date (T2). Figure 12 As shown, the sampling time period may include 5 periodic statistical nodes, for example, Figure 12 Shown are the period statistics node T11, the period statistics node T12, the period statistics node T13, the period statistics node T14 and the period statistics node T15.

[0308] In addition, if Figure 12As shown, the quality inspection analysis report also includes the watermark ratio of each data set under the corresponding data type (for example, the watermark ratio of each data set at different period statistical nodes can be highlighted). For example, when the data type is data type V1, the data set under data type V1 can include Figure 12 Therefore, after the server performs data quality inspection on the data to be inspected under data type V1 through the target quality inspection model, it can analyze and count the data quality inspection results to obtain Figure 12 The watermark ratio of the data set V11, the watermark ratio of the data set V12, the watermark ratio of the data set V13, and the watermark ratio of the data set V14 are shown. Figure 12 As shown, since the watermark proportion of dataset V11, the watermark proportion of dataset V12, the watermark proportion of dataset V13, and the watermark proportion of dataset V14 are all less than the watermark proportion threshold (for example, 10%) indicated by the watermark quality inspection strategy, it can be determined that the watermark proportion of dataset V11, the watermark proportion of dataset V12, the watermark proportion of dataset V13, and the watermark proportion of dataset V14 all meet the quality inspection data strategy, so that the text data in datasets V11, dataset V12, dataset V13, and dataset V14 that meet the quality inspection data strategy can be used as the above-mentioned model training text data and added to the full database (that is, the above-mentioned second business database), so that the target business model (for example, the above-mentioned mixed-yuan large model) can be trained subsequently through the training samples obtained in the full database.

[0309] Similarly, with this type, Figure 12 As shown, considering that the watermark ratio of dataset V21 and the watermark ratio of dataset V22 under dataset V2 are also less than the watermark ratio threshold (for example, 10%) indicated by the watermark quality inspection strategy, it can be determined that the watermark ratio of dataset V21 and the watermark ratio of dataset V22 both meet the quality inspection data strategy, so that the text data in dataset V21 and dataset V22 that meet the quality inspection data strategy can be used as the above-mentioned model training text data and added to the full database (that is, the above-mentioned second business database) to ensure the sample data quality of the training samples obtained during the training process of the target business model.

[0310] Thus, the embodiment of the present application can, after training the target quality inspection model, perform data quality inspection on the input text data obtained from the text data to be processed through the target quality inspection model, and then obtain the text fragment associated with the current cleaning task as the text quality inspection fragment from the text quality inspection result obtained by the data quality inspection, and then obtain the training data carrying the text quality inspection fragment from the input text data as the sample data to be cleaned for training the initial recall model and the initial cleaning model, and then the sample data to be cleaned and the text quality inspection fragment in the sample data to be cleaned can be used to train the initial recall model and the initial cleaning model respectively to obtain the target recall model for data recall and the target cleaning model for data cleaning. In this way, the embodiment of the present application can construct a data cleaning model for two-stage data cleaning through the target recall model and the target cleaning model, and can use the data cleaning model to perform two-stage data cleaning on the text data to be cleaned associated with the current data set (i.e., the target data set) in units of data sets to ensure the data security and reliability of the data to be cleaned for input into the target quality inspection model. In addition, when the embodiment of the present application performs a quality inspection task on the data to be quality inspected through the target quality inspection model, the target quality inspection model can be used to identify whether the data to be quality inspected carries text quality inspection fragments that require data cleaning in units of data sets. If so, the quality inspection ratio of the data to be quality inspected carrying text quality inspection fragments in this data set to all the data to be quality inspected in this data set will be counted in units of data sets. This can allow all the data to be quality inspected in the data set whose quality inspection ratio meets the quality inspection data strategy indicated by the quality inspection task to be used as model training sample data for training the target business model. This can then allow the model training sample data to be added to the business database (for example, the second business database) used to train the target business model to ensure the quality inspection quality and quality inspection accuracy of the training sample data participating in the training of the model.

[0311] See Figure 13 , Figure 13 This is a structural diagram of a text data processing device provided by an embodiment of the present application. Figure 13 As shown, the text data processing device 1 can be a computer device (for example, the above Figure 1 A computer program (including program code) of the server 100a) in the embodiment of the present application, for example, the text data processing device 1 is an application software; it can be understood that the text data processing device 1 can be used to execute the corresponding steps of the text data processing method provided in the embodiment of the present application. Figure 13As shown, the text data processing device 1 may include: a definition information determination module 11, a sample data acquisition module 12, a tag information determination module 13, a sample data pair determination module 14, an initial business model training module 15 and a quality inspection task execution module 16;

[0312] The definition information determination module 11 is used to determine the text definition information of the text data to be processed based on the quality inspection task associated with the text data to be processed; the text definition information is determined based on the text key fields in the text data to be processed and the text attributes of the text key fields;

[0313] The sampling data acquisition module 12 is used to obtain the sampling text data from the text data to be processed;

[0314] The tag information determining module 13 is used to determine the text tag information of the sampled text data based on the text definition information; the text tag information is determined based on the sampling key field in the sampled text data and the text attributes of the sampling key field; the sampling key field is a field in the text key field;

[0315] a sample data pair determination module 14 for obtaining a text prompt template associated with the text definition information, determining a tag prompt template corresponding to the text prompt template based on the text tag information, and using the text tag information and the sampled text data as a training sample data pair for training an initial business model;

[0316] The initial business model training module 15 is used to train the initial business model by using training sample data pairs and labeling prompt templates to obtain a target quality inspection model for performing quality inspection tasks;

[0317] The quality inspection task execution module 16 is used to obtain input text data for inputting into the target quality inspection model from the text data to be processed, and the target quality inspection model executes the quality inspection task on the input text data to obtain the text quality inspection result of the input text data.

[0318] The specific implementation of the definition information determination module 11, the sampling data acquisition module 12, the label information determination module 13, the sample data pair determination module 14, the initial business model training module 15 and the quality inspection task execution module 16 can be found in the above Figure 3 The description of steps S101 to S105 in the corresponding embodiment will not be repeated here.

[0319] The definition information determination module 11 includes: a data set acquisition unit 111, a quality inspection task acquisition unit 112, a quality inspection definition unit 113 and a definition information determination unit 114;

[0320] The data set acquisition unit 111 is used to acquire a training text data set for training a target business model, and use each training text data in the training text data set as text data to be processed; the target business model is a business model associated with the target quality inspection model;

[0321] A quality inspection task acquisition unit 112 is used to acquire a quality inspection task associated with the training text dataset;

[0322] A quality inspection definition unit 113 is used to perform quality inspection definition on the text data to be processed based on the quality inspection task and the data type of the text data to be processed, and obtain text key fields and text attributes of the text key fields of the text data to be processed;

[0323] The definition information determining unit 114 is configured to use the text key field and the text attributes of the text key field as text definition information of the text data to be processed.

[0324] The specific implementation of the data set acquisition unit 111, the quality inspection task acquisition unit 112, the quality inspection definition unit 113 and the definition information determination unit 114 can be found in the above Figure 3 The description of the specific process of obtaining text definition information in the corresponding embodiment will not be repeated here.

[0325] Among them, the quality inspection task at least includes the watermark quality inspection task;

[0326] The quality inspection definition unit 113 is specifically configured to obtain a watermark quality inspection strategy associated with the watermark quality inspection task, and obtain a first target data type to be watermark defined from the data type of the to-be-processed text data based on the watermark quality inspection strategy;

[0327] The quality inspection definition unit 113 is further specifically configured to use, in the text data to be processed, the text data to be processed that matches the first target data type as first target text data, perform watermark definition on the first target text data, obtain a watermark definition field of the first target text data, and use the watermark definition field as a text key field of the text data to be processed;

[0328] The quality inspection definition unit 113 is further specifically configured to classify the watermark definition field into a watermark type indicated by the watermark quality inspection strategy, and determine the watermark type as a text attribute of the text key field.

[0329] Optionally, the quality inspection task includes at least a noise quality inspection task;

[0330] The quality inspection definition unit 113 is further specifically configured to obtain a noise quality inspection strategy associated with the noise quality inspection task, and obtain a second target data type to be subjected to noise definition from the data type of the to-be-processed text data based on the noise quality inspection strategy;

[0331] The quality inspection definition unit 113 is further specifically configured to use, in the text data to be processed, the text data to be processed that matches the second target data type as second target text data, perform noise definition on the second target text data to obtain a noise definition field of the second target text data, and use the noise definition field as a text key field of the text data to be processed;

[0332] The quality inspection definition unit 113 is further specifically configured to classify the noise definition field into the noise type indicated by the noise quality inspection strategy, and determine the noise type as a text attribute of the text key field.

[0333] The sampling data acquisition module 12 includes: a sampling processing unit 121 and a sampling text determination unit 122;

[0334] The sampling processing unit 121 is used to perform sampling processing on the text data to be processed based on the data type of the text data to be processed, so as to obtain a sampling data set that matches the data type of the text data to be processed;

[0335] The sampled text determining unit 122 is configured to determine sampled text data based on a sampled data set that matches the data type of the text data to be processed.

[0336] The specific implementation of the sampling processing unit 121 and the sampling text determination unit 122 can be found in the above Figure 3 The description of the specific process of obtaining the sampled text data in the corresponding embodiment will not be repeated here.

[0337] The data types of the text data to be processed include at least a first data type and a second data type;

[0338] The sampling processing unit 121 is specifically configured to determine a first data type from the data types of the text data to be processed, extract N text data matching the first data type from the text data to be processed, and use the extracted N text data matching the first data type as a first sampling data set; N is a positive integer;

[0339] The sampling processing unit 121 is further specifically configured to determine a second data type from the data types of the text data to be processed, extract N text data matching the second data type from the text data to be processed, and use the extracted N text data matching the second data type as a second sampling data set;

[0340] The sampling processing unit 121 is further specifically configured to determine a sampling data set that matches the data type of the text data to be processed based on the first sampling data set and the second sampling data set.

[0341] The tag information determination module 13 includes: a sample text distribution unit 131, an annotation information receiving unit 132 and a tag information determination unit 133;

[0342] The sampled text distribution unit 131 is used to distribute the sampled text data to multiple quality inspection marking terminals associated with the quality inspection task, and each quality inspection marking terminal is used to perform text annotation on the sampled text data based on the text definition information, thereby obtaining a text annotation information for the sampled text data by the quality inspection marking terminal;

[0343] The annotation information receiving unit 132 is used to receive the text annotation information returned by each quality inspection and marking terminal for the sampled text data, cross-check the received text annotation information, and obtain a cross-check result for the sampling key fields in the sampled text data and the text attributes of the sampling key fields;

[0344] The tag information determining unit 133 is configured to determine text tag information of the sampled text data based on the sampling key fields and text attributes of the sampling key fields if the cross-check result indicates that the received text annotation information is consistent.

[0345] The specific implementation of the sampling text distribution unit 131, the annotation information receiving unit 132 and the tag information determination unit 133 can be found in the above Figure 3 The description of the specific process of obtaining text markup information in the corresponding embodiment will not be repeated here.

[0346] Optionally, the quality inspection task includes at least a watermark quality inspection task, the text key field includes a watermark definition field, and the text attribute of the text key field includes the watermark type to which the watermark definition field belongs;

[0347] The tag information determination module 13 includes: a prompt template construction unit 134 and a reply information output unit 135;

[0348] The prompt template construction unit 134 is configured to obtain watermark definition information including a watermark definition field and a watermark type from the text definition information based on the watermark quality inspection task, and determine the prompt template constructed based on the watermark definition information as the watermark prompt template;

[0349] The reply information output unit 135 is used to input the sampled text data into the watermark prompt template, and the watermark prompt template outputs the marked reply information of the sampled text data based on the watermark definition information;

[0350] The sampled text distribution unit 131 is further configured to send the sampled text data to a watermark quality inspection terminal associated with the watermark quality inspection task, so that the watermark quality inspection terminal performs text marking on the sampled text data based on the watermark definition information to obtain watermark marking information of the sampled text data;

[0351] The marking information receiving unit 132 is further used to receive the watermark marking information returned by the watermark quality inspection terminal;

[0352] The marking information determining unit 133 is further configured to determine text marking information of the sampled text data based on the watermark marking information and the marking reply information.

[0353] The marking information determining unit 133 is specifically configured to use the watermark text segment of the sampled text data carried in the watermark marking information as the first watermark segment, and use the watermark text segment of the sampled text data carried in the marking reply information as the second watermark segment;

[0354] The marking information determining unit 133 is further specifically configured to, if the first watermark segment contains a field different from the second watermark segment, correct the second watermark segment in the marking reply information to be the first watermark segment, use the watermark text field indicated by the first watermark segment as a sampling key field in the sampled text data, and use the watermark type to which the watermark text field indicated by the first watermark segment belongs as a text attribute of the sampling key field;

[0355] The tag information determining unit 133 is further specifically configured to determine the text tag information of the sampled text data based on the sampled key fields and the text attributes of the sampled key fields.

[0356] It should be understood that the embodiment of the present application can jointly determine the text tag information through the prompt template construction unit 134, the reply information output unit 135 and the sample text distribution unit 131, the annotation information receiving unit 132 and the tag information determination unit 133, and the specific implementation method of jointly determining the text tag information through the prompt template construction unit 134, the reply information output unit 135 and the sample text distribution unit 131, the annotation information receiving unit 132 and the tag information determination unit 133 can be referred to above. Figure 3 The description of the specific process of calibrating the mark reply information through the watermark mark information in the corresponding embodiment will not be repeated here.

[0357] The initial business model training module 15 includes: a prediction mark output unit 151, a loss value determination unit 152 and a parameter optimization unit 153;

[0358] The prediction tag output unit 151 is used to input the sampled text data in the training sample data pair into the initial business model, and the initial business model predicts and outputs the text tag information of the sampled text data, and uses the predicted output text tag information as the prediction tag information of the sampled text data;

[0359] A loss value determining unit 152 is configured to determine a model loss value of the initial business model based on the text tag information and the prediction tag information in the training sample data pair;

[0360] The parameter optimization unit 153 is used to optimize the model parameters of the initial business model if the model loss value of the initial business model does not meet the model optimization conditions indicated by the quality inspection task, until the model loss value of the business model after parameter optimization meets the model optimization conditions, and use the business model whose model loss value meets the model optimization conditions as the target quality inspection model for executing the quality inspection task; the model optimization conditions are used to indicate that the business model after parameter optimization predicts and outputs predicted tag information that matches the text tag information according to the tag prompt template.

[0361] The specific implementation of the prediction mark output unit 151, the loss value determination unit 152 and the parameter optimization unit 153 can be found in the above Figure 3 The description of the specific process of obtaining the target quality inspection model in the corresponding embodiment will not be repeated here.

[0362] The quality inspection task execution module 16 includes: a target data type determination unit 161, a training data extraction unit 162, a quality inspection task execution unit 163 and a tag information parsing unit 164;

[0363] A target data type determining unit 161 is configured to obtain a target data type from the data type of the text data to be processed, and obtain K data sets deployed under the target data type; one data set corresponds to one data identifier, K is a positive integer, and each of the K data sets includes at least k1 training data;

[0364] The training data extraction unit 162 is configured to obtain k2 training data associated with the target data type based on the k1 training data extracted from each data set, and use the k2 training data as input text data for inputting into the target quality inspection model; k2 = k1*K;

[0365] The quality inspection task execution unit 163 is configured to input the input text data into the target quality inspection model, and when the target quality inspection model performs the quality inspection task on the input text data based on the tag prompt template, predict and output predicted tag information that matches the text output style of the tag prompt template, and use the predicted tag information as the target tag information of the input text data;

[0366] The tag information parsing unit 164 is used to parse the target tag information to obtain text parsing information corresponding to the target tag information. Based on the data identifier of each data set, the text parsing information associated with the training data in each data set is searched in the text parsing information. The text parsing information associated with the training data in each data set found is determined as the data set quality inspection result of each data set, and the data set quality inspection result of each data set is used as the text quality inspection result of the input text data.

[0367] The specific implementation of the target data type determination unit 161, the training data extraction unit 162, the quality inspection task execution unit 163 and the tag information parsing unit 164 can be found in the above Figure 3 The description of the specific process of obtaining the text quality inspection results in the corresponding embodiment will not be repeated here.

[0368] Among them, the quality inspection task at least includes the watermark quality inspection task;

[0369] The device 1 further comprises: a target data set acquisition module 17, a watermark detection quantity determination module 18, a watermark proportion determination module 19 and a watermark proportion curve display module 20;

[0370] The target dataset acquisition module 17 is configured to acquire a target dataset from the K datasets and obtain a dataset quality inspection result that matches the data identifier of the target dataset from the text quality inspection result; the total number of quality inspections corresponding to the dataset quality inspection result of the target dataset is k1;

[0371] The watermark detection quantity determination module 18 is configured to search for data quality inspection results associated with the watermark quality inspection task in the data set quality inspection results of the target data set, use the data quality inspection results associated with the watermark quality inspection task found as the watermark detection results, and use the cumulative number of watermark detection results as the number of watermark detections corresponding to the target data set;

[0372] The watermark ratio determination module 19 is used to determine the watermark ratio corresponding to the target data set based on the number of watermark detections and the total number of quality inspections. When each of the K data sets is used as the target data set, the watermark ratio corresponding to each data set is obtained;

[0373] The watermark proportion curve display module 20 is used to fit the watermark proportion curve corresponding to each data set based on the watermark proportion corresponding to each data set, and display the watermark proportion curve on the quality inspection display platform associated with the watermark quality inspection task.

[0374] The specific implementation of the target data set acquisition module 17, the watermark detection quantity determination module 18, the watermark ratio determination module 19 and the watermark ratio curve display module 20 can be found in the above Figure 7 The description of the specific process of displaying the watermark ratio curve on the quality inspection display platform in the corresponding embodiment will not be repeated here.

[0375] Optionally, the apparatus 1 further comprises: a cleaning task acquisition module 21, a to-be-cleaned sample determination module 22, a recall model training module 23 and a cleaning model training module 24;

[0376] A cleaning task acquisition module 21 is used to acquire a cleaning task associated with a quality inspection task, and to acquire a text quality inspection segment associated with the cleaning task from a text quality inspection result;

[0377] The to-be-cleaned sample determination module 22 is configured to obtain an initial recall model and an initial cleaning model associated with the cleaning task, and use the training data carrying the text quality inspection fragments obtained from the input text data as the to-be-cleaned sample data associated with the initial recall model and the initial cleaning model;

[0378] The recall model training module 23 is used to train the initial recall model using the sample data to be cleaned and the text quality inspection fragments in the sample data to be cleaned, so as to obtain a target recall model for data recall;

[0379] The cleaning model training module 24 is used to train the initial cleaning model using the sample data to be cleaned and the text quality inspection fragments in the sample data to be cleaned, so as to obtain a target cleaning model for data cleaning.

[0380] The specific implementation of the cleaning task acquisition module 21, the sample to be cleaned determination module 22, the recall model training module 23 and the cleaning model training module 24 can be found in the above Figure 7 The description of the specific process of training the recall model and the cleaning model in the corresponding embodiment will not be repeated here.

[0381] Optionally, the device 1 further includes: a data cleaning module 25;

[0382] The data cleaning module 25 is configured to perform a model cascade process on the target recall model and the target cleaning model to obtain a data cleaning model associated with the cleaning task, and obtain text data to be cleaned associated with a target dataset from a first business database associated with the cleaning task; the target dataset is any dataset deployed under the data type of the text data to be processed;

[0383] The data cleaning module 25 is further configured to input the text data to be cleaned into the target recall model in the data cleaning model, and the target recall model selects target text data to be cleaned that carries text quality inspection segments from the text data to be cleaned, performs text marking on the text quality inspection segments in the target text data to be cleaned, uses the marked text quality inspection segments as cleaning marked segments of the target text data to be cleaned, and outputs the target text data to be cleaned that carries the cleaning marked segments from the target text data to be cleaned;

[0384] The data cleaning module 25 is further configured to input the target text data to be cleaned into the target cleaning model in the data cleaning model, and the target cleaning model identifies cleaning mark segments from the target text data to be cleaned, performs text cleaning on the cleaning mark segments in the target text data to be cleaned, and uses the cleaned target text data as the data to be quality inspected for input into the target quality inspection model.

[0385] The specific implementation of the data cleaning module 25 can be found in the above Figure 7 The description of the specific process of obtaining the data to be quality inspected by text cleaning in the corresponding embodiment will not be repeated here.

[0386] Optionally, the quality inspection task execution module 16 is further configured to input the data to be quality inspected into the target quality inspection model, and the target quality inspection model executes the quality inspection task on the data to be quality inspected to obtain the data quality inspection result of the data to be quality inspected;

[0387] The quality inspection task execution module 16 is further used to determine the quality inspection ratio corresponding to the target data set based on the data quality inspection results;

[0388] The quality inspection task execution module 16 is also used to use the data to be quality inspected associated with the target data set as model training sample data for training the target business model if the quality inspection ratio meets the quality inspection data strategy indicated by the quality inspection task, and add the model training sample data to the second business database corresponding to the target business model.

[0389] The specific implementation method of the quality inspection task execution module 16 for performing data quality inspection on the quality inspection data can be found in the above Figure 3 or Figure 7 The description of the text quality inspection result of the input text data in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here either.

[0390] See Figure 14 , Figure 14 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 14As shown, the computer device 1000 may include: a processor 1001, a network interface 1004 and a memory 1005. In addition, the above-mentioned computer device 1000 may also include: a user interface 1003, and at least one communication bus 1002. The communication bus 1002 is used to realize the connection and communication between these components. The optional user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The memory 1005 may optionally be at least one storage device located away from the aforementioned processor 1001. As Figure 14 As shown, the memory 1005 as a computer-readable storage medium may include an operating system, a network communication module, a user interface module, and a device control application.

[0391] In such Figure 14 In the computer device 1000 shown, the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the device control application stored in the memory 1005 to execute the above Figure 3 or Figure 7 The description of the text data processing method provided in any embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here either.

[0392] In addition, it should be noted that: the embodiment of the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the text data processing device 1 mentioned above, and the computer program includes program instructions. When the processor executes the program instructions, it can execute the above-mentioned text data processing device 1. Figure 3 or Figure 7 The description of the text data processing method provided in any embodiment will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated here. For technical details not disclosed in the computer-readable storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application.

[0393] The computer-readable storage medium may be the text data processing device provided in any of the aforementioned embodiments or the internal storage unit of the computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Furthermore, the computer-readable storage medium may include both the internal storage unit of the computer device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been output or is about to be output.

[0394] In addition, it should be noted that the present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above-mentioned Figure 3 or Figure 7 The method provided in any embodiment. In addition, the description of the beneficial effects of adopting the same method will not be repeated. For technical details not disclosed in the computer program product or computer program embodiment involved in this application, please refer to the description of the method embodiment of this application.

[0395] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0396] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0397] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A text data processing method, characterized in that: The method comprises: Determining text definition information of the text data to be processed based on a quality inspection task associated with the text data to be processed; the text definition information is determined based on text key fields in the text data to be processed and text attributes of the text key fields; Acquire sampled text data from the text data to be processed, and determine text tag information of the sampled text data based on the text definition information; the text tag information is determined based on a sampling key field in the sampled text data and a text attribute of the sampling key field; the sampling key field is a field in the text key field; Acquire a text prompt template associated with the text definition information, determine a tag prompt template corresponding to the text prompt template based on the text tag information, and use the text tag information and the sampled text data as a training sample data pair for training an initial business model; Performing model training on the initial business model using the training sample data pair and the marking prompt template to obtain a target quality inspection model for performing the quality inspection task; Input text data for inputting into the target quality inspection model is acquired from the text data to be processed, and the target quality inspection model performs the quality inspection task on the input text data to obtain a text quality inspection result of the input text data.

2. The method according to claim 1, characterized in that The determining of text definition information of the text data to be processed based on the quality inspection task associated with the text data to be processed includes: Acquire a training text data set for training a target business model, and use each training text data in the training text data set as text data to be processed; the target business model is a business model associated with the target quality inspection model; Obtaining a quality inspection task associated with the training text dataset; Based on the quality inspection task and the data type of the text data to be processed, a quality inspection definition is performed on the text data to be processed, and a text key field of the text data to be processed and a text attribute of the text key field are obtained; The text key field and the text attribute of the text key field are used as text definition information of the text data to be processed.

3. The method according to claim 2, characterized in that The quality inspection task at least includes a watermark quality inspection task; The method of performing quality inspection on the text data to be processed based on the quality inspection task and the data type of the text data to be processed, and obtaining text key fields of the text data to be processed and text attributes of the text key fields, includes: Acquire a watermark quality inspection strategy associated with the watermark quality inspection task, and acquire a first target data type to be watermarked from the data type of the to-be-processed text data based on the watermark quality inspection strategy; In the text data to be processed, the text data to be processed that matches the first target data type is used as the first target text data, the watermark definition is performed on the first target text data to obtain a watermark definition field of the first target text data, and the watermark definition field is used as the text key field of the text data to be processed; The watermark definition field is divided into a watermark type indicated by the watermark quality inspection strategy, and the watermark type is determined as a text attribute of the text key field.

4. The method according to claim 2, characterized in that The quality inspection task at least includes a noise quality inspection task; The method of performing quality inspection on the text data to be processed based on the quality inspection task and the data type of the text data to be processed, and obtaining text key fields of the text data to be processed and text attributes of the text key fields, includes: Acquire a noise quality inspection strategy associated with the noise quality inspection task, and acquire a second target data type to be subjected to noise definition from the data type of the to-be-processed text data based on the noise quality inspection strategy; In the text data to be processed, the text data to be processed that matches the second target data type is used as the second target text data, the noise definition is performed on the second target text data to obtain a noise definition field of the second target text data, and the noise definition field is used as the text key field of the text data to be processed; The noise definition field is classified into a noise type indicated by the noise quality inspection strategy, and the noise type is determined as a text attribute of the text key field.

5. The method according to claim 1, wherein The step of obtaining sampled text data from the text data to be processed includes: Based on the data type of the text data to be processed, sampling the text data to be processed to obtain a sample data set that matches the data type of the text data to be processed; The sampled text data is determined based on a sampling data set that matches the data type of the text data to be processed.

6. The method according to claim 5, characterized in that The data types of the text data to be processed include at least a first data type and a second data type; The method of sampling the text data to be processed based on the data type of the text data to be processed to obtain a sample data set that matches the data type of the text data to be processed includes: Determining the first data type from the data types of the text data to be processed, extracting N text data matching the first data type from the text data to be processed, and using the extracted N text data matching the first data type as a first sampling data set; N is a positive integer; Determining the second data type from the data types of the text data to be processed, extracting N text data matching the second data type from the text data to be processed, and using the extracted N text data matching the second data type as a second sampling data set; Based on the first sampling data set and the second sampling data set, a sampling data set matching the data type of the text data to be processed is determined.

7. The method according to claim 1, characterized in that The determining of the text tag information of the sampled text data based on the text definition information includes: Distributing the sampled text data to a plurality of quality inspection marking terminals associated with the quality inspection task, wherein one quality inspection marking terminal is used to perform text annotation on the sampled text data based on the text definition information, thereby obtaining text annotation information for the sampled text data by one quality inspection marking terminal; Receiving text annotation information returned by each quality inspection mark terminal for the sampled text data, cross-checking the received text annotation information, and obtaining a cross-check result for the sampled key fields in the sampled text data and the text attributes of the sampled key fields; If the cross-check result indicates that the received text annotation information is consistent, the text tag information of the sampled text data is determined based on the sampling key field and the text attribute of the sampling key field.

8. The method according to claim 1, characterized in that The quality inspection task at least includes a watermark quality inspection task, the text key field includes a watermark definition field, and the text attribute of the text key field includes the watermark type to which the watermark definition field belongs; The determining of the text tag information of the sampled text data based on the text definition information includes: Based on the watermark quality inspection task, obtaining watermark definition information including the watermark definition field and the watermark type from the text definition information, and determining a prompt template constructed based on the watermark definition information as a watermark prompt template; Inputting the sampled text data into the watermark prompt template, and having the watermark prompt template output the marked reply information of the sampled text data based on the watermark definition information; Sending the sampled text data to a watermark quality inspection terminal associated with the watermark quality inspection task, so that the watermark quality inspection terminal performs text marking on the sampled text data based on the watermark definition information to obtain watermark marking information of the sampled text data; The watermark marking information returned by the watermark quality inspection terminal is received, and the text marking information of the sampled text data is determined based on the watermark marking information and the marking reply information.

9. The method according to claim 8, characterized in that The determining of the text marking information of the sampled text data based on the watermark marking information and the marking reply information includes: Using the watermark text segment of the sampled text data carried in the watermark marking information as the first watermark segment, and using the watermark text segment of the sampled text data carried in the marking reply information as the second watermark segment; If the first watermark segment contains a field different from that of the second watermark segment, the second watermark segment in the mark reply information is corrected to be the first watermark segment, and the watermark text field indicated by the first watermark segment is used as a sampling key field in the sampled text data, and the watermark type to which the watermark text field indicated by the first watermark segment belongs is used as a text attribute of the sampling key field; Based on the sampling key field and the text attribute of the sampling key field, text tag information of the sampling text data is determined.

10. The method according to claim 1, characterized in that The initial business model is trained using the training sample data pair and the marking prompt template to obtain a target quality inspection model for performing the quality inspection task, including: Inputting the sampled text data in the training sample data pair into the initial business model, having the initial business model predict and output text tag information of the sampled text data, and using the predicted output text tag information as predicted tag information of the sampled text data; Determining a model loss value of the initial business model based on the text tag information and the prediction tag information in the training sample data pair; If the model loss value of the initial business model does not meet the model optimization conditions indicated by the quality inspection task, the model parameters of the initial business model are optimized until the model loss value of the business model after parameter optimization meets the model optimization conditions. The business model whose model loss value meets the model optimization conditions is used as the target quality inspection model for executing the quality inspection task; the model optimization conditions are used to indicate that the business model after parameter optimization predicts and outputs predicted tag information that matches the text tag information according to the tag prompt template.

11. The method according to claim 1, wherein The step of obtaining input text data for inputting into the target quality inspection model from the text data to be processed, and having the target quality inspection model perform the quality inspection task on the input text data to obtain a text quality inspection result of the input text data includes: Obtain a target data type from the data type of the text data to be processed, and obtain K data sets deployed under the target data type; one data set corresponds to one data identifier, K is a positive integer, and each of the K data sets includes at least k1 training data; Based on the k1 training data extracted from each data set, k2 training data associated with the target data type are obtained, and the k2 training data are used as input text data for inputting the target quality inspection model; k2=k1*K; Inputting the input text data into the target quality inspection model, and having the target quality inspection model predict and output predicted tag information that matches the text output style of the tag prompt template when performing the quality inspection task on the input text data based on the tag prompt template, and using the predicted tag information output as the target tag information of the input text data; Information parsing is performed on the target tag information to obtain text parsing information corresponding to the target tag information; based on the data identifier of each data set, text parsing information associated with the training data in each data set is searched in the text parsing information; the text parsing information associated with the training data in each data set that is found is determined as a data set quality inspection result of each data set; and the data set quality inspection result of each data set is used as the text quality inspection result of the input text data.

12. The method according to claim 11, characterized in that The quality inspection task at least includes a watermark quality inspection task; The method further comprises: A target data set is obtained from the K data sets, and a data set quality inspection result that matches the data identifier of the target data set is obtained from the text quality inspection result; the total number of quality inspections corresponding to the data set quality inspection result of the target data set is k1; Searching for data quality inspection results associated with the watermark quality inspection task in the data set quality inspection results of the target data set, using the data quality inspection results associated with the watermark quality inspection task found as watermark detection results, and using the cumulative number of the watermark detection results as the number of watermark detections corresponding to the target data set; Determine the watermark ratio corresponding to the target data set based on the number of watermark detections and the total number of quality inspections, and when each of the K data sets is used as the target data set, obtain the watermark ratio corresponding to each data set; Based on the watermark proportion corresponding to each data set, a watermark proportion curve corresponding to each data set is obtained by fitting, and the watermark proportion curve is displayed on a quality inspection display platform associated with the watermark quality inspection task.

13. The method according to claim 1, wherein The method further comprises: Acquire a cleaning task associated with the quality inspection task, and acquire a text quality inspection segment associated with the cleaning task from the text quality inspection result; Acquire an initial recall model and an initial cleaning model associated with the cleaning task, and use the training data carrying the text quality inspection segment acquired from the input text data as sample data to be cleaned associated with the initial recall model and the initial cleaning model; The initial recall model is trained using the sample data to be cleaned and the text quality inspection fragments in the sample data to be cleaned to obtain a target recall model for data recall; The initial cleaning model is trained using the sample data to be cleaned and the text quality inspection segments in the sample data to be cleaned to obtain a target cleaning model for data cleaning.

14. The method according to claim 13, characterized in that The method further comprises: Performing model cascade processing on the target recall model and the target cleaning model to obtain a data cleaning model associated with the cleaning task, and obtaining text data to be cleaned associated with a target data set from a first business database associated with the cleaning task; the target data set refers to any data set deployed under the data type of the text data to be processed; Inputting the text data to be cleaned into the target recall model in the data cleaning model, having the target recall model screen target text data to be cleaned that carries the text quality inspection segment from the text data to be cleaned, marking the text quality inspection segment in the target text data to be cleaned, using the marked text quality inspection segment as a cleaning mark segment of the target text data to be cleaned, and having the target recall model output the target text data to be cleaned that carries the cleaning mark segment; The target text data to be cleaned is input into the target cleaning model in the data cleaning model, the target cleaning model identifies the cleaning mark segments from the target text data to be cleaned, and the cleaning mark segments in the target text data to be cleaned are subjected to text cleaning. The target text data to be cleaned after text cleaning is used as the data to be quality inspected for input into the target quality inspection model.

15. The method according to claim 14, characterized in that The method further comprises: Inputting the data to be quality inspected into the target quality inspection model, and having the target quality inspection model perform the quality inspection task on the data to be quality inspected to obtain a data quality inspection result of the data to be quality inspected; Determining a quality inspection ratio corresponding to the target data set based on the data quality inspection result; If the quality inspection ratio meets the quality inspection data strategy indicated by the quality inspection task, the data to be quality inspected associated with the target data set will be used as model training sample data for training the target business model, and the model training sample data will be added to the second business database corresponding to the target business model.

16. A text data processing device, characterized in that: The device comprises: A definition information determination module is configured to determine text definition information of the text data to be processed based on a quality inspection task associated with the text data to be processed; the text definition information is determined based on text key fields in the text data to be processed and text attributes of the text key fields; a tag information determining module, configured to obtain sampled text data from the text data to be processed, and determine text tag information of the sampled text data based on the text definition information; the text tag information is determined based on a sampling key field in the sampled text data and text attributes of the sampling key field; the sampling key field is a field in the text key field; a sample data pair determination module, configured to obtain a text prompt template associated with the text definition information, determine a tag prompt template corresponding to the text prompt template based on the text tag information, and use the text tag information and the sampled text data as a training sample data pair for training an initial business model; An initial business model training module is used to perform model training on the initial business model using the training sample data pairs and the marking prompt template, so as to obtain a target quality inspection model for performing the quality inspection task; The quality inspection task execution module is used to obtain input text data for inputting the target quality inspection model from the text data to be processed, and the target quality inspection model executes the quality inspection task on the input text data to obtain a text quality inspection result of the input text data.

17. A computer device, characterized in that: including memory and processor; The memory is connected to the processor, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1 to 15.

18. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which is suitable for being loaded and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1 to 15.

19. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 15 when executed by a processor.