Data processing method and device and electronic equipment

By obtaining the data quality matrix and utilizing knowledge distillation technology of the student model and the teacher model, computing resources are dynamically allocated, which solves the problems of low data screening efficiency and resource waste, improves the robustness and generalization ability of the model, and reduces the cost of model training.

CN120653993APending Publication Date: 2025-09-16CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511102474.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-09-16

Smart Images

  • Figure CN120653993A_ABST
    Figure CN120653993A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device and electronic equipment. The method comprises the steps that a data quality matrix is obtained, and the data quality matrix comprises original data used for model training and quality evaluation scores obtained after the original data are evaluated through multi-dimensional data quality indexes; grouping the data quality matrix according to the quality evaluation score and a preset evaluation threshold to obtain grouped data of different quality levels; calculating power resources needed by the grouped data are determined through the student model, the grouped data are distributed to the corresponding calculating power equipment for processing according to the calculating power resources, the student model is trained by minimizing the total loss between the student model and the teacher model, and the training efficiency is improved. The teacher model takes the first data exceeding a preset quality threshold value in the data quality matrix and the corresponding annotation data as input. The technical problems of low data screening efficiency, computing power resource waste and high model training cost existing in related technologies are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data processing method, device and electronic device. Background Art

[0002] With the rapid development of artificial intelligence technology, enterprises are increasingly demanding efficient processing of large-scale data and model training. However, enterprises still face many challenges in data management and model training:

[0003] (1) Data quality and screening efficiency issues: With the explosive growth of data volume, data types are diverse and of varying quality (e.g., insufficient accuracy, lack of completeness, redundancy, etc.). Data screening technologies in related technologies are unable to efficiently integrate multi-source heterogeneous data, resulting in complex and costly data governance processes. Furthermore, erroneous or unreliable inference results generated by models trained on low-quality data can also reduce the reliability and generalization capabilities of model inference.

[0004] (2) Model training and computing power waste: Large model training usually relies on high-computing hardware (such as GPU clusters), but the resource allocation methods in related technologies often lack a mechanism for coordinating the optimization of data quality and hardware resources. On the one hand, unfiltered low-quality data takes up a large amount of computing power resources, resulting in low training efficiency; on the other hand, the computing power and computing efficiency (computing power per unit power consumption) of heterogeneous hardware resources (such as CPU, GPU, NPU) vary significantly, making it difficult to achieve dynamic resource allocation and optimization, resulting in low hardware utilization and high energy consumption.

[0005] (3) Insufficient model compression and robustness: Model compression techniques, such as knowledge distillation, aim to reduce the computational burden and accelerate the inference process by transferring complex knowledge from large models to smaller models. However, the data processing methods in related technologies do not combine data quality stratification and dynamic training strategies, resulting in insufficient robustness of the compressed models in complex scenarios. In addition, the existing distillation process lacks a dynamic adaptation mechanism for data quality, making it difficult to balance model performance and resource consumption, further limiting efficiency improvements in practical applications.

[0006] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0007] The embodiments of the present application provide a data processing method, device, and electronic device to at least solve the technical problems of low data screening efficiency, waste of computing resources, and high model training costs existing in related technologies.

[0008] According to one aspect of an embodiment of the present application, a data processing method is provided, including: obtaining a data quality matrix, wherein the data quality matrix includes original data used for model training and a quality assessment score obtained by evaluating the original data using multi-dimensional data quality indicators; grouping the data quality matrix according to the quality assessment score and a preset assessment threshold to obtain grouped data of different quality levels; determining the computing power resources required for the grouped data through a student model, and allocating the grouped data to corresponding computing power devices for processing based on the computing power resources, wherein the student model is trained by minimizing the total loss between the student model and the teacher model, and the teacher model takes the first data in the data quality matrix that exceeds the preset quality threshold and the corresponding labeled data as input.

[0009] Optionally, the quality assessment score is determined in the following manner: respectively determine a first assessment score, a second assessment score and a third assessment score of the original data, wherein the first assessment score is used to reflect the accuracy of the original data, the second assessment score is used to reflect the integrity of the original data, and the third assessment score is used to reflect the uniqueness of the original data; respectively determine a first weight corresponding to the first assessment score, a second weight corresponding to the second assessment score, and a third weight corresponding to the third assessment score; and weight the first assessment score, the second assessment score and the third assessment score according to the first weight, the second weight and the third weight to obtain the quality assessment score of the original data.

[0010] Optionally, the data quality matrix is ​​grouped according to the quality assessment score and the preset assessment threshold to obtain grouped data of different quality levels, including: determining the interval greater than 0 and less than the first assessment threshold as the first grouping interval, determining the interval greater than the first assessment threshold and less than the second assessment threshold as the second grouping interval, and determining the interval greater than the second assessment threshold and less than 1 as the third grouping interval; in the data quality matrix, determining the data with the quality assessment score in the first grouping interval as the first grouping data, determining the data with the quality assessment score in the second grouping interval as the second grouping data, and determining the data with the quality assessment score in the third grouping interval as the third grouping data, wherein the data quality of the first grouping data is less than the data quality of the second grouping data, and the quality of the second grouping data is less than the data quality of the third grouping data.

[0011] Optionally, the method also includes: determining an operation instruction corresponding to the second data, wherein the second data is any data in the data quality matrix, and the operation instruction is used to query the abnormality of the second data; processing the second data and the operation instruction through a data screening model to obtain a data processing result; when the data processing result indicates that the second data is abnormal data, removing the second data from the data quality matrix.

[0012] Optionally, the student model is trained in the following manner: determining a target loss function for knowledge distillation of the student model; determining the total loss between the student model and the teacher model based on the target loss function, wherein the total loss includes at least the cross entropy loss between the student model and the labeled data, the knowledge distillation loss between the student model and the teacher model, and the sequence loss between the student model and the teacher model on a specific hidden layer; adjusting the loss weight corresponding to the target loss function until the total loss is less than the preset loss, thereby obtaining the student model.

[0013] Optionally, determining a target loss function for knowledge distillation of the student model includes: determining a first probability distribution and a second probability distribution corresponding to the first data through the teacher model and the student model respectively based on the first data, the labeled data and the preset temperature coefficient, wherein the preset temperature coefficient is used to soften the first probability distribution output by the teacher model; determining the sequence length of the first data, and determining the knowledge distillation loss function based on the sequence length, the first probability distribution and the second probability distribution, wherein the knowledge distillation loss function is used to reflect the probability distribution difference between the student model and the teacher model; determining the first hidden state and the second hidden state of the teacher model and the student model in a specific hidden layer respectively, and determining the sequence loss function based on the sequence length, the first hidden state and the second hidden state, wherein the sequence loss function is used to reflect the hidden state difference between the student model and the teacher model; determining the cross entropy loss function corresponding to the student model based on the first data and the labeled data; determining the target loss function based on the knowledge distillation loss function, the sequence loss function, the cross entropy loss function and the loss weight.

[0014] Optionally, the computing power resources required for the grouped data are determined through the student model, and the grouped data are allocated to the corresponding computing power equipment for processing based on the computing power resources, including: mapping the grouped data to the corresponding computing power tasks through the student model; determining the computing performance indicators corresponding to the computing power tasks; and allocating the grouped data to the computing power equipment corresponding to the computing performance indicators for processing.

[0015] According to another aspect of an embodiment of the present application, a data processing device is also provided, including: an acquisition module for acquiring a data quality matrix, wherein the data quality matrix includes the original data used for model training and the quality assessment score obtained after evaluating the original data using multi-dimensional data quality indicators; a grouping module for grouping the data quality matrix according to the quality assessment score and a preset assessment threshold to obtain grouped data of different quality levels; an allocation module for determining the computing power resources required for the grouped data through a student model, and allocating the grouped data to corresponding computing power equipment for processing based on the computing power resources, wherein the student model is trained by minimizing the total loss between the student model and the teacher model, and the teacher model takes the first data in the data quality matrix that exceeds the preset quality threshold and the corresponding labeled data as input.

[0016] According to another aspect of the embodiments of the present application, an electronic device is provided, including: a memory and a processor, wherein the memory is used to store program instructions; the processor is connected to the memory and is used to execute the above-mentioned data processing method.

[0017] According to another aspect of the embodiments of the present application, a non-volatile storage medium is provided. The non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the above-mentioned data processing method by running the computer program.

[0018] According to another aspect of the embodiments of the present application, a computer program product is provided, including computer instructions, which implement the above-mentioned data processing method when executed by a processor.

[0019] In an embodiment of the present application, a data quality matrix is ​​obtained, wherein the data quality matrix includes the original data used for model training and the quality assessment score obtained by evaluating the original data using multi-dimensional data quality indicators; the data quality matrix is ​​grouped according to the quality assessment score and the preset assessment threshold to obtain grouped data of different quality levels; the computing power resources required for the grouped data are determined by the student model, and the grouped data are allocated to the corresponding computing power equipment for processing based on the computing power resources, wherein the student model is trained by minimizing the total loss between the student model and the teacher model, and the teacher model uses the first data in the data quality matrix that exceeds the preset quality threshold and the corresponding labeled data as input, thereby achieving the purpose of efficient data screening and computing power resource allocation, thereby realizing the optimal utilization of model training resources, and at the same time enhancing the technical effect of the robustness and generalization ability of the model, thereby solving the technical problems of low data screening efficiency, waste of computing power resources and high model training cost existing in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0021] Figure 1 is a hardware structure diagram of a computer terminal for implementing a data processing method according to an embodiment of the present application;

[0022] Figure 2 is a flow chart of a data processing method according to an embodiment of the present application;

[0023] Figure 3 is a structural diagram of a data processing system according to an embodiment of the present application;

[0024] Figure 4 It is a structural diagram of a data processing device according to an embodiment of the present application. DETAILED DESCRIPTION

[0025] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0027] First, some nouns or terms that appear in the process of explaining the embodiments of this application are subject to the following explanations:

[0028] Data quality refers to the comprehensive performance of data's accuracy, completeness, consistency, timeliness, credibility, and relevance. It is a key indicator of data suitability for a specific purpose. High-quality data is the foundation of data analysis, decision-making, and business operations, directly impacting the effectiveness of data-driven decisions and business outcomes.

[0029] Computational Power (CP): refers to the ability of a server to output results after processing data. It is a comprehensive indicator to measure the computing power of a data center. The larger the value, the stronger the comprehensive computing power.

[0030] Computational Efficiency (CE): This refers to the ratio of a data center's computing power to its power, specifically the computing power generated per watt of power. This efficiency factor considers both data center computing performance and power consumption. A larger value indicates greater computing power per unit of power, and therefore higher efficiency.

[0031] Knowledge Distillation (KD): refers to the transfer of knowledge from a larger or more complex model (called the teacher) to a smaller, simpler model (called the student).

[0032] In order to solve the problem of low efficiency of data screening and distribution in related technologies, the present invention provides a data processing method that can be run on Figure 1 Among the computer terminals shown, the computer terminal will be described below.

[0033] The data processing method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG1 shows a hardware structure block diagram of a computer terminal for implementing a data processing method. Figure 1 As shown, the computer terminal 10 may include one or more (illustrated by 102a, 102b, ..., 102n in the figure) processors (the processor may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions connected via a wired and / or wireless network. In addition, it may also include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, and a BUS bus. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0034] It should be noted that the one or more processors and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10. As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0035] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the data processing method in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned data processing method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0036] The transmission module 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission module 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission module 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0037] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 .

[0038] It should be noted that, in some optional embodiments, the above Figure 1 The computer terminal shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the computer terminal described above.

[0039] In the above-mentioned operating environment, an embodiment of the present application provides an embodiment of a data processing method. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0040] Figure 2 is a flow chart of a data processing method according to an embodiment of the present application, such as Figure 2 As shown, the method includes the following steps:

[0041] Step S202: Obtain a data quality matrix, wherein the data quality matrix includes the original data used for model training and a quality assessment score obtained by evaluating the original data using multi-dimensional data quality indicators.

[0042] Step S204 : grouping the data quality matrix according to the quality evaluation score and the preset evaluation threshold to obtain grouped data of different quality levels.

[0043] Step S206: Determine the computing power resources required for the grouped data through the student model, and allocate the grouped data to the corresponding computing power equipment for processing based on the computing power resources. The student model is trained by minimizing the total loss between the student model and the teacher model. The teacher model takes the first data in the data quality matrix that exceeds the preset quality threshold and the corresponding labeled data as input.

[0044] Through steps S202 to S206 above, efficient data screening and computing resource allocation are achieved, thereby achieving optimal utilization of model training resources and enhancing the robustness and generalization capabilities of the model. This resolves the technical issues of low data screening efficiency, waste of computing resources, and high model training costs that exist in related technologies. This is described in detail below.

[0045] In the above step S202, it is first necessary to perform quality assessment on the original data participating in the large model training using multiple dimensional data quality indicators (such as accuracy, completeness, uniqueness, etc.), and normalize the quality assessment scores to form a data quality matrix, where the higher the quality score, the higher the data quality of the original data.

[0046] Optionally, the quality assessment score is determined in the following manner: respectively determine a first assessment score, a second assessment score and a third assessment score of the original data, wherein the first assessment score is used to reflect the accuracy of the original data, the second assessment score is used to reflect the integrity of the original data, and the third assessment score is used to reflect the uniqueness of the original data; respectively determine a first weight corresponding to the first assessment score, a second weight corresponding to the second assessment score, and a third weight corresponding to the third assessment score; and weight the first assessment score, the second assessment score and the third assessment score according to the first weight, the second weight and the third weight to obtain the quality assessment score of the original data.

[0047] In the embodiments of the present application, the initial screening of data quality mainly involves the evaluation of accuracy, completeness, and uniqueness. Among them, the accuracy evaluation is used to determine whether the original data truly reflects the objective facts and whether there are errors or deviations; the completeness evaluation is used to determine whether the original data is comprehensive and whether key information is missing; the uniqueness evaluation is used to determine whether the original data is unique and whether there are duplications or redundancies. The calculation process of the data quality evaluation score is as follows:

[0048] The specific expression is as follows:

[0049] F=αF1+βF2+γF3

[0050] α+β+γ=1

[0051] Where F1 represents the first evaluation score, α represents the first weight corresponding to F1;

[0052] F2 represents the second evaluation score, β represents the second weight corresponding to F2;

[0053] F3 represents the third evaluation score, γ represents a third weight corresponding to F3.

[0054] The three evaluation scores are combined and processed using a weighted summation method, combining their weights to produce a comprehensive quality assessment score for each data item. The introduction of this scoring mechanism not only ensures the comprehensiveness and accuracy of data quality assessments, but also provides a solid foundation for subsequent data screening, model training, and efficient allocation of computing resources. It effectively addresses the problems of uneven data quality and inefficient resource utilization in traditional data processing processes.

[0055] Furthermore, the data quality matrix can be re-screened, and relevant business parameters can be combined with small model training to filter out abnormal data, including: determining the operation instruction corresponding to the second data, wherein the second data is any data in the data quality matrix, and the operation instruction is used to query the abnormal situation of the second data; processing the second data and the operation instruction through the data screening model to obtain the data processing result; when the data processing result indicates that the second data is abnormal data, the second data is eliminated from the data quality matrix.

[0056] In an embodiment of the present application, in order to ensure the purity of the data quality matrix, a strategy for rescreening the data after preliminary screening is proposed to identify and eliminate abnormal data therein. This process involves determining an operating instruction associated with any one of the data (i.e., the second data) in the data quality matrix, and the operating instruction specifically refers to a specific task for querying and judging data abnormalities. Then, the second data and the corresponding operating instruction are processed using a data screening model to generate a data processing result, which includes the judgment of data abnormality. If the data processing result shows that the second data does have anomalies, such as unreasonable values, statistical outliers, or features that do not conform to the expected pattern, the data will be removed from the data quality matrix to maintain the overall quality of the data set and the accuracy of subsequent model training.

[0057] Taking the knowledge graph as an example, the query parameter value is matched with the corresponding "data quality matrix" for correlation, and the queried business includes information such as business type, business name, and business weight.

[0058] The prompts for data screening model queries are shown in Table 1:

[0059] Table 1 Prompt example for data screening model query

[0060]

[0061] During the inference process of the data screening model, an indicator query is performed. The queried indicator and related indicators are used to answer questions. The queried relationships include the indicator name and the indicator error rate. The prompts used for re-screening are shown in Table 2:

[0062] Table 2 Prompt example for rescreening data screening model

[0063]

[0064]

[0065] In the above process, the re-screening mechanism makes up for the possible omissions in the initial data screening. At the same time, through dynamic adjustment and continuous optimization, the reliability and applicability of the data quality matrix are further improved, ensuring that the data used for model training is not only sufficient in quantity but also of high quality.

[0066] In the above step S204, the data quality matrix can be weighted and grouped according to the quality assessment score of each data and the preset assessment threshold to obtain grouped data of different quality levels, including: determining the interval greater than 0 and less than the first assessment threshold as the first grouping interval, determining the interval greater than the first assessment threshold and less than the second assessment threshold as the second grouping interval, and determining the interval greater than the second assessment threshold and less than 1 as the third grouping interval; in the data quality matrix, determining the data with the quality assessment score in the first grouping interval as the first grouping data, determining the data with the quality assessment score in the second grouping interval as the second grouping data, and determining the data with the quality assessment score in the third grouping interval as the third grouping data, wherein the data quality of the first grouping data is less than the data quality of the second grouping data, and the quality of the second grouping data is less than the data quality of the third grouping data.

[0067] In an embodiment of the present application, the data quality matrix is ​​further subdivided into grouped data of different quality levels. Specifically, a first evaluation threshold (such as 0.5) and a second evaluation threshold (such as 0.75) are first set, corresponding to the lower limits of the intermediate weight and the advanced weight, respectively. Subsequently, the low-scoring data whose quality assessment score is in the interval of (0,0.5] (i.e., the first grouping interval) is marked as a low-level weight and classified into the low-quality data group, i.e., the first grouping data; the medium-scoring data whose quality assessment score is in the interval of (0.5,0.75] (i.e., the second grouping interval) is marked as an intermediate weight and classified into the medium-quality data group, i.e., the second grouping data; the high-scoring data whose quality assessment score is in the interval of (0.75,1] (i.e., the third grouping interval) is marked as an advanced weight and classified into the high-quality data group, i.e., the third grouping data.

[0068] Through this precise quality grouping, computing resources can be allocated more intelligently, ensuring that high-quality data is processed more efficiently, while low-quality data is used reasonably, avoiding resource waste and improving the efficiency and effectiveness of overall data processing and model training.

[0069] In the above step S206, the computing resources required for data of different quality levels, such as computing power and computing efficiency, can be determined by the student model, and allocated to the corresponding hardware computing resource equipment for processing. Among them, the training of the student model is achieved by minimizing the total loss between the student model and the teacher model. The teacher model uses high-quality data that has been strictly screened and labeled as input, and transfers knowledge by training the student model, while ensuring that the student model can still maintain similar performance to the teacher model after compression. During the model distillation process, the system will focus on using high-scoring data and improve the quality of the student model through mixed training (including manually labeled data). Finally, according to the quantitative indicators of computing power and efficiency, the data is matched with the hardware equipment to achieve the collaborative work of multiple hardware resources, thereby improving resource utilization and overall computing efficiency.

[0070] Optionally, the student model is trained in the following manner: determining a target loss function for knowledge distillation of the student model; determining the total loss between the student model and the teacher model based on the target loss function, wherein the total loss includes at least the cross entropy loss between the student model and the labeled data, the knowledge distillation loss between the student model and the teacher model, and the sequence loss between the student model and the teacher model on a specific hidden layer; adjusting the loss weight corresponding to the target loss function until the total loss is less than the preset loss, thereby obtaining the student model.

[0071] Among them, determining the target loss function for knowledge distillation of the student model includes: determining the first probability distribution and the second probability distribution corresponding to the first data through the teacher model and the student model respectively based on the first data, the labeled data and the preset temperature coefficient, wherein the preset temperature coefficient is used to soften the first probability distribution output by the teacher model; determining the sequence length of the first data, and determining the knowledge distillation loss function based on the sequence length, the first probability distribution and the second probability distribution, wherein the knowledge distillation loss function is used to reflect the probability distribution difference between the student model and the teacher model; determining the first hidden state and the second hidden state of the teacher model and the student model in a specific hidden layer respectively, and determining the sequence loss function based on the sequence length, the first hidden state and the second hidden state, wherein the sequence loss function is used to reflect the hidden state difference between the student model and the teacher model; determining the cross entropy loss function corresponding to the student model based on the first data and the labeled data; determining the target loss function based on the knowledge distillation loss function, the sequence loss function, the cross entropy loss function and the loss weight.

[0072] In the embodiment of this application, the training of the student model is a comprehensive optimization process involving multiple loss functions. The purpose is to ensure that the student model can efficiently and accurately inherit the knowledge and performance of the teacher model through the knowledge distillation mechanism, while achieving better resource utilization under specific conditions. The specific process can be as follows:

[0073] First, define the target loss function to guide the learning and optimization of the student model. This target loss function consists of at least the cross entropy loss between the student model and the labeled data, the knowledge distillation loss between the student model and the teacher model, and the sequence loss between the two at a specific hidden layer. The specific expression is as follows:

[0074] L=L CLM +L logits +α·L is

[0075] Where L represents the target loss function, which is used to calculate the total loss between the student model and the teacher model; L logitsrepresents the knowledge distillation loss function, which is used to reflect the difference in probability distribution between the student model and the teacher model; L is represents the sequence loss function, which is used to reflect the difference in hidden states between the student model and the teacher model; L CLM Represents the cross entropy loss function, which is used to calculate the cross entropy loss between the student model and the teacher model; α represents the loss weight, which is used to balance L logits and L is significant size differences.

[0076] Specifically, for the knowledge distillation loss function L logits :By comparing the output probability distributions of the teacher and student models on high-quality data, we can construct a knowledge distillation loss function to quantify the output difference between the two, thereby guiding the parameter adjustment of the student model and simulating the decision logic of the teacher model as much as possible. The specific expression is as follows:

[0077]

[0078] Where, and Represent the first probability distribution and second probability distribution of the teacher model and the student model on the kth token respectively, and l represents the sequence length of the high-quality data (i.e., the first data mentioned above) in the data quality matrix.

[0079]

[0080] Where τ represents the temperature coefficient, which is used to control the smoothness of the output probability distribution; |V| represents the vocabulary size, that is, the total number of different tokens that the model can predict; x i 、x j Represents the unnormalized logarithmic probability (logit) of the model for the i-th and j-th token respectively; exp represents the exponential function, which is used to convert logit to probability.

[0081] It is worth noting that in this process, a temperature coefficient τ is introduced to soften the probability distribution of the teacher model output, making it more suitable for student model learning.

[0082] For the sequence loss function L is :By comparing the hidden states of the teacher model and the student model at a specific hidden layer, we can determine the sequence loss associated with the sequence length of high-quality data, ensuring that the student model not only imitates the teacher model at the output layer but also remains consistent with the teacher in the learned representation of the intermediate layers, thereby enhancing the model's generalization ability and ability to handle complex scenarios. The specific expression is as follows:

[0083]

[0084] Where, and denote the first hidden state and the second hidden state of the teacher model and the student model of the i-th token in the k-th hidden layer, respectively, l denotes the sequence length of the high-quality data (i.e., the first data mentioned above) in the data quality matrix, and H denotes the selected intermediate state set.

[0085] For the cross entropy loss function L CLM : The cross entropy loss between the student model and the labeled data is used to measure the error of the student model output relative to the true labeled result, ensuring that the student model can accurately perform classification or prediction tasks and improve the accuracy of the model.

[0086] Then, based on the target loss function, iterative training is performed. By adjusting the loss weight α, the optimal balance point is found so that the total loss reaches the preset minimization standard. A student model with performance comparable to that of the uncompressed model (teacher model) can be obtained, but its size is much smaller than that of the teacher model, thereby achieving model compression and efficient use of resources.

[0087] Optionally, the computing power resources required for the grouped data are determined through the student model, and the grouped data are allocated to the corresponding computing power equipment for processing based on the computing power resources, including: mapping the grouped data to the corresponding computing power tasks through the student model; determining the computing performance indicators corresponding to the computing power tasks; and allocating the grouped data to the computing power equipment corresponding to the computing performance indicators for processing.

[0088] In the embodiment of this application, a comprehensive data-driven computing power optimization strategy can be formed by building a unified operator library, supporting cross-architecture compilation, abstracting the computing power call interface, and coordinating the student model with heterogeneous hardware resources. The specific process can be as follows:

[0089] 1. Unified large-model operator library: By providing a unified standard operator interface, it allows upper-layer reasoning and training frameworks to call operators in a consistent manner. This not only simplifies the implementation of large models, but also supports operator optimization and acceleration, improving overall computing efficiency.

[0090] 2. Cross-architecture compilation: The use of cross-architecture compilation technology ensures that the same code can run on different hardware devices, such as CPUs, GPUs, TPUs, and FPGAs, without the need for additional adaptation and optimization. This greatly improves code reusability and system flexibility, reduces development and deployment costs, and also provides technical support for the dynamic allocation of computing power tasks.

[0091] 3. Unified operator interface: supports unified abstraction and encapsulation of the underlying heterogeneous computing power calling interface, supports interfaces such as calculation graphs and runtime management operator development, and the underlying unified mapping adaptation layer supports unified abstraction and encapsulation of the calling interfaces of different hardware devices.

[0092] 4. Through the collaborative work of the student model and multiple hardware resources, dynamic matching and optimization based on data quality and computing power and efficiency are achieved. The specific process can be as follows:

[0093] First, the student model analyzes a data quality matrix with different weights and maps different grouped data into a series of computing tasks. This mapping process essentially quantifies data processing requirements, estimating the basic computational effort and complexity required to complete data processing based on different data quality levels.

[0094] Secondly, based on the mapping of computing tasks, we further determine the corresponding computing performance indicators. This computing performance indicator comprehensively considers the computing output capacity and computing efficiency (i.e., computing power output per unit power consumption) to evaluate the hardware device's ability to handle specific tasks. By quantifying computing performance, we can more intuitively compare the advantages and disadvantages of different hardware devices, providing a scientific basis for resource scheduling.

[0095] Finally, based on the computing power tasks and computing performance indicators, different grouped data are allocated to the most suitable computing power devices for processing. Specifically, the system uses hardware resource pooling technology, treating various heterogeneous computing powers (such as CPU, GPU, TPU, FPGA, etc.) as a unified resource pool, and dynamically scheduling them according to the computing power requirements of the task and the computing efficiency characteristics of the device. Through intelligent matching, high-quality data is preferentially allocated to devices with high computing efficiency, such as GPUs, so that complex data processing and model training can be completed quickly and efficiently. For low-quality or low-complexity data, devices with lower computing efficiency but better cost-effectiveness, such as CPUs, may be used to save resources.

[0096] In the above process, computing power (CP) refers to the ability of the server to output results after processing data. It is a comprehensive indicator to measure the computing power of the data center. The larger the value, the stronger the comprehensive computing power. Computing power (CP) includes general computing power represented by the central processing unit (CPU) and high-performance computing power represented by the graphics processing unit (GPU). That is, CP = CP 通用 +CP 高效能 , the commonly used unit of measurement is EELOPS (10 18 FLOPS).

[0097] Computing efficiency (CE) refers to the ratio of data center computing power to power, that is, the computing power generated per watt of power in the data center. It is an efficiency that considers both the computing performance and power of the data center. The larger the value, the stronger the computing power per unit power and the higher the efficiency. The specific expression is as follows:

[0098]

[0099] Where, PC IT Indicates the overall power consumption of IT equipment in the data center, in W.

[0100] Figure 3 This is a structural diagram of a data processing system according to an embodiment of the present application. The data processing system includes a data fusion and screening module, a data distillation model training module, and a resource collaboration module. The specific analysis is as follows:

[0101] (1) Data fusion and screening module: First, the original data involved in the large model training is subjected to multi-dimensional data quality assessment and normalization processing to obtain a data quality matrix. Then, the data quality matrix is ​​re-screened using a small model (such as a data screening model) to filter out abnormal data. Finally, the data quality matrix is ​​grouped according to the quality assessment score and the preset assessment threshold to obtain grouped data of different quality levels.

[0102] (2) Data distillation model training module: Through model distillation technology, the total loss between the model and the teacher model is minimized to obtain the student model. Subsequently, the trained student model is used to map different grouped data to corresponding computing tasks and send them to the corresponding computing equipment for processing.

[0103] (3) Resource collaboration module: It is used to emphasize the synergy of heterogeneous hardware resources during model training. Different types of computing infrastructure are intelligently allocated according to task requirements to jointly complete model training and data processing.

[0104] In the embodiments of the present application, a new framework integrating data quality assessment and resource collaborative optimization is proposed. This framework realizes multi-level screening and value maximization of data by constructing a sophisticated data quality assessment system. Specifically, the present application adopts knowledge distillation technology, combined with data quality matrix and model training, to innovatively solve the problem of decreased robustness in the process of model compression. At the same time, through the quantification and intelligent scheduling of computing power and efficiency, it ensures the efficient utilization and performance optimization of heterogeneous hardware resources. This multi-dimensional collaborative solution from data, model to computing power not only significantly improves the efficiency of data screening and model training, reduces resource consumption, but also opens up a new path for building high-performance, low-cost, green and energy-saving computing systems. It is an important technological breakthrough in the fields of artificial intelligence, big data processing and cloud computing.

[0105] According to an embodiment of the present application, a data processing device is provided. It should be noted that the data processing device of the embodiment of the present application can be used to execute the data processing method provided in the embodiment of the present application. The data processing device provided in the embodiment of the present application is introduced below.

[0106] Figure 4This is a structural diagram of a data processing device provided according to an embodiment of the present application. Figure 4 As shown, the device includes:

[0107] An acquisition module 40 is configured to acquire a data quality matrix, wherein the data quality matrix includes the original data used for model training and a quality assessment score obtained by evaluating the original data using multi-dimensional data quality indicators;

[0108] A grouping module 42 is used to group the data quality matrix according to the quality assessment score and a preset assessment threshold to obtain grouped data of different quality levels;

[0109] The allocation module 44 is used to determine the computing power resources required for the grouped data through the student model, and allocate the grouped data to the corresponding computing power equipment for processing based on the computing power resources, wherein the student model is trained by minimizing the total loss between it and the teacher model, and the teacher model takes the first data in the data quality matrix that exceeds the preset quality threshold and the corresponding labeled data as input.

[0110] Through the acquisition module, grouping module and allocation module in the above-mentioned data processing device, the purpose of efficient data screening and computing resource allocation is achieved, thereby realizing the optimal utilization of model training resources, while enhancing the technical effect of the robustness and generalization ability of the model, and thus solving the technical problems of low data screening efficiency, waste of computing resources and high model training cost existing in related technologies.

[0111] In the data processing device provided in an embodiment of the present application, the acquisition module is also used to respectively determine a first evaluation score, a second evaluation score and a third evaluation score of the original data, wherein the first evaluation score is used to reflect the accuracy of the original data, the second evaluation score is used to reflect the integrity of the original data, and the third evaluation score is used to reflect the uniqueness of the original data; respectively determine a first weight corresponding to the first evaluation score, a second weight corresponding to the second evaluation score, and a third weight corresponding to the third evaluation score; and perform weighted processing on the first evaluation score, the second evaluation score and the third evaluation score according to the first weight, the second weight and the third weight to obtain a quality evaluation score of the original data.

[0112] In the data processing device provided in an embodiment of the present application, the acquisition module is also used to determine an operation instruction corresponding to the second data, wherein the second data is any data in the data quality matrix, and the operation instruction is used to query the abnormality of the second data; the second data and the operation instruction are processed by the data screening model to obtain a data processing result; when the data processing result indicates that the second data is abnormal data, the second data is eliminated from the data quality matrix.

[0113] In the data processing device provided in an embodiment of the present application, the grouping module is also used to determine an interval greater than 0 and less than a first evaluation threshold as a first grouping interval, determine an interval greater than the first evaluation threshold and less than a second evaluation threshold as a second grouping interval, and determine an interval greater than the second evaluation threshold and less than 1 as a third grouping interval; in the data quality matrix, data whose quality evaluation score is located in the first grouping interval is determined as first grouped data, data whose quality evaluation score is located in the second grouping interval is determined as second grouped data, and data whose quality evaluation score is located in the third grouping interval is determined as third grouped data, wherein the data quality of the first grouped data is less than the data quality of the second grouped data, and the quality of the second grouped data is less than the data quality of the third grouped data.

[0114] In the data processing device provided in the embodiment of the present application, the allocation module is also used to map the grouped data to corresponding computing tasks through the student model; determine the computing performance indicators corresponding to the computing tasks; and allocate the grouped data to the computing equipment corresponding to the computing performance indicators for processing.

[0115] The data processing device provided in the embodiment of the present application also includes a training module 46, which is used to determine the target loss function for knowledge distillation of the student model; determine the total loss between the student model and the teacher model based on the target loss function, wherein the total loss includes at least the cross entropy loss between the student model and the labeled data, the knowledge distillation loss between the student model and the teacher model, and the sequence loss between the student model and the teacher model on a specific hidden layer; adjust the loss weight corresponding to the target loss function until the total loss is less than the preset loss, and obtain the student model.

[0116] In the data processing device provided in an embodiment of the present application, the training module is also used to determine the first probability distribution and the second probability distribution corresponding to the first data through the teacher model and the student model respectively based on the first data, the labeled data and the preset temperature coefficient, wherein the preset temperature coefficient is used to soften the first probability distribution output by the teacher model; determine the sequence length of the first data, and determine the knowledge distillation loss function based on the sequence length, the first probability distribution and the second probability distribution, wherein the knowledge distillation loss function is used to reflect the probability distribution difference between the student model and the teacher model; determine the first hidden state and the second hidden state of the teacher model and the student model in a specific hidden layer respectively, and determine the sequence loss function based on the sequence length, the first hidden state and the second hidden state, wherein the sequence loss function is used to reflect the hidden state difference between the student model and the teacher model; determine the cross entropy loss function corresponding to the student model based on the first data and the labeled data; determine the target loss function based on the knowledge distillation loss function, the sequence loss function, the cross entropy loss function and the loss weight.

[0117] An embodiment of the present application further provides an electronic device, comprising: a memory and a processor, wherein the memory is used to store program instructions; the processor is connected to the memory and is used to execute and implement the above-mentioned data processing method.

[0118] It should be noted that the above electronic equipment is used to perform Figure 2 The data processing method shown, therefore the relevant explanations in the above data processing method are also applicable to the electronic device and will not be repeated here.

[0119] An embodiment of the present application further provides a non-volatile storage medium, which includes a stored computer program, wherein a device where the non-volatile storage medium is located executes the above-mentioned data processing method by running the computer program.

[0120] It should be noted that the above non-volatile storage medium is used to execute Figure 2 The data processing method shown, therefore the relevant explanations in the above data processing method are also applicable to the non-volatile storage medium, and will not be repeated here.

[0121] An embodiment of the present application also provides a computer program product, including computer instructions, which implement the above-mentioned data processing method when executed by a processor.

[0122] It should be noted that the above-mentioned computer program product is used to execute Figure 2 The data processing method shown, therefore the relevant explanations in the above data processing method are also applicable to the computer program product and will not be repeated here.

[0123] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0124] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0125] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0126] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0127] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0128] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0129] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A data processing method, characterized in that: include: Obtaining a data quality matrix, wherein the data quality matrix includes original data used for model training and quality assessment scores obtained by evaluating the original data using multi-dimensional data quality indicators; Grouping the data quality matrix according to the quality assessment score and a preset assessment threshold to obtain grouped data of different quality levels; The computing power resources required for the grouped data are determined through the student model, and the grouped data are allocated to corresponding computing power devices for processing based on the computing power resources, wherein the student model is trained by minimizing the total loss between the student model and the teacher model, and the teacher model takes the first data exceeding the preset quality threshold in the data quality matrix and the corresponding labeled data as input.

2. The method according to claim 1, characterized in that The quality assessment score is determined as follows: respectively determining a first evaluation score, a second evaluation score, and a third evaluation score for the original data, wherein the first evaluation score is used to reflect the accuracy of the original data, the second evaluation score is used to reflect the completeness of the original data, and the third evaluation score is used to reflect the uniqueness of the original data; respectively determining a first weight corresponding to the first evaluation score, a second weight corresponding to the second evaluation score, and a third weight corresponding to the third evaluation score; The first evaluation score, the second evaluation score, and the third evaluation score are weighted according to the first weight, the second weight, and the third weight to obtain a quality evaluation score of the original data.

3. The method according to claim 1, characterized in that The data quality matrix is ​​grouped according to the quality assessment score and a preset assessment threshold to obtain grouped data of different quality levels, including: Determine an interval greater than 0 and less than a first evaluation threshold as a first grouping interval, determine an interval greater than the first evaluation threshold and less than a second evaluation threshold as a second grouping interval, and determine an interval greater than the second evaluation threshold and less than 1 as a third grouping interval; In the data quality matrix, data whose quality assessment score is in the first grouping interval is determined as first grouping data, data whose quality assessment score is in the second grouping interval is determined as second grouping data, and data whose quality assessment score is in the third grouping interval is determined as third grouping data, wherein the data quality of the first grouping data is less than the data quality of the second grouping data, and the quality of the second grouping data is less than the data quality of the third grouping data.

4. The method according to claim 1, wherein The method further comprises: Determining an operation instruction corresponding to second data, wherein the second data is any item of data in the data quality matrix, and the operation instruction is used to query an abnormality of the second data; Processing the second data and the operation instruction through a data screening model to obtain a data processing result; If the data processing result indicates that the second data is abnormal data, the second data is removed from the data quality matrix.

5. The method according to claim 1, characterized in that The student model is trained in the following way: determining a target loss function for performing knowledge distillation on the student model; Determining a total loss between the student model and the teacher model according to the objective loss function, wherein the total loss includes at least a cross entropy loss between the student model and the labeled data, a knowledge distillation loss between the student model and the teacher model, and a sequence loss between the student model and the teacher model on a specific hidden layer; The loss weight corresponding to the target loss function is adjusted until the total loss is less than the preset loss, thereby obtaining the student model.

6. The method according to claim 5, characterized in that Determine the target loss function for performing knowledge distillation on the student model, including: Determining a first probability distribution and a second probability distribution corresponding to the first data using the teacher model and the student model respectively based on the first data, the labeled data, and a preset temperature coefficient, wherein the preset temperature coefficient is used to soften the first probability distribution output by the teacher model; Determining a sequence length of the first data, and determining a knowledge distillation loss function based on the sequence length, the first probability distribution, and the second probability distribution, wherein the knowledge distillation loss function is used to reflect the difference in probability distribution between the student model and the teacher model; Determining a first hidden state and a second hidden state of the teacher model and the student model in the specific hidden layer, respectively, and determining a sequence loss function according to the sequence length, the first hidden state, and the second hidden state, wherein the sequence loss function is used to reflect the difference in hidden states between the student model and the teacher model; Determining a cross entropy loss function corresponding to the student model based on the first data and the labeled data; The target loss function is determined based on the knowledge distillation loss function, the sequence loss function, the cross entropy loss function and the loss weight.

7. The method according to claim 1, characterized in that Determining the computing resources required for the grouped data using the student model, and allocating the grouped data to corresponding computing devices for processing based on the computing resources, including: Mapping the grouped data to corresponding computing tasks through the student model; Determine the computing performance index corresponding to the computing task; The grouped data is distributed to a computing device corresponding to the computing performance indicator for processing.

8. A data processing device, characterized in that: include: An acquisition module is used to obtain a data quality matrix, wherein the data quality matrix includes the original data used for model training and a quality assessment score obtained by evaluating the original data using multi-dimensional data quality indicators; A grouping module, configured to group the data quality matrix according to the quality assessment score and a preset assessment threshold to obtain grouped data of different quality levels; An allocation module is used to determine the computing power resources required for the grouped data through a student model, and allocate the grouped data to corresponding computing power devices for processing based on the computing power resources, wherein the student model is trained by minimizing the total loss between the student model and the teacher model, and the teacher model takes the first data in the data quality matrix that exceeds a preset quality threshold and the corresponding labeled data as input.

9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory is used to store program instructions; The processor is connected to the memory and is used to execute the data processing method described in any one of claims 1 to 7.

10. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the data processing method according to any one of claims 1 to 7 by running the computer program.

11. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the data processing method according to any one of claims 1 to 7 is implemented.