Model training data generation
By obtaining and updating the model processing labels of the target domain data features and generating high-quality model training datasets, the problem of insufficient performance of large language models in professional vertical fields is solved, and more accurate professional field processing is achieved.
Patent Information
- Application Number
- PCT/CN2025/085413
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-27
- Filing Date
- 2025-03-27
- Publication Date
- 2025-10-02
AI Technical Summary
Large language models perform poorly in professional vertical fields, mainly due to the lack of high-quality professional field data, resulting in insufficient model training data to accurately process complex professional knowledge.
By acquiring the target domain data, determining its domain data features and model processing labels, performing label update processing, generating the target domain model training dataset, and using the dataset to train the generative processing model.
The quality of training samples of the target domain data feature set is improved, a high-quality processing model for the target domain is generated, and the processing capability of the model in professional vertical fields is enhanced.
Smart Images

Figure CN2025085413_02102025_PF_FP_ABST
Abstract
Description
Model training data generation Technical Field
[0001] This specification relates to the field of data processing technology, and in particular to methods, devices, storage media, and electronic devices for generating model training data. Background Art
[0002] In recent years, large language models have performed well in general language processing, but have performed poorly in specialized verticals, which require complex expertise. To improve the performance of large language models in these areas, they need to be fine-tuned using precise domain-specific data, but obtaining high-quality domain-specific data is difficult. Summary of the Invention
[0003] This specification provides a method, device, storage medium and electronic device for generating model training data, and the technical solution is as follows.
[0004] In a first aspect, the present specification provides a method for generating model training data, the method comprising: obtaining a generative processing model and target domain data; determining domain data features corresponding to the target domain data and model processing labels of the domain data features; performing label update processing on the model processing labels of the domain data features based on the domain data features to obtain a target domain model training data set; and using the target domain model training data set to perform model training on the generative processing model.
[0005] In the second aspect, this specification provides a model training data generation device, which includes: an acquisition module, suitable for acquiring a generative processing model and target domain data; a determination module, suitable for determining the domain data features corresponding to the target domain data and the model processing labels of the domain data features; an update module, suitable for performing label update processing on the model processing labels of the domain data features based on the domain data features to obtain a target domain model training data set; a training module, suitable for using the target domain model training data set to perform model training on the generative processing model.
[0006] In a third aspect, this specification provides a computer storage medium, wherein the computer storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the above-mentioned method steps.
[0007] In a fourth aspect, this specification provides an electronic device that may include a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the above-mentioned method steps.
[0008] In a fifth aspect, this specification provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded by a processor to execute any one of the above method steps.
[0009] The beneficial effects brought about by the technical solutions provided in some embodiments of this specification include at least: by determining the domain data features corresponding to the target domain data and the model processing labels of the domain data features, the domain data features with large differences from other domain data features are determined based on the domain data features, and the model processing labels corresponding to the domain data features with large differences from other domain data features are often prone to labeling errors. Therefore, based on the domain data features, the model processing labels of the domain data features that are prone to labeling errors in the domain data features are updated, thereby obtaining a target domain model training data set, improving the training sample quality of the target domain data feature set, and finally using the obtained target domain model training data set to train the generative processing model, thereby obtaining a high-quality processing model for the target domain. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] FIG1 is a schematic diagram of a scenario of a model training data generation system provided in this specification.
[0011] FIG2 is a flow chart of a method for generating model training data provided in an embodiment of this specification.
[0012] FIG3 is a schematic diagram of a process for determining a target domain model training dataset according to an embodiment of this specification.
[0013] FIG4 is a schematic diagram of a process for determining domain-specific data features according to an embodiment of this specification.
[0014] FIG5 is a schematic diagram of another process for determining domain-specific data features according to an embodiment of this specification.
[0015] FIG6 is a schematic diagram of another process for determining domain-specific data features according to an embodiment of this specification.
[0016] FIG7 is a schematic diagram of a process for determining a new specific domain data feature in a domain data feature set according to an embodiment of this specification.
[0017] FIG8 is a schematic diagram of a process for determining domain data features corresponding to target domain data according to an embodiment of this specification.
[0018] FIG9 is a flow chart of a generative processing model and target domain data provided by an embodiment of this specification.
[0019] FIG10 is a schematic diagram of a process for updating a target domain model training dataset provided in an embodiment of this specification.
[0020] FIG11 is a model training data generating device provided in an embodiment of this specification.
[0021] FIG12 is a structural block diagram of an electronic device provided in an embodiment of this specification.
[0022] FIG13 is a schematic diagram of the structure of an operating system and user space provided in an embodiment of this specification.
[0023] FIG14 is an architecture diagram of the Android operating system in FIG13 provided in an embodiment of this specification.
[0024] FIG15 is an architectural diagram of the IOS operating system in FIG13 provided in an embodiment of this specification. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in this specification in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments in this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.
[0026] In the description of this specification, it should be understood that the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In the description of this specification, it should be noted that, unless otherwise expressly specified and limited, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units inherent to these processes, methods, products or devices. For those of ordinary skill in the art, the specific meanings of the above terms in this specification can be understood according to the specific circumstances. In addition, in the description of this specification, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.
[0027] The present specification is described in detail below with reference to specific embodiments.
[0028] Please refer to Figure 1, which is a schematic diagram of a scenario of a model training data generation system provided in this specification. As shown in Figure 1, the model training data generation system may include at least a client cluster and a service platform 100.
[0029] The client cluster may include at least one client, as shown in FIG1 , specifically including client 1 corresponding to user 1, client 2 corresponding to user 2, ..., client n corresponding to user n, where n is an integer greater than 0.
[0030] Each client in the client cluster can be an electronic device with communication capabilities, including but not limited to wearable devices, handheld devices, personal computers, tablet computers, in-vehicle devices, smartphones, computing devices, or other processing devices connected to a wireless modem. Electronic devices may be called different names in different networks, such as user equipment, access terminal, subscriber unit, subscriber station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), electronic devices in 5G network or future evolution network, etc.
[0031] The service platform 100 can be a separate server device, such as a rack-mounted, blade, tower, or cabinet-mounted server device, or a workstation, mainframe computer, or other hardware device with strong computing capabilities; it can also be a server cluster composed of multiple servers. The servers in the service cluster can be symmetrically composed, wherein each server has equivalent functions and status in the transaction link, and each server can provide services to the outside world independently. The independent service can be understood as not requiring the assistance of other servers.
[0032] In one or more embodiments of the present specification, the service platform 100 may establish a communication connection with at least one client in the client cluster, and complete data interaction during the model training data generation process based on the communication connection, such as online transaction data interaction. For example, the service platform 100 may implement model training for the generative processing model based on the target domain model training data set obtained based on the model training data generation method of the present specification.
[0033] The service platform 100 establishes a communication connection with at least one client in the client cluster through a network for interactive communication, wherein the network can be a wireless network or a wired network, the wireless network includes but is not limited to a cellular network, a wireless local area network, an infrared network or a Bluetooth network, and the wired network includes but is not limited to an Ethernet, a universal serial bus (USB) or a controller area network. In one or more embodiments of the specification, technologies and / or formats including Hypertext Markup Language (HTML) and Extensible Markup Language (XML) are used to represent data (such as a target compressed package) exchanged over the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can also be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above-mentioned data communication technologies.
[0034] The model training data generation system embodiments provided in this specification share the same concept as the model training data generation method described in one or more embodiments. The execution entity corresponding to the model training data generation method described in one or more embodiments of the specification may be the aforementioned service platform 100; the execution entity corresponding to the model training data generation method described in one or more embodiments of the specification may also be the electronic device corresponding to the client, specifically determined based on the actual application environment. The implementation process of the model training data generation system embodiment can be found in the following method embodiment, which will not be detailed here.
[0035] Based on the scenario diagram shown in FIG1 , the model training data generation method provided by one or more embodiments of this specification is introduced in detail below.
[0036] Please refer to Figure 2, which is a flow chart illustrating a method for generating model training data according to an embodiment of this specification. This method can be implemented using a computer program and can be run on a model training data generation device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone tool application. The model training data generation device can be a service platform.
[0037] Specifically, the model training data generation method includes the following steps.
[0038] S202: Obtain a generative processing model and target domain data.
[0039] Among them, the generative processing model can be a generative large language model, which is mainly constructed using deep learning technologies such as neural networks. This type of model can learn based on large amounts of text data and combine complex algorithms to understand, interpret and generate human language.
[0040] Here, generative processing models can be applied to target domains to process data. Generally, generative processing models are effective at general language processing tasks, but their performance in specialized verticals is limited, as these domains are often filled with complex terminology and unique operational rules.
[0041] The target domain refers to a specific problem domain or knowledge domain. It can be a collection of domain-specific terminology, concepts, rules, processes, and entities. The target domain can be any specific knowledge domain, such as finance, healthcare, education, or e-commerce, as well as any specific problem domain, specifically an application domain, such as natural language processing, image recognition, or machine learning.
[0042] Therefore, in addition to obtaining a generative processing model for the target domain, it is also necessary to obtain target domain data for the target domain. The target domain data can be of various types, such as structured data, unstructured data, and semi-structured data. Structured data refers to data with a clear format and rules, such as tabular data and relational data in a database; unstructured data refers to data without a clear format and rules, such as text data, image data, and audio data; and semi-structured data is data between structured and unstructured data, such as XML (Extensible Markup Language) documents and JSON (JavaScript Object Notation) data.
[0043] Target domain data can also be from finance, healthcare, education, or transportation. For example, financial data in the financial sector can include securities market data, exchange rate data, financial institution data, and financial product data. Medical data can include patient data, drug data, medical device data, and disease diagnosis data. Data from different domains has its own unique characteristics and application scenarios, so the specificity of each domain must be considered during data processing and analysis.
[0044] S204: Determine domain data features corresponding to the target domain data and model processing labels of the domain data features.
[0045] After obtaining the target domain data, data features are extracted from the target domain data to obtain domain data features. Here, the domain data features include both general domain dimension features and domain vertical dimension features.
[0046] It should be understood that the presentation forms of domain data features include but are not limited to domain data feature vectors, domain data feature matrices, domain data feature graphs, domain data feature values, etc. The selection of these presentation forms depends on the specific application scenario.
[0047] A domain data feature vector is a concise representation of domain data features, representing data attributes or characteristics. For example, in text classification tasks, a domain data feature vector can represent keywords, phrases, or themes in a text. A domain data feature matrix is a more detailed representation of domain data features. A domain data feature graph is an intuitive representation of domain data features, often used to visualize the relationships and network structure between domain data features.
[0048] Specifically, when extracting data features from target domain data, data features of two dimensions can be extracted from the target domain data, namely, the general domain dimension of the target domain data and the domain vertical dimension corresponding to the target domain. The data features of these two dimensions are then fused to obtain domain data features. When the domain data features are domain data feature vectors, data feature extraction includes a data vectorization process; when the domain data features are domain data feature matrices, data feature extraction includes a data matrixization process; when the domain data features are domain data feature graphs, data feature extraction includes a data visualization process; and when the domain data features are domain data feature values, data feature extraction includes a numeration process.
[0049] Since the characteristics of domain data have both the characteristics of the general domain dimension and the characteristics of the domain vertical dimension, the generative processing model can process general processing tasks more efficiently and accurately after learning the characteristics of the general domain dimension; and, after learning the characteristics of the domain vertical dimension, the generative processing model can process the processing tasks of the target domain more efficiently and accurately based on the characteristics of the general domain dimension.
[0050] To train the generative processing model based on domain data features, the domain data features can be automatically labeled based on the task type handled by the generative processing model. For example, the task type handled by the generative processing model may be to identify whether text in the target domain is true or false. In this case, the result of the automated labeling of the domain data features can be a model processing label that automatically labels each domain data feature as true or false based on labeling rules or a pre-trained model.
[0051] Determining the model processing label of the domain data feature may include: obtaining a data labeling rule set, labeling the domain data feature based on the data labeling rule set, and obtaining the model processing label corresponding to the domain data feature.
[0052] The data annotation rule set can be determined based on expert experience and the professional knowledge base of the target field.
[0053] Optionally, determining the model processing label of the domain data feature includes: labeling the domain data feature based on a pre-trained model to obtain the model processing label corresponding to the domain data feature.
[0054] Here, the pre-trained model can be a pre-trained model that can annotate domain data features. Of course, in other embodiments, automatic annotation processing can also be performed based on the pre-trained model and the data annotation rule set. Specifically, the process of automatic annotation processing can be expressed by the following formula: pre (D)=AutoAnnotate(D;R,M).
[0055] Among them, L pre (D) represents the model processing label of the domain data feature D; R represents the data annotation rule set; M represents the pre-trained model; AutoAnnotate represents the automatic annotation function.
[0056] S206: performing label update processing on the model processing labels of the domain data features based on the domain data features to obtain a target domain model training dataset.
[0057] After determining the domain data features, domain data features that differ significantly from other domain data features are identified based on feature similarity between the domain data features. Since model processing labels corresponding to domain data features that differ significantly from other domain data features are often prone to mislabeling, model processing labels for domain data features that are prone to mislabeling can be updated based on the domain data features.
[0058] Therefore, the model processing labels corresponding to the domain data features that are significantly different from other domain data features can be updated, thereby achieving label update processing of the model processing labels of the domain data features based on the domain data features, and obtaining the updated domain data features corresponding to the target domain data and the model processing labels of the domain data features. Afterwards, the target domain model training dataset is obtained based on the domain data features corresponding to the target domain data and the model processing labels of the domain data features.
[0059] S208: Using the target domain model training dataset to perform model training on the generative processing model.
[0060] Among them, after obtaining the target domain model training dataset, the target domain model training dataset is used to perform high-quality model training on the generative processing model, so that the generative processing model can accurately learn the professional domain features corresponding to the target domain data for the target domain.
[0061] In the embodiments provided in this specification, by determining the domain data features corresponding to the target domain data and the model processing labels of the domain data features, the domain data features that are more different from other domain data features in the domain data features are determined based on the domain data features. The model processing labels corresponding to the domain data features that are more different from other domain data features are often prone to labeling errors. Therefore, based on the domain data features, the model processing labels of the domain data features that are prone to labeling errors in the domain data features are updated, thereby obtaining a target domain model training data set, improving the training sample quality of the target domain data feature set, and finally using the obtained target domain model training data set to train the generative processing model, thereby obtaining a high-quality processing model for the target domain.
[0062] Please refer to Figure 3, which is a schematic diagram of a process for determining a target domain model training dataset according to an embodiment of this specification. As shown in Figure 3, in S206, the model processing labels of the domain data features are updated based on the domain data features to obtain the target domain model training dataset, including the following steps.
[0063] S302: Determine specific domain data features based on the domain data features, perform label correction processing on specific model processing labels corresponding to the specific domain data features, and obtain corrected model processing labels.
[0064] After obtaining domain data features with model processing labels, their accuracy is difficult to guarantee because they are typically obtained through automated annotation. Therefore, label correction is necessary for the model processing labels of domain data features. However, this requires extensive expertise in the target domain. Therefore, to improve label correction efficiency and save costs, selective label correction can be performed on domain data features.
[0065] Therefore, based on the feature similarity between domain data features, domain data features that are more different from other domain data features can be determined, that is, specific domain data features. However, the specific model processing labels corresponding to specific domain data features that are more different from other domain data features are often prone to labeling errors.
[0066] Therefore, the specific model processing labels corresponding to the specific domain data features can be modified to obtain modified model processing labels. Here, the label modification process can be based on the training model; or, the label modification process can be performed by expert intervention.
[0067] S304: performing label update processing on the model processing labels of the domain data features based on the modified model processing labels to obtain a target domain model training dataset.
[0068] Among them, the features corresponding to the specific domain data features in the domain data features corresponding to the target domain data are determined, and then the labels of the features corresponding to the specific domain data features are updated based on the modified model processing labels corresponding to the specific domain data features, that is, label replacement is performed, thereby obtaining the domain data features corresponding to the target domain data and the model processing labels of the domain data features corresponding to the updated target domain data. Based on the domain data features corresponding to the target domain data and the model processing labels of the domain data features corresponding to the updated target domain data, a target domain model training dataset is obtained. At this time, the target domain data feature set has a high guarantee in terms of domain expertise for the target domain and the accuracy of the domain features.
[0069] In the embodiments provided in this specification, domain data features that are more different from other domain data features in the domain data features are determined based on domain data features, namely specific domain data features. Compared with ordinary domain data features, specific domain data features are more prone to errors when labeling. Therefore, it is necessary to perform label correction processing on the specific model processing labels corresponding to the specific domain data features to obtain corrected model processing labels. Then, the model processing labels of the domain data features are updated with the corrected model processing labels corresponding to the specific domain data features to obtain the target domain model training data set, thereby effectively improving the training sample quality of the target domain data feature set, avoiding the intervention of human experts, saving costs, and improving the generation quality and efficiency of the target domain data feature set.
[0070] In an embodiment provided in this specification, determining specific domain data features based on domain data features in S302 includes: clustering the domain data features to obtain at least two types of domain data feature sets, and determining specific domain data features based on the domain data feature sets.
[0071] During the selective label correction process for domain data features, the domain data features are first clustered, such as Cluster(D;K), where K is the number of clusters (K is greater than or equal to 2), D is the domain data feature, and Cluster(D;K) is the clustering function of D based on the number of clusters K. Clustering can be performed based on the overall similarity between domain data features, or based on the similarity of one or more dimensions between domain data features.
[0072] After clustering, at least two domain data feature sets are obtained. The domain data features in each domain data feature set are similar in one or more dimensions. Therefore, specific domain data features can be determined based on the degree of similarity between the domain data features in the domain data feature set. Specific domain data features can be features in the domain data feature set that have low similarity to other domain data features. In other words, specific domain data features can be understood as features that are critical to improving the performance of the generative processing model.
[0073] Since the similarity between specific domain data features and other domain data features in the domain data feature set is low, there is a high possibility that the specific model processing labels of specific domain data features contain errors. At this time, the specific model processing labels corresponding to the specific domain data features in each domain data feature set can be corrected to obtain corrected model processing labels for the specific domain data features.
[0074] In the embodiments provided herein, since model processing labels are often obtained through automated annotation, their accuracy is difficult to guarantee. Therefore, label correction processing is required for the model processing labels of domain data features. Before label correction processing, clustering processing is first performed. Since the domain data features in each domain data feature set obtained after clustering processing are similar in one or more dimensions, the specific domain data features have a low similarity with other domain data features in the domain data feature set.
[0075] Therefore, there is a high possibility that the specific model processing labels of specific domain data features are erroneous. At this time, the specific model processing labels corresponding to the specific domain data features in each domain data feature set can be corrected to improve the training sample quality of the target domain data feature set. Finally, the target domain model training data set is used to train the generative processing model to obtain a high-quality processing model for the target domain.
[0076] Please refer to Figure 4, which is a schematic diagram of a process for determining specific domain data features according to an embodiment of this specification. Specifically, in the above embodiment, the domain data features are clustered to obtain at least two domain data feature sets. Determining specific domain data features based on the domain data feature sets includes the following steps.
[0077] S402: Calculate feature similarity between two domain data features, perform clustering based on the feature similarity, and obtain at least two categories of domain data feature sets.
[0078] The feature similarity between each pair of domain data features can be calculated, for example, based on a similarity calculation formula or a Euclidean distance similarity calculation formula. After obtaining the feature similarity between each pair of domain data features, the domain data features are divided based on the feature similarity, and domain data features with higher feature similarity are grouped into the same domain data feature set, thereby obtaining at least two categories of domain data feature sets.
[0079] In other embodiments, the coordinates of the data points corresponding to the domain data features may also be determined. For example, when the domain data feature is (1, 1, 1, 1), the coordinates of the data points may be (1, 1, 1, 1). After obtaining the coordinates of the data points corresponding to each domain data feature, the distribution of the data points corresponding to each domain data feature in space is determined. The domain data features corresponding to the data points clustered together in space are then grouped into the same domain data feature set, ultimately obtaining at least two categories of domain data feature sets.
[0080] S404: Determine feature set similarity distribution information based on all feature similarities of the domain data feature set.
[0081] Among them, all feature similarities of each field data feature set are obtained, that is, the feature similarities between each two field data features in each field data feature set, and then statistical analysis and processing are performed based on all the obtained feature similarities to obtain feature set similarity distribution information.
[0082] Specifically, the feature set similarity distribution information may be domain data feature pairs distributed in each feature similarity interval. For example, domain data feature pairs distributed in a feature similarity interval of 40% to 50% and their corresponding feature similarities, and domain data feature pairs distributed in a feature similarity interval of 70% to 80% and their corresponding feature similarities.
[0083] S406: Determine specific domain data features from the domain data feature set based on the feature set similarity distribution information.
[0084] The feature set similarity distribution information records the frequency of occurrence of data features in each field in each feature similarity interval.
[0085] It is easy to understand that the domain data features that appear more frequently in the low feature similarity area can be considered as vectors with lower similarity to other domain data features, and therefore can be used as specific domain data features.
[0086] In the embodiments provided in this specification, clustering is performed by the feature similarity between two domain data features to obtain at least two types of domain data feature sets, and then the feature set similarity distribution information in each domain data feature set is determined, so as to quickly determine specific domain data features with low similarity to other domain data features through the feature set similarity distribution information.
[0087] Please refer to Figure 5, which is another flowchart of determining specific domain data features according to an embodiment of this specification. Specifically, in S406, determining specific domain data features from the domain data feature set based on feature set similarity distribution information includes the following steps.
[0088] S502: Based on the feature set similarity distribution information, query the target feature similarity whose feature similarity is less than or equal to the feature similarity threshold from the domain data feature set, determine the domain data feature pairs corresponding to the target feature similarity, and determine the domain data feature pair set based on the domain data feature pairs.
[0089] The feature similarity threshold can be flexibly set based on the feature set similarity distribution information. For example, if the minimum feature similarity interval in the feature set similarity distribution information is 30% to 40%, the feature similarity threshold can be set to 40% or 50%. Of course, the feature similarity threshold can also be manually pre-set. The feature similarity thresholds corresponding to different domain data feature sets can be the same or different.
[0090] After determining the feature similarity threshold, the target feature similarities whose feature similarity is less than or equal to the feature similarity threshold in the feature set similarity distribution information are searched from the domain data feature set. Each target feature similarity corresponds to a pair of domain data features, i.e., a domain data feature pair. The domain data feature pairs corresponding to each target feature similarity are put into the same set to obtain the domain data feature pair set corresponding to the domain data feature set.
[0091] S504: Determine the feature repetition frequency of the domain data feature for each reference domain data feature in the set.
[0092] Among them, the feature similarity of each domain data feature pair in the domain data feature pair set is less than or equal to the feature similarity threshold. Therefore, the vector similarity of each domain data feature pair in the domain data feature pair set is low. At this time, the feature repetition frequency of each reference domain data feature in the domain data feature pair set can be counted.
[0093] The higher the feature repetition frequency of the reference domain data feature, the greater the difference between the reference domain data feature and other reference domain data features; the lower the feature repetition frequency of the reference domain data feature, the smaller the difference between the reference domain data feature and other reference domain data features.
[0094] S506: Determine the specific domain data features in the domain data feature set based on the feature repetition frequency of each reference domain data feature.
[0095] Among them, when the feature repetition frequency of the reference domain data feature is higher, the reference domain data feature is more different from other reference domain data features. In this case, the reference domain data feature can be considered as a specific domain data feature in the domain data feature set. Here, the specific domain data feature can be one or more.
[0096] Specifically, the feature repetition frequencies of each reference domain data feature may be sorted, such as in descending order, and the reference domain data feature corresponding to the feature repetition frequency ranked first may be used as the specific domain data feature in the domain data feature set.
[0097] In the embodiments provided in this specification, domain data feature pairs with higher similarity in the domain data feature set are filtered out by using a feature similarity threshold to obtain domain data feature pairs with lower similarity in the domain data feature set, i.e., a set of domain data feature pairs. Then, the feature repetition frequency of each reference domain data feature in the set of domain data feature pairs is counted. When the feature repetition frequency of the reference domain data feature is higher, it indicates that the reference domain data feature is more different from other reference domain data features. At this time, the reference domain data feature can be considered as a specific domain data feature in the domain data feature set.
[0098] Please refer to Figure 6, which is a flowchart illustrating another method for determining specific domain data features according to an embodiment of this specification. Specifically, in S506, determining specific domain data features in the domain data feature set based on the feature repetition frequency of each reference domain data feature includes the following steps.
[0099] S602: Obtain an indicator ratio mapping relationship between a model training indicator and a label correction ratio, determine a target model training indicator for the generative processing model, and determine a target label correction ratio based on the target model training indicator and the indicator ratio mapping relationship.
[0100] The model training metric is the expected performance metric after generative processing model training. Different label correction ratios correspond to different model training metrics. A higher label correction ratio results in a higher performance metric after generative processing model training; a lower label correction ratio results in a lower performance metric after generative processing model training. The label correction ratio is the proportion of specific domain data features in the domain data feature set that undergo label correction. Specifically, the target label correction ratio can be expressed as Stratify(D;P), where P is the target model training metric and Stratify can be a metric ratio mapping function.
[0101] After obtaining the indicator ratio mapping relationship, the target model training indicator of the generative processing model is obtained. For example, when processing the target domain task in the target domain, the accuracy is 95%. The indicator ratio mapping relationship can be used to determine the target label correction ratio corresponding to the target model training indicator.
[0102] S604: Determine the number of specific domain data features based on the target label correction ratio and the number of domain data features in the domain data feature set.
[0103] Among them, after obtaining the target label correction ratio, the target label correction ratio is multiplied by the number of domain data features in the domain data feature set to calculate the number of features that need to be label corrected, that is, the number of specific domain data features.
[0104] S606: Determine the specific domain data features in the domain data feature set based on the number of specific domain data features and the feature repetition frequency of each reference domain data feature.
[0105] In step S404, the feature repetition frequencies of each reference domain data feature have been obtained, and these feature repetition frequencies can be sorted in descending order. When the number of specific domain data features is X, the reference domain data features corresponding to the first X feature repetition frequencies in the descending order can be used as the specific domain data features, thereby obtaining the specific domain data features in the domain data feature set. The specific domain data features can be represented as the intersection of Cluster(D;K) and Stratify(D;P).
[0106] The embodiments of this specification determine the target label correction ratio through the mapping relationship between the target model training indicators and the indicator ratio, thereby determining the number of specific domain data features, and then screening the specific domain data features based on the feature repetition frequency of each reference domain data feature and the number of specific domain data features, thereby ensuring the training quality of the generative processing model while reducing the workload of label correction processing and improving the efficiency of model training data generation.
[0107] In one embodiment provided in this specification, in S208, a target domain model training dataset is used to perform model training on a generative processing model, including: performing a first model training process on the generative processing model based on the target domain model training dataset to obtain a first generative processing model; detecting whether the first generative processing model meets a model training end condition; if the first generative processing model does not meet the model training end condition, determining to add a new specific domain data feature to the domain data feature set, performing label correction on the specific model processing label carried by the new specific domain data feature to obtain a new specific domain data feature after label correction, updating the target domain model training dataset based on the new specific domain data feature to obtain an updated target domain model training dataset, performing a second model training process on the first generative processing model based on the target domain model training dataset to obtain a second generative processing model, and using the second generative processing model as the first generative processing model to execute in parallel the step of detecting whether the first generative processing model meets the model training end condition; if the first generative processing model meets the model training end condition, determining to end the model training to obtain the target generative processing model.
[0108] After obtaining the target domain model training dataset, the target domain model training dataset can be used to perform a first model training process on the generative processing model to obtain a first generative processing model. The first generative processing model may not meet the model end training conditions, so the vector portion of the target domain model training dataset excluding the specific domain data features can be label-corrected again to train the first generative processing model again. This process is then repeated, that is, incrementally training the model obtained in the previous round using each round of labeled data until the obtained model meets the model end training conditions. This determines that the model training is terminated and the target generative processing model is obtained.
[0109] Specifically, it can be expressed as follows: M i+1 =Train(M i , S i ).
[0110] Among them, M i+1 is the model after the i+1th iteration, i is an integer greater than or equal to 0; M i is the model of round i; S i It is the target domain model training dataset after the i-th round of update; Train is the ongoing training process.
[0111] By implementing the aforementioned strategies, this solution not only significantly reduces the annotation burden of label correction processing, but also improves the accuracy of domain data annotation through multiple rounds of model training. This optimizes the quality and availability of target domain model training datasets, laying a solid foundation for domain-specific model training. This enables the trained target generative processing model to accurately and efficiently handle target domain tasks in the target domain. Furthermore, through continuous iterative cycles, the model is continuously refined using newly identified annotated data, gradually reducing the need for high-quality annotations and optimizing the entire annotation process.
[0112] Please refer to Figure 7, which is a schematic diagram of a process for determining a new specific domain data feature in a domain data feature set according to an embodiment of this specification. In the above embodiment, determining a new specific domain data feature in a domain data feature set includes the following steps.
[0113] S702: Determine the current model training index of the first generative processing model, and determine the number of newly added specific domain data features based on the index difference between the current model training index and the target model training index.
[0114] Among them, the current model training index of the first generative processing model is the training effect index of the first generative processing model obtained after training.
[0115] Generally, there will be a certain difference between the current model training index and the target model training index. When the current model training index is smaller than the target model training index, the number of newly added specific domain data features can be determined based on the index difference between the current model training index and the target model training index; when the current model training index is greater than the target model training index, it indicates that the model meets the model end training conditions, and it is determined to end the model training to obtain the target generative processing model.
[0116] The number of newly added specific domain data features determined by the indicator difference may be determined based on a pre-established mapping relationship or a pre-determined model.
[0117] S704: Determine the new specific domain data features in the domain data feature set based on the number of new specific domain data features and the feature repetition frequency of each reference domain data feature.
[0118] Among them, after determining the number of newly added specific field data features, the newly added specific field data features can be selected again from the feature repetition frequencies that have not been determined as specific field data features based on the feature repetition frequencies of each reference field data feature.
[0119] Here, the method of selecting the newly added specific domain data features is similar to step S506, so it will not be repeated here.
[0120] The embodiment of this specification determines the number of newly added specific domain data features by the indicator difference between the current model training indicator and the target model training indicator, and thus determines the newly added specific domain data features in the domain data feature set by the number of newly added specific domain data features and the feature repetition frequency of each reference domain data feature. While ensuring the training quality of the generative processing model, it also reduces the workload of label correction processing and improves the efficiency of model training data generation.
[0121] Please refer to Figure 8, which is a flow chart of determining domain data features corresponding to target domain data according to an embodiment of this specification. Determining the domain data features corresponding to the target domain data in S204 includes the following steps.
[0122] S802: Extracting domain comprehensive features and domain attribute features of target domain data, and extracting common domain features of target domain data.
[0123] Among them, the target domain data is extracted from two dimensions of data features, namely, the general domain dimension of the target domain data and the domain vertical dimension corresponding to the target domain, so as to obtain the domain comprehensive features and domain attribute features corresponding to the domain vertical dimension corresponding to the target domain, as well as the general domain features corresponding to the general domain dimension.
[0124] The domain comprehensive features integrate the attribute characteristics of each domain. The domain attribute characteristics can be understood as the attribute characteristics of each sub-domain obtained by further dividing the target domain. The domain comprehensive features summarize the characteristics of the target domain from a macro perspective, and the domain attribute characteristics split the characteristics of the target domain from the attribute characteristics of specific sub-domains.
[0125] S804: Determine the domain vertical characteristics based on the domain comprehensive characteristics and domain attribute characteristics, and determine the domain data characteristics based on the domain data characteristics and general domain data characteristics.
[0126] Among them, since the domain comprehensive characteristics summarize the characteristics of the target domain from a macro perspective, and the domain attribute characteristics split the characteristics of the target domain from the attribute characteristics of specific sub-domains, based on the domain comprehensive characteristics and domain attribute characteristics, the domain vertical characteristics can be comprehensively determined from both the macro and specific sub-domain perspectives, which can effectively avoid the loss of domain characteristics.
[0127] The domain data features are determined based on the domain data features and the general domain data features, so that after the generative processing model learns the general domain data features, the model can process the general processing tasks more efficiently and accurately; and, after the generative processing model learns the domain data features, the model can process the processing tasks of the target domain more efficiently and accurately based on the learning of the general domain data features.
[0128] Specifically, S804 may include the following steps: determining a first weight coefficient corresponding to the domain comprehensive feature, determining a second weight coefficient corresponding to the domain attribute feature; applying a first calculation formula to the domain comprehensive feature, the domain attribute feature, the first weight coefficient, and the second weight coefficient to obtain the domain vertical feature; the first calculation formula satisfies the following formula: F L =α*f c +β*f b .
[0129] Among them, F L is the domain vertical feature, α is the first weight coefficient, β is the second weight coefficient, f c is the comprehensive feature of the domain, f b It is the domain attribute feature; determine the third weight coefficient corresponding to the general domain data feature, and the sum of the first weight coefficient, the second weight coefficient and the third weight coefficient is 1; use the second calculation formula to obtain the domain data feature by applying the domain vertical feature, the general domain data feature and the third weight coefficient, and determine the domain data feature based on the domain data feature; the second calculation formula satisfies the following formula: F D =F L +γ*f i .
[0130] Among them, F D is the domain data feature, f i is a general domain data feature, and γ is the third weight coefficient.
[0131] It should be understood that the first weight coefficient, the second weight coefficient and the third weight coefficient can be determined based on the feature importance of the domain comprehensive feature, the domain attribute feature and the general domain data feature.
[0132] Please refer to Figure 9, which is a schematic diagram of a process flow of a generative processing model and target domain data according to an embodiment of this specification. Specifically, obtaining the generative processing model and target domain data in S202 includes the following steps.
[0133] S902: Obtain a generative processing model for the target domain.
[0134] Among them, the acquisition of the generative processing model for the target domain in S902 can refer to the relevant description in S202, which will not be repeated here.
[0135] S904: Acquire source data and a target domain identifier for the target domain, and identify target domain source data from the source data based on the target domain identifier.
[0136] The target domain identifier may be a domain keyword in the target domain, and the data in the corresponding target domain may be identified based on the domain keyword. Therefore, the target domain source data may be identified from the source data based on the target domain identifier.
[0137] S906: Obtain target domain entity definition information, and extract entity information from the target domain source data based on the target domain entity definition information to obtain target domain entity information.
[0138] Among them, the target domain entity definition information can be the definition information corresponding to professional terms or professional terminology in the target domain. Entity information is extracted from the target domain source data through the target domain entity definition information to obtain the target domain entity information. The target domain entity information may include descriptive statements such as professional terms or professional terminology.
[0139] S908: Acquire data extraction logic corresponding to the target domain entity information, and extract data from the target domain entity information based on the data extraction logic to obtain target domain data.
[0140] The target domain entity information includes descriptions of professional terms or terminology. Therefore, different target domain entity information has different corresponding descriptions. By obtaining the data extraction logic corresponding to the target domain entity information, data features can be extracted from the target domain entity information to obtain the target domain data.
[0141] Please refer to Figure 10, which is a flow chart of updating a target domain model training dataset according to an embodiment of this specification. As shown in Figure 9, the method includes the following steps.
[0142] S1002: Determine the feature similarity between the features of each domain data in the target domain model training dataset.
[0143] Among them, the feature similarity calculation formula can be
[0144] Among them, x i and y i These two vectors represent the domain data features in the target domain model training dataset. The higher the vector similarity between the two vectors, the stronger the feature similarity between the two features. In other words, the stronger the correlation between the two features.
[0145] S1004: Determine reference feature similarities whose feature similarities are greater than or equal to a similarity threshold and reference feature pairs corresponding to the reference feature similarities.
[0146] The similarity threshold can be preset manually, such as 98%. Reference feature pairs greater than or equal to the similarity threshold are highly similar. In order to save data storage space, data fusion processing can be performed on the two features corresponding to the reference feature pairs.
[0147] S1006: Determine two similar domain data features corresponding to the reference feature pair, and perform data fusion processing on the similar domain data features to update the target domain model training data set.
[0148] Among them, the reference features in the target domain model training dataset can be used to extract the same parts of the two similar domain data features, and the different parts can be spliced, so as to realize data fusion processing of the corresponding two similar domain data features to update the target domain model training dataset.
[0149] Furthermore, after obtaining the target domain model training dataset, it can be stored in a vector database optimized for efficient retrieval and computation. This vector database utilizes sophisticated indexing mechanisms, such as the approximate nearest neighbor search algorithm, and advanced data representation techniques to facilitate query and analysis across multiple dimensions, such as unique data identifiers, feature similarity, and Gaussian probability distribution. Furthermore, the Gaussian probability distribution can be used to estimate the distribution characteristics of the data. If the distribution characteristics do not meet the requirements, the parameters can be adjusted to regenerate the target domain model training dataset.
[0150] In a specific embodiment provided in this specification, the model training data generation method may specifically include four stages: a data introduction stage, a data processing stage, a data annotation stage, and a domain knowledge base stage. In the data introduction stage, domain data management is performed on the source data. The source data may include text, charts, audio and video, etc. Domain data management may include determining the target domain identifier, determining the target domain entity definition information, determining the data extraction logic, and verifying the data extraction logic for the source data. Then, the target domain data is obtained. In the process of extracting the target domain data, the extraction methods may include web crawlers, optical character recognition, and direct acquisition of document information. High-precision data extraction is performed according to the data extraction logic; finally, the collected data will be strictly quality controlled according to the data verification rules, thereby ensuring the accuracy and completeness of the introduced data.
[0151] During the data processing phase, the target domain data is first processed through a standard process, then through a custom process for the domain. Finally, a universal large model is used to intelligently correct the common knowledge components in the data. This mechanism significantly reduces reliance on manual annotation, thereby lowering annotation costs and improving data processing efficiency. The standard processing process includes data segmentation, data cleaning, standard formatting, and data enhancement.
[0152] The domain-customization process includes domain classification, general dimension processing, domain attribute processing, and attribute enhancement. Considering the diverse processing requirements for data across different domains, this solution builds a loosely coupled, highly scalable, component-based data processing architecture. This architecture supports user-defined domain processing flows, including but not limited to domain classification algorithms, multi-dimensional general processing, attribute-specific processing, and attribute augmentation. Finally, data is output through a universal large model and vectorized to generate domain data feature vectors corresponding to domain data features.
[0153] The data labeling stage includes automated labeling processing, label correction processing, and cyclic label correction processing for newly added specific domain data features.
[0154] In the domain knowledge base stage, the target domain model training data set can be stored. The domain knowledge base stage includes calculating data similarity, calculating data probability distribution information, the target domain model training data set and the corresponding model processing labels.
[0155] The following will be combined with Figure 11, which shows a model training data generation device provided in an embodiment of this specification. The model training data generation device provided in this specification is described in detail. It should be noted that the model training data generation device shown in Figure 11 is used to execute the method of the embodiment shown in Figures 1 to 10 of this specification. For ease of explanation, only the parts related to this specification are shown. For specific technical details not disclosed, please refer to the embodiment shown in Figures 1 to 10 of this specification.
[0156] Please refer to Figure 11, which shows a structural diagram of the model training data generation device of this specification. The model training data generation device 1 can be implemented as all or part of the user terminal through software, hardware or a combination of both. According to some embodiments, the model training data generation device 1 includes an acquisition module 11, a determination module 12, an update module 13 and a training module 14, wherein: the acquisition module 11 is suitable for acquiring a generative processing model and target domain data; the determination module 12 is suitable for determining the domain data features corresponding to the target domain data and the model processing labels of the domain data features; the update module 13 is suitable for performing label update processing on the model processing labels of the domain data features based on the domain data features to obtain a target domain model training data set; the training module 14 is suitable for using the target domain model training data set to perform model training on the generative processing model.
[0157] Optionally, the update module 13 includes: a correction unit, suitable for determining specific domain data features based on domain data features, performing label correction processing on specific model processing labels corresponding to specific domain data features, and obtaining corrected model processing labels; an update unit, suitable for performing label update processing on model processing labels of domain data features based on corrected model processing labels, and obtaining a target domain model training data set.
[0158] Optionally, the correction unit is further adapted to perform clustering processing on the domain data features to obtain at least two types of domain data feature sets, and determine the specific domain data features based on the domain data feature sets.
[0159] Optionally, the correction unit includes: a calculation subunit, suitable for calculating the feature similarity between two domain data features, clustering based on the feature similarity, and obtaining at least two categories of domain data feature sets; a first determination subunit, suitable for determining the feature set similarity distribution information based on all feature similarities of the domain data feature set; a second determination subunit, suitable for determining specific domain data features from the domain data feature set based on the feature set similarity distribution information, and determining specific domain data features based on the specific domain data features.
[0160] Optionally, the second determination subunit includes: a domain data feature pair set determination subunit, adapted to query target feature similarities whose feature similarities are less than or equal to a feature similarity threshold from the domain data feature set based on feature set similarity distribution information, determine the domain data feature pairs corresponding to the target feature similarities, and determine the domain data feature pair set based on the domain data feature pairs; a feature repetition frequency determination subunit, adapted to determine the feature repetition frequency of each reference domain data feature in the domain data feature pair set; and a domain data feature determination subunit, adapted to determine specific domain data features in the domain data feature set based on the feature repetition frequency of each reference domain data feature.
[0161] Optionally, the domain data feature determination subunit includes: an acquisition subunit, suitable for obtaining the indicator ratio mapping relationship between the model training indicator and the label correction ratio, determining the target model training indicator for the generative processing model, and determining the target label correction ratio based on the target model training indicator and the indicator ratio mapping relationship; a specific domain data feature quantity determination subunit, suitable for determining the number of specific domain data features based on the target label correction ratio and the number of domain data features in the domain data feature set; a specific domain data feature determination subunit, suitable for determining the specific domain data features in the domain data feature set based on the number of specific domain data features and the feature repetition frequency of each reference domain data feature.
[0162] Optionally, the training module 14 includes: a first model training unit, suitable for performing a first model training process on the generative processing model based on the target domain model training data set to obtain a first generative processing model; a detection unit, suitable for detecting whether the first generative processing model meets the model end training condition; a first judgment unit, suitable for determining to add new specific domain data features to the domain data feature set if the first generative processing model does not meet the model end training condition, performing label correction on the specific model processing label carried by the new specific domain data feature to obtain the new specific domain data feature after label correction, updating the target domain model training data set based on the new specific domain data feature to obtain an updated target domain model training data set, performing a second model training process on the first generative processing model based on the target domain model training data set to obtain a second generative processing model, and using the second generative processing model as the first generative processing model to execute in parallel the step of detecting whether the first generative processing model meets the model end training condition; a second judgment unit, suitable for determining to end the model training to obtain the target generative processing model if the first generative processing model meets the model end training condition.
[0163] Optionally, the first judgment unit includes: a subunit for determining the number of newly added specific domain data features, suitable for determining the current model training index of the first generative processing model, and determining the number of newly added specific domain data features based on the index difference between the current model training index and the target model training index; a subunit for determining newly added specific domain data features, suitable for determining the newly added specific domain data features in the domain data feature set based on the number of newly added specific domain data features and the feature repetition frequency of each reference domain data feature.
[0164] Optionally, the determination module 12 includes: an extraction unit, suitable for extracting the domain comprehensive features and domain attribute features of the target domain data, and extracting the general domain features of the target domain data; a domain data feature determination unit, suitable for determining the domain vertical features based on the domain comprehensive features and domain attribute features, and determining the domain data features based on the domain data features and the general domain data features.
[0165] Optionally, the domain data feature determination unit includes: a weight determination subunit, adapted to determine a first weight coefficient corresponding to the domain comprehensive feature, and a second weight coefficient corresponding to the domain attribute feature; a domain vertical feature determination subunit, adapted to obtain the domain vertical feature by using a first calculation formula to obtain the domain comprehensive feature, the domain attribute feature, the first weight coefficient, and the second weight coefficient; the first calculation formula satisfies the following formula: F L =α*f c +β*f b .
[0166] Among them, F L is the domain vertical feature, α is the first weight coefficient, β is the second weight coefficient, f c is the comprehensive feature of the domain, f b It is a domain attribute feature; the weight coefficient determination subunit is suitable for determining the third weight coefficient corresponding to the general domain data feature, and the sum of the first weight coefficient, the second weight coefficient and the third weight coefficient is 1; the domain data feature subunit is suitable for using the second calculation formula to obtain the domain data feature by applying the domain vertical feature, the general domain data feature and the third weight coefficient, and determining the domain data feature based on the domain data feature; the second calculation formula satisfies the following formula: F D =F L +γ*f i .
[0167] Among them, F D is the domain data feature, f i is a general domain data feature, and γ is the third weight coefficient.
[0168] Optionally, the acquisition module 11 includes: a generative processing model acquisition unit, suitable for acquiring a generative processing model for the target domain; an identification unit, suitable for acquiring source data and a target domain identifier for the target domain, and identifying the target domain source data from the source data based on the target domain identifier; a target domain entity information determination unit, suitable for acquiring target domain entity definition information, and extracting entity information from the target domain source data based on the target domain entity definition information to obtain target domain entity information; a target domain data extraction unit, suitable for acquiring data extraction logic corresponding to the target domain entity information, and extracting data from the target domain entity information based on the data extraction logic to obtain target domain data.
[0169] Optionally, the model training data generating device 1 also includes: a feature similarity determination module, suitable for determining the feature similarity between pairwise domain data features in the target domain model training data set; a reference feature pair determination module, suitable for determining a reference feature similarity whose feature similarity is greater than or equal to a similarity threshold and a reference feature pair corresponding to the reference feature similarity; a data fusion module, suitable for determining two similar domain data features corresponding to the reference feature pair, and performing data fusion processing on the similar domain data features to update the target domain model training data set.
[0170] It should be noted that the model training data generation device provided in the above embodiment only uses the division of the above functional modules as an example when executing the model training data generation method. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the model training data generation device provided in the above embodiment and the model training data generation method embodiment are of the same concept. The implementation process is detailed in the method embodiment and will not be repeated here.
[0171] This specification also provides a computer storage medium, which can store multiple instructions. The instructions are suitable for being loaded by a processor and executed by the model training data generation method of the embodiments shown in Figures 1 to 10 above. The specific execution process can be found in the specific description of the embodiments shown in Figures 1 to 10, which will not be repeated here.
[0172] This specification also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded by the processor and executed by the model training data generation method of the embodiments shown in Figures 1 to 10 above. The specific execution process can be found in the specific description of the embodiments shown in Figures 1 to 10, and will not be repeated here.
[0173] Please refer to Figure 12, which is a block diagram of the structure of an electronic device provided in an embodiment of this specification. The electronic device described in this specification may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, the memory 120, the input device 130, and the output device 140 may be connected via the bus 150.
[0174] The processor 110 may include one or more processing cores. The processor 110 utilizes various interfaces and circuits to connect various components within the electronic device. It executes instructions, programs, code sets, or instruction sets stored in the memory 120, as well as accesses data stored in the memory 120, to perform various functions of the electronic device and process data. Optionally, the processor 110 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 110 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 110 and may be implemented separately via a communications chip.
[0175] The memory 120 may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory 120 includes a non-transitory computer-readable storage medium. The memory 120 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The operating system may be an Android system, including a system deeply developed based on the Android system, an IOS system developed by Apple, including a system deeply developed based on the IOS system or other systems. The data storage area may also store data created by the electronic device during use, such as a phone book, audio and video data, chat record data, etc.
[0176] Refer to Figure 13, which is a structural diagram of an operating system and user space provided in an embodiment of this specification. The memory 120 can be divided into an operating system space and a user space. The operating system runs in the operating system space, and native and third-party applications run in the user space. In order to ensure that different third-party applications can achieve better operating results, the operating system allocates corresponding system resources to different third-party applications. However, different application scenarios in the same third-party application also have different requirements for system resources. For example, in the local resource loading scenario, the third-party application has higher requirements for disk reading speed; in the animation rendering scenario, the third-party application has higher requirements for GPU performance. The operating system and the third-party application are independent of each other, and the operating system often cannot perceive the current application scenario of the third-party application in a timely manner, resulting in the operating system being unable to perform targeted system resource adaptation according to the specific application scenario of the third-party application.
[0177] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to open up data communication between third-party applications and the operating system so that the operating system can obtain the current scenario information of third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.
[0178] Referring to FIG14 , FIG14 is an architectural diagram of the Android operating system in FIG13 provided in an embodiment of this specification. Taking the Android operating system as an example, the programs and data stored in the memory 120 are shown in FIG14 . The memory 120 may store a Linux kernel layer 320, a system runtime library layer 340, an application framework layer 360, and an application layer 380. The Linux kernel layer 320, the system runtime library layer 340, and the application framework layer 360 belong to the operating system space, and the application layer 380 belongs to the user space. The Linux kernel layer 320 provides underlying drivers for various hardware of electronic devices, such as display drivers, audio drivers, camera drivers, Bluetooth drivers, Wi-Fi drivers, power management, etc. The system runtime library layer 340 provides the main feature support for the Android system through some C / C++ libraries. For example, the SQLite library provides database support, the OpenGL / ES library provides 3D drawing support, and the Webkit library provides browser kernel support. The system runtime layer 340 also includes the Android runtime library (Android runtime), which primarily provides core libraries that allow developers to write Android applications using the Java language. The application framework layer 360 provides various APIs that may be used when building applications. Developers can also use these APIs to build their own applications, such as activity management, window management, view management, notification management, content provider management, package management, call management, resource management, and location management. The application layer 380 runs at least one application. These applications can be native applications that come with the operating system, such as contacts, SMS, clock, and camera applications, or third-party applications developed by third-party developers, such as game applications, instant messaging programs, and photo beautification programs.
[0179] Referring to Figure 15, Figure 15 is an architectural diagram of the iOS operating system in Figure 13 provided in an embodiment of this specification. Taking the iOS operating system as an example, the programs and data stored in the memory 120 are shown in Figure 14. The iOS system includes: a core operating system layer 420 (Core OS layer), a core service layer 440 (Core Services layer), a media layer 460 (Media layer), and a touchable layer 480 (Cocoa Touch Layer). The core operating system layer 420 includes the operating system kernel, drivers, and underlying program frameworks. These underlying program frameworks provide functions closer to the hardware for use by the program framework located in the core service layer 440. The core service layer 440 provides system services and / or program frameworks required by the application, such as the foundation framework, account framework, advertising framework, data storage framework, network connection framework, geographic location framework, motion framework, etc. The media layer 460 provides audio-visual interfaces for the application, such as graphics and image-related interfaces, audio technology-related interfaces, video technology-related interfaces, and wireless playback (AirPlay) interfaces for audio and video transmission technology. The touchable layer 480 provides various commonly used interface-related frameworks for application development. It is responsible for user touch interaction operations on electronic devices, such as local notification services, remote push services, advertising frameworks, game tool frameworks, message user interface (UI) frameworks, UIKit frameworks, and map frameworks.
[0180] In the frameworks shown in FIG15 , those relevant to most applications include, but are not limited to, the Foundation framework in the core services layer 440 and the UIKit framework in the touchable layer 480. The Foundation framework provides many basic object classes and data types, offering fundamental system services for all applications and having nothing to do with the UI. The classes provided by the UIKit framework are the foundational UI class library for creating touch-based user interfaces. iOS applications can use the UIKit framework to provide their UIs, providing the application infrastructure for building user interfaces, drawing, handling user interaction events, responding to gestures, and so on.
[0181] Among them, the method and principle of implementing data communication between third-party applications and the operating system in the IOS system can be referred to the Android system, and this manual will not go into details here.
[0182] Among them, the input device 130 is used to receive input instructions or data, and the input device 130 includes but is not limited to a keyboard, a mouse, a camera, a microphone or a touch device. The output device 140 is used to output instructions or data, and the output device 140 includes but is not limited to a display device and a speaker. In one example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 are touch screen displays, which are used to receive touch operations on or near the touch screen using any suitable object such as a finger or a touch pen, and to display the user interface of each application. The touch screen display is usually provided on the front panel of the electronic device. The touch screen display can be designed as a full screen, a curved screen or a special-shaped screen. The touch screen display can also be designed as a combination of a full screen and a curved screen, or a combination of a special-shaped screen and a curved screen, which is not limited in this specification.
[0183] In addition, those skilled in the art will understand that the structures of the electronic devices shown in the above figures do not limit the electronic devices. The electronic devices may include more or fewer components than shown, or may combine certain components, or arrange the components differently. For example, the electronic devices may also include radio frequency circuits, input units, sensors, audio circuits, wireless fidelity (WiFi) modules, power supplies, Bluetooth modules, and other components, which are not described in detail here.
[0184] In this specification, the execution entity of each step can be the electronic device described above. Optionally, the execution entity of each step is the operating system of the electronic device. The operating system can be Android, iOS, or other operating systems, which is not limited in this specification.
[0185] The electronic device of this specification may also be equipped with a display device, which may be any device capable of realizing a display function, such as a cathode ray tube display (CR), a light-emitting diode display (LED), an electronic ink screen, a liquid crystal display (LCD), a plasma display panel (PDP), etc. The user may use the display device on the electronic device 101 to view displayed text, images, videos and other information. The electronic device may be a smart phone, a tablet computer, a gaming device, an AR (Augmented Reality) device, a car, a data storage device, an audio playback device, a video playback device, a notebook, a desktop computing device, a wearable device such as an electronic watch, electronic glasses, an electronic helmet, an electronic bracelet, an electronic necklace, electronic clothing and the like.
[0186] In the electronic device shown in Figure 12, which can be a terminal, the processor 110 can be used to call the model training data generation program stored in the memory 120, and specifically perform the following operations: obtain the generative processing model and the target domain data; determine the domain data features corresponding to the target domain data and the model processing labels of the domain data features; update the model processing labels of the domain data features based on the domain data features to obtain the target domain model training data set; and use the target domain model training data set to perform model training on the generative processing model.
[0187] Optionally, the processor 110 performs label update processing on the model processing labels of the domain data features based on the domain data features, and when obtaining the target domain model training data set, specifically performs: determining the specific domain data features based on the domain data features, and performing label correction processing on the specific model processing labels corresponding to the specific domain data features to obtain the corrected model processing labels; and performing label update processing on the model processing labels of the domain data features based on the corrected model processing labels to obtain the target domain model training data set.
[0188] Optionally, when the processor 110 determines the specific domain data feature based on the domain data feature, it specifically performs: clustering the domain data feature to obtain at least two types of domain data feature sets, and determining the specific domain data feature based on the domain data feature sets.
[0189] Optionally, the processor 110 performs clustering processing on the domain data features to obtain at least two categories of domain data feature sets. When determining the specific domain data features based on the domain data feature sets, the following specifically executes: calculating the feature similarity between each pair of domain data features, clustering based on the feature similarity to obtain at least two categories of domain data feature sets; determining the feature set similarity distribution information based on all feature similarities of the domain data feature sets; determining the specific domain data features from the domain data feature sets based on the feature set similarity distribution information, and determining the specific domain data features based on the specific domain data features.
[0190] Optionally, when the processor 110 executes determining specific domain data features from the domain data feature set based on the feature set similarity distribution information, it specifically performs: based on the feature set similarity distribution information, querying the target feature similarity whose feature similarity is less than or equal to the feature similarity threshold from the domain data feature set, determining the domain data feature pairs corresponding to the target feature similarity, and determining the domain data feature pair set based on the domain data feature pairs; determining the feature repetition frequency of each reference domain data feature in the domain data feature pair set; and determining the specific domain data features in the domain data feature set based on the feature repetition frequency of each reference domain data feature.
[0191] Optionally, when the processor 110 determines the specific domain data features in the domain data feature set based on the feature repetition frequency of each reference domain data feature, it specifically performs the following: obtaining the indicator ratio mapping relationship between the model training indicator and the label correction ratio, determining the target model training indicator for the generative processing model, and determining the target label correction ratio based on the target model training indicator and the indicator ratio mapping relationship; determining the number of specific domain data features based on the target label correction ratio and the number of domain data features in the domain data feature set; determining the specific domain data features in the domain data feature set based on the number of specific domain data features and the feature repetition frequency of each reference domain data feature.
[0192] Optionally, when the processor 110 executes model training of the generative processing model using the target domain model training dataset, it specifically executes: performing a first model training process on the generative processing model based on the target domain model training dataset to obtain a first generative processing model; detecting whether the first generative processing model meets the model training end condition; if the first generative processing model does not meet the model training end condition, determining to add new specific domain data features to the domain data feature set, performing label correction on the specific model processing labels carried by the new specific domain data features to obtain the new specific domain data features after label correction, updating the target domain model training dataset based on the new specific domain data features to obtain an updated target domain model training dataset, performing a second model training process on the first generative processing model based on the target domain model training dataset to obtain a second generative processing model, and using the second generative processing model as the first generative processing model to execute in parallel the step of detecting whether the first generative processing model meets the model training end condition; if the first generative processing model meets the model training end condition, determining to end the model training to obtain the target generative processing model.
[0193] Optionally, when the processor 110 executes the determination of adding new specific domain data features to the domain data feature set, it specifically performs: determining the current model training index of the first generative processing model, and determining the number of new specific domain data features based on the index difference between the current model training index and the target model training index; determining the new specific domain data features in the domain data feature set based on the number of new specific domain data features and the feature repetition frequency of each reference domain data feature.
[0194] Optionally, when the processor 110 determines the domain data features corresponding to the target domain data, it specifically performs: extracting the domain comprehensive features and domain attribute features of the target domain data, extracting the general domain features of the target domain data; determining the domain vertical features based on the domain comprehensive features and domain attribute features, and determining the domain data features based on the domain data features and the general domain data features.
[0195] Optionally, when the processor 110 determines the domain vertical category features based on the domain comprehensive features and the domain attribute features, and determines the domain data features based on the domain data features and the general domain data features, the processor 110 specifically performs the following steps: determining the first weight coefficient corresponding to the domain comprehensive features, and determining the second weight coefficient corresponding to the domain attribute features; applying the domain comprehensive features, the domain attribute features, the first weight coefficient, and the second weight coefficient to the first calculation formula to obtain the domain vertical category features; the first calculation formula satisfies the following formula: F L =α*f c +β*f b ;
[0196] Among them, F L is the domain vertical feature, α is the first weight coefficient, β is the second weight coefficient, f c is the comprehensive feature of the domain, f b It is the domain attribute feature; determine the third weight coefficient corresponding to the general domain data feature, and the sum of the first weight coefficient, the second weight coefficient and the third weight coefficient is 1; use the second calculation formula to obtain the domain data feature by applying the domain vertical feature, the general domain data feature and the third weight coefficient, and determine the domain data feature based on the domain data feature; the second calculation formula satisfies the following formula: F D =F L +γ*f i ;
[0197] Among them, F D is the domain data feature, f i is a general domain data feature, and γ is the third weight coefficient.
[0198] Optionally, when the processor 110 executes the acquisition of the generative processing model and the target domain data, it specifically performs the following: acquiring the generative processing model for the target domain; acquiring the source data and the target domain identifier for the target domain, and identifying the target domain source data from the source data based on the target domain identifier; acquiring the target domain entity definition information, and extracting entity information from the target domain source data based on the target domain entity definition information to obtain the target domain entity information; acquiring the data extraction logic corresponding to the target domain entity information, and extracting data from the target domain entity information based on the data extraction logic to obtain the target domain data.
[0199] Optionally, the processor 110 is also suitable for executing: determining the feature similarity between pairwise domain data features in the target domain model training data set; determining a reference feature similarity whose feature similarity is greater than or equal to a similarity threshold and a reference feature pair corresponding to the reference feature similarity; determining two similar domain data features corresponding to the reference feature pair, and performing data fusion processing on the similar domain data features to update the target domain model training data set.
[0200] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory, or a random access memory.
[0201] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the object characteristics, interactive behavior characteristics, and user information involved in this specification are all obtained with full authorization.
[0202] The above disclosure is only a preferred embodiment of this specification, and certainly cannot be used to limit the scope of rights of this specification. Therefore, equivalent changes made according to the claims of this specification are still within the scope covered by this specification.
Claims
1. A method for generating model training data, the method comprising: Obtain generative processing models and target domain data; Determining domain data features corresponding to the target domain data and model processing labels of the domain data features; Performing label updating processing on the model processing labels of the domain data features based on the domain data features to obtain a target domain model training data set; The target domain model training dataset is used to perform model training on the generative processing model.
2. The method according to claim 1, wherein the step of performing label updating processing on the model processing labels of the domain data features based on the domain data features to obtain a target domain model training dataset comprises: Determine a specific domain data feature based on the domain data feature, and perform label correction processing on the specific model processing label corresponding to the specific domain data feature to obtain a corrected model processing label; Based on the modified model processing label, the model processing label of the domain data feature is updated to obtain a target domain model training data set.
3. The method according to claim 2, wherein determining the specific domain data feature based on the domain data feature comprises: Clustering is performed on the domain data features to obtain at least two types of domain data feature sets, and specific domain data features are determined based on the domain data feature sets.
4. The method according to claim 3, wherein clustering the domain data features to obtain at least two domain data feature sets, and determining specific domain data features based on the domain data feature sets, comprises: Calculating feature similarities between each of the domain data features, and performing clustering based on the feature similarities to obtain at least two categories of domain data feature sets; Determining feature set similarity distribution information based on all of the feature similarities of the domain data feature set; Based on the feature set similarity distribution information, specific domain data features are determined from the domain data feature set, and specific domain data features are determined based on the specific domain data features.
5. The method according to claim 4, wherein determining the specific domain data features from the domain data feature set based on the feature set similarity distribution information comprises: Based on the feature set similarity distribution information, querying target feature similarities whose feature similarity is less than or equal to a feature similarity threshold from the domain data feature set, determining domain data feature pairs corresponding to the target feature similarities, and determining a domain data feature pair set based on the domain data feature pairs; determining a frequency of feature repetition of the domain data feature for each reference domain data feature in the set; Based on the feature repetition frequency of each reference domain data feature, the specific domain data feature in the domain data feature set is determined.
6. The method according to claim 5, wherein determining the specific domain data features in the domain data feature set based on the feature repetition frequency of each reference domain data feature comprises: Obtaining an indicator ratio mapping relationship between a model training indicator and a label correction ratio, determining a target model training indicator for the generative processing model, and determining a target label correction ratio based on the target model training indicator and the indicator ratio mapping relationship; determining the number of specific domain data features based on the target label correction ratio and the number of domain data features in the domain data feature set; Based on the number of the specific domain data features and the feature repetition frequency of each reference domain data feature, the specific domain data features in the domain data feature set are determined.
7. The method according to any one of claims 3 to 6, wherein the step of using the target domain model training dataset to perform model training on the generative processing model comprises: Performing a first model training process on the generative processing model based on the target domain model training data set to obtain a first generative processing model; Detecting whether the first generative processing model meets the model end training condition; If the first generative processing model does not meet the model end training condition, then determine to add a new specific domain data feature to the domain data feature set, perform label correction on the specific model processing label carried by the new specific domain data feature to obtain the new specific domain data feature after label correction, update the target domain model training data set based on the new specific domain data feature to obtain the updated target domain model training data set, perform a second model training process on the first generative processing model based on the target domain model training data set to obtain a second generative processing model, and use the second generative processing model as the first generative processing model to execute the step of detecting whether the first generative processing model meets the model end training condition in parallel; If the first generative processing model meets the model training end condition, it is determined to end the model training to obtain the target generative processing model.
8. The method according to claim 7, wherein determining to add a new specific domain data feature to the domain data feature set comprises: Determining a current model training index of the first generative processing model, and determining a number of newly added specific domain data features based on an index difference between the current model training index and the target model training index; Based on the number of the newly added specific domain data features and the feature repetition frequency of each reference domain data feature, the newly added specific domain data features in the domain data feature set are determined.
9. The method according to claim 1, wherein determining the domain data features corresponding to the target domain data comprises: Extracting domain comprehensive features and domain attribute features of the target domain data, and extracting common domain features of the target domain data; The domain vertical characteristics are determined based on the domain comprehensive characteristics and the domain attribute characteristics, and the domain data characteristics are determined based on the domain data characteristics and the general domain data characteristics.
10. The method according to claim 9, wherein determining the domain vertical category feature based on the domain comprehensive feature and the domain attribute feature, and determining the domain data feature based on the domain data feature and the general domain data feature, comprises: Determine a first weight coefficient corresponding to the domain comprehensive feature, and determine a second weight coefficient corresponding to the domain attribute feature; The field comprehensive feature, the field attribute feature, the first weight coefficient, and the second weight coefficient are calculated using a first formula to obtain a field vertical feature; The first calculation formula satisfies the following formula: F L =α*f c +β*f b ; Among them, F L is the vertical feature of the field, α is the first weight coefficient, β is the second weight coefficient, f c is the comprehensive characteristic of the field, f b is the domain attribute feature; Determining a third weight coefficient corresponding to the general domain data feature, where the sum of the first weight coefficient, the second weight coefficient, and the third weight coefficient is 1; The domain vertical category feature, the general domain data feature, and the third weight coefficient are calculated using a second formula to obtain a domain data feature, and a domain data feature is determined based on the domain data feature; The second calculation formula satisfies the following formula: F D =F L +γ*f i ; Among them, F D is the domain data feature, f i is the general domain data feature, and γ is the third weight coefficient.
11. The method according to claim 1, wherein obtaining the generative processing model and target domain data comprises: Obtain a generative processing model for the target domain; Acquire source data and a target domain identifier for a target domain, and identify target domain source data from the source data based on the target domain identifier; Acquire target domain entity definition information, and extract entity information from the target domain source data based on the target domain entity definition information to obtain target domain entity information; The data extraction logic corresponding to the target domain entity information is obtained, and data extraction is performed on the target domain entity information based on the data extraction logic to obtain target domain data.
12. The method according to claim 1, after obtaining the target domain model training dataset, the method further comprises: Determining the feature similarity between each pair of domain data features in the target domain model training dataset; Determining reference feature similarities having feature similarities greater than or equal to a similarity threshold and reference feature pairs corresponding to the reference feature similarities; Two similar domain data features corresponding to the reference feature pair are determined, and data fusion processing is performed on the similar domain data features to update the target domain model training data set.
13. A device for generating model training data, the device comprising: An acquisition module, suitable for acquiring generative processing models and target domain data; a determination module adapted to determine domain data features corresponding to the target domain data and a model processing label of the domain data features; An updating module, adapted to perform label updating processing on the model processing labels of the domain data features based on the domain data features to obtain a target domain model training data set; The training module is adapted to perform model training on the generative processing model using the target domain model training dataset.
14. A computer storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the method steps according to any one of claims 1 to 12.
15. A computer program product, wherein the computer program product stores at least one instruction, wherein the at least one instruction is loaded by a processor and executes the method steps according to any one of claims 1 to 12.
16. An electronic device comprising: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method steps according to any one of claims 1 to 12.
Citation Information
Patent Citations
Classification model generation method and device, storage medium and electronic equipment
CN113378895A
Training method and device of an image recognition model, equipment and a storage medium
CN113705554A
Vertical class model training method and device, electronic equipment and readable storage medium
CN117172328A
Model training data generation method and device, storage medium and electronic equipment
CN118070923A