Enterprise label prediction optimization method and device based on large model, equipment and medium
By using a large-model-based enterprise label prediction optimization method, and leveraging pre-trained expert sub-models to extract features and iteratively optimize enterprise data, the problem of low enterprise data utilization is solved, resulting in more accurate label prediction and precise customer outreach.
Patent Information
- Application Number
- CN202311868230.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2043-12-29
AI Technical Summary
In corporate finance scenarios, due to the diversity of corporate data and the scarcity of labeled data, existing technologies cannot fully and accurately utilize corporate data, resulting in low prediction accuracy of the trained labeled prediction models and difficulty in achieving precise customer outreach.
A large-model-based enterprise label prediction optimization method is adopted. By allocating enterprise sample data to corresponding expert sub-models for feature extraction and combining the prediction sub-models for iterative optimization, the pre-trained expert sub-models are used to comprehensively understand and extract features from enterprise data, thereby improving the model's cognitive accuracy.
It improves the utilization rate of enterprise data and the accuracy of tag prediction models, enabling a more comprehensive and accurate understanding of enterprise customers and enhancing the precise reach of message recommendations.
Smart Images

Figure CN117649241B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology in financial technology (Fintech), and in particular to a method, apparatus, device and medium for enterprise label prediction optimization based on a large model. Background Technology
[0002] With the continuous development of fintech, especially internet fintech, more and more technologies (such as distributed systems and artificial intelligence) are being applied in the financial field, but the financial industry is also placing higher demands on technology.
[0003] Artificial intelligence is increasingly being applied in the financial industry, and model training often requires a large amount of training sample data and labels corresponding to the prediction tasks. In corporate finance scenarios, a wealth of corporate data is available, but it covers many dimensions, such as business registration information, tax information, legal information, and financial information. This data also takes many forms, including textual data like business scope information and penalty information in legal documents, as well as numerical data in tax and financial information. In this context, manual feature engineering is insufficient to accurately and comprehensively utilize this complex and diverse corporate data, potentially missing certain dimensions. Furthermore, the proportion of labeled corporate data is relatively small, which may also lead to missing or insufficient data in certain dimensions. Consequently, the model inherently lacks understanding of certain dimensions of corporate customers, resulting in a clear performance ceiling and making it difficult to achieve accurate customer targeting during message recommendation processes. Summary of the Invention
[0004] The main purpose of this application is to provide a method, apparatus, device and medium for optimizing enterprise label prediction based on a large model, which aims to solve the technical problem that the prediction accuracy of the trained label prediction model is low due to the low utilization rate of enterprise data in related technologies.
[0005] To achieve the above objectives, this application provides a large-model-based enterprise label prediction optimization method. This method is applied to a first device, on which an enterprise label prediction model to be trained is deployed. The enterprise label prediction model to be trained includes a prediction sub-model and at least one pre-trained expert sub-model. The large-model-based enterprise label prediction optimization method includes the following steps:
[0006] Obtain training enterprise sample data and enterprise sample labels, wherein the training enterprise sample data includes at least one set of training enterprise sample sub-data;
[0007] Based on the mapping relationship between the preset grouping and the expert sub-model, each of the training enterprise sample sub-data is assigned to its corresponding target expert sub-model for feature extraction, thereby obtaining at least one training enterprise sample data sub-feature of the training enterprise sample.
[0008] The prediction sub-model predicts enterprise labels based on the sub-features of each training enterprise sample data, thereby determining the training label prediction results for the training enterprise samples.
[0009] Based on the training label prediction results and the enterprise sample labels, the prediction sub-model and each of the target expert sub-models are iteratively optimized.
[0010] This application also provides a method for predicting enterprise labels, which is applied to a second device. The second device is equipped with an enterprise label prediction model, which includes a prediction sub-model and at least one expert sub-model. The enterprise label prediction model is trained using the large-model-based enterprise label prediction optimization method described above. The enterprise label prediction method includes the following steps:
[0011] Obtain the sample data of the enterprises to be predicted, wherein the sample data of the enterprises to be predicted includes at least one set of sample sub-data of the enterprises to be predicted;
[0012] Based on the mapping relationship between the preset grouping and the expert sub-model, each of the enterprise sample sub-data to be predicted is assigned to its corresponding target expert sub-model for feature extraction, thereby obtaining at least one enterprise sample data sub-feature of the enterprise sample to be predicted.
[0013] The enterprise label of the enterprise sample is determined by predicting the enterprise label based on the sub-features of the sample data of each enterprise to be predicted using the prediction sub-model.
[0014] This application also provides a large-model-based enterprise label prediction optimization device, which is applied to a first device. The first device has an enterprise label prediction model to be trained deployed on it. The enterprise label prediction model to be trained includes at least one expert sub-model and a prediction sub-model. The large-model-based enterprise label prediction optimization device includes:
[0015] The first acquisition module is used to acquire training enterprise sample data and enterprise sample labels of training enterprise samples, wherein the training enterprise sample data includes at least one set of training enterprise sample sub-data.
[0016] The first feature extraction module is used to allocate each of the training enterprise sample sub-data to its corresponding target expert sub-model for feature extraction based on the mapping relationship between the preset group and the expert sub-model, so as to obtain at least one training enterprise sample data sub-feature of the training enterprise sample.
[0017] The optimization module is used to predict enterprise labels based on the sub-features of each of the training enterprise sample data through the prediction sub-model, determine the training label prediction result of the training enterprise sample, and iteratively optimize the prediction sub-model and each of the target expert sub-models based on the training label prediction result and the enterprise sample label.
[0018] This application also provides a large-model-based enterprise label prediction optimization device, which is applied to a second device. The second device is equipped with an enterprise label prediction model, which includes a prediction sub-model and at least one expert sub-model. The enterprise label prediction model is trained using the large-model-based enterprise label prediction optimization method described above. The large-model-based enterprise label prediction optimization device includes:
[0019] The second acquisition module is used to acquire the sample data of the enterprises to be predicted, wherein the sample data of the enterprises to be predicted includes at least one set of sample sub-data of the enterprises to be predicted.
[0020] The second feature extraction module is used to allocate each of the enterprise sample sub-data to its corresponding target expert sub-model for feature extraction based on the mapping relationship between the preset group and the expert sub-model, so as to obtain at least one enterprise sample data sub-feature of the enterprise sample to be predicted.
[0021] The prediction module is used to predict enterprise labels based on the sub-features of the sample data of each enterprise to be predicted through the prediction sub-model, and to determine the enterprise labels of the enterprise samples to be predicted.
[0022] This application also provides an electronic device, which is a physical device, comprising: a memory, a processor, and a program of the enterprise label prediction optimization method based on a large model stored in the memory and executable on the processor. When the program of the enterprise label prediction optimization method based on a large model is executed by the processor, it can implement the steps of the enterprise label prediction optimization method based on a large model as described above.
[0023] This application also provides a medium, which is a computer-readable storage medium, on which a program implementing a large-model-based enterprise label prediction optimization method is stored. When the program is executed by a processor, it implements the steps of the large-model-based enterprise label prediction optimization method as described above.
[0024] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described enterprise label prediction optimization method based on a large model.
[0025] This application provides a method, apparatus, device, and medium for enterprise label prediction optimization based on a large model. The method is applied to a first device, on which an enterprise label prediction model to be trained is deployed. The enterprise label prediction model to be trained includes a prediction sub-model and at least one pre-trained expert sub-model. Since the enterprise data used for pre-training can be related to or unrelated to the enterprise label prediction task, and can be labeled or unlabeled, and is not limited by the downstream prediction model using labeled data for training, the expert sub-model can utilize a large amount of enterprise data from various dimensions for pre-training during the pre-training stage, thereby improving the accuracy of each expert sub-model's understanding of some enterprise data. Furthermore, in the process of model training combined with prediction tasks, firstly, by acquiring training enterprise sample data and enterprise sample labels, wherein the training enterprise sample data includes at least one set of training enterprise sample sub-data, the labeled data of the training enterprise samples is grouped; then, based on the mapping relationship between the preset grouping and the expert sub-model, each training enterprise sample sub-data is assigned to its corresponding target expert sub-model for feature extraction, obtaining at least one training enterprise sample data sub-feature of the training enterprise sample, thus achieving a comprehensive understanding of the training enterprise sample; then, the prediction sub-model performs enterprise label prediction based on each training enterprise sample data sub-feature, determining the training label prediction result of the training enterprise sample; based on the training label prediction result and the enterprise sample label, the prediction sub-model and each of the target expert sub-models are iteratively optimized, thus achieving the training of the prediction sub-model, and the fine-tuning of each of the target expert sub-models in conjunction with downstream prediction tasks. In this way, on the one hand, using multiple different expert sub-models can improve the comprehensiveness of the understanding of the training enterprise samples and the accuracy of the understanding from each perspective. On the other hand, since the expert sub-models have been pre-trained and already have a relatively accurate understanding of some enterprise data, fine-tuning them in conjunction with the label prediction task allows the expert sub-models to extract features that better meet the needs of the prediction task, thereby further improving the accuracy of the understanding from each perspective. Therefore, this overcomes the technical shortcomings of models that inherently lack understanding of certain dimensions of enterprise customers, resulting in a clear upper limit to model performance and making it difficult to achieve precise customer reach in the message recommendation process. This allows the enterprise label prediction model to utilize more enterprise data, improving the utilization rate of enterprise data and achieving a more comprehensive and accurate understanding of enterprise customers, thereby improving the subsequent message recommendation effect and enabling precise message reach. Attached Figure Description
[0026] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0027] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 This is a flowchart illustrating the first embodiment of the enterprise label prediction optimization method based on a large model in this application.
[0029] Figure 2 This is a flowchart illustrating the second embodiment of the enterprise label prediction optimization method based on a large model in this application.
[0030] Figure 3 This is a flowchart illustrating the third embodiment of the enterprise label prediction optimization method based on a large model in this application;
[0031] Figure 4 This is a schematic diagram of the structure of the enterprise label prediction and optimization device based on a large model in the embodiments of this application;
[0032] Figure 5 This is a schematic diagram of the hardware operating environment involved in the enterprise label prediction optimization method based on a large model in the embodiments of this application.
[0033] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0034] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0035] Example 1
[0036] This application provides an enterprise label prediction optimization method based on a large model, referring to... Figure 1 In the first embodiment of the enterprise label prediction optimization method based on a large model in this application, the enterprise label prediction optimization method based on a large model is applied to a first device, on which an enterprise label prediction model to be trained is deployed. The enterprise label prediction model to be trained includes a prediction sub-model and at least one pre-trained expert sub-model. The enterprise label prediction optimization method based on a large model includes the following steps:
[0037] Step S10: Obtain the training enterprise sample data and enterprise sample labels of the training enterprise samples, wherein the training enterprise sample data includes at least one set of training enterprise sample sub-data.
[0038] The execution subject of the method in this embodiment can be a large-model-based enterprise label prediction and optimization device, or a large-model-based enterprise label prediction and optimization terminal device or server. This embodiment takes a large-model-based enterprise label prediction and optimization device as an example. This large-model-based enterprise label prediction and optimization device can be integrated into terminal devices such as smartphones and computers with data processing functions.
[0039] In this embodiment, it should be noted that the enterprise label prediction optimization method based on a large model is applied to a first device. The first device deploys an enterprise label prediction model to be trained. This enterprise label prediction model is a multi-expert pre-trained large model. The model structure of the multi-expert pre-trained large model can include multiple expert sub-models. The model training process of the multi-expert pre-trained large model includes an independent pre-training process for each expert sub-model, and a fine-tuning process where the pre-trained expert sub-models are trained together with the downstream prediction model. Thus, each expert sub-model, after independent pre-training, possesses the cognitive accuracy for the data it is responsible for. Each expert sub-model can be used by one or more downstream prediction tasks, thereby reducing a significant amount of time and resources spent on repetitive training. Furthermore, for downstream prediction tasks, since each expert sub-model has already been pre-trained and possesses good cognitive accuracy for the data it is responsible for, only fine-tuning is needed in conjunction with the specific prediction task. Therefore, the training time is greatly shortened, and the training efficiency is significantly improved. The enterprise label prediction model to be trained includes prediction sub-models and at least one pre-trained expert sub-model. The prediction sub-model is used to perform enterprise label prediction tasks. Each expert sub-model is used to extract features from a set of enterprise data. By using multiple expert sub-models, feature extraction is achieved across all enterprise data, ensuring comprehensive feature extraction and thus comprehensive understanding of the enterprise samples. Simultaneously, each expert sub-model is only responsible for a portion of the enterprise data, allowing the complex and diverse enterprise data to be divided into multiple groups, reducing the differences in features extracted by each expert sub-model. This enables each expert sub-model to more accurately understand the portion of enterprise data it is responsible for. Each expert sub-model is pre-trained. After pre-training, it is trained together with the prediction sub-model. Fine-tuning the pre-trained expert sub-models in conjunction with the label prediction task allows them to extract features that better meet the requirements of the prediction task, further improving cognitive accuracy.
[0040] The training enterprise sample data may contain one or more sets of training enterprise sample sub-data. In one feasible embodiment, the training enterprise sample data can be grouped from multiple perspectives such as dimension, region, and industry, and a corresponding expert sub-model can be matched for each group of training enterprise sample data. Each training enterprise sample data can be matched with one or more expert sub-models. For example, the training enterprise sample sub-data can be divided into different dimensions, and each dimension of enterprise sample dimension data can be matched with a corresponding expert sub-model; and / or, one or more dimensions can be further subdivided, and the training enterprise sample sub-data can be divided into different groups according to the subdivision method, and each group can be matched with a corresponding expert sub-model; and / or, multiple dimensions can be combined, and the training enterprise sample sub-data can be divided into different groups according to the combination method, and each group can be matched with a corresponding expert sub-model. For example, assuming there are two dimensions: industry and region, we can assign corresponding expert sub-models to each dimension, and then group them according to different industries, assigning a corresponding expert sub-model to each industry group. Alternatively, we can combine industry and region into basic information groups, assign corresponding expert sub-models to these basic information groups, and then further group them according to different industries and regions, assigning a corresponding expert sub-model to each industry group and each region group. In this way, each expert sub-model can comprehensively extract features from enterprise data from multiple perspectives, and pre-training can ensure the accuracy of each expert sub-model's understanding of the enterprise data within its assigned group.
[0041] The pre-training method for each of the expert sub-models is not limited and can include at least one of supervised training, unsupervised training, and self-supervised training. Therefore, the enterprise data used for pre-training can be related to or unrelated to the enterprise label prediction task, and can be labeled or unlabeled. Since it is not limited by the downstream prediction model using labeled data for training, the expert sub-models can utilize a large amount of enterprise data during the pre-training stage. Using a large amount of enterprise data for pre-training can effectively improve cognitive accuracy. Furthermore, since it is not necessary to cognitively process all enterprise data, but only needs to focus on enterprise data of one category, it is less susceptible to interference, and both training efficiency and training effect can be effectively improved.
[0042] As an example, step S10 includes: identifying labeled training enterprise samples from a large number of training samples, and obtaining the training enterprise sample data and its enterprise sample labels. The training enterprise sample data includes a large number of training enterprise sample sub-data, which are grouped according to a preset grouping rule to determine at least one group of training enterprise sample sub-data. The preset grouping rule can be determined according to actual needs, and this embodiment does not impose any limitations on it.
[0043] In one feasible implementation, a gating network can be used to determine the training enterprise sample sub-data assigned to each expert sub-model.
[0044] Optionally, before the step of allocating each of the training enterprise sample sub-data to its corresponding target expert sub-model for feature extraction based on the mapping relationship between preset groups and expert sub-models, to obtain at least one training enterprise sample data sub-feature of the training enterprise sample, the method further includes:
[0045] Step A10: Obtain the pre-training corpus corresponding to each of the expert sub-models;
[0046] Step B10: Perform self-supervised pre-training on each of the expert sub-models based on the pre-training corpus.
[0047] In this embodiment, it should be noted that the pre-training process of each expert sub-model can be carried out independently without relying on the downstream prediction task, and the pre-training is performed in a self-supervised manner. The pre-training data does not need to be restricted by labels. In this way, the amount of enterprise data that can be used for pre-training is greatly increased, which can effectively improve the model training effect and significantly improve the model performance.
[0048] As an example, steps A10-A20 include: for each expert sub-model, pre-training corpora can be selected from a large amount of enterprise data according to pre-matched groupings. For example, as much enterprise data as possible is obtained first. Then, for each regional industry expert sub-model, target enterprises that meet the target regional industry grouping rules are identified from among the enterprises based on their corresponding target regional industry grouping rules, and all enterprise data of the target enterprises are used as the pre-training corpora for that regional industry expert sub-model. For the data domain expert sub-model, the enterprise data of each enterprise can be first compiled into a general enterprise data table, with each row representing the enterprise data of one enterprise and each column representing the enterprise data of different enterprises in the same dimension. Therefore, for each data domain expert sub-model, all enterprise data corresponding to one or more target dimensions of its corresponding data domain can be used as the pre-training corpora for that data domain expert model. Then, each of the pre-training corpora is input into its corresponding expert sub-model for self-supervised training to obtain the pre-trained expert sub-model.
[0049] In one feasible implementation, each expert sub-model can be a BERT model with 110 million parameters. The method of self-supervised pre-training of each expert sub-model based on pre-training corpus can include: first, initializing each expert sub-model. In one feasible implementation, the parameters of BERT-base-chinese can be used to initialize each expert sub-model; then, using pre-training corpus prepared for each expert sub-model, each expert sub-model is trained to obtain multiple different pre-trained expert sub-models. The method of using pre-training corpus to train the expert sub-model can be: masking some fields in the pre-training corpus so that each expert sub-model learns to recover the masked fields.
[0050] Step S20: Based on the mapping relationship between the preset group and the expert sub-model, each of the training enterprise sample sub-data is assigned to its corresponding target expert sub-model for feature extraction, thereby obtaining at least one training enterprise sample data sub-feature of the training enterprise sample.
[0051] In this embodiment, it should be noted that after determining the grouping rules, a corresponding expert sub-model can be determined for each group, so that each expert sub-model can focus on the enterprise data in the group it is responsible for, thereby improving the accuracy of its understanding of the enterprise data in the group it is responsible for.
[0052] As an example, step S20 includes: after determining the grouping of each training enterprise sample sub-data, the training enterprise sample sub-data belonging to the same group can be spliced into a sub-data sequence, and then based on the mapping relationship between the preset grouping and the expert sub-model, the target expert sub-model corresponding to each sub-data sequence is determined, and each sub-data sequence is assigned to its corresponding target expert sub-model for feature extraction, thereby obtaining at least one training enterprise sample data sub-feature of the training enterprise sample.
[0053] Optionally, the preset grouping includes at least one regional industry grouping and at least one data domain grouping; the expert sub-model includes at least one regional industry expert sub-model and at least one data domain expert sub-model; the training enterprise sample data sub-features include regional industry sub-features and at least one data domain sub-feature.
[0054] The step of allocating each training enterprise sample sub-data to its corresponding target expert sub-model for feature extraction based on the mapping relationship between preset groups and expert sub-models, and obtaining at least one training enterprise sample data sub-feature of the training enterprise samples, includes:
[0055] Step B10: Determine the target region and target industry corresponding to the training enterprise sample based on the training enterprise sample data; determine the target region and industry group to which the training enterprise sample data belongs based on the target region and target industry; and allocate the training enterprise sample data to the target region and industry expert sub-model corresponding to the target region and industry group for feature extraction based on the mapping relationship between the preset group and the expert sub-model, so as to obtain the region and industry sub-features of the training enterprise sample.
[0056] Step B20: Determine the target data domain group to which each of the training enterprise sample sub-data belongs based on the dimensional information corresponding to each of the training enterprise sample sub-data. Based on the mapping relationship between the preset group and the expert sub-model, determine the target data domain expert sub-model corresponding to each of the training enterprise sample sub-data. Assign each of the training enterprise sample sub-data to its corresponding target data domain expert sub-model for feature extraction to obtain at least one data domain sub-feature of the training enterprise sample.
[0057] In this embodiment, it should be noted that enterprises are distributed across the country and various industries. Different regions and industries are important factors affecting the enterprise's attributes, operating conditions, and risk performance. If enterprise data from different regions and industries are mixed together for model training, the trained model will find it difficult to accurately recognize each region and industry, resulting in low accuracy in the final predicted enterprise labels. Therefore, based on region and industry, different expert sub-models are assigned to different regional and industry groups. Each expert sub-model can focus on the enterprise data of the region and industry it is responsible for, improving the accuracy of the recognition of enterprise data in each region and industry. At the same time, multiple expert sub-models can ensure comprehensive coverage of all regional industries, thereby ensuring comprehensive recognition. One method for grouping by region and industry is as follows: First, obtain the geographical and industry distribution of enterprises. Then, divide the geographical distribution into multiple geographical areas according to a preset geographical division method, and divide the industry distribution into multiple industry categories according to a preset industry division method. Combine each geographical area with each industry category to obtain multiple regional industry groups. Then, input the enterprise data of the target region and target industry corresponding to each regional industry group into the corresponding expert sub-model for separate training to obtain the expert sub-model corresponding to each regional industry group. For example, the geographical area can be divided into nine categories based on geographical location and business situation: Northeast, Northwest, East, South, Southwest, Central, North, Jiangsu-Zhejiang-Shanghai, and Guangdong. The industry category can be divided into ten categories based on industry type and business situation: Unaffiliated, Resources, Public Facilities, Livelihood, Technology, Construction, Manufacturing, Commerce, Retail, and Wholesale. By combining each geographical region with each industry category, 90 geographical and industry groups can be obtained. Then, the enterprise data of each geographical and industry group is input into the corresponding expert sub-model for pre-training, thereby obtaining the expert sub-model corresponding to each of the 90 geographical and industry groups.
[0058] A data domain refers to a combination of data dimensions that have a certain correlation. All data dimensions can be divided into multiple data domains based on the correlation between them to ensure that each data domain covers all data dimensions. These data domains can include at least one of the following: basic information, financial data, tax data, corporate relationships, tag ratings, operational information, operational anomalies, and legal data. For example, if the data dimensions include registered capital, region, and industry, then these three data dimensions can be considered as belonging to the company's basic information, and therefore all three dimensions are classified into the basic information data domain.
[0059] Data across different data domains exhibits varying forms of representation and underlying logic. Representation can include numerical, textual, and graphical formats. For example, the underlying logic differs between registered capital and risk factors. Higher registered capital generally indicates stronger company strength and lower risk, while more risk factors suggest higher risk – the two are essentially opposite. Therefore, grouping data based on data domains and assigning different expert sub-models to each domain allows for focused attention on company data within their assigned domain. This improves the accuracy of understanding company data across each domain. Furthermore, multiple expert sub-models ensure comprehensive coverage of all data domains, guaranteeing a holistic understanding.
[0060] The training enterprise sample data is grouped according to region and industry, and a corresponding regional industry expert sub-model is assigned to each regional industry group. At the same time, it is also grouped according to data domain, and a corresponding data domain expert sub-model is assigned to each data domain group. In this way, features are extracted from multiple perspectives for each enterprise data, so the features of enterprise data can be mined from multiple angles, improving the comprehensiveness of feature extraction and thus improving the comprehensiveness of the understanding of enterprise data.
[0061] Step B10 is the process of extracting regional industry sub-features, and step B20 is the process of extracting data domain sub-features. Since each expert sub-model extracts features independently, this embodiment does not limit the order of steps B10 and B20, and they can be executed simultaneously or in any order.
[0062] As an example, step B10 includes: after obtaining the training enterprise sample data, determining the target region and target industry corresponding to the training enterprise sample based on the training enterprise sample data; searching for a preset regional industry group based on the target region and target industry; determining the regional industry group that includes both the target region and the target industry as the target regional industry group to which the training enterprise sample data belongs; then concatenating each of the training enterprise sample sub-data into a regional industry sub-data sequence; determining the target regional industry expert sub-model corresponding to the regional industry sub-data sequence based on the mapping relationship between the preset group and the expert sub-model; and assigning the regional industry sub-data sequence to the target industry expert sub-model for feature extraction to obtain the regional industry sub-features of the training enterprise sample.
[0063] As an example, step B20 includes: after obtaining the training enterprise sample data, the dimensional information corresponding to each training enterprise sample sub-data can be found from the training enterprise sample data; based on the data domain to which the dimensional information belongs, the target data domain group to which each training enterprise sample sub-data belongs is determined; then, the training enterprise sample sub-data belonging to the same target data domain group is concatenated together to obtain at least one data domain sub-data sequence; based on the mapping relationship between the preset group and the expert sub-model, the target data domain expert sub-model corresponding to each data domain sub-data sequence is determined; and each data domain sub-data sequence is assigned to its corresponding target data domain expert sub-model for feature extraction to obtain at least one data domain sub-feature of the training enterprise sample.
[0064] Step S30: The prediction sub-model predicts enterprise labels based on the sub-features of each of the training enterprise sample data to determine the training label prediction results of the training enterprise samples. Based on the training label prediction results and the enterprise sample labels, the prediction sub-model and each of the target expert sub-models are iteratively optimized.
[0065] As an example, step S30 includes: merging and concatenating the sub-features of the training enterprise sample data extracted by each target expert sub-model into a feature vector or feature matrix, inputting it into the prediction sub-model to predict enterprise labels, obtaining the training label prediction result of the training enterprise sample, calculating the label prediction loss based on the difference between the training label prediction result and the enterprise sample label, and iteratively optimizing the prediction sub-model and each of the target expert sub-models through backpropagation based on the label prediction loss.
[0066] This allows for one round of iterative optimization of the predictive sub-model and one round of fine-tuning of each of the target expert sub-models. During multiple rounds of iterative optimization, each round obtains different training enterprise samples. These different training enterprise samples may belong to different groups. Therefore, each round of iterative optimization will perform one iteration of optimization on the predictive sub-model and one round of fine-tuning on some expert sub-models. Thus, as long as the training enterprise sample data can cover all groups, multiple rounds of iterative optimization can achieve the goal of training the predictive sub-model and simultaneously fine-tuning all expert sub-models.
[0067] In one feasible implementation, during model training, each iteration can acquire a batch of training enterprise sample data and enterprise sample labels. Therefore, after feature extraction of this batch of training enterprise sample data, batch normalization can be performed on the output of this batch of expert sub-models to align them to a uniform distribution. Then, a two-layer multilayer perceptron can be used. The first layer is used to merge the outputs of all expert sub-models, and the second layer is used to map to labels. That is, the output of the second layer is the training label prediction result, which should be consistent with the enterprise sample labels.
[0068] In one feasible implementation, the outputs of the various expert sub-models can be aggregated using a gating network.
[0069] In this embodiment, the enterprise label prediction optimization method based on a large model is applied to a first device. The first device is equipped with an enterprise label prediction model to be trained. The enterprise label prediction model to be trained includes a prediction sub-model and at least one pre-trained expert sub-model. Since the enterprise data used for pre-training can be related to or unrelated to the enterprise label prediction task, and can be labeled or unlabeled, and is not subject to the limitation of downstream prediction models using labeled data for training, the expert sub-model can utilize a large amount of enterprise data from various dimensions for pre-training during the pre-training stage, thereby improving the accuracy of each expert sub-model's understanding of some enterprise data. Furthermore, in the process of model training combined with prediction tasks, firstly, by acquiring training enterprise sample data and enterprise sample labels, wherein the training enterprise sample data includes at least one set of training enterprise sample sub-data, the labeled data of the training enterprise samples is grouped; then, based on the mapping relationship between the preset grouping and the expert sub-model, each training enterprise sample sub-data is assigned to its corresponding target expert sub-model for feature extraction, obtaining at least one training enterprise sample data sub-feature of the training enterprise sample, thus achieving a comprehensive understanding of the training enterprise sample; then, the prediction sub-model performs enterprise label prediction based on each training enterprise sample data sub-feature, determining the training label prediction result of the training enterprise sample; based on the training label prediction result and the enterprise sample label, the prediction sub-model and each of the target expert sub-models are iteratively optimized, thus achieving the training of the prediction sub-model, and the fine-tuning of each of the target expert sub-models in conjunction with downstream prediction tasks. In this way, on the one hand, using multiple different expert sub-models can improve the comprehensiveness of the understanding of the training enterprise samples and the accuracy of the understanding from each perspective. On the other hand, since the expert sub-models have been pre-trained and already have a relatively accurate understanding of some enterprise data, fine-tuning them in conjunction with the label prediction task allows the expert sub-models to extract features that better meet the needs of the prediction task, thereby further improving the accuracy of the understanding from each perspective. Therefore, this overcomes the technical shortcomings of models that inherently lack understanding of certain dimensions of enterprise customers, resulting in a clear upper limit to model performance and making it difficult to achieve precise customer reach in the message recommendation process. This allows the enterprise label prediction model to utilize more enterprise data, improving the utilization rate of enterprise data and achieving a more comprehensive and accurate understanding of enterprise customers, thereby improving the subsequent message recommendation effect and enabling precise message reach.
[0070] Example 2
[0071] Furthermore, in the second embodiment of this application, content that is the same as or similar to that in the above embodiments can be referred to the above description, and will not be repeated hereafter. Based on this, refer to... Figure 2The training enterprise sample data includes training enterprise sample table data; the step of assigning each training enterprise sample sub-data to its corresponding target expert sub-model for feature extraction based on the mapping relationship between preset groups and expert sub-models, and obtaining at least one training enterprise sample data sub-feature of the training enterprise samples, includes:
[0072] Step C10: Transcribe the training enterprise sample table data into training enterprise sample sequence data, wherein the training enterprise sample sequence data includes at least one sub-data element, and the sub-data element includes dimensional information and training enterprise sample sub-data.
[0073] In this embodiment, it's important to note that in corporate finance marketing scenarios, there is a wealth of corporate data, encompassing information across multiple dimensions, such as business registration information, tax information, legal information, and financial information. This corporate data is typically stored in tabular form in a distributed database. A significant portion of this data is text-based, such as the business scope in business registration information and penalty information in legal information, while also including a large amount of numerical data. This complex and diverse corporate data, scattered across tables, makes it difficult to manually incorporate features into a model. Consequently, the model inherently lacks understanding of these customer dimensions, resulting in a clear performance ceiling. This makes it difficult to achieve precise customer reach during subsequent marketing and customer acquisition processes, and also poses significant challenges to developing fine-grained customer operation strategies.
[0074] The training enterprise sample table data includes enterprise identifiers, training enterprise sample sub-data, and dimensional information. In one feasible implementation, the first column of the training enterprise sample table data is the enterprise identifier, and the first row is the dimensional information. That is, each row of the training enterprise sample table data contains data belonging to different dimensions of the same enterprise, and each column of the training enterprise sample table data contains data belonging to the same dimension but to different enterprises.
[0075] As an example, step C10 includes: after obtaining the training enterprise sample table data, sequentially obtaining training enterprise sample sub-data from the training enterprise sample table data, determining their corresponding dimensional information, combining each training enterprise sample sub-data and its corresponding dimensional information into a sub-data element, until the sub-data elements corresponding to all training enterprise sample sub-data are determined, thereby completing the transcription of the training enterprise sample table data and obtaining training enterprise sample sequence data composed of sub-data elements.
[0076] Optionally, the step of transcribing the training enterprise sample table data into training enterprise sample sequence data includes:
[0077] Step C11: Extract at least one training enterprise sample sub-data from the training enterprise sample table data;
[0078] Step C12: Concatenate the sub-data of each training enterprise sample with its corresponding dimensional information to form at least one sub-data element;
[0079] Step C13: Concatenate the sub-data elements to form the training enterprise sample sequence data.
[0080] As an example, steps C11-C13 include: after obtaining the training enterprise sample table data, extracting training enterprise sample sub-data from the training enterprise sample table data, determining the dimensional information corresponding to each of the training enterprise sample sub-data, concatenating each training enterprise sample sub-data and its corresponding dimensional information to obtain at least one sub-data element; and then concatenating the sub-data elements to obtain training enterprise sample sequence data.
[0081] Optionally, the training enterprise sample sub-data includes at least one of training enterprise sample text sub-data and training enterprise sample numerical sub-data; the sub-data element includes at least one of dimensional text data element and dimensional numerical data element;
[0082] The step of concatenating the sub-data of each training enterprise sample with its corresponding dimensional information to form at least one sub-data element includes:
[0083] Step C121: Concatenate the text sub-data of each training enterprise sample with their respective dimensional information to form at least one dimensional text data element;
[0084] And / or, in step C122, the numerical sub-data of each training enterprise sample is binned according to the preset binning rules, the bin name corresponding to each of the numerical sub-data of each training enterprise sample is determined, and the bin name corresponding to each of the numerical sub-data of each training enterprise sample is concatenated with the corresponding dimension information to form at least one dimension numerical data element.
[0085] In this embodiment, it should be noted that enterprise data typically contains both numerical and textual data. When predicting enterprise tags, the model primarily analyzes and predicts by understanding the meaning within the enterprise data. Textual data itself can readily convey its meaning; therefore, it can be directly concatenated with dimensional information to serve as dimensional textual data elements for the model's use. However, the meaning of numerical data is often difficult to represent through a single numerical value. For example, for a total amount of ten million, ten thousand might be a very small number, but if the total amount is twenty thousand, ten thousand would represent an average level. Therefore, by using a binning approach, the actual meaning of each numerical value can be better represented, thus facilitating the model's understanding of the meaning of numerical data.
[0086] Step C121 describes the processing method for the text sub-data of the training enterprise samples, and step C122 describes the processing method for the numerical sub-data of the training enterprise samples. These two steps are performed independently, and this embodiment does not limit the order of steps C121 and C122. For the text sub-data of the training enterprise samples, step C121 is executed to obtain its corresponding dimensional text data elements; for the numerical sub-data of the training enterprise samples, step C122 is executed to obtain its corresponding dimensional numerical data elements.
[0087] As an example, step C121 includes: if the training enterprise sample sub-data includes at least one training enterprise sample text sub-data, then each training enterprise sample text sub-data is concatenated with its corresponding dimension information to obtain at least one dimension text data element.
[0088] As an example, step C122 includes: if the training enterprise sample sub-data includes at least one training enterprise sample numerical sub-data, then the training enterprise sample numerical sub-data is binned according to a preset binning rule, the target bin to which each training enterprise sample numerical sub-data belongs is determined, and the bin name corresponding to each target bin is determined. The bin name of the target bin corresponding to each training enterprise sample numerical sub-data is concatenated with its corresponding dimension information to obtain at least one dimension numerical data element. The preset binning rules for different dimensions can be the same or different, and can be specifically determined based on some or all of the training enterprise sample numerical sub-data corresponding to each dimension. This embodiment does not impose any restrictions on this. For example, Company A has a registered capital of 500,000 yuan, and Company B has a registered capital of 1,000,000 yuan. The rule for dividing registered capital is that less than or equal to 500,000 yuan is designated as "Registered Capital_1", and greater than 500,000 yuan but less than or equal to 1,000,000 yuan is designated as "Registered Capital_2". Therefore, the bin name for Company A is "Registered Capital_1", and the corresponding dimension data element for Company A is "Registered Capital: Registered Capital_1". The bin name for Company B is "Registered Capital_2", and the corresponding dimension data element for Company B is also "Registered Capital: Registered Capital_2". In this way, compared to the two separate texts of 500,000 yuan and 1,000,000 yuan, "Registered Capital_1" and "Registered Capital_2" are easier for the model to understand as similar yet different, and thus more comparable.
[0089] Step C20: Based on the mapping relationship between the preset group and the expert sub-model, each of the sub-data elements is assigned to its corresponding target expert sub-model for feature extraction, thereby obtaining at least one training enterprise sample data sub-feature of the training enterprise sample.
[0090] As an example, step C20 includes: after determining the grouping of each training enterprise sample sub-data, the sub-data elements corresponding to the training enterprise sample sub-data belonging to the same group can be concatenated into a sub-data sequence, and then, based on the mapping relationship between the preset grouping and the expert sub-model, the target expert sub-model corresponding to each sub-data sequence is determined, and each sub-data sequence is assigned to its corresponding target expert sub-model for feature extraction, thereby obtaining at least one training enterprise sample data sub-feature of the training enterprise sample.
[0091] In this embodiment, by transcribing the training enterprise sample table data into training enterprise sample sequence data, the model can adapt to the storage method of enterprise data, enabling the model to effectively utilize enterprise data stored in the distributed database in tabular form. By improving the utilization rate of enterprise data, the model can avoid the situation where some data cannot be included in the model, resulting in a lack of understanding of certain dimensions of customers. This improves the comprehensiveness of the model's understanding of enterprise customers, thereby improving the accuracy of enterprise label prediction.
[0092] Example 3
[0093] Furthermore, in the third embodiment of this application, a method for predicting enterprise labels is also provided. Contents identical or similar to those in the above embodiments can be referred to the above description and will not be repeated hereafter. Based on this, refer to... Figure 3 The enterprise label prediction optimization based on a large model is applied to a second device, on which an enterprise label prediction model is deployed. The enterprise label prediction model includes a prediction sub-model and at least one expert sub-model. The enterprise label prediction model is trained using the enterprise label prediction optimization method based on a large model as described above. The enterprise label prediction optimization method based on a large model includes the following steps:
[0094] Step D10: Obtain the sample data of the enterprises to be predicted, wherein the sample data of the enterprises to be predicted includes at least one set of sample sub-data of the enterprises to be predicted.
[0095] Step D20: Based on the mapping relationship between the preset group and the expert sub-model, each of the enterprise sample sub-data to be predicted is assigned to its corresponding target expert sub-model for feature extraction, thereby obtaining at least one enterprise sample data sub-feature of the enterprise sample to be predicted.
[0096] Step D30: Based on the sub-features of the data of each enterprise sample to be predicted, the prediction sub-model is used to predict the enterprise label, and the enterprise label of the enterprise sample to be predicted is determined.
[0097] In this embodiment, it should be noted that after the enterprise label prediction model is trained, it can be deployed on a second device for enterprise label prediction. The second device may be the same as or different from the first device. The trained enterprise label prediction model includes a prediction sub-model and at least one expert sub-model. For each enterprise sample to be predicted, based on the available enterprise sample data, some or all of the expert sub-models can be used to extract features from the enterprise sample's sub-data. The extracted features are then merged and input into the prediction sub-model for enterprise label prediction.
[0098] As an example, steps D10-D30 include: acquiring the enterprise sample data to be predicted, wherein the enterprise sample data consists of at least one sub-data of the enterprise sample to be predicted; grouping these sub-data of the enterprise sample to be predicted according to a preset grouping rule to determine at least one set of training enterprise sample sub-data, wherein the preset grouping rule is determined in advance during the training stage of each expert sub-model; after determining the grouping of each sub-data of the enterprise sample to be predicted, the sub-data of the enterprise sample to be predicted belonging to the same group can be concatenated into a sequence of sub-data to be predicted; then, based on the mapping relationship between the preset grouping and the expert sub-model, the target expert sub-model corresponding to each sequence of sub-data to be predicted is determined, and each sequence of sub-data to be predicted is assigned to its corresponding target expert sub-model for feature extraction to obtain at least one sub-feature of the enterprise sample to be predicted; then, the sub-features of the enterprise sample to be predicted extracted by each target expert sub-model are merged and concatenated into a feature vector or feature matrix, which is input into the prediction sub-model for enterprise label prediction to determine the enterprise label of the enterprise sample to be predicted.
[0099] The enterprise label prediction method provided by this invention employs the enterprise label prediction optimization method based on a large model as described in the above embodiments, solving the technical problem in related technologies where the prediction accuracy of the trained label prediction model is low due to the low utilization rate of enterprise data. Compared with related technologies, the enterprise label prediction method provided by this invention has the same advantages as the enterprise label prediction optimization method based on a large model provided in the above embodiments, and other technical features in this enterprise label prediction method are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0100] Example 4
[0101] Furthermore, embodiments of this application also provide an enterprise label prediction optimization device based on a large model, referring to... Figure 4The enterprise label prediction optimization method based on a large model is applied to a first device, on which a first-party feature extraction model is deployed; the enterprise label prediction optimization device based on a large model includes:
[0102] The first acquisition module 10 is used to acquire training enterprise sample data and enterprise sample labels of training enterprise samples, wherein the training enterprise sample data includes at least one set of training enterprise sample sub-data.
[0103] The first feature extraction module 20 is used to allocate each of the training enterprise sample sub-data to its corresponding target expert sub-model for feature extraction based on the mapping relationship between the preset group and the expert sub-model, so as to obtain at least one training enterprise sample data sub-feature of the training enterprise sample.
[0104] The optimization module 30 is used to predict enterprise labels based on the sub-features of each of the training enterprise sample data through the prediction sub-model, determine the training label prediction result of the training enterprise sample, and iteratively optimize the prediction sub-model and each of the target expert sub-models based on the training label prediction result and the enterprise sample label.
[0105] Furthermore, the training enterprise sample data includes training enterprise sample table data; the first feature extraction module 20 is also used for:
[0106] The training enterprise sample table data is transcribed into training enterprise sample sequence data, wherein the training enterprise sample sequence data includes at least one sub-data element, and the sub-data element includes dimensional information and training enterprise sample sub-data.
[0107] Based on the mapping relationship between the preset grouping and the expert sub-model, each of the sub-data elements is assigned to its corresponding target expert sub-model for feature extraction, thereby obtaining at least one training enterprise sample data sub-feature of the training enterprise sample.
[0108] Furthermore, the first feature extraction module 20 is also used for:
[0109] Extract at least one sub-data of training enterprise samples from the training enterprise sample table data;
[0110] Each of the training enterprise sample sub-datas is concatenated with its corresponding dimensional information to form at least one sub-data element;
[0111] The sub-data elements are concatenated to form the training enterprise sample sequence data.
[0112] Furthermore, the first feature extraction module 20 is also used for:
[0113] Each of the training enterprise sample text sub-data and its corresponding dimensional information are concatenated to form at least one dimensional text data element;
[0114] And / or, according to the preset bucketing rules, the numerical data of each training enterprise sample is bucketed, the bucket name corresponding to each numerical data of each training enterprise sample is determined, and the bucket name corresponding to each numerical data of each training enterprise sample is concatenated with the corresponding dimension information to form at least one dimension numerical data element.
[0115] Further, the preset grouping includes at least one regional industry grouping and at least one data domain grouping; the expert sub-model includes at least one regional industry expert sub-model and at least one data domain expert sub-model; the training enterprise sample data sub-features include regional industry sub-features and at least one data domain sub-features; the first feature extraction module 20 is also used for:
[0116] Based on the training enterprise sample data, the target region and target industry corresponding to the training enterprise sample are determined. Based on the target region and target industry, the target region and industry group to which the training enterprise sample data belongs is determined. Based on the mapping relationship between the preset group and the expert sub-model, the training enterprise sample data is assigned to the target region and industry expert sub-model corresponding to the target region and industry group for feature extraction, thereby obtaining the region and industry sub-features of the training enterprise sample.
[0117] Based on the dimensional information corresponding to each of the training enterprise sample sub-data, the target data domain group to which each of the training enterprise sample sub-data belongs is determined. Based on the mapping relationship between the preset group and the expert sub-model, the target data domain expert sub-model corresponding to each of the training enterprise sample sub-data is determined. Each of the training enterprise sample sub-data is assigned to its corresponding target data domain expert sub-model for feature extraction, thereby obtaining at least one data domain sub-feature of the training enterprise sample.
[0118] Furthermore, before the operation of assigning each training enterprise sample sub-data to its corresponding target expert sub-model for feature extraction based on the mapping relationship between preset groups and expert sub-models to obtain at least one training enterprise sample data sub-feature of the training enterprise sample, the enterprise label prediction optimization device based on the large model further includes a pre-training module, which is used for:
[0119] Obtain the pre-training corpus corresponding to each of the expert sub-models;
[0120] Each of the expert sub-models is pre-trained under self-supervised conditions based on the pre-training corpus.
[0121] The enterprise label prediction optimization device based on a large model provided by this invention employs the enterprise label prediction optimization method based on a large model in the above embodiments, solving the technical problem in related technologies where the prediction accuracy of the trained label prediction model is low due to the low utilization rate of enterprise data. Compared with related technologies, the advantages of the enterprise label prediction optimization device based on a large model provided in this invention are the same as those of the enterprise label prediction optimization method based on a large model provided in the above embodiments, and other technical features in this enterprise label prediction optimization device based on a large model are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0122] Example 5
[0123] Furthermore, this application embodiment also provides an enterprise label prediction device. The enterprise label prediction optimization device based on a large model is applied to a second device, on which an enterprise label prediction model is deployed. The enterprise label prediction model includes a prediction sub-model and at least one expert sub-model. The enterprise label prediction model is obtained by training the model using the enterprise label prediction optimization method based on a large model as described above. The enterprise label prediction optimization device based on a large model includes:
[0124] The second acquisition module is used to acquire the sample data of the enterprises to be predicted, wherein the sample data of the enterprises to be predicted includes at least one set of sample sub-data of the enterprises to be predicted.
[0125] The second feature extraction module is used to allocate each of the enterprise sample sub-data to its corresponding target expert sub-model for feature extraction based on the mapping relationship between the preset group and the expert sub-model, so as to obtain at least one enterprise sample data sub-feature of the enterprise sample to be predicted.
[0126] The prediction module is used to predict enterprise labels based on the sub-features of the sample data of each enterprise to be predicted through the prediction sub-model, and to determine the enterprise labels of the enterprise samples to be predicted.
[0127] The enterprise label prediction device provided by this invention employs the enterprise label prediction method described in the above embodiments, solving the technical problem in related technologies where the prediction accuracy of the trained label prediction model is low due to the low utilization rate of enterprise data. Compared with related technologies, the advantages of the enterprise label prediction device provided in this invention are the same as those of the enterprise label prediction method described in the above embodiments, and other technical features in this enterprise label prediction device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0128] Example 6
[0129] Furthermore, embodiments of the present invention provide an electronic device, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to execute the enterprise label prediction optimization method based on a large model as described above.
[0130] The following is for reference. Figure 5 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of the present disclosure. The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as Bluetooth headsets, mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0131] like Figure 5 As shown, an electronic device may include a processing unit (such as a central processing unit, graphics processing unit, etc.) that can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) or a program loaded from a storage device into random access memory (RAM). The RAM also stores various programs and arrays required for the operation of the electronic device. The processing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0132] Typically, the following systems can be connected to the I / O interface: input devices including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices including, for example, magnetic tapes, hard disks, etc.; and communication devices. Communication devices allow electronic devices to communicate wirelessly or wiredly with other devices to exchange arrays. Although electronic devices with various systems are shown in the figures, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems may be implemented alternatively.
[0133] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processing device, it performs the functions defined above in the methods of embodiments of this disclosure.
[0134] The electronic device provided by this invention employs the enterprise label prediction optimization method based on a large model as described in the above embodiments, solving the technical problem in related technologies where the prediction accuracy of the trained label prediction model is low due to the low utilization rate of enterprise data. Compared with related technologies, the electronic device provided by this invention has the same advantages as the enterprise label prediction optimization method based on a large model provided in the above embodiments, and other technical features of this electronic device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0135] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.
[0136] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
[0137] Example 7
[0138] Furthermore, this embodiment provides a computer-readable storage medium having computer-readable program instructions stored thereon, the computer-readable program instructions being used to execute the enterprise label prediction optimization method based on a large model in the above embodiment.
[0139] The computer-readable storage medium provided in this embodiment of the invention may be, for example, a USB flash drive, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0140] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.
[0141] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the electronic device to: acquire training enterprise sample data and enterprise sample labels of training enterprise samples, wherein the training enterprise sample data includes at least one set of training enterprise sample sub-data; based on the mapping relationship between a preset group and an expert sub-model, assign each of the training enterprise sample sub-data to its corresponding target expert sub-model for feature extraction, thereby obtaining at least one training enterprise sample data sub-feature of the training enterprise sample; predict enterprise labels based on each of the training enterprise sample data sub-features through the prediction sub-model, determine the training label prediction result of the training enterprise sample, and iteratively optimize the prediction sub-model and each of the target expert sub-models based on the training label prediction result and the enterprise sample labels.
[0142] Alternatively, the aforementioned computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the electronic device to: acquire the enterprise sample data to be predicted, wherein the enterprise sample data to be predicted includes at least one set of enterprise sample sub-data; based on the mapping relationship between a preset group and an expert sub-model, assign each of the enterprise sample sub-data to its corresponding target expert sub-model for feature extraction, thereby obtaining at least one enterprise sample data sub-feature of the enterprise sample to be predicted; and determine the enterprise label of the enterprise sample to be predicted by predicting enterprise label based on each of the enterprise sample data sub-features of the enterprise sample to be predicted through the prediction sub-model.
[0143] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0144] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0145] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0146] The computer-readable storage medium provided by this invention stores computer-readable program instructions for executing the above-described enterprise label prediction optimization method based on a large model, thus solving the technical problem in related technologies where the low utilization rate of enterprise data leads to low prediction accuracy of the trained label prediction model. Compared with related technologies, the advantages of the computer-readable storage medium provided in this invention are the same as those of the enterprise label prediction optimization method based on a large model provided in the above-described embodiments, and will not be repeated here.
[0147] Example 8
[0148] Furthermore, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the enterprise label prediction optimization method based on the large model described above.
[0149] The computer program product provided in this application solves the technical problem in related technologies where the prediction accuracy of the trained label prediction model is low due to the low utilization rate of enterprise data. Compared with related technologies, the advantages of the computer program product provided in this embodiment are the same as those of the enterprise label prediction optimization method based on a large model provided in the above embodiments, and will not be repeated here.
[0150] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.
Claims
1. A method for enterprise label prediction optimization based on a large model, characterized in that, The large-model-based enterprise label prediction optimization method is applied to a first device, on which an enterprise label prediction model to be trained is deployed. The enterprise label prediction model to be trained includes a prediction sub-model and at least one pre-trained expert sub-model. The large-model-based enterprise label prediction optimization method includes the following steps: Obtain training enterprise sample data and enterprise sample labels, wherein the training enterprise sample data includes at least one set of training enterprise sample sub-data; Based on the mapping relationship between the preset grouping and the expert sub-model, each of the training enterprise sample sub-data is assigned to its corresponding target expert sub-model for feature extraction, thereby obtaining at least one training enterprise sample data sub-feature of the training enterprise sample. The prediction sub-model predicts enterprise labels based on the sub-features of each of the training enterprise sample data, determines the training label prediction results of the training enterprise samples, and iteratively optimizes the prediction sub-model and each of the target expert sub-models based on the training label prediction results and the enterprise sample labels. The training enterprise sample data includes training enterprise sample table data; the step of assigning each training enterprise sample sub-data to its corresponding target expert sub-model for feature extraction based on the mapping relationship between preset groups and expert sub-models, and obtaining at least one training enterprise sample data sub-feature of the training enterprise samples, includes: Extract at least one sub-data of training enterprise samples from the training enterprise sample table data; The training enterprise sample sub-data includes at least one of training enterprise sample text sub-data and training enterprise sample numerical sub-data; the sub-data element includes at least one of dimensional text data element and dimensional numerical data element. Each of the training enterprise sample text sub-data and its corresponding dimensional information are concatenated to form at least one dimensional text data element; And / or, according to the preset bucketing rules, the numerical sub-data of each training enterprise sample is bucketed, the bucket name corresponding to each numerical sub-data of each training enterprise sample is determined, and the bucket name corresponding to each numerical sub-data of each training enterprise sample is concatenated with the corresponding dimension information to form at least one dimension numerical data element. The sub-data elements are concatenated to form training enterprise sample sequence data, wherein the training enterprise sample sequence data includes at least one sub-data element, and the sub-data element includes dimensional information and training enterprise sample sub-data. Based on the mapping relationship between the preset grouping and the expert sub-model, each of the sub-data elements is assigned to its corresponding target expert sub-model for feature extraction, thereby obtaining at least one training enterprise sample data sub-feature of the training enterprise sample.
2. The enterprise label prediction optimization method based on a large model as described in claim 1, characterized in that, The preset grouping includes at least one regional industry grouping and at least one data domain grouping; the expert sub-model includes at least one regional industry expert sub-model and at least one data domain expert sub-model; the training enterprise sample data sub-features include regional industry sub-features and at least one data domain sub-feature. The step of allocating each training enterprise sample sub-data to its corresponding target expert sub-model for feature extraction based on the mapping relationship between preset groups and expert sub-models, and obtaining at least one training enterprise sample data sub-feature of the training enterprise samples, includes: Based on the training enterprise sample data, the target region and target industry corresponding to the training enterprise sample are determined. Based on the target region and target industry, the target region and industry group to which the training enterprise sample data belongs is determined. Based on the mapping relationship between the preset group and the expert sub-model, the training enterprise sample data is assigned to the target region and industry expert sub-model corresponding to the target region and industry group for feature extraction, thereby obtaining the region and industry sub-features of the training enterprise sample. Based on the dimensional information corresponding to each of the training enterprise sample sub-data, the target data domain group to which each of the training enterprise sample sub-data belongs is determined. Based on the mapping relationship between the preset group and the expert sub-model, the target data domain expert sub-model corresponding to each of the training enterprise sample sub-data is determined. Each of the training enterprise sample sub-data is assigned to its corresponding target data domain expert sub-model for feature extraction, thereby obtaining at least one data domain sub-feature of the training enterprise sample.
3. The enterprise label prediction optimization method based on a large model as described in claim 1, characterized in that, Before the step of allocating each training enterprise sample sub-data to its corresponding target expert sub-model for feature extraction based on the mapping relationship between preset groups and expert sub-models, and obtaining at least one training enterprise sample data sub-feature of the training enterprise samples, the method further includes: Obtain the pre-training corpus corresponding to each of the expert sub-models; Each of the expert sub-models is pre-trained under self-supervised conditions based on the pre-training corpus.
4. A method for predicting enterprise labels, characterized in that, The enterprise label prediction method is applied to a second device, on which an enterprise label prediction model is deployed. The enterprise label prediction model includes a prediction sub-model and at least one expert sub-model. The enterprise label prediction model is trained using the enterprise label prediction optimization method based on a large model as described in any one of claims 1-3. The enterprise label prediction method includes the following steps: Obtain the sample data of the enterprises to be predicted, wherein the sample data of the enterprises to be predicted includes at least one set of sample sub-data of the enterprises to be predicted; Based on the mapping relationship between the preset grouping and the expert sub-model, each of the enterprise sample sub-data to be predicted is assigned to its corresponding target expert sub-model for feature extraction, thereby obtaining at least one enterprise sample data sub-feature of the enterprise sample to be predicted. The enterprise label of the enterprise sample is determined by predicting the enterprise label based on the sub-features of the sample data of each enterprise to be predicted using the prediction sub-model.
5. A large-model-based enterprise label prediction and optimization device, characterized in that, The enterprise label prediction optimization device based on a large model is applied to a first device, on which an enterprise label prediction model to be trained is deployed. The enterprise label prediction model to be trained includes at least one expert sub-model and a prediction sub-model. The enterprise label prediction and optimization device based on the large model includes: The first acquisition module is used to acquire training enterprise sample data and enterprise sample labels of training enterprise samples, wherein the training enterprise sample data includes at least one set of training enterprise sample sub-data. The first feature extraction module is used to allocate each training enterprise sample sub-data to its corresponding target expert sub-model for feature extraction based on the mapping relationship between preset groups and expert sub-models, thereby obtaining at least one training enterprise sample data sub-feature of the training enterprise sample; the training enterprise sample data includes training enterprise sample table data; the first feature extraction module is further used to: extract at least one training enterprise sample sub-data from the training enterprise sample table data; the training enterprise sample sub-data includes at least one of training enterprise sample text sub-data and training enterprise sample numerical sub-data; the sub-data element includes at least one of dimensional text data element and dimensional numerical data element; and concatenate each training enterprise sample text sub-data with its corresponding dimensional information to form at least one dimension. Text data elements; and / or, according to a preset bucketing rule, the numerical sub-data of each training enterprise sample is bucketed, and the bucket name corresponding to each numerical sub-data of the training enterprise sample is determined. The bucket name corresponding to each numerical sub-data of the training enterprise sample is concatenated with the corresponding dimension information to form at least one dimension numerical data element; the sub-data elements are concatenated to form training enterprise sample sequence data, wherein the training enterprise sample sequence data includes at least one sub-data element, and the sub-data element includes dimension information and training enterprise sample sub-data; based on the mapping relationship between preset grouping and expert sub-model, each sub-data element is assigned to its corresponding target expert sub-model for feature extraction to obtain at least one training enterprise sample data sub-feature of the training enterprise sample; The optimization module is used to predict enterprise labels based on the sub-features of each of the training enterprise sample data through the prediction sub-model, determine the training label prediction result of the training enterprise sample, and iteratively optimize the prediction sub-model and each of the target expert sub-models based on the training label prediction result and the enterprise sample label.
6. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to at least one of the processors; wherein, The memory stores instructions that can be executed by at least one of the processors to enable the at least one processor to perform the steps of the method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program implementing a large-model-based enterprise label prediction optimization method, the program being executed by a processor to implement the steps of the method as described in any one of claims 1 to 4.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Task execution method, device and equipment and computer readable storage medium
CN112148952A
Industrial classification model training method and device, industry classification model using method and device, equipment and medium
CN112417150A