Model training and service execution

By using the two-stage feature extraction model training method in the large language model, processing structured data and generating business text, the problem of poor accuracy in structured data business execution is solved, and higher business execution accuracy and cost-effectiveness are achieved.

WO2025123985A1PCT designated stage expired Publication Date: 2025-06-19ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Patent Information

Application Number
PCT/CN2024/128662
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-10-30
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

When the existing large language model processes structured data, the accuracy of business execution is poor, resulting in the unsatisfactory results of some business execution.

Method used

By obtaining historical structured business data in the specified business scenario, using the two-stage feature extraction model training method, the structured feature representation is first extracted, and then interacted with the preset supplementary feature representation to generate business text. The second feature extraction model is trained to minimize the deviation between the business text and the actual business text as the optimization goal.

Benefits of technology

Improve the accuracy of results of business execution through large language models, reduce the cost of executing various types of businesses, and improve the interpretability of text generation models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024128662_19062025_PF_FP_ABST
    Figure CN2024128662_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present description are a model training method and apparatus, a service execution method and apparatus, and a storage medium and a device. A server can obtain a second feature extraction model by means of two-stage training, so as to adjust, by means of the second feature extraction model, a structured feature representation, which is output by a first feature extraction model, thereby reducing the difference between the structured feature representation, which is output by the first feature extraction model, and a text feature representation corresponding to text data required by a text generation model, and thus the accuracy of a result of service execution performed by means of the text generation model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Model training and business execution Technical Field

[0001] This specification relates to the field of computer technology, and in particular to model training and business execution. Background Art

[0002] With the development of artificial intelligence technology, large language models (Generative Pre-trained Transformer, ChatGPT) have been widely used in various fields. By converting various types of downstream businesses into language generation businesses (for example, converting risk control businesses based on user data into language generation businesses that generate text based on user data to determine whether a user is a risky user, so that business risk control can be performed based on the language text generated by the converted language generation business to protect the user's personal privacy data), and uniformly executing them through large language models, the method of reducing the execution cost of downstream businesses has also received increasing attention.

[0003] However, since the data used in some businesses is usually complex structured data, and the data that can be processed by large language models is usually text data, the accuracy of the results of executing these businesses through large language generation models is poor.

[0004] Therefore, how to improve the accuracy of business execution results through large language models is an urgent problem that needs to be solved.

[0005] Summary of the Invention

[0006] This specification provides a model training method, including: obtaining historical structured business data in a specified business scenario as first sample data; inputting the first sample data into a pre-trained first feature extraction model to determine the structured feature representation corresponding to the first sample data through the first feature extraction model; inputting the structured feature representation into a pre-trained second feature extraction model to determine the interactive feature representation between the structured feature representation and a preset supplementary feature representation through the second feature extraction model, the preset supplementary feature representation being used to characterize the descriptive text used to describe the first sample data; inputting the interactive feature representation into a preset text generation model to generate business text according to the interactive feature representation through the text generation model; and training the second feature extraction model with the optimization goal of minimizing the deviation between the business text generated by the text generation model and the business text actually corresponding to the first sample data.

[0007] In some embodiments, training the first feature extraction model includes: inputting the first sample data into the first feature extraction model to be trained, so as to determine the field feature representation corresponding to each type of data field contained in the first sample data through the first feature extraction model to be trained, and determining the attention field feature representation corresponding to the field feature representation based on the correlation between the field feature representation and other field feature representations; determining the sample structured feature representation corresponding to the first sample data based on each attention field feature representation; taking other sample data whose sample labels match the first sample data as reference sample data; and training the first feature extraction model to be trained with the optimization goal being that the similarity between the sample structured feature representation corresponding to the first sample data and the sample structured feature representation corresponding to the reference sample data is greater than the similarity between the sample structured feature representation corresponding to the first sample data and the sample structured feature representation corresponding to other sample data except the first sample data and the reference sample data.

[0008] In some embodiments, training the first feature extraction model includes: inputting the first sample data into the first feature extraction model to be trained, so as to determine the field feature representation corresponding to each type of data field contained in the first sample data through the first feature extraction model to be trained, and determining the attention field feature representation corresponding to the field feature representation based on the correlation between the field feature representation and other field feature representations; determining the sample structured feature representation corresponding to the first sample data based on each attention field feature representation; inputting the sample structured feature representation into a preset classifier model, so as to predict the sample label corresponding to the sample data corresponding to the sample structured feature representation based on the sample structured feature representation through the classifier model as the predicted sample label; and training the first feature extraction model to be trained with the optimization goal of minimizing the deviation between the predicted sample label and the sample label actually corresponding to the first sample data.

[0009] In some embodiments, pre-training a second feature extraction model includes: obtaining second sample data; inputting the second sample data into the first feature extraction model to obtain a sample structured feature representation; inputting the sample structured feature representation and a preset initial supplementary feature representation into the second feature extraction model to be trained to determine a sample interaction feature representation between the sample structured feature representation and the initial supplementary feature representation; and pre-training the second feature extraction model based on the sample interaction feature representation and the descriptive text corresponding to the second sample data.

[0010] In some embodiments, a feature interaction network and a feature extraction network are provided in the second feature extraction model; the second feature extraction model is pre-trained according to the sample interaction feature representation and the descriptive text corresponding to the second sample data, including: inputting the descriptive text corresponding to the second sample data into the feature extraction network in the second feature extraction model to be trained, so as to determine the text feature representation of the descriptive text corresponding to the second sample data through the feature extraction network, as the text feature representation corresponding to the sample structured feature representation; the second feature extraction model is pre-trained with the optimization goal of the greater similarity between the sample interaction feature representation and the text feature representation than the similarity between the sample interaction feature representation and the text feature representation corresponding to other sample structured feature representations.

[0011] In some embodiments, a feature interaction network and a feature extraction network are provided in the second feature extraction model; the second feature extraction model is pre-trained according to the sample interaction feature representation and the descriptive text corresponding to the second sample data, including: inputting the descriptive text corresponding to the second sample data into the feature extraction network in the second feature extraction model to be trained, so as to determine the text feature representation of the descriptive text corresponding to the second sample data through the feature extraction network, as the text feature representation corresponding to the sample structured feature representation; inputting the sample interaction feature representation and the text feature representation into a preset classification model, so that the classification model determines the classification result of the second sample data according to the sample interaction feature representation and the text feature representation; and pre-training the second feature extraction model with the optimization goal of minimizing the deviation between the classification result of the second sample data and the actual classification result of the second sample data.

[0012] In some embodiments, the second feature extraction model is pre-trained based on the sample interaction feature representation and the descriptive text corresponding to the second sample data, including: inputting the sample structured feature representation into a preset generation model to generate the descriptive text corresponding to the second sample data as predicted description text through the generation model; and pre-training the second feature extraction model with the optimization goal of minimizing the deviation between the predicted description text and the description text corresponding to the second sample data.

[0013] In some embodiments, the second feature extraction model is pre-trained based on the sample interaction feature representation and the descriptive text corresponding to the second sample data, including: adjusting the initial supplementary feature representation based on the sample interaction feature representation and the descriptive text corresponding to the second sample data to obtain an adjusted supplementary feature representation, and inputting the adjusted supplementary feature representation and the sample structured feature representation into the second feature extraction model to pre-train the second feature extraction model.

[0014] This specification provides a business execution method, including: receiving business data sent by a user, wherein the business data is structured data; inputting the business data into a preset first feature extraction model to determine the structured feature representation corresponding to the business data through the first feature extraction model; inputting the structured feature representation into a pre-trained second feature extraction model to determine the interactive feature representation between the structured feature representation and a preset supplementary feature representation through the second feature extraction model, wherein the second feature extraction model is trained using the above-mentioned model training method; inputting the interactive feature representation into a preset text generation model to generate a business text according to the interactive feature representation through the text generation model, and performing business execution according to the business text.

[0015] This specification provides a model training device, comprising: an acquisition module for acquiring historical structured business data in a specified business scenario as first sample data; a first feature extraction module for inputting the first sample data into a pre-trained first feature extraction model to determine the structured feature representation corresponding to the first sample data through the first feature extraction model; a second feature extraction module for inputting the structured feature representation into a pre-trained second feature extraction model to determine the interactive feature representation between the structured feature representation and a preset supplementary feature representation through the second feature extraction model, wherein the preset supplementary feature representation is used to characterize the descriptive text used to describe the first sample data; an interaction module for inputting the interactive feature representation into a preset text generation model to generate business text according to the interactive feature representation through the text generation model; and a training module for training the second feature extraction model with the optimization goal of minimizing the deviation between the business text generated by the text generation model and the business text actually corresponding to the first sample data.

[0016] This specification provides a business execution device, including: a data receiving module, used to receive business data sent by a user, wherein the business data is structured data; a structured feature extraction module, used to input the business data into a preset first feature extraction model, so as to determine the structured feature representation corresponding to the business data through the first feature extraction model; an interactive feature extraction module, used to input the structured feature representation into a pre-trained second feature extraction model, so as to determine the interactive feature representation between the structured feature representation and a preset supplementary feature representation through the second feature extraction model, wherein the second feature extraction model is trained by the above-mentioned model training method; an execution module, used to input the interactive feature representation into a preset text generation model, so as to generate a business text according to the interactive feature representation through the text generation model, and perform business execution according to the business text.

[0017] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned model training and business execution methods.

[0018] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, the above-mentioned model training and business execution methods are implemented.

[0019] At least one of the above-mentioned technical solutions adopted in this specification can achieve the following beneficial effects: in the model training method provided in this specification, historical structured business data in a specified business scenario is obtained as the first sample data, and the first sample data is input into a pre-trained first feature extraction model to determine the structured feature representation corresponding to the first sample data through the first feature extraction model, and the structured feature representation is input into a pre-trained second feature extraction model to determine the interactive feature representation between the structured feature representation and the preset supplementary feature representation through the second feature extraction model, where the preset supplementary feature representation is used to characterize the descriptive text used to describe the first sample data, and the interactive feature representation is input into a preset text generation model to generate business text according to the interactive feature representation through the text generation model, and the second feature extraction model is trained with the optimization goal of minimizing the deviation between the business text generated by the text generation model and the business text actually corresponding to the first sample data.

[0020] It can be seen from the above method that the server can obtain the second feature extraction model through two-stage training, so that the structured feature representation output by the first feature extraction model can be adjusted through the second feature extraction model to reduce the difference between the structured feature representation output by the first feature extraction model and the text feature representation corresponding to the text data required by the text generation model, thereby improving the accuracy of the results of business execution through the text generation model. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings described herein are used to provide further understanding of this specification and constitute a part of this specification. The illustrative embodiments of this specification and their descriptions are used to explain this specification and do not constitute improper limitations on this specification.

[0022] FIG1 is a flow chart of a model training method provided in this specification.

[0023] FIG2 is a schematic diagram of the structure of the first feature extraction model provided in this specification.

[0024] FIG3 is a schematic diagram of the structure of the second feature extraction model provided in this specification.

[0025] FIG4 is a flow chart of a service execution method provided in this specification.

[0026] FIG5 is a schematic diagram of a model training device provided in this specification.

[0027] FIG6 is a schematic diagram of a service execution device provided in this specification.

[0028] FIG7 is a schematic diagram of an electronic device provided in this specification corresponding to FIG1 . DETAILED DESCRIPTION

[0029] To make the objectives, technical solutions, and advantages of this specification more clear, the following will clearly and completely describe the technical solutions of this specification in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.

[0030] The technical solutions provided by the embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0031] FIG1 is a flow chart of a model training method provided in this specification, including steps S100 to S108.

[0032] S100: Acquire historical structured business data in a specified business scenario as first sample data.

[0033] In this specification, when the business platform needs to convert a target business that uses structured data into a text generation business that is executed by a text generation model, it can use a first feature extraction model to extract the corresponding structured feature representation for the structured data used in the target business, so that the structured feature representation extracted by the first feature extraction model can be adjusted by a second feature extraction model to avoid the situation where the text generation model has poor accuracy when executing the text generation business after converting the target business due to the difference between the structured feature representation extracted by the first feature extraction model and the traditional text feature representation used by the text generation model. The above-mentioned text generation model can be a large natural language processing model (Chat Generative Pre-trained Transformer, ChatGPT).

[0034] The above-mentioned first feature extraction model and second feature extraction model need to be trained before they can be deployed to the business platform to be used to execute the text generation business after the conversion of the target business. Before training the above-mentioned first feature extraction model and second feature extraction model, the business platform can first obtain historical structured business data under the specified business scenario as the first sample data, and then use the first sample data to train the above-mentioned first feature extraction model and second feature extraction model. The specified business scenario here can be determined according to actual needs, such as: risk control business scenario, data classification business scenario, etc.

[0035] For example, in a risk control business scenario, the risk control business that detects whether a user's transaction behavior is risky based on the user's structured data can be converted into a text generation business that generates text based on the user's structured data through a text generation model to determine whether the user's transaction behavior is risky, and a text generation business that generates text about the reasons why the user's transaction behavior is risky. In this case, the historical structured business data in the risk control business scenario is the user's structured data.

[0036] In this specification, the execution entity used to implement the model training method can refer to a designated device such as a server set up on the business platform, or it can refer to a terminal device such as a desktop computer, a laptop computer, etc. For the sake of convenience of description, the model training method provided in this specification is explained below using the server as the execution entity as an example.

[0037] S102: Input the first sample data into a pre-trained first feature extraction model to determine a structured feature representation corresponding to the first sample data through the first feature extraction model.

[0038] In some embodiments, the server can input the first sample data into a pre-trained first feature extraction model, so that through the first feature extraction model, for each type of data field contained in the first sample data, the data field of that type and the type name of that type are separately encoded to obtain the field feature representation corresponding to the data field of that type, and then the structured feature representation corresponding to the first sample data can be determined based on the feature representation of each field, where the structure of the first feature extraction model is shown in Figure 2.

[0039] FIG2 is a schematic diagram of the structure of the first feature extraction model provided in this specification.

[0040] As can be seen from Figure 2, the first feature extraction model includes: an input processing module and a feature extraction module, wherein the server can use the input processing module to separately encode the data field of each type (for example, Cat, Bin, Num) contained in the first sample data and the type name of the type, and obtain the field feature representation corresponding to the data field of the type. Then, the feature extraction module can be used to determine the attention field feature representation corresponding to the field feature representation based on the correlation between each field feature representation and other field feature representations, and then the sample structured feature representation corresponding to the first sample data can be determined based on each attention field feature representation.

[0041] Among them, the method for the server to determine the attention field feature representation corresponding to the field feature representation through the feature extraction module according to the correlation between each field feature representation and other field feature representations can be to determine the sub-attention field feature representations corresponding to the field feature representation through a multi-head attention mechanism, and then the sub-attention field feature representations corresponding to the field feature representation can be spliced ​​and input into the linear layer, so as to perform linear conversion on the sub-attention field feature representations corresponding to the spliced ​​field feature representation through the linear layer to obtain the attention field feature representation corresponding to the field feature representation.

[0042] It should be noted that there are two methods for training the first feature extraction model. The following describes in detail these two methods for training the first feature extraction model.

[0043] The first method may be to input the first sample data into a first feature extraction model to be trained, so as to determine, through the first feature extraction model to be trained, for each type of data field contained in the first sample data, the field feature representation corresponding to the data field of that type, and determine, based on the correlation between the field feature representation and other field feature representations, the attention field feature representation corresponding to the field feature representation.

[0044] Therefore, the sample structured feature representation corresponding to the first sample data can be determined based on the feature representation of each attention field, and then other sample data with sample labels matching the first sample data can be used as reference sample data. Then, the first feature extraction model to be trained is trained with the optimization goal being that the greater the similarity between the sample structured feature representation corresponding to the first sample data and the sample structured feature representation corresponding to the reference sample data is compared with the similarity between the sample structured feature representation corresponding to the first sample data and the sample structured feature representation corresponding to other sample data except the first sample data and the reference sample data.

[0045] The second method may be to input the first sample data into a first feature extraction model to be trained, so as to determine, through the first feature extraction model to be trained, for each type of data field contained in the first sample data, a field feature representation corresponding to the data field of that type, and determine, based on the correlation between the field feature representation and other field feature representations, the attention field feature representation corresponding to the field feature representation.

[0046] In some embodiments, the sample structured feature representation corresponding to the first sample data can be determined based on the feature representation of each attention field, and the sample structured feature representation can be input into a preset classifier model to predict the sample label corresponding to the sample data corresponding to the sample structured feature representation based on the sample structured feature representation through the classifier model. The sample label is used as the predicted sample label, and the first feature extraction model to be trained is trained with the optimization goal of minimizing the deviation between the predicted sample label and the sample label actually corresponding to the first sample data.

[0047] The above-mentioned sample labels can be determined according to the specified business scenario described in the sample data. For example, in a risk control business scenario, the result of whether the sample data is risky can be used as the above-mentioned sample label.

[0048] It should be noted that the above two methods can be used separately or simultaneously. In order to improve the accuracy of the structured feature representation of the sample data extracted by the trained first feature extraction model, the server can first train the first feature extraction model using the above first method to obtain the first feature extraction model after the first training, and then use the above second method to train the first feature extraction model after the first training for the second time to obtain the final trained first feature extraction model.

[0049] In addition, in some business scenarios, the amount of first sample data that the server can obtain for training the first feature extraction model may be small. At this time, in order to avoid the situation where the training effect of the first feature extraction model is poor due to the small amount of first sample data for training the first feature extraction model, the server can also split the acquired first sample data (such as: split the first sample data according to the various types of data fields contained in the first sample data) to obtain each split first sample data, and then the first feature extraction model can be trained based on each split first sample data through the above-mentioned method of training the first feature extraction model to obtain a trained first feature extraction model.

[0050] S104: Input the structured feature representation into a pre-trained second feature extraction model to determine, through the second feature extraction model, an interactive feature representation between the structured feature representation and a preset supplementary feature representation, where the preset supplementary feature representation is used to characterize the descriptive text used to describe the first sample data.

[0051] In this specification, the server can input the structured feature representation extracted by the first feature extraction model into a pre-trained second feature extraction model to determine the interactive feature representation between the structured feature representation and the preset supplementary feature representation through the second feature extraction model, wherein the second feature extraction model is shown in Figure 3.

[0052] FIG3 is a schematic diagram of the structure of the second feature extraction model provided in this specification.

[0053] As can be seen from Figure 3, a feature interaction network and a feature extraction network are provided in the second feature extraction model. The server can input the structured feature representation extracted by the first feature extraction model into the feature interaction network of the pre-trained second feature extraction model to determine the interactive feature representation between the structured feature representation and the preset supplementary feature representation through the feature interaction network of the second feature extraction model.

[0054] Among them, the pre-training method of the above-mentioned second feature extraction model can be to obtain second sample data, input the second sample data into the first feature extraction model to obtain a sample structured feature representation, input the sample structured feature representation and the preset initial supplementary feature representation into the second feature extraction model to be trained to determine the sample interaction feature representation between the sample structured feature representation and the initial supplementary feature representation, and pre-train the second feature extraction model according to the sample interaction feature representation and the descriptive text corresponding to the second sample data.

[0055] In the above content, there are three methods for the server to pre-train the second feature extraction model based on the sample interaction feature representation and the descriptive text corresponding to the second sample data. These three methods can be used individually or together. The following describes these three methods in detail.

[0056] The first method can be to input the descriptive text corresponding to the second sample data into the feature extraction network in the second feature extraction model to be trained, so as to determine the text feature representation of the descriptive text corresponding to the second sample data through the feature extraction network, as the text feature representation corresponding to the sample structured feature representation, and then pre-train the second feature extraction model with the optimization goal being that the similarity between the sample interaction feature representation and the text feature representation is greater than the similarity between the sample interaction feature representation and the text feature representation corresponding to other sample structured feature representations, and adjust the initial supplementary feature representation to obtain the adjusted supplementary feature representation, and input the adjusted supplementary feature representation and the sample structured feature representation into the second feature extraction model to pre-train the second feature extraction model.

[0057] The second method can be to input the descriptive text corresponding to the second sample data into the feature extraction network in the second feature extraction model to be trained, so as to determine the text feature representation of the descriptive text corresponding to the second sample data through the feature extraction network, as the text feature representation corresponding to the sample structured feature representation, and then the sample interaction feature representation and the text feature representation can be input into a preset classification model, so that the classification model determines the classification result of the second sample data according to the sample interaction feature representation and the text feature representation, and pre-train the second feature extraction model with minimizing the deviation between the classification result of the second sample data and the actual classification result of the second sample data as the optimization goal, and adjust the initial supplementary feature representation to obtain the adjusted supplementary feature representation, and input the adjusted supplementary feature representation and the sample structured feature representation into the second feature extraction model to pre-train the second feature extraction model.

[0058] A third method is to input the sample structured feature representation into a preset generation model to generate a description text corresponding to the second sample data through the generation model as the predicted description text, pre-train the second feature extraction model with minimizing the deviation between the predicted description text and the description text corresponding to the second sample data as the optimization goal, and adjust the initial supplementary feature representation to obtain an adjusted supplementary feature representation, and input the adjusted supplementary feature representation and the sample structured feature representation into the second feature extraction model to pre-train the second feature extraction model.

[0059] It should be noted that the above-mentioned descriptive text can be determined based on the second sample data. For example, if the second sample data is user data in a risk control business scenario, the above-mentioned descriptive text can be an evaluation text of whether there are high-frequency abnormal transactions in the user's historical transaction behavior, an evaluation text of the user's current network environment, etc.

[0060] It is worth noting that there may be only one supplementary feature representation or a set of multiple supplementary feature representations.

[0061] S106: Inputting the interaction feature representation into a preset text generation model, so as to generate a business text according to the interaction feature representation through the text generation model.

[0062] S108: The second feature extraction model is trained with minimizing the deviation between the business text generated by the text generation model and the business text actually corresponding to the first sample data as an optimization goal.

[0063] In some embodiments, the server can input the interactive feature representation into a preset text generation model to generate business text based on the interactive feature representation through the text generation model, and then train the second feature extraction model with the optimization goal of minimizing the deviation between the business text generated by the text generation model and the business text actually corresponding to the first sample data. Then, the trained first feature extraction model, second feature extraction model and text generation model can be used to execute the text generation business obtained by converting the target business.

[0064] From the above content, it can be seen that the server can obtain the second feature extraction model through two-stage training, so that the structured feature representation output by the first feature extraction model can be adjusted through the second feature extraction model to reduce the difference between the structured feature representation output by the first feature extraction model and the text feature representation corresponding to the text data required by the text generation model, thereby improving the accuracy of the results of business execution through the text generation model.

[0065] In order to illustrate the above content in some embodiments, the following describes in detail the method for executing business using the first feature extraction model, the second feature extraction model, and the text generation model trained by the above method, as shown in FIG4 .

[0066] FIG4 is a flowchart of a service execution method provided in this specification, including steps S400 to S406 .

[0067] S400: receiving business data sent by a user, where the business data is structured data.

[0068] S402: Input the business data into a preset first feature extraction model to determine a structured feature representation corresponding to the business data through the first feature extraction model.

[0069] S404: Input the structured feature representation into a pre-trained second feature extraction model to determine the interactive feature representation between the structured feature representation and the preset supplementary feature representation through the second feature extraction model, where the second feature extraction model is trained using the above-mentioned model training method.

[0070] S406: Inputting the interaction feature representation into a preset text generation model, so as to generate a business text according to the interaction feature representation through the text generation model, and perform business execution according to the business text.

[0071] In this specification, when the business platform receives structured data sent by a user as business data, it can input the received business data into a preset first feature extraction model to determine the structured feature representation corresponding to the business data through the first feature extraction model.

[0072] In some embodiments, the server can input the determined structured feature representation into a pre-trained second feature extraction model to determine the interactive feature representation between the structured feature representation and the preset supplementary feature representation through the second feature extraction model, and then input the interactive feature representation into a preset text generation model to generate business text according to the interactive feature representation through the text generation model, and perform business execution based on the business text.

[0073] Among them, the above-mentioned business execution can be determined according to the business scenario to which the input business data belongs. For example: in the risk control business scenario, a text generation model can be used to generate text based on the interactive feature representation to characterize whether the input business data is at risk, and to generate text corresponding to the reasons for determining that the business data is at risk.

[0074] From the above content, it can be seen that the server can perform text generation tasks obtained by business conversion involving structured data through the trained first feature extraction model, second feature extraction model and text generation model, and can also generate corresponding explanatory text to improve the interpretability of the text generation model, and can reduce the cost of executing various types of businesses by uniformly converting various types of businesses into text generation businesses.

[0075] The above are model training and business execution methods provided in one or more embodiments of this specification. Based on the same idea, this specification also provides corresponding model training and business execution devices, as shown in Figures 5 and 6.

[0076] Figure 5 is a schematic diagram of a model training device provided in this specification, including: an acquisition module 501, used to acquire historical structured business data in a specified business scenario as first sample data; a first feature extraction module 502, used to input the first sample data into a pre-trained first feature extraction model, so as to determine the structured feature representation corresponding to the first sample data through the first feature extraction model; a second feature extraction module 503, used to input the structured feature representation into a pre-trained second feature extraction model, so as to determine the interactive feature representation between the structured feature representation and a preset supplementary feature representation through the second feature extraction model, and the preset supplementary feature representation is used to characterize the descriptive text used to describe the first sample data; an interaction module 504, used to input the interactive feature representation into a preset text generation model, so as to generate business text according to the interactive feature representation through the text generation model; a training module 505, used to train the second feature extraction model with the optimization goal of minimizing the deviation between the business text generated by the text generation model and the business text actually corresponding to the first sample data.

[0077] In some embodiments, the training module 505 is used to input the first sample data into the first feature extraction model to be trained, so as to determine the field feature representation corresponding to each type of data field contained in the first sample data through the first feature extraction model to be trained, and determine the attention field feature representation corresponding to the field feature representation based on the correlation between the field feature representation and other field feature representations; determine the sample structured feature representation corresponding to the first sample data based on each attention field feature representation; use other sample data whose sample labels match the first sample data as reference sample data; and train the first feature extraction model to be trained with the optimization goal being that the similarity between the sample structured feature representation corresponding to the first sample data and the sample structured feature representation corresponding to the reference sample data is greater than the similarity between the sample structured feature representation corresponding to the first sample data and the sample structured feature representation corresponding to other sample data except the first sample data and the reference sample data.

[0078] In some embodiments, the training module 505 is used to input the first sample data into the first feature extraction model to be trained, so as to determine the field feature representation corresponding to each type of data field contained in the first sample data through the first feature extraction model to be trained, and determine the attention field feature representation corresponding to the field feature representation based on the correlation between the field feature representation and other field feature representations; determine the sample structured feature representation corresponding to the first sample data based on each attention field feature representation; input the sample structured feature representation into a preset classifier model, so as to predict the sample label corresponding to the sample data corresponding to the sample structured feature representation based on the sample structured feature representation through the classifier model, as the predicted sample label; and train the first feature extraction model to be trained with the optimization goal of minimizing the deviation between the predicted sample label and the sample label actually corresponding to the first sample data.

[0079] In some embodiments, the training module 505 is used to obtain second sample data; input the second sample data into the first feature extraction model to obtain a sample structured feature representation; input the sample structured feature representation and a preset initial supplementary feature representation into the second feature extraction model to be trained to determine a sample interaction feature representation between the sample structured feature representation and the initial supplementary feature representation; and pre-train the second feature extraction model based on the sample interaction feature representation and the descriptive text corresponding to the second sample data.

[0080] In some embodiments, a feature interaction network and a feature extraction network are provided in the second feature extraction model; the training module 505 is used to input the description text corresponding to the second sample data into the feature extraction network in the second feature extraction model to be trained, so as to determine the text feature representation of the description text corresponding to the second sample data through the feature extraction network as the text feature representation corresponding to the sample structured feature representation; the second feature extraction model is pre-trained with the optimization goal being that the similarity between the sample interaction feature representation and the text feature representation is greater than the similarity between the sample interaction feature representation and the text feature representation corresponding to other sample structured feature representations.

[0081] In some embodiments, a feature interaction network and a feature extraction network are provided in the second feature extraction model; the training module 505 is used to input the descriptive text corresponding to the second sample data into the feature extraction network in the second feature extraction model to be trained, so as to determine the text feature representation of the descriptive text corresponding to the second sample data through the feature extraction network, as the text feature representation corresponding to the sample structured feature representation; the sample interaction feature representation and the text feature representation are input into a preset classification model, so that the classification model determines the classification result of the second sample data according to the sample interaction feature representation and the text feature representation; and the second feature extraction model is pre-trained with the optimization goal of minimizing the deviation between the classification result of the second sample data and the actual classification result of the second sample data.

[0082] In some embodiments, the training module 505 is used to input the sample structured feature representation into a preset generation model to generate a description text corresponding to the second sample data as a predicted description text through the generation model; and pre-train the second feature extraction model with the optimization goal of minimizing the deviation between the predicted description text and the description text corresponding to the second sample data.

[0083] In some embodiments, the training module 505 is used to adjust the initial supplementary feature representation according to the sample interaction feature representation and the descriptive text corresponding to the second sample data to obtain an adjusted supplementary feature representation, and input the adjusted supplementary feature representation and the sample structured feature representation into the second feature extraction model to pre-train the second feature extraction model.

[0084] Figure 6 is a schematic diagram of a business execution device provided in this specification, including: a data receiving module 601, used to receive business data sent by a user, wherein the business data is structured data; a structured feature extraction module 602, used to input the business data into a preset first feature extraction model, so as to determine the structured feature representation corresponding to the business data through the first feature extraction model; an interactive feature extraction module 603, used to input the structured feature representation into a pre-trained second feature extraction model, so as to determine the interactive feature representation between the structured feature representation and the preset supplementary feature representation through the second feature extraction model, wherein the second feature extraction model is trained by the above-mentioned model training method; an execution module 604, used to input the interactive feature representation into a preset text generation model, so as to generate a business text according to the interactive feature representation through the text generation model, and perform business execution according to the business text.

[0085] This specification also provides a computer-readable storage medium, which stores a computer program. The computer program can be used to execute a model training method provided in Figure 1 above.

[0086] This specification also provides a schematic structural diagram of an electronic device corresponding to Figure 1, as shown in Figure 7. As shown in Figure 7, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and of course may also include hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the model training method of Figure 1 above. Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0087] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages ​​and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.

[0088] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.

[0089] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0090] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0091] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0092] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0093] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0094] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0095] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0096] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0097] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0098] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0099] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.

[0100] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.

[0101] The above are some embodiments of this specification and are not intended to limit this specification. For those skilled in the art, various modifications and variations of this specification are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this specification should be included within the scope of the claims of this specification.

Claims

1. A model training method, comprising: Obtain historical structured business data under a specified business scenario as first sample data; Inputting the first sample data into a pre-trained first feature extraction model to determine a structured feature representation corresponding to the first sample data through the first feature extraction model; Inputting the structured feature representation into a pre-trained second feature extraction model to determine, through the second feature extraction model, an interactive feature representation between the structured feature representation and a preset supplementary feature representation, wherein the preset supplementary feature representation is used to characterize a description text used to describe the first sample data; Inputting the interaction feature representation into a preset text generation model, so as to generate a business text according to the interaction feature representation through the text generation model; The second feature extraction model is trained with the optimization goal of minimizing the deviation between the business text generated by the text generation model and the business text actually corresponding to the first sample data.

2. The method of claim 1, wherein: Training the first feature extraction model includes: Inputting the first sample data into a first feature extraction model to be trained, so as to determine, for each type of data field contained in the first sample data, a field feature representation corresponding to the data field of the type through the first feature extraction model to be trained, and determining, according to a correlation between the field feature representation and other field feature representations, an attention field feature representation corresponding to the field feature representation; Determining, according to each attention field feature representation, a sample structured feature representation corresponding to the first sample data; Using other sample data whose sample labels match the first sample data as reference sample data; The first feature extraction model to be trained is trained with the optimization goal being that the similarity between the sample structured feature representation corresponding to the first sample data and the sample structured feature representation corresponding to the reference sample data is greater than the similarity between the sample structured feature representation corresponding to the first sample data and the sample structured feature representation corresponding to other sample data except the first sample data and the reference sample data.

3. The method of claim 1, wherein: Training the first feature extraction model includes: Inputting the first sample data into a first feature extraction model to be trained, so as to determine, for each type of data field contained in the first sample data, a field feature representation corresponding to the data field of the type through the first feature extraction model to be trained, and determining, according to a correlation between the field feature representation and other field feature representations, an attention field feature representation corresponding to the field feature representation; Determining, according to each attention field feature representation, a sample structured feature representation corresponding to the first sample data; The sample structured feature representation is input into a preset classifier model, so that the classifier model predicts the sample data corresponding to the sample structured feature representation according to the sample structured feature representation. This label is used as the predicted sample label; The first feature extraction model to be trained is trained with minimizing the deviation between the predicted sample label and the sample label actually corresponding to the first sample data as an optimization goal.

4. The method of claim 1, wherein: Pre-training the second feature extraction model includes: Acquire second sample data; Inputting the second sample data into the first feature extraction model to obtain a sample structured feature representation; Inputting the sample structured feature representation and the preset initial supplementary feature representation into a second feature extraction model to be trained to determine a sample interaction feature representation between the sample structured feature representation and the initial supplementary feature representation; The second feature extraction model is pre-trained according to the sample interaction feature representation and the description text corresponding to the second sample data.

5. The method of claim 4, wherein: The second feature extraction model is provided with a feature interaction network and a feature extraction network; Pre-training the second feature extraction model according to the sample interaction feature representation and the description text corresponding to the second sample data includes: Inputting the description text corresponding to the second sample data into the feature extraction network in the second feature extraction model to be trained, so as to determine the text feature representation of the description text corresponding to the second sample data through the feature extraction network as the text feature representation corresponding to the sample structured feature representation; The second feature extraction model is pre-trained with the optimization goal of increasing the similarity between the sample interaction feature representation and the text feature representation compared to the similarity between the sample interaction feature representation and the text feature representations corresponding to other sample structured feature representations.

6. The method of claim 4, wherein: The second feature extraction model is provided with a feature interaction network and a feature extraction network; Pre-training the second feature extraction model according to the sample interaction feature representation and the description text corresponding to the second sample data includes: Inputting the description text corresponding to the second sample data into the feature extraction network in the second feature extraction model to be trained, so as to determine the text feature representation of the description text corresponding to the second sample data through the feature extraction network as the text feature representation corresponding to the sample structured feature representation; Inputting the sample interaction feature representation and the text feature representation into a preset classification model, so that the classification model determines a classification result of the second sample data according to the sample interaction feature representation and the text feature representation; The second feature extraction model is pre-trained with minimizing the deviation between the classification result of the second sample data and the actual classification result of the second sample data as an optimization goal.

7. The method of claim 4, wherein: Pre-training the second feature extraction model according to the sample interaction feature representation and the description text corresponding to the second sample data includes: Inputting the sample structured feature representation into a preset generation model to generate a description text corresponding to the second sample data as a predicted description text through the generation model; The second feature extraction model is pre-trained with minimizing the deviation between the predicted description text and the description text corresponding to the second sample data as an optimization goal.

8. The method according to any one of claims 4 to 7, wherein: Pre-training the second feature extraction model according to the sample interaction feature representation and the description text corresponding to the second sample data includes: According to the sample interaction feature representation and the description text corresponding to the second sample data, the initial supplementary feature representation is adjusted to obtain an adjusted supplementary feature representation, and the adjusted supplementary feature representation and the sample structured feature representation are input into a second feature extraction model to pre-train the second feature extraction model.

9. A business execution method, comprising: Receiving business data sent by a user, wherein the business data is structured data; Inputting the business data into a preset first feature extraction model to determine a structured feature representation corresponding to the business data through the first feature extraction model; Inputting the structured feature representation into a pre-trained second feature extraction model to determine an interactive feature representation between the structured feature representation and a preset supplementary feature representation through the second feature extraction model, wherein the second feature extraction model is trained by the method of any one of claims 1 to 8 above; The interactive feature representation is input into a preset text generation model, so that a business text is generated according to the interactive feature representation through the text generation model, and business execution is performed according to the business text.

10. A model training device, comprising: An acquisition module, used to acquire historical structured business data in a specified business scenario as first sample data; A first feature extraction module, used for inputting the first sample data into a pre-trained first feature extraction model, so as to determine a structured feature representation corresponding to the first sample data through the first feature extraction model; A second feature extraction module, used for inputting the structured feature representation into a pre-trained second feature extraction model, so as to determine an interactive feature representation between the structured feature representation and a preset supplementary feature representation through the second feature extraction model, wherein the preset supplementary feature representation is used to characterize a description text used to describe the first sample data; An interaction module, used for inputting the interaction feature representation into a preset text generation model, so as to generate a business text according to the interaction feature representation through the text generation model; A training module is used to train the second feature extraction model with the optimization goal of minimizing the deviation between the business text generated by the text generation model and the business text actually corresponding to the first sample data.

11. A service execution device, comprising: A data receiving module, used to receive business data sent by a user, wherein the business data is structured data; A structured feature extraction module, used to input the business data into a preset first feature extraction model to determine a structured feature representation corresponding to the business data through the first feature extraction model; an interactive feature extraction module, configured to input the structured feature representation into a pre-trained second feature extraction model, so as to determine an interactive feature representation between the structured feature representation and a preset supplementary feature representation through the second feature extraction model, wherein the second feature extraction model is trained by the method according to any one of claims 1 to 8; An execution module is used to input the interaction feature representation into a preset text generation model, so as to generate a business text according to the interaction feature representation through the text generation model, and perform business execution according to the business text.

12. A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 9.

13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 9 when executing the computer program.

Citation Information

Patent Citations

  • Model training method, business risk control method and device

    CN114997472A

  • Fault diagnosis method and device

    CN115408190A

  • Model training method, business risk control method and device

    CN115618962A

  • Model training method and device, service execution method and device, storage medium and equipment

    CN117743824A

  • Domain knowledge based feature extraction for enhanced text representation

    US20220129625A1

Cited By

  • Acquisition method and device of training data and related equipment

    CN121071482A