Transaction processing method and device, storage medium and electronic equipment
Patent Information
- Application Number
- CN202510080600.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
Smart Images

Figure CN119989052A_ABST
Abstract
Description
Technical Field
[0001] The present specification relates to the field of computer technology, and in particular to a transaction processing method, device, storage medium and electronic device. Background Art
[0002] With the rapid development of computer technology, transaction scenarios such as financial credit decision-making, insurance underwriting and claims, risk management and anti-fraud detection, service recommendation, and e-commerce recommendation often involve the use of large (language) models (LLM) for daily transaction processing. In practical applications, large models show a high dependence on transaction data features. At present, it is found that some factors that affect the performance of large models directly affect the transaction task reasoning ability of large models. Sparse features or insufficient quality in model input data will lead to a significant decrease in model classification performance. However, in many real scenarios, due to data collection conditions, user authorization restrictions and other issues, key feature information in model input data is often missing or insufficient, resulting in the inability of large models to fully utilize potential information resources for transaction processing. Summary of the invention
[0003] This specification provides a transaction processing method, device, storage medium and electronic device, and the technical solution is as follows:
[0004] In a first aspect, this specification provides a transaction processing method, the method comprising:
[0005] Obtaining a large model processing capability impact factor for a target transaction, and performing classification task modeling processing on the large model processing capability impact factor to obtain a task classifier;
[0006] Based on the big model base network and the task classifier, create an initial transaction processing big model for the target transaction, and determine transaction training data for the initial transaction processing big model;
[0007] The transaction training data is input into the initial transaction processing large model for at least one round of model training. During the model training process, the classification coding boundary is determined by the task classifier, and the sample transaction semantic vector of the large model base network for the transaction training data is obtained. The model parameters of the large model base network are adjusted based on the classification coding boundary and the sample transaction semantic vector to obtain the transaction processing large model.
[0008] In a second aspect, this specification provides a transaction processing method, the method comprising:
[0009] Obtain target transaction data for a target transaction;
[0010] Input the target transaction data into the transaction processing model and output the transaction processing result;
[0011] Target transaction processing is performed based on the transaction processing result.
[0012] In a third aspect, this specification provides a transaction processing device, the device comprising:
[0013] A modeling module, used to obtain a large model processing capability impact factor for a target transaction, and to perform classification task modeling processing on the large model processing capability impact factor to obtain a task classifier;
[0014] A creation module, used for creating an initial transaction processing big model for the target transaction based on the big model base network and the task classifier, and determining transaction training data for the initial transaction processing big model;
[0015] A training module is used to input the transaction training data into the initial transaction processing large model for at least one round of model training. During the model training process, the classification coding boundary is determined by the task classifier, and the sample transaction semantic vector of the large model base network for the transaction training data is obtained. The model parameters of the large model base network are adjusted based on the classification coding boundary and the sample transaction semantic vector to obtain the transaction processing large model.
[0016] In a fourth aspect, this specification provides a transaction processing device, the device comprising:
[0017] A data acquisition module, used for acquiring target transaction data for a target transaction;
[0018] A model processing module, used for inputting the target transaction data into a transaction processing large model and outputting a transaction processing result, wherein the transaction processing large model is obtained by the method steps according to any one of claims 1 to 9;
[0019] The transaction processing module is used to perform target transaction processing based on the transaction processing result.
[0020] In a fifth aspect, the present specification provides a computer storage medium, wherein the computer storage medium stores at least one instruction, wherein the instruction is suitable for being loaded by a processor and executing the method steps of one or more embodiments of the present specification.
[0021] In a sixth aspect, the present specification provides a computer program product, wherein the computer program product stores at least one instruction, wherein the instruction is suitable for being loaded by a processor and executing the method steps of one or more embodiments of the present specification.
[0022] In a seventh aspect, the present specification provides an electronic device, which may include: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method steps of one or more embodiments of the present specification.
[0023] The beneficial effects brought by the technical solutions provided by some embodiments of this specification include at least:
[0024] In one or more embodiments of the present specification, the electronic device determines the influencing factors of the large model processing capability and constructs a task classifier in a generative adversarial manner to accurately identify the key performance bottlenecks of the transaction processing model under the target transaction. During the model training process, the task classifier is combined with the large model base network to establish an initial transaction processing model and determine targeted training data. The classification coding boundary determined by the task classifier after training is used as a guide. Based on the classification coding boundary, the semantic vector distribution of the large model base network is optimized to gradually bring low-quality feature samples closer to high-quality feature samples, effectively avoiding the limitation that the performance of the large model is degraded due to feature sparsity and quality imbalance in actual transaction scenarios. The overall transaction processing process not only improves the robustness of the large model to sparse and low-quality data under specific target transactions, but also explicitly controls the semantic vector representation tendency of the large model, so that the large model moves closer to the feature encoding in the direction of good influencing factors, thereby improving the feature representation capability of the large model according to actual needs, thereby improving the performance of the large model in downstream target transaction processing tasks, meeting the diverse needs in complex scenarios, and significantly enhancing the generalization ability and accuracy of the large model in complex transaction scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in this specification or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0026] Figure 1 It is a scenario diagram of a transaction processing system provided in this specification;
[0027] Figure 2 It is a flowchart of a transaction processing method provided in this specification;
[0028] Figure 3 It is a flowchart of another transaction processing method provided in this specification;
[0029] Figure 4 It is a schematic diagram of a model training scenario of a large transaction processing model provided in this specification;
[0030] Figure 5 It is a scene diagram of a classification coding boundary provided in this manual;
[0031] Figure 6 This is a flowchart of a task classifier training process provided in this manual;
[0032] Figure 7 It is a flow chart of determining a network gradient adjustment error provided in this specification;
[0033] Figure 8 It is a flow chart determined by a large transaction processing model provided in this specification;
[0034] Fig. 9 It is a flow chart of model parameter adjustment provided in this manual;
[0035] Fig.10 It is a flowchart of a transaction processing method provided in this specification;
[0036] Fig.11 It is a structural schematic diagram of a transaction processing device provided in this specification;
[0037] Fig.12 It is a structural schematic diagram of a transaction processing device provided in this specification;
[0038] Fig.13 It is a structural schematic diagram of an electronic device provided in this manual. DETAILED DESCRIPTION
[0039] The following will be combined with the drawings in this specification to clearly and completely describe the technical solutions in this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.
[0040] In the description of this specification, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In the description of this specification, it should be noted that, unless otherwise clearly specified and limited, "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices. For those of ordinary skill in the art, the specific meanings of the above terms in this specification can be understood in specific circumstances. In addition, in the description of this specification, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are an "or" relationship.
[0041] At present, it is found that some factors that affect the performance of large models directly affect the transaction task reasoning ability of large models. Sparse features or insufficient quality in the model input data will lead to a significant decrease in the model classification performance. However, in many real scenarios, due to problems such as data collection conditions and user authorization restrictions, key feature information in the model input data is often missing or insufficient, resulting in the inability of large models to fully utilize potential information resources for transaction processing;
[0042] Furthermore, through creative work, it is found that the reasoning ability of the large model is weakened under the influence of some influencing factors (feature coverage, feature quality). In the actual modeling process, some influencing factors will greatly affect the reasoning ability of the large model. For example, taking the two influencing factors of feature coverage and feature quality as examples:
[0043] 1) In terms of feature coverage, the recognition accuracy of users with important features is higher than that of users without important features. In user driving recognition, users with trajectory and interest features have higher recognition accuracy than users without such information; in wholesaler identity recognition, users with rich features have higher accuracy than users with only store names.
[0044] 2) In terms of feature quality, the richer the feature information, the higher the recognition accuracy. In driver recognition, users whose trajectories have more than 30 points in the past month have a higher accuracy rate than users with less than 10 points;
[0045] Furthermore, through creative work, it was found that data features were seriously missing in transaction scenarios.
[0046] In the actual user identity prediction transaction scenario, due to objective reasons such as user authorization, a large number of key factors are missing. For example, in the driver identity prediction transaction, the trajectory data with users only accounts for a small part of the target customer group, and even users with more than 30 points only account for 14%. In the wholesaler identity restoration, the user's business scope and industry coverage are all below 50%, and only 3% have sales product information.
[0047] The existing solutions do not have a modeling method for these influencing factors. Therefore, one or more embodiments of this specification propose a large transaction processing model based on the idea of generative adversarial learning, namely GanLLM, to model these influencing factors, thereby improving the model recognition performance.
[0048] The present specification is described in detail below with reference to specific embodiments.
[0049] See also Figure 1 , is a scenario diagram of a transaction processing system provided in this specification. Figure 1 As shown, the transaction processing system may include at least a client cluster and a service platform 100 .
[0050] The client cluster may include at least one client, such as Figure 1 As shown, it specifically includes client 1 corresponding to user 1, client 2 corresponding to user 2, ..., client n corresponding to user n, where n is an integer greater than 0.
[0051] Each client in the client cluster may be an electronic device with communication function, including but not limited to: wearable device, handheld device, personal computer, tablet computer, vehicle-mounted device, smart phone, computing device or other processing device connected to wireless modem, etc. Electronic devices may be called different names in different networks, such as: user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, electronic device, wireless communication device, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), electronic device in 5G network or future evolution network, etc.
[0052] The service platform 100 can be a separate server device, such as a rack-mounted, blade, tower, or cabinet-mounted server device, or a workstation, mainframe computer, or other hardware device with strong computing power; it can also be a server cluster composed of multiple servers, and the servers in the service cluster can be composed in a symmetrical manner, wherein each server has equivalent functions and status in the transaction link, and each server can provide services to the outside independently, and the independent service can be understood as not requiring the assistance of other servers.
[0053] In one or more embodiments of the present specification, the service platform 100 may establish a communication connection with at least one client in the client cluster, and complete data interaction during the transaction processing based on the communication connection, such as online transaction data interaction. For example, the service platform 100 may implement service recommendations to the client based on the transaction processing model obtained by the transaction processing method of the present specification;
[0054] It should be noted that the service platform 100 and at least one client in the client cluster establish a communication connection through a network for interactive communication, wherein the network can be a wireless network or a wired network, the wireless network includes but is not limited to a cellular network, a wireless local area network, an infrared network or a Bluetooth network, and the wired network includes but is not limited to Ethernet, a universal serial bus (USB) or a controller local area network. In one or more embodiments of the specification, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent data (such as a target compressed package) exchanged through a network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can also be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.
[0055] The transaction processing system embodiment provided in this specification and the transaction processing method in one or more embodiments belong to the same concept. The execution subject corresponding to the transaction processing method involved in one or more embodiments of the specification can be the above-mentioned service platform 100; the execution subject corresponding to the transaction processing method involved in one or more embodiments of the specification can also be the electronic device corresponding to the client, which is determined based on the actual application environment. The embodiment of the transaction processing system and its implementation process can be detailed in the following method embodiment, which will not be repeated here.
[0056] based on Figure 1 The scenario diagram is shown, and the transaction processing method provided by one or more embodiments of this specification is introduced in detail below.
[0057] See also Figure 2 , a flowchart of a transaction processing method is provided for one or more embodiments of the present specification, the method can be implemented by a computer program and can be run on a transaction processing device based on the von Neumann system. The computer program can be integrated into an application or run as an independent tool application. The transaction processing device can be an electronic device.
[0058] Specifically, the transaction processing method includes:
[0059] S102: Obtain a large model processing capability impact factor for a target transaction, and perform classification task modeling processing on the large model processing capability impact factor to obtain a task classifier;
[0060] Target transaction: It can be understood as a specific task in a user transaction scenario. Target transactions can be financial transactions, shopping transactions, instant messaging transactions, content recommendation transactions, and other transactions.
[0061] Large model processing capability influencing factors: refers to the influencing factor characteristics corresponding to the key factors that affect the processing performance of the large model on the target transaction, such as the feature coverage of the data, feature quality, sample distribution and other influencing factors. These large model processing capability influencing factors affect the reasoning effect of the transaction processing large model applied under the target transaction. The large model processing capability influencing factors can be the factors that affect the prediction effect of the large model. Taking the feature coverage rate as an example, it is found that when the feature coverage rate is generally lower than a certain proportion, the prediction effect of the large model is poor, and vice versa. In this case, the feature coverage rate can be set as the large model processing capability influencing factor of the transaction processing large model.
[0062] Classification task modeling: Classify and encode feature vectors of "big model processing power influence factors" according to their different states (such as high or low), and use machine learning models to build task classifiers. Classifier tasks: Each task classifier predicts an "influencing factor" (IF) that affects the prediction effect of the big model.
[0063] Task classifier: A model specifically used to classify influencing factors (such as an FNN classifier), which determines the category of influencing factors by learning the characteristics of sample data. Optionally, the task classifier includes a multi-layer neural network and at least one softmax layer. For example, a multi-layer neural network and a softmax layer can be used to form a task classifier.
[0064] Indicatively, factors that may affect the performance of the trained transaction processing large model are mined in the target transaction, and based on the factors, the classifier task of the large model processing capacity influencing factor is defined, and then the task classifier corresponding to the classifier task is constructed. The output of the task classifier is the classification result corresponding to the "large model processing capacity influencing factor", and the classification result can usually be an influencing factor encoding feature vector. The classifier task corresponding to the large model processing capacity influencing factor can guide the feature vector encoding of the large model processing capacity influencing factor dimension of the model input data during the large model processing process;
[0065] In a feasible implementation, the factors affecting the processing capability of large models can be quickly defined manually by expert-side data mining to define some clear factors affecting the processing capability of large models (such as feature coverage, number of trajectory points, etc.) based on transaction knowledge.
[0066] In a feasible implementation, the factors affecting the processing capacity of large models can adopt a data-driven factor mining method, which automatically finds the factors affecting the processing capacity of large models that may affect the performance of the model by analyzing the contribution of each feature in the transaction data.
[0067] For example, use an importance evaluation model (such as decision tree, XGBoost, etc.) to calculate the importance parameter of each feature. Sort the features according to the importance parameter, and select the factors that may affect the large model processing capability.
[0068] S104: creating an initial transaction processing large model for the target transaction based on the large model base network and the task classifier, and determining transaction training data for the initial transaction processing large model;
[0069] Large model base network: a core machine learning network used to process input data, a pre-trained large model. In the field of natural language processing, the large model base network is a pre-trained large language model upgraded based on the transformer architecture (trained on a large amount of public data and has a strong understanding of text). The number of parameters can usually reach hundreds of millions. Large model base networks include the Tongyi Qianwen large model, the Vicuna-7B model, the ChatGPT large model, the Wenxin Yiyan large model, Bert, GPT, LLAMA, etc.
[0070] Initial transaction processing big model: A big model built at least based on a big model base network and a task classifier, serving as a neural network for transaction processing tasks.
[0071] Optionally, the initial transaction processing big model may include a first task classifier, a second task classifier and a big model base network, wherein the first task classifier is obtained by performing classification task modeling processing on the target scenario task corresponding to the target transaction, and the second task classifier is obtained by performing classification task modeling processing on the big model processing capability influencing factor of the target transaction, and the target scenario task corresponding to the first task classifier can be understood as a downstream transaction processing task;
[0072] Optionally, the initial transaction processing big model may include a task classifier and a big model base network. In some embodiments, considering the different scenario processing tasks of downstream transaction processing tasks, the initial transaction processing big model may include a target transaction processing scenario adaptation network.
[0073] Transaction training data: A dedicated model training data set used in advance to optimize the initial transaction processing large model. The transaction training data carries the classifier labels of the factors that affect the processing capability of the large model for each task classifier.
[0074] The "large model processing capability influencing factor" can be any factor found that affects the prediction effect of the large model. Taking feature coverage as an example, it is found that for transaction training data, the feature coverage is generally lower than 20%, and the model prediction effect is poor, otherwise it is better. In order to enable the task classifier to guide the large model to accurately encode the "large model processing capability influencing factor", the classification task target (IF) of a certain task classifier can be set as the feature coverage. The transaction training data with a feature coverage rate lower than 20% has a corresponding label of 0 in the classification coding feature vector of the classifier, and the samples with a feature coverage rate higher than 20% have a corresponding label of 1 in the classification coding feature vector of the classifier. For each "large model processing capability influencing factor", refer to the above method to set the classifier label of the task classifier for the transaction training data one by one.
[0075] S106: Input the transaction training data into the initial transaction processing large model for at least one round of model training. During the model training process, the classification coding boundary is determined by the task classifier, and the sample transaction semantic vector of the large model base network for the transaction training data is obtained. The model parameters of the large model base network are adjusted based on the classification coding boundary and the sample transaction semantic vector to obtain the transaction processing large model.
[0076] Classification encoding boundary: The classification rules learned by the task classifier based on the "large model processing capability influencing factor" are used to distinguish the semantic vectors of samples of different categories.
[0077] Example: Taking feature coverage as an example, the task classifier can divide the sample transaction semantic vector into "high coverage area" and "low coverage area" by classifying the coding boundary to achieve the coding feature coverage classification feature;
[0078] Sample transaction semantic vector: A high-dimensional vector representation generated by the large model base network processing the transaction training data, which represents the semantic features of the transaction training data.
[0079] During the model forward propagation training process, the transaction training data is input into the initial transaction processing large model. The transaction training data is passed through the large model base network to generate a high-dimensional semantic vector V LLM , that is, the sample transaction semantic vector, which is used as the input of the task classifier (FNN) corresponding to each "large model processing capability influencing factor" to classify the "large model processing capability influencing factor" and obtain the classification result;
[0080] During the reverse training of the model, the model parameters (such as model weight parameters) of the task classifier are adjusted based on the classification results and classifier labels of the corresponding task classifier until the task classifier is trained; and the classification coding boundary is determined based on the trained task classifier, and the error loss is determined based on the classification coding boundary and the sample transaction semantic vector. The model weight parameters of the large model base network are optimized through back propagation through the error loss, so that the distribution of the semantic vector output by it is adjusted to meet the classification coding boundary, until the model training is completed and the transaction processing large model is obtained.
[0081] The optimization goal of the large model base network can be understood as adjusting the semantic vectors of low-coverage samples (transaction training data) so that they are gradually distributed closer to the high-coverage area of the classifier, and ultimately making the semantic vectors of all samples as consistent as possible with the classification encoding boundary of the classifier.
[0082] In one or more embodiments of the present specification, the electronic device determines the influencing factors of the large model processing capability and constructs a task classifier in a generative adversarial manner to accurately identify the key performance bottlenecks of the transaction processing model under the target transaction. During the model training process, the task classifier is combined with the large model base network to establish an initial transaction processing model and determine targeted training data. The classification coding boundary determined by the task classifier after training is used as a guide. Based on the classification coding boundary, the semantic vector distribution of the large model base network is optimized to gradually bring low-quality feature samples closer to high-quality feature samples, effectively avoiding the limitation that the performance of the large model is degraded due to feature sparsity and quality imbalance in actual transaction scenarios. The overall transaction processing process not only improves the robustness of the large model to sparse and low-quality data under specific target transactions, but also explicitly controls the semantic vector representation tendency of the large model, so that the large model moves closer to the feature encoding in the direction of good influencing factors, thereby improving the feature representation capability of the large model according to actual needs, thereby improving the performance of the large model in downstream target transaction processing tasks, meeting the diverse needs in complex scenarios, and significantly enhancing the generalization ability and accuracy of the large model in complex transaction scenarios.
[0083] See also Figure 3 , Figure 3 This is a flowchart of another embodiment of a transaction processing method proposed in one or more embodiments of this specification. Specifically:
[0084] S202: Obtain a large model processing capability impact factor for a target transaction, and perform classification task modeling processing on the large model processing capability impact factor to obtain a task classifier;
[0085] Target business: large model task scenarios that need to be optimized, such as text classification, user identity recognition, and recommendation systems.
[0086] Factors affecting large model processing capabilities: key features that affect the performance of large models, such as feature coverage, data quality, and data distribution deviation.
[0087] Task classifier: A small neural network (such as FNN) used to determine the factor category to which the sample belongs.
[0088] For details, please refer to the method steps in other embodiments of this specification, which will not be repeated here.
[0089] S204: creating an initial transaction processing large model for the target transaction based on the large model base network and the task classifier, and determining transaction training data for the initial transaction processing large model;
[0090] Optionally, the initial transaction processing big model may include a first task classifier (a classifier for downstream transaction processing tasks), a second task classifier (a task classifier focusing on factors affecting the processing capability of the big model) and a big model base network, wherein the first task classifier is obtained by performing classification task modeling processing on the target scenario tasks corresponding to the target transaction, and the second task classifier is obtained by performing classification task modeling processing on the factors affecting the processing capability of the big model of the target transaction. The target scenario tasks corresponding to the first task classifier may be understood as downstream transaction processing tasks;
[0091] Optionally, the initial transaction processing big model may include a task classifier and a big model base network. In some embodiments, considering the different scenario processing tasks of downstream transaction processing tasks, the initial transaction processing big model may include a target transaction processing scenario adaptation network.
[0092] S206: Input the transaction training data into the initial transaction processing large model to perform at least one round of model training.
[0093] The model training process includes a first model training process and a second model training process; Figure 4 As shown, Figure 4 This is a schematic diagram of a model training scenario for a large transaction processing model. Figure 4 In the first model training process, Figure 4 Step 1: large model multi-task learning stage, the second model training process is Step 2: freeze the influencing factor tasks and retrain the model. Figure 4 The initial transaction processing model in the figure shows part of the neural network structure. Figure 4 The neural network structure consists of a large model base network ( Figure 4 LLM) and several task classifiers ( Figure 4 The classifier tasks corresponding to each task classifier group are at least not all the same. Classifier tasks: Each classifier predicts a factor that affects the prediction effect of the large model.
[0094] S208: In the first model training process, the transaction training data is encoded by the large model base network to obtain a first sample transaction semantic vector, and the task classifier is trained on a classification coding task for a classification coding boundary based on the first sample transaction semantic vector and a first classifier label carried by the transaction training data, until the task classifier completes training;
[0095] The first sample transaction semantic vector: a high-dimensional vector generated by the large model base on the training data, representing the semantic features of the sample.
[0096] Classification encoding boundary: The classification rules learned by the task classifier are used to distinguish samples from different categories of "large model processing capability influencing factors".
[0097] A series of transaction training data is obtained in advance, including training features (a series of texts) and the first classification label of each corresponding task classifier.
[0098] In the first model training process of training, transaction training data is input, and the large model is trained using the transaction training data. When the gradient is returned, the weight update of the neurons (feature detector) of the large model base network is turned off, and only the weight update of the neurons of all classifiers is turned on (such as Figure 4 As shown in the left figure). At this time, the large model with weight update turned off (frozen) is equivalent to a feature encoder that does not update. The encoding vector (embedding) of the training feature is output through the general knowledge of the large model, that is, the first sample transaction semantic vector, while the task classifiers with weight update turned on can learn the classification coding boundary (decision boundary) during the first model training process. Each task classifier can learn different classification coding boundaries under different influencing factors. Figure 5 As shown, Figure 5 is a scene diagram of a classification coding boundary. The classification coding boundary learned in the first model training process is as follows Figure 5 The subgraph corresponding to step 1 in .
[0099] Furthermore, in the first model training process, the task classifier FNN learns how to distinguish the categories of input samples on a certain influencing factor (large model processing capability influencing factor) (e.g., low feature coverage vs. high feature coverage) based on the first sample transaction semantic vector representation provided by the large model base network to generate a binary classification result (the binary classification result is Class 0 or Class 1). The first sample transaction semantic vector and the first classifier label carried by the transaction training data are back-propagated to optimize the weight parameters of the task classifier to gradually adjust the learned classification boundaries, such as Figure 5 The sub-figure corresponding to step 1 in is shown to enable the use of classification encoding boundaries ( Figure 5 The task classifier accurately classifies the task within the boundary shown in the figure until the task classifier reaches the end training condition of the first model and completes the training; the FNN feeds back the learned classification coding boundary to the entire training process (especially in the second model training process), helping the large model base network to update its coding transaction semantic vector and optimize its adaptability to low feature quality samples.
[0100] Optionally, the first model training end condition may include, for example, the value of the loss function is less than or equal to a preset loss function threshold, the number of iterations reaches a preset number threshold, etc. The specific model training end condition may be determined based on actual conditions and is not specifically limited here.
[0101] exist Figure 5 In the sub-graph corresponding to step 1: the black dots and asterisks represent the classification results of the two types of semantic vectors output by the task classifier for the transaction semantic vector: Class 0 (e.g., low coverage or low-quality feature samples) and Class 1 (e.g., high coverage or high-quality feature samples). The red solid line is the classification coding boundary (Boundary) learned by the task classifier (FNN) during the model training process, which is used to distinguish Class 0 from Class 1. In S208, the first sample transaction semantic vector is generated by inputting the transaction training data into the large model base network. The task classifier uses these first sample transaction semantic vectors and the first classifier label of the sample (Class 0 or Class 1) for classification task training, and the classifier gradually learns the classification coding boundary that separates Class 0 and Class 1 samples. In Step 1, the task classifier successfully learns the classification coding boundary, but there is still a significant separation in the distribution of semantic vectors of Class 0 and Class 1. The semantic vector of the low-quality sample (Class 0) is close to the classification boundary and is easily misclassified.
[0102] Exemplarily, the network structure of the task classifier may include an input layer, a hidden layer, and an output layer. The input layer receives the high-dimensional feature vector output from the large model base (LLM). The hidden layer includes several fully connected layers, each of which is followed by an activation function (such as ReLU, Tanh, etc.). The output layer outputs the classification result through the Softmax or Sigmoid function, which is specifically designed according to the classification task (such as binary classification or multi-classification) that affects the processing power of the large model.
[0103] like Figure 4 As shown, each task classifier is used as an independent classifier, and its role can be summarized as follows:
[0104] Classification task: For a specific large model processing capacity influencing factor (IF), such as feature coverage, FNN determines whether the input sample belongs to low coverage (0) or high coverage (1) to generate the first classification result.
[0105] Multi-task learning: Each FNN is a subtask, and multiple FNNs jointly learn the feature boundaries of multiple influencing factors.
[0106] Phased optimization:
[0107] In the first model training process, the FNN learns the classification encoding boundary and the large model base network is frozen.
[0108] In the second model training process, the large model base adjusts the output vector to adapt to the classification encoding boundary of the FNN.
[0109] S210: In the second model training process, based on the task classifier, the transaction training data is label-modified to obtain reference transaction training data carrying the second classifier label, the reference transaction training data is encoded through the large model base network to obtain a second sample transaction semantic vector, based on the second sample transaction semantic vector, the task classifier and the classification encoding boundary, the network gradient adjustment information is determined, and the model weight of the large model base network is adjusted based on the network gradient adjustment information, until the large model base network completes the training, and the transaction processing large model is obtained;
[0110] Label modification: After the second model training process is completed, the classification results output by the task classifier for the transaction training data are used as the new training target to generate the second classifier label. At this time, the first classification label of the transaction training data is reset by the trained task classifier at this stage according to the learned classification coding boundary, and the labels of all influencing factor classifiers (FNN) are all set to 1. For example, the classification result of the task classifier is used to replace the original first classification label to generate the second classifier label, indicating that the second-stage optimization goal is to bring all samples (transaction training data) "close" to the distribution of high-quality feature samples.
[0111] Reference transaction training data: training data set after modifying the classifier labels, used to optimize the large model base network.
[0112] Network gradient adjustment information: Gradients generated based on classification errors are used to optimize the base weights of the large model.
[0113] Second sample transaction semantic vector: a sample transaction semantic vector generated by inputting reference transaction training data into the large model base network;
[0114] During the second model training process: optimize the weights of the large model base network (LLM) so that the encoding vector it outputs can adapt to the impact factor classification boundary determined by the FNN classifier in the first stage. The following is an explanation of the second model training process:
[0115] Input data: Use the same transaction training data (text or other features) as in the first stage. The first classification label of the transaction training data is reset by the trained task classifier in this stage according to the learned classification coding boundary. The labels of all influencing factor classifiers (FNN) are all configured to 1, indicating that the second-stage optimization goal is to move all samples (transaction training data) "close" to the distribution of high-quality feature samples.
[0116] Fix the classifier during backpropagation training and unfreeze the large model
[0117] Classifier freezing: In this stage, the weights of all FNN classifiers are frozen and the classification encoding boundaries remain unchanged. The frozen classifier uses the classification encoding boundaries as fixed rules to evaluate the transaction semantic vectors output by the large model base network.
[0118] Large model unfreezing: Unfreeze the weights of the large model base network so that it can adjust its output vector according to the feedback of the task classifier during back propagation.
[0119] The second model training process is as follows:
[0120] 1. After the first model training process is completed, the transaction training data is labeled based on the task classifier to obtain reference transaction training data with the second classifier label
[0121] 2. The reference transaction training data (such as text, features) is input into the large model base network, and the large model base processes the data and outputs the second sample transaction semantic vector;
[0122] 3. Determine the network gradient adjustment information based on the second sample transaction semantic vector, the task classifier and the classification coding boundary, and adjust the model weight of the large model base network based on the network gradient adjustment information until the large model base network completes training to obtain a transaction processing large model.
[0123] The second sample transaction semantic vector output by the large model base network is input into the frozen FNN classifier. The FNN classifier evaluates whether the second sample transaction semantic vector falls into the correct category area according to the classification coding boundary determined in the first stage. If the second sample transaction semantic vector cannot pass the boundary of the task classifier (such as being judged as a low-quality feature category), an error signal is generated during the back propagation process. The weights of the large model base network are updated according to the error signal, so that its output vector adjusts its direction and gradually approaches the distribution of the target category (high-quality feature category) of the FNN classifier. The goal is to make the semantic vector distribution of low-quality samples close to the distribution of high-quality samples, and finally achieve the distribution fusion of the two.
[0124] like Figure 5 As shown, Figure 5The middle and right sub-figures in the figure describe the second model training process, Step 2: Optimize the large model base network, in Figure 5 In the middle sub-graph and the right sub-graph in , the sample distribution of the black dots (Class 0) begins to gradually approach the distribution of Class 1. Through the optimization of the large model base network, some Class 0 samples have successfully crossed the classification boundary. The corresponding model training process: In S210, the task classifier corrects the labels of the transaction training data to generate reference training data carrying the second classifier label. The reference transaction training data is input into the large model base network to generate the second sample transaction semantic vector. The gradient adjustment error information of the large model base network is calculated by combining the task classifier and the classification encoding boundary. The large model base network is optimized through back propagation, so that the semantic vector of Class 0 samples gradually approaches the distribution of Class 1 samples. The final result is as follows: Figure 5 Middle right sub-figure: low-quality samples (Class 0) gradually move closer to high-quality samples (Class 1), and the semantic vectors output by the large model base network are more concentrated, weakening the impact of feature quality and coverage. Figure 5 In the middle right sub-figure: the sample distributions of black dots and asterisks are highly overlapped, the sample distributions of Class 0 and Class 1 are no longer significantly separated, the classification boundary is still retained, but most samples have been correctly classified. The corresponding process: At the end of the S210 model training, the weights of the large model base network have fully adapted to the classification boundary of the classifier after multiple rounds of training. The semantic vectors of low-quality samples successfully cross the classification boundary and are close to the distribution of high-quality samples. The optimized large model base network can generate more stable and robust semantic vectors in scenarios with sparse features or low-quality data, and obtain a trained large model base network;
[0125] The effect of model training application in actual scenarios: Through the learning of classification boundaries and the optimization of semantic vectors, the problem of performance degradation of large models in scenarios with sparse features and uneven quality is effectively solved. For example, in user identity recognition, low coverage samples (such as drivers with missing historical trajectory data) no longer significantly affect recognition accuracy. For example, in recommendation systems, sparse user data can generate more reasonable recommendation results through optimized semantic vectors.
[0126] In one or more embodiments of this specification, the above method is used to gradually identify influencing factors, train task classifiers, and then optimize the large model base network, and finally generate a transaction processing large model that adapts to the target transaction. Each step combines the powerful semantic representation ability of the large model and the targeted rules of the classifier, significantly improving the robustness and accuracy of the model for sparse and low-quality feature data.
[0127] Optional, see Figure 6 , Figure 6This is a flowchart of a task classifier training process proposed in one or more embodiments of this specification. Specifically, the classification coding task training of the task classifier for the classification coding boundary based on the first sample transaction semantic vector and the first classifier label carried by the transaction training data, until the task classifier is trained, until the large model base network is trained, can be performed in the following manner:
[0128] S3002: Inputting the first sample transaction semantic vector into the task classifier to obtain a first classification result, and determining classifier gradient adjustment information based on the first classifier label and the first classification result;
[0129] First classification result: The category prediction (such as 0 or 1) output by the task classifier based on the semantic vector of the first sample transaction. Example: For the feature coverage task classifier, the output may be: 0: low feature coverage. 1: high feature coverage.
[0130] First classifier label: the true category label of the sample, manually annotated based on the impact factor.
[0131] Example: The label of a driver whose sample trajectory points are less than 20 is 0, and the label of a driver whose trajectory points are greater than or equal to 20 is 1.
[0132] Classifier gradient adjustment information: The gradient generated by back-propagation of the classification error of the task classifier is used to update the classifier weights.
[0133] Schematically, the first sample transaction semantic vector is input into the task classifier, and the task classifier outputs the predicted category probability distribution, i.e., the first classification result, through the fully connected layer and the activation function; the classifier output result is compared with the true label to calculate the classification error. In a feasible implementation, the classifier gradient adjustment information is determined based on the first classifier label and the first classification result, including:
[0134] A loss calculation is performed based on the first classifier label and the first classification result to obtain a classifier processing loss, and classifier gradient adjustment information is generated based on the classifier processing loss.
[0135] For example, a first loss function may be used, the first classifier label and the first classification result may be input into the first loss function, the first loss function may be used to calculate the classification error, and based on the classification error, the gradient adjustment information of the task classifier weight may be calculated through back propagation.
[0136] Optionally, the first loss function may be a cross entropy loss function, a Euclidean distance loss function, a hinge loss function, etc.
[0137] S3004: Freeze the network weight parameters of the large model base network, and adjust the classifier weight of the task classifier based on the classifier gradient adjustment information to update the classification coding boundary of the task classifier through the classifier weight adjustment until the task classifier completes training.
[0138] Freeze the large model base network: The weight parameters of the large model base remain unchanged, and it only generates semantic vectors as a feature extractor. This ensures that the output features of the large model base network are not changed when the task classifier is trained, and focuses on optimizing the classifier.
[0139] Classifier weight adjustment: Update the weight of the task classifier through classifier gradient adjustment information to gradually optimize its classification boundary.
[0140] Classification encoding boundary: The sample category segmentation rule learned by the classifier is used to distinguish low-quality and high-quality samples.
[0141] Indicatively, the gradient adjustment information is used to update the classifier weights through an optimization algorithm (such as SGD, Adam). The semantic vector is repeatedly input, the classification error is calculated and the classifier weights are adjusted, and the classification boundary is gradually optimized until the task classifier reaches the first end training condition to complete the training.
[0142] In one or more embodiments of the present specification, the classifier (task classifier) focuses on optimizing its own classification boundaries while freezing the weights of the large model base network. The semantic vector generated by the large model is input into the classifier, and the gradient adjustment information is generated by calculating the classification error to clarify the optimization direction of the classifier; and the weights of the classifier are iteratively optimized using the gradient adjustment information, and the classification boundaries are gradually adjusted until the classifier can accurately distinguish between sample categories. The entire process decouples feature extraction and classification tasks, ensuring that the classifier can clearly learn the classification rules of the influencing factors, laying the foundation for the subsequent weight adjustment of the large model base network.
[0143] Optional, see Figure 7 , Figure 7 The flowchart of determining a network gradient adjustment error proposed in one or more embodiments of the present specification is as follows. Specifically, the method of determining the network gradient adjustment information based on the second sample transaction semantic vector, the task classifier and the classification coding boundary includes:
[0144] S4002: inputting the second sample transaction semantic vector into the task classifier, and determining a second classification result through the classification coding boundary of the task classifier;
[0145] Classification encoding boundary of task classifier: The task classifier classifies the input semantic vector according to the classification boundary learned during the first stage model training process (for example, the boundary between high and low feature coverage).
[0146] Example: If the classifier boundary is based on feature coverage (less than 20% is Class 0, more than 20% is Class 1), the classifier judges the input semantic vector based on this boundary.
[0147] Second classification result: The task classifier outputs a classification prediction result, usually a category label (such as 0 or 1), based on the second sample transaction semantic vector and the learned classification boundary.
[0148] Example: If the feature coverage represented by the second sample transaction semantic vector is higher than 20%, the task classifier outputs Class 1.
[0149] Schematically, the semantic vector of the second sample transaction is input into the task classifier, and the task classifier determines which category the semantic vector belongs to based on its internal classification coding boundary, and outputs a second classification result, which is a predicted category based on the classification coding boundary.
[0150] S4004: Determine network gradient adjustment information based on the second classification result and the second classifier label;
[0151] Network gradient adjustment information: By calculating the difference between the task classifier output result (second classification result) and the true label (second classifier label), the error information is back-propagated, and the weights of the large model are updated based on this error information. The error is calculated using a loss function (such as the cross entropy loss function), and then the gradient is calculated based on the error through the back-propagation algorithm. The calculated gradient information is used to update the weights of the large model base network.
[0152] In a feasible implementation manner, determining the network gradient adjustment information based on the second classification result and the second classifier label includes:
[0153] A loss is calculated based on the second classification result and the second classifier label to obtain a network processing loss, and network gradient adjustment information is generated based on the network processing loss.
[0154] For example, a second loss function can be used, the second classifier label and the second classification result can be input into the second loss function, the second loss function can be used to calculate the network processing loss, and the network gradient adjustment information can be calculated through back propagation based on the network processing loss.
[0155] Optionally, the second loss function may be a cross entropy loss function, a Euclidean distance loss function, a hinge loss function, etc.
[0156] S4006: Unfreeze the network weight parameters of the large model base network and freeze the classifier weight parameters of the task classifier, and adjust the network weight parameters of the large model base network based on the network gradient adjustment information until the large model base network completes training.
[0157] Unfreezing the network weight parameters of the large model base network can be understood as: restoring the previously frozen weight parameters of the large model base network to a trainable state, which means that the neurons of the large model base network (such as the parameters of the Transformer layer) will start to participate in training in this step and be updated according to the gradient adjustment information of the task classifier.
[0158] Freeze the classifier weight parameters of the task classifier: Freezing is to fix the weight of the task classifier and no longer update it. The weight of the task classifier will not change at this stage. This ensures that the classifier has learned the appropriate classification boundary and will not be destroyed during the training of the large model base network.
[0159] Network weight parameter adjustment: By calculating the gradient and loss function, the weight parameters in the large model base network are adjusted to reduce the encoding error, so that the large model base network can better meet the needs of the task classifier. Network weight parameter adjustment usually uses optimization algorithms such as gradient descent (SGD), Adam or other variants to perform weight updates.
[0160] Schematically, during the training of the first model, the weights of the large model base network are frozen, and only feature extraction is performed. In S4006, - the parameters of the large model base network are unfrozen so that it can participate in training and accept network gradient adjustment information. And the weights of the task classifier are frozen, the weights of the task classifier are fixed, and the classification boundary of the task classifier no longer changes. This ensures that the training of the large model base network will not affect what the classifier has learned; further, the weights of the large model base network are updated using the calculated network gradient adjustment information. The goal is to allow the semantic vector of the large model base network to better pass the classification boundary of the task classifier. In other words, the large model base network needs to "adapt" to the classification rules of the task classifier to ensure that the feature vectors it generates can be accurately classified; repeat this process until the large model base network meets the second model training end conditions to complete the training, the weight parameters of the large model base network converge, the error of the classifier is minimized, and the large model base network can generate semantic vectors that better meet the requirements of the classifier.
[0161] Optionally, the second model training end condition may include, for example, the value of the loss function is less than or equal to a preset loss function threshold, the number of iterations reaches a preset number threshold, etc. The specific model training end condition can be determined based on actual conditions and is not specifically limited here.
[0162] In this specification, through the above method, the task classifier and the large model base network achieve efficient collaborative training. The task classifier helps the large model adjust its generated feature representation by classifying semantic vectors, ensuring that the large model can adapt to the requirements of different feature categories. By unfreezing the weights of the large model and freezing the weights of the classifier, the entire training process effectively separates the training tasks of the classifier and the large model, ensuring the stability of the classification boundary, and optimizing the feature learning ability of the large model through gradient information. This process significantly improves the adaptability and classification accuracy of the large model on specific tasks.
[0163] Optional, see Figure 8 , Figure 8 This is a flow chart of determining a transaction processing large model proposed in one or more embodiments of this specification. Specifically, the model parameters of the large model base network are adjusted based on the classification coding boundary and the sample transaction semantic vector to obtain the transaction processing large model, which can be referred to in the following manner:
[0164] S5002: Adjusting the model parameters of the large model base network based on the classification coding boundary and the sample transaction semantic vector to obtain the trained large model base network and the task classifier;
[0165] For details, please refer to the method steps in other embodiments of this specification, which will not be repeated here;
[0166] In this specification, the initial transaction processing big model may include a task classifier and a big model base network. In some embodiments, considering the different scenario processing tasks of downstream transaction processing tasks, the initial transaction processing big model may include a scenario adaptation network corresponding to the target transaction processing.
[0167] S5004: Determine the target scenario task corresponding to the target transaction, and use the transaction training data to perform target scenario task adaptation processing on the initial transaction processing big model to obtain the target scenario task adapted transaction processing big model.
[0168] Target scenario task: The specific task faced by the transaction processing model in actual application according to the task requirements under the specific target transaction. For example, in the driver identification scenario, the target scenario task is to identify whether the driver's identity matches the preset conditions.
[0169] Target scenario task adaptation processing: Since the trained large model base network and task classifier are obtained through model training, the transaction processing large model including the large model base network and task classifier has the ability to weaken / eliminate the various influencing factors mentioned in the classifier, and can adapt to any downstream classification / generation tasks and apply to any large model scenario. Then, in order to better handle the task requirements of the target transaction, scenario adaptation is required, that is, the transaction processing large model can be adapted to the target scenario task. The target scenario task can be understood as the downstream transaction processing task. The target scenario task refers to the optimization of specific transactions in actual applications, so that the transaction processing large model can accurately handle the task requirements related to the transaction. For example, the target transaction may involve scenarios such as driver identity identification and wholesaler identity authentication. These tasks require the model to make accurate predictions based on specific data features.
[0170] Through the adaptation of target scenario tasks, the transaction processing big model's capabilities are further enhanced, enabling it to provide customized processing effects for different downstream tasks and scenarios. This process ensures the flexibility and efficiency of the big model in actual transaction applications.
[0171] In an illustrative manner, the target transaction and the corresponding target scenario task are first determined. Then, the initial transaction processing model is fine-tuned using the transaction training data to make it better adaptable to the target scenario task. The adaptation process ensures that the large model can generate more accurate prediction and classification results for specific tasks.
[0172] In a feasible implementation manner, the using of the transaction training data to perform target scenario task adaptation processing on the initial transaction processing big model to obtain the target scenario task adapted transaction processing big model includes:
[0173] The task classifier in the initial transaction processing large model is filtered out, and the initial transaction processing large model is adapted to the target scenario task using the transaction training data to obtain the transaction processing large model adapted to the target scenario task.
[0174] The filtered task classifier is a component used to learn and identify various factors that affect the performance of the large model during the early training process. As the training of the large model base network is completed, the main function of the task classifier is no longer necessary, and filtering can be considered when the model is applied to specific target scenarios in the future.
[0175] Based on this, the task classifier can be filtered out, which means removing the task classifier from the initial transaction processing model and at least retaining the large model base network, so that the large model can focus on adapting to specific target transaction tasks. After filtering out the classifier, the remaining model is the model that at least contains the large model base network, which already has the ability to adapt to different characteristics and influencing factors.
[0176] After filtering out the task classifiers, the initial transaction processing model is further trained using the transaction training data to adapt it to the needs of the target scenario tasks. The target scenario tasks usually refer to the tasks that the model needs to perform in actual applications (such as driver identification, wholesaler identity verification, etc.).
[0177] By training the model with specific transaction training data, the model parameters are further optimized so that it can make accurate predictions for specific data features in specific scenarios. For example, if the target scenario task is driver identification, then the transaction training data may contain the driver's historical trajectory information, POI data, etc., and the large model is optimized through these data to better identify the driver's identity.
[0178] After the adaptation process, the output of the model will be more in line with the needs of the target transaction and can accurately complete the target task. At this point, the large transaction processing model adapted to the target scenario task already has the ability to process transactions in a specific scenario.
[0179] For example, after adapting to the target scenario task, the large model can determine the driver's identity based on his trajectory data, or verify his identity through the wholesaler's business information and product sales data.
[0180] In this specification, by filtering out the task classifier and using transaction training data to adapt the target scenario task, the originally more general large model can be adjusted to a transaction processing model specifically for a certain target scenario. This adaptation process enables the large model to not only handle general tasks, but also accurately complete tasks in specific fields or transaction scenarios, thereby improving the practicality and accuracy of the large model.
[0181] Optional, see Fig. 9 , Fig. 9 This is a flow chart of a model parameter adjustment proposed in one or more embodiments of this specification. Specifically, the model parameter adjustment of the large model base network based on the classification coding boundary and the sample transaction semantic vector is performed to obtain a transaction processing large model, which can be referred to in the following manner:
[0182] S6002: Obtain the large model processing capability impact factor for the target transaction, determine the target scenario task corresponding to the target transaction, perform classification task modeling processing on the target scenario task to obtain a first task classifier, obtain the large model processing capability impact factor for the target transaction, perform classification task modeling processing on the large model processing capability impact factor to obtain a second task classifier, and determine a task classifier based on the first task classifier and the second task classifier;
[0183] In this specification, the initial transaction processing big model may include a first task classifier, a second task classifier and a big model base network. The first task classifier is obtained by performing classification task modeling on the target scenario task corresponding to the target transaction based on the target scenario task. The second task classifier is obtained by performing classification task modeling on the big model processing capability influencing factor of the target transaction. The target scenario task corresponding to the first task classifier can be understood as a downstream transaction processing task. An initial transaction processing big model is created and trained through the synergy of the task classifiers combined with the big model base network, and finally adapted to the specific task requirements of the target transaction.
[0184] For example, assuming that the target task is driver identification, a first task classifier is needed, whose task is to identify the driver based on trajectory data and other features. Then, based on influencing factors (such as feature coverage), a second task classifier is created, which will optimize the performance of the large model in this scenario based on whether the feature data is complete. Finally, a comprehensive task classifier is determined by combining these two classifiers to simultaneously optimize the processing power of the large model and adapt to the needs of the target task.
[0185] In addition, in the execution of some embodiments of "in the second model training process, the transaction training data is labeled based on the task classifier to obtain reference transaction training data carrying the second classifier label", only the transaction training data is labeled based on the second task classifier to obtain the reference transaction training data carrying the second classifier label, that is, the annotated labels of the first task classifier in the transaction training data are not modified, and the original notes are retained.
[0186] S6004: Creating an initial transaction processing large model for the target transaction based on the large model base network and the task classifier, and determining transaction training data for the initial transaction processing large model;
[0187] For details, please refer to the method steps in other embodiments of this specification, which will not be repeated here;
[0188] S6006: Input the transaction training data into the initial transaction processing large model for at least one round of model training. During the model training process, the classification coding boundary is determined by the task classifier to obtain the sample transaction semantic vector of the large model base network for the transaction training data;
[0189] For details, please refer to the method steps in other embodiments of this specification, which will not be repeated here;
[0190] S6008: Adjusting the model parameters of the large model base network based on the classification coding boundary and the sample transaction semantic vector to obtain the trained large model base network and the task classifier;
[0191] For details, please refer to the method steps in other embodiments of this specification, which will not be repeated here;
[0192] S6010: Obtain a transaction processing large model based on the large model base network and the first task classifier.
[0193] Through the synergy of the large model base network and the task classifier, a transaction processing large model is finally obtained based on at least the large model base network and the first task classifier. That is, the second task classifier can be filtered out after the model training is completed. The transaction processing large model has the ability to handle the target scenario tasks and has optimized the adaptability of the influencing factors.
[0194] In this specification, by combining the first task classifier and the second task classifier, as well as the large model base network, a transaction processing large model is created that can simultaneously optimize task adaptability and processing capabilities. In this way, the large model can not only show stronger adaptability when processing target transactions (for example, for driver identification tasks), but also effectively cope with the challenges of influencing factors such as data quality, thereby improving the performance and accuracy of the model in practical applications.
[0195] Optional, see Fig.10 , Fig.10 This is a flowchart of a transaction processing method proposed in one or more embodiments of this specification. The method can be implemented by a computer program and can be run on a transaction processing device based on the von Neumann architecture. The computer program can be integrated into an application or run as an independent tool application. The transaction processing device can be an electronic device.
[0196] Specifically,
[0197] S7002: Obtain target transaction data for the target transaction;
[0198] Target transaction data refers to transaction processing data related to the target transaction, which usually contains feature data used for decision-making, classification or generation.
[0199] In an illustrative manner, data related to the target transaction is first obtained. For example, if the target transaction is driver identification, the target transaction data may include driver trajectory information, vehicle information, POI data, etc. The target transaction data can be collected from different data sources, including databases, real-time data streams, or manually uploaded data.
[0200] S7004: Input the target transaction data into the transaction processing model and output the transaction processing result;
[0201] In an illustrative manner, the target transaction data is input into a trained transaction processing model. The model will output transaction processing results, such as classification labels, predicted values, or generated text, based on the characteristics of the data and transaction requirements.
[0202] For example, if it is driver identification, the output transaction result can be the driver identity (such as driver A or B).
[0203] For example, if it is a wholesaler identity verification, the output transaction processing result may be whether the wholesaler meets the predetermined identity verification standard.
[0204] S7006: Perform target transaction processing based on the transaction processing result.
[0205] Target transaction processing refers to subsequent transaction decisions or process operations based on the transaction processing results. This step will further promote the execution of transaction logic, such as identity authentication, classification result application, task execution, etc., based on the output results of the big model.
[0206] Indicatively, the target transaction is processed specifically based on the transaction processing result. For example, if the target transaction is identity authentication, the output of the model may be a label of "qualified" or "unqualified", and the subsequent system will perform corresponding operations based on this label, such as approving or rejecting identity authentication, activating an account, etc.
[0207] In one or more embodiments of the present specification, the trained transaction processing model is applied to actual tasks, and can generate transaction processing results based on the input target transaction data, and process the target transaction based on these results. This process ensures that the output of the model can not only provide accurate predictions and classifications, but also be closely integrated with the subsequent transaction process, thereby achieving effective management and processing of specific transactions.
[0208] The following will be combined Fig.11 , the transaction processing device provided in this specification is introduced in detail. It should be noted that, Fig.11 The transaction processing device shown is used to execute this specification Figures 1 to 10 For the convenience of explanation, only the parts related to this specification are shown. For the specific technical details not disclosed, please refer to this specification. Figures 1 to 10 The embodiment shown.
[0209] See also Fig.11, which shows a schematic diagram of the structure of the transaction processing device of this specification. The transaction processing device 1 can be implemented as all or part of the device through software, hardware or a combination of both. According to some embodiments, the transaction processing device 1 includes a modeling module 11, a creation module 12 and a training module 13, which are specifically used to:
[0210] Modeling module 11, used to obtain the large model processing capability impact factor for the target transaction, and perform classification task modeling processing on the large model processing capability impact factor to obtain a task classifier;
[0211] A creation module 12, configured to create an initial transaction processing big model for the target transaction based on the big model base network and the task classifier, and determine transaction training data for the initial transaction processing big model;
[0212] The training module 13 is used to input the transaction training data into the initial transaction processing large model for at least one round of model training. During the model training process, the classification coding boundary is determined by the task classifier, and the sample transaction semantic vector of the large model base network for the transaction training data is obtained. The model parameters of the large model base network are adjusted based on the classification coding boundary and the sample transaction semantic vector to obtain the transaction processing large model.
[0213] Optionally, the model training process includes a first model training process and a second model training process, and the training module 13 is used to:
[0214] In the first model training process, the transaction training data is encoded by the large model base network to obtain a first sample transaction semantic vector, and the task classifier is trained on a classification coding task for a classification coding boundary based on the first sample transaction semantic vector and a first classifier label carried by the transaction training data until the task classifier completes training;
[0215] During the training process of the second model, the transaction training data is labeled modified based on the task classifier to obtain reference transaction training data carrying the second classifier label, the reference transaction training data is encoded through the large model base network to obtain a second sample transaction semantic vector, network gradient adjustment information is determined based on the second sample transaction semantic vector, the task classifier and the classification coding boundary, and model weights of the large model base network are adjusted based on the network gradient adjustment information until the large model base network completes training.
[0216] Optionally, the training module 13 is used to:
[0217] Inputting the first sample transaction semantic vector into the task classifier to obtain a first classification result, and determining classifier gradient adjustment information based on the first classifier label and the first classification result;
[0218] Freeze the network weight parameters of the large model base network, and adjust the classifier weight of the task classifier based on the classifier gradient adjustment information to update the classification encoding boundary of the task classifier through the classifier weight adjustment until the task classifier completes training.
[0219] Optionally, the training module 13 is used to:
[0220] A loss calculation is performed based on the first classifier label and the first classification result to obtain a classifier processing loss, and classifier gradient adjustment information is generated based on the classifier processing loss.
[0221] Optionally, the training module 13 is used to:
[0222] Inputting the two sample transaction semantic vectors into the task classifier, and determining a second classification result through the classification coding boundary of the task classifier;
[0223] Determine network gradient adjustment information based on the second classification result and the second classifier label;
[0224] Unfreeze the network weight parameters of the large model base network and freeze the classifier weight parameters of the task classifier, and adjust the network weight parameters of the large model base network based on the network gradient adjustment information until the large model base network completes training.
[0225] Optionally, the training module 13 is used to:
[0226] A loss is calculated based on the second classification result and the second classifier label to obtain a network processing loss, and network gradient adjustment information is generated based on the network processing loss.
[0227] Optionally, the training module 13 is used to:
[0228] Adjust the model parameters of the large model base network based on the classification coding boundary and the sample transaction semantic vector to obtain the trained large model base network and the task classifier;
[0229] Determine the target scenario task corresponding to the target transaction, use the transaction training data to perform target scenario task adaptation processing on the initial transaction processing big model, and obtain the target scenario task adapted transaction processing big model.
[0230] Optionally, the training module 13 is used to:
[0231] The task classifier in the initial transaction processing large model is filtered out, and the initial transaction processing large model is adapted to the target scenario task using the transaction training data to obtain the transaction processing large model adapted to the target scenario task.
[0232] Optionally, the modeling module 11 is used to:
[0233] Determine a target scenario task corresponding to the target transaction, perform classification task modeling processing on the target scenario task to obtain a first task classifier, obtain a large model processing capability impact factor for the target transaction, perform classification task modeling processing on the large model processing capability impact factor to obtain a second task classifier, and determine a task classifier based on the first task classifier and the second task classifier;
[0234] The step of adjusting the model parameters of the large model base network based on the classification coding boundary and the sample transaction semantic vector to obtain a large transaction processing model includes:
[0235] Adjust the model parameters of the large model base network based on the classification coding boundary and the sample transaction semantic vector to obtain the trained large model base network and the task classifier;
[0236] A transaction processing big model is obtained based on the big model base network and the first task classifier.
[0237] See also Fig.12 , which shows a schematic diagram of the structure of the transaction processing device of this specification. The transaction processing device 1 can be implemented as all or part of the device through software, hardware or a combination of both. According to some embodiments, the transaction processing device 2 includes a data acquisition module 21, a model processing module 22 and a transaction processing module 23, which are specifically used to:
[0238] A data acquisition module 21 is used to acquire target transaction data for a target transaction;
[0239] A model processing module 22, used for inputting the target transaction data into a transaction processing model and outputting a transaction processing result, wherein the transaction processing model is obtained by the method steps according to any one of claims 1 to 9;
[0240] The transaction processing module 23 is used to perform target transaction processing based on the transaction processing result.
[0241] It should be noted that the transaction processing device provided in the above embodiment only uses the division of the above functional modules as an example when executing the transaction processing method. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the transaction processing device provided in the above embodiment and the transaction processing method embodiment belong to the same concept, and the implementation process thereof is detailed in the method embodiment, which will not be repeated here.
[0242] The above serial numbers in this specification are for description only and do not represent the advantages or disadvantages of the embodiments.
[0243] The present specification also provides a computer storage medium, which can store multiple instructions, which are suitable for being loaded and executed by a processor as described above. Figures 1 to 10 The transaction processing method of the embodiment shown in the figure can be found in the specific execution process. Figures 1 to 10 The specific description of the illustrated embodiment will not be repeated here.
[0244] The present specification also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor as described above. Figures 1 to 10 The transaction processing method of the embodiment shown in the figure can be found in the specific execution process. Figures 1 to 10 The specific description of the illustrated embodiment will not be repeated here.
[0245] Please refer to Fig.13 , is a block diagram of a structure of an electronic device provided in an embodiment of this specification. The electronic device in this specification may include one or more of the following components: a processor 1010, a memory 1020, an input device 1030, an output device 1040, and a bus 1050. The processor 1010, the memory 1020, the input device 1030, and the output device 1040 may be connected via a bus 1050.
[0246] The processor 1010 may include one or more processing cores. The processor 1010 uses various interfaces and lines to connect various parts of the entire electronic device, and executes various functions of the electronic device and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 1020, and calling data stored in the memory 1020. Optionally, the processor 1010 can be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 1010 can integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 1010, but may be implemented separately through a communication chip.
[0247] The memory 1020 may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory 1020 includes a non-transitory computer-readable storage medium. The memory 1020 may be used to store instructions, programs, codes, code sets, or instruction sets.
[0248] The input device 1030 is used to receive input instructions or data, and the input device 1030 includes but is not limited to a keyboard, a mouse, a camera, a microphone, or a touch device. The output device 1040 is used to output instructions or data, and the output device 1040 includes but is not limited to a display device and a speaker. In the embodiment of this specification, the input device 1030 can be a temperature sensor for obtaining the operating temperature of the electronic device. The output device 1040 can be a speaker for outputting audio information.
[0249] In addition, those skilled in the art will appreciate that the structure of the electronic device shown in the above drawings does not constitute a limitation on the electronic device, and the electronic device may include more or fewer components than shown, or combine certain components, or arrange the components differently. For example, the electronic device also includes a radio frequency circuit, an input unit, a sensor, an audio circuit, a wireless fidelity (WIFI) module, a power supply, a Bluetooth module and other components, which will not be described in detail here.
[0250] In the embodiments of this specification, the execution subject of each step may be the electronic device described above. Optionally, the execution subject of each step is the operating system of the electronic device. The operating system may be an Android system, an IOS system, or other operating systems, which is not limited in the embodiments of this specification.
[0251] exist Fig.13 In the electronic device, the processor 1010 can be used to call the program stored in the memory 1020 and execute it to implement the transaction processing method as described in the various method embodiments of this specification.
[0252] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only storage memory, or a random access memory, etc.
[0253] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and information involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the target transaction data, transaction training data, etc. involved in this specification are all obtained with full authorization.
[0254] The above disclosure is only the preferred embodiment of this specification, which certainly cannot be used to limit the scope of rights of this specification. Therefore, equivalent changes made according to the claims of this specification are still within the scope covered by this specification.
Claims
1. A transaction processing method, the method comprising: Obtaining a large model processing capability impact factor for a target transaction, and performing classification task modeling processing on the large model processing capability impact factor to obtain a task classifier; Based on the big model base network and the task classifier, create an initial transaction processing big model for the target transaction, and determine transaction training data for the initial transaction processing big model; The transaction training data is input into the initial transaction processing large model for at least one round of model training. During the model training process, the classification coding boundary is determined by the task classifier, and the sample transaction semantic vector of the large model base network for the transaction training data is obtained. The model parameters of the large model base network are adjusted based on the classification coding boundary and the sample transaction semantic vector to obtain the transaction processing large model.
2. The method according to claim 1, wherein the model training process comprises a first model training process and a second model training process. In the model training process, the classification coding boundary is determined by the task classifier, a sample transaction semantic vector of the large model base network for the transaction training data is obtained, and model parameters of the large model base network are adjusted based on the classification coding boundary and the sample transaction semantic vector, including: In the first model training process, the transaction training data is encoded by the large model base network to obtain a first sample transaction semantic vector, and the task classifier is trained on a classification coding task for a classification coding boundary based on the first sample transaction semantic vector and a first classifier label carried by the transaction training data until the task classifier completes training; During the training process of the second model, the transaction training data is labeled modified based on the task classifier to obtain reference transaction training data carrying the second classifier label, the reference transaction training data is encoded through the large model base network to obtain a second sample transaction semantic vector, network gradient adjustment information is determined based on the second sample transaction semantic vector, the task classifier and the classification coding boundary, and model weights of the large model base network are adjusted based on the network gradient adjustment information until the large model base network completes training.
3. The method according to claim 2, wherein the task classifier is trained on the classification coding task for the classification coding boundary based on the first sample transaction semantic vector and the first classifier label carried by the transaction training data, until the task classifier is trained and the large model base network is trained, comprising: Inputting the first sample transaction semantic vector into the task classifier to obtain a first classification result, and determining classifier gradient adjustment information based on the first classifier label and the first classification result; Freeze the network weight parameters of the large model base network, and adjust the classifier weight of the task classifier based on the classifier gradient adjustment information to update the classification encoding boundary of the task classifier through the classifier weight adjustment until the task classifier completes training.
4. The method according to claim 2, wherein determining classifier gradient adjustment information based on the first classifier label and the first classification result comprises: A loss calculation is performed based on the first classifier label and the first classification result to obtain a classifier processing loss, and classifier gradient adjustment information is generated based on the classifier processing loss.
5. The method according to claim 2, wherein determining the network gradient adjustment information based on the second sample transaction semantic vector, the task classifier and the classification coding boundary comprises: Inputting the two sample transaction semantic vectors into the task classifier, and determining a second classification result through the classification coding boundary of the task classifier; Determine network gradient adjustment information based on the second classification result and the second classifier label; Unfreeze the network weight parameters of the large model base network and freeze the classifier weight parameters of the task classifier, and adjust the network weight parameters of the large model base network based on the network gradient adjustment information until the large model base network completes training.
6. The method according to claim 5, wherein determining the network gradient adjustment information based on the second classification result and the second classifier label comprises: A loss is calculated based on the second classification result and the second classifier label to obtain a network processing loss, and network gradient adjustment information is generated based on the network processing loss.
7. According to the method of claim 1, the adjusting the model parameters of the large model base network based on the classification coding boundary and the sample transaction semantic vector to obtain the transaction processing large model comprises: Adjust the model parameters of the large model base network based on the classification coding boundary and the sample transaction semantic vector to obtain the trained large model base network and the task classifier; Determine the target scenario task corresponding to the target transaction, use the transaction training data to perform target scenario task adaptation processing on the initial transaction processing big model, and obtain the target scenario task adapted transaction processing big model.
8. The method according to claim 7, wherein the step of using the transaction training data to perform target scenario task adaptation processing on the initial transaction processing big model to obtain the target scenario task adapted transaction processing big model comprises: The task classifier in the initial transaction processing large model is filtered out, and the initial transaction processing large model is adapted to the target scenario task using the transaction training data to obtain the transaction processing large model adapted to the target scenario task.
9. The method according to claim 1, wherein obtaining the large model processing capability impact factor for the target transaction and performing classification task modeling processing on the large model processing capability impact factor to obtain a task classifier comprises: Determine a target scenario task corresponding to the target transaction, perform classification task modeling processing on the target scenario task to obtain a first task classifier, obtain a large model processing capability impact factor for the target transaction, perform classification task modeling processing on the large model processing capability impact factor to obtain a second task classifier, and determine a task classifier based on the first task classifier and the second task classifier; The step of adjusting the model parameters of the large model base network based on the classification coding boundary and the sample transaction semantic vector to obtain a large transaction processing model includes: Adjust the model parameters of the large model base network based on the classification coding boundary and the sample transaction semantic vector to obtain the trained large model base network and the task classifier; A transaction processing big model is obtained based on the big model base network and the first task classifier.
10. A transaction processing method, the method comprising: Obtain target transaction data for a target transaction; Inputting the target transaction data into a transaction processing model and outputting a transaction processing result, wherein the transaction processing model is obtained by the method steps according to any one of claims 1 to 9; Target transaction processing is performed based on the transaction processing result.
11. A transaction processing device, comprising: A modeling module, used to obtain a large model processing capability impact factor for a target transaction, and to perform classification task modeling processing on the large model processing capability impact factor to obtain a task classifier; A creation module, used for creating an initial transaction processing big model for the target transaction based on the big model base network and the task classifier, and determining transaction training data for the initial transaction processing big model; A training module is used to input the transaction training data into the initial transaction processing large model for at least one round of model training. During the model training process, the classification coding boundary is determined by the task classifier, and the sample transaction semantic vector of the large model base network for the transaction training data is obtained. The model parameters of the large model base network are adjusted based on the classification coding boundary and the sample transaction semantic vector to obtain the transaction processing large model.
12. A transaction processing device, comprising: A data acquisition module, used for acquiring target transaction data for a target transaction; A model processing module, used for inputting the target transaction data into a transaction processing large model and outputting a transaction processing result, wherein the transaction processing large model is obtained by the method steps according to any one of claims 1 to 9; The transaction processing module is used to perform target transaction processing based on the transaction processing result.
13. A computer storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the method steps as claimed in any one of claims 1 to 9 or 10.
14. A computer program product, the computer program product storing at least one instruction, wherein the at least one instruction is loaded by a processor and executes the method steps according to any one of claims 1 to 9 or 10.
15. An electronic device, comprising: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method steps as claimed in any one of claims 1 to 9 or 10.