Big model-based accounting test verification method, apparatus and device, and medium
By using a large-model-based accounting testing and verification method, which automatically processes accounting test data using repayment prediction and journal entry prediction models, the problem of low efficiency and insufficient accuracy in traditional accounting testing and verification is solved, and efficient and accurate accounting test results are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 湖南长银五八消费金融股份有限公司
- Filing Date
- 2025-12-01
- Publication Date
- 2026-04-21
AI Technical Summary
Traditional accounting testing and verification processes suffer from inefficiency and inaccuracy, especially in complex accounting transaction scenarios. Manual comparison is prone to errors, and the maintenance costs are high when business rules change, making it difficult for testing to detect problems in a timely manner.
A large-scale model-based accounting testing and verification method is adopted. By obtaining relevant data on loan receipts from the accounting testing database, a feature set is constructed, and automatic verification is performed using repayment prediction models and journal entry prediction models, including feature extraction, model training, and result comparison, to achieve efficient and accurate accounting testing.
It enables efficient acquisition and accurate verification of intermediate data for accounting tests, improving testing efficiency and accuracy, reducing manual intervention, and increasing flexibility to adapt to changes in business rules.
Smart Images

Figure CN121901904A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, computer device, storage medium, and computer program product for accounting testing and verification based on a large model. Background Technology
[0002] Ensuring the accuracy and stability of accounting transactions is crucial during the testing of various accounting systems. As accounting systems become increasingly complex and business volume continues to rise, higher demands are placed on the efficiency and accuracy of accounting testing and verification. Currently, in many accounting system testing scenarios, the test execution phase has gradually become automated. With the help of various testing tools and scripts, a large number of transaction operations can be executed quickly, effectively improving the efficiency of the test execution phase.
[0003] However, traditional techniques still have significant shortcomings in post-transaction verification. For example, many accounting systems face the situation where testers need to manually query relevant data from multiple database tables during verification. In scenarios involving complex accounting transactions, this might involve querying information from multiple tables such as the loan agreement table, account table, and accounting transaction log. Furthermore, for each transaction, the actual results must be compared one by one with the expected results preset in the test cases. This process is not only repetitive and tedious but also consumes a significant amount of time, accounting for a substantial proportion of the overall testing time and severely hindering the improvement of testing efficiency.
[0004] More importantly, traditional manual comparison methods have numerous drawbacks, leading to inaccurate testing and verification. On the one hand, for accounting scenarios involving complex interest calculation rules and multiple accounting entries, manual comparison is prone to errors due to fatigue, negligence, and other factors, causing some potential defects to be overlooked and making it impossible to detect problems in the accounting system in a timely manner. On the other hand, when business rules change or test data changes, the pre-set expected results in the test cases need to be updated and maintained synchronously. This is not only costly to maintain, but also, if maintenance is not timely or errors occur, the pre-set results will become invalid, causing the test to be unable to continue normally. Manual recalculation and intervention are required, further slowing down the testing progress and reducing the accuracy and reliability of testing and verification. Therefore, how to improve the efficiency and accuracy of accounting testing and verification has become a key issue that urgently needs to be addressed in the current field of accounting system testing. Summary of the Invention
[0005] Therefore, it is necessary to provide an efficient and accurate accounting testing and verification method, apparatus, computer equipment, computer-readable storage medium, and computer program product based on a large model to address the aforementioned technical problems.
[0006] Firstly, this application provides a method for accounting testing and verification based on a large-scale model. The method includes:
[0007] Obtain relevant data on loan receipts from the accounting test database, and obtain the corresponding feature set based on the relevant data on loan receipts. The feature set includes numerical features and categorical features.
[0008] Input the feature set into the preset repayment prediction model to obtain multiple repayment prediction features;
[0009] Multiple repayment prediction features and feature sets are input into a preset journal entry prediction model to obtain a set of predicted accounting item labels;
[0010] Obtain the actual accounting entries after the repayment transaction occurs, compare the actual accounting entries with the predicted set of accounting item labels, and obtain the accounting test verification results.
[0011] In one embodiment, before inputting the feature set into a preset repayment prediction model to obtain multiple repayment prediction features, the method further includes:
[0012] Obtain the accounting system's business rules;
[0013] Based on the accounting system's business rules, numerical features, and categorical features, derived features are constructed.
[0014] Add the derived features to the feature set.
[0015] In one embodiment, the feature set is input into a preset repayment prediction model to obtain multiple repayment prediction features, including:
[0016] The numerical features in the feature set are standardized or normalized, and the categorical features in the feature set are one-hot encoded to obtain the processed feature set.
[0017] The processed feature set is input into a preset repayment prediction model to obtain multiple repayment prediction features.
[0018] In one embodiment, before inputting the feature set into a preset repayment prediction model to obtain multiple repayment prediction features, the method further includes:
[0019] Obtain training sample data for repayment prediction. The training sample data includes the feature vector of the loan agreement before the repayment transaction and multiple repayment labels after the repayment transaction.
[0020] We construct a total loss function by weighting the loss values of each repayment tag;
[0021] The initial LightGBM model is trained based on the loan feature vector, multiple repayment labels, and the total loss function. During the model training process, the gradient descent algorithm is used to iteratively optimize the model parameters to obtain the preset repayment prediction model.
[0022] In one embodiment, an initial LightGBM model is trained based on the loan feature vector, multiple repayment tags, and a total loss function. During model training, a gradient descent algorithm is used to iteratively optimize the model parameters, resulting in a pre-defined repayment prediction model, including:
[0023] Based on multiple repayment labels and the total loss function, the hyperparameters of the initial LightGBM model are configured. The initial LightGBM model is built based on the LightGBM multi-output regressor.
[0024] The initial LightGBM model is iteratively trained based on the loan receipt feature vector and multiple repayment tags. In each iteration, the negative gradient of the loss function between the current model's predicted repayment tags and the actual repayment tags is calculated, and a new decision tree is constructed to fit the negative gradient of the loss function. The prediction results corresponding to the new decision tree are weighted and superimposed into the current ensemble tree model. During the construction of each tree in the model, all features in the loan receipt feature vector are traversed, and the best split point is selected based on the principle of minimizing the squared error. The nonlinear relationship between the loan receipt feature vector and multiple repayment tags is automatically learned.
[0025] When the preset iteration stop condition is reached, the final generated ensemble tree model will be used as the preset repayment prediction model.
[0026] In one embodiment, the above-described accounting test verification method based on a large model further includes:
[0027] Obtain the accounting system's business rules;
[0028] Extract loan receipt information prior to repayment transactions from the relevant loan receipt data;
[0029] Based on the accounting system's business rules and the loan receipt information before the repayment transaction, verify whether the repayment prediction characteristics meet the constraints corresponding to the accounting system's business rules.
[0030] If the verification passes, the process proceeds to inputting the repayment prediction features and feature set into the preset journal entry prediction model to obtain the predicted accounting item label set.
[0031] In one embodiment, before inputting the repayment prediction features and feature set into a preset journal entry prediction model to obtain the predicted accounting item label set, the method further includes:
[0032] Obtain historical repayment transaction data;
[0033] Extract the accounting entries corresponding to each repayment transaction from the historical repayment transaction data to determine the sample accounting subject labels;
[0034] The sample accounting item labels are converted into multi-label 0-1 vectors to form the sample label matrix Y; the multi-label 0-1 vectors are used to represent whether each accounting item appears in each repayment transaction data;
[0035] Extract the sample feature matrix X corresponding to the sample label matrix Y from the historical repayment transaction data. The sample feature matrix X is constructed based on the sample repayment prediction features and sample features corresponding to the historical repayment transaction data.
[0036] Based on the sample feature matrix X and the corresponding sample label matrix Y, an initial multi-label classification model is trained using weighted multi-label binary cross-entropy as the loss function to obtain the preset entry prediction model.
[0037] Secondly, this application also provides an accounting testing and verification device based on a large model. The device includes:
[0038] The feature acquisition module is used to obtain relevant data on loan receipts from the accounting test database and obtain the corresponding feature set based on the relevant data on loan receipts. The feature set includes numerical features and categorical features.
[0039] The dual-model processing module is used to input the feature set into a preset repayment prediction model to obtain multiple repayment prediction features; and to input the multiple repayment prediction features and feature set into a preset journal entry prediction model to obtain a set of predicted accounting item labels.
[0040] The verification module is used to obtain the actual accounting entries after the repayment transaction occurs, compare the actual accounting entries with the predicted accounting subject label set, and obtain the accounting test verification results.
[0041] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0042] Obtain relevant data on loan receipts from the accounting test database, and obtain the corresponding feature set based on the relevant data on loan receipts. The feature set includes numerical features and categorical features.
[0043] Input the feature set into the preset repayment prediction model to obtain multiple repayment prediction features;
[0044] Multiple repayment prediction features and feature sets are input into a preset journal entry prediction model to obtain a set of predicted accounting item labels;
[0045] Obtain the actual accounting entries after the repayment transaction occurs, compare the actual accounting entries with the predicted set of accounting item labels, and obtain the accounting test verification results.
[0046] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0047] Obtain relevant data on loan receipts from the accounting test database, and obtain the corresponding feature set based on the relevant data on loan receipts. The feature set includes numerical features and categorical features.
[0048] Input the feature set into the preset repayment prediction model to obtain multiple repayment prediction features;
[0049] Multiple repayment prediction features and feature sets are input into a preset journal entry prediction model to obtain a set of predicted accounting item labels;
[0050] Obtain the actual accounting entries after the repayment transaction occurs, compare the actual accounting entries with the predicted set of accounting item labels, and obtain the accounting test verification results.
[0051] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0052] Obtain relevant data on loan receipts from the accounting test database, and obtain the corresponding feature set based on the relevant data on loan receipts. The feature set includes numerical features and categorical features.
[0053] Input the feature set into the preset repayment prediction model to obtain multiple repayment prediction features;
[0054] Multiple repayment prediction features and feature sets are input into a preset journal entry prediction model to obtain a set of predicted accounting item labels;
[0055] Obtain the actual accounting entries after the repayment transaction occurs, compare the actual accounting entries with the predicted set of accounting item labels, and obtain the accounting test verification results.
[0056] The aforementioned accounting test verification method, apparatus, computer equipment, storage medium, and computer program product based on a large model obtains relevant data on loan receipts from an accounting test database and acquires corresponding feature sets based on this data. These feature sets include both numerical and categorical features. The feature sets are then input into a pre-defined repayment prediction model to obtain multiple repayment prediction features. These multiple repayment prediction features and the feature set are then input into a pre-defined journal entry prediction model to obtain a set of predicted accounting subject labels. The actual accounting entries after the repayment transaction occur are obtained, and the actual accounting entries are compared with the set of predicted accounting subject labels to obtain the accounting test verification results. Throughout the process, the pre-defined repayment prediction model and journal entry prediction model are used sequentially to process the feature sets, resulting in multiple repayment prediction features and sets of predicted accounting subject labels. Leveraging the powerful analytical and predictive capabilities of the models, efficient acquisition of intermediate accounting test data is achieved. Finally, by comparing the predicted accounting subject labels with the actual journal entries, efficient and accurate accounting test verification is achieved. Attached Figure Description
[0057] Figure 1 This is a diagram illustrating the application environment of an accounting test and verification method based on a large model in one embodiment.
[0058] Figure 2 This is a flowchart illustrating an accounting test and verification method based on a large model in one embodiment.
[0059] Figure 3 This is a flowchart illustrating an accounting test and verification method based on a large model, as described in another embodiment.
[0060] Figure 4 This is an interactive diagram illustrating an accounting test and verification method based on a large model in a specific application example.
[0061] Figure 5 This is a structural block diagram of an accounting test and verification device based on a large model in one embodiment;
[0062] Figure 6 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0063] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0064] The accounting testing and verification method based on a large model provided in this application can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on a cloud or other network server. Terminal 102 sends an accounting test verification request to server 104. Server 104 responds to the request by retrieving relevant data from the accounting test database and obtaining a corresponding feature set based on the data. The feature set includes numerical and categorical features. The feature set is input into a preset repayment prediction model to obtain multiple repayment prediction features. The multiple repayment prediction features and the feature set are input into a preset journal entry prediction model to obtain a predicted accounting subject label set. The actual accounting entry after the repayment transaction occurs is obtained, and the actual accounting entry is compared with the predicted accounting subject label set to obtain the accounting test verification result. Furthermore, server 104 can feed back the accounting test verification result to terminal 102. The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle systems. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0065] In one embodiment, such as Figure 2 As shown, an accounting testing and verification method based on a large model is provided, which is then applied to... Figure 1 Taking server 104 as an example, the following steps are included:
[0066] S200: Obtain relevant data on loan receipts from the accounting test database, and obtain the corresponding feature set based on the relevant data on loan receipts. The feature set includes numerical features and categorical features.
[0067] Server 104 establishes a connection with the accounting test database, which stores a large amount of data related to IOUs. This data covers various aspects of the IOUs, such as IOU number, loan amount, loan term, interest rate, and repayment method. Through query statements, the required IOU-related data can be accurately retrieved from the accounting test database.
[0068] After obtaining the relevant data of the loan agreement, feature extraction is performed. The purpose of feature extraction is to extract information from the raw data that has a significant impact on subsequent model predictions. Numerical features are those that can be represented by specific numerical values, such as loan amount, loan term (specific values in months or years), and interest rate (specific percentage values). These features directly reflect the quantitative information of the loan agreement. Categorical features are those used to represent different categories or states, such as repayment methods (equal principal and interest payments, equal principal payments, interest-only payments, etc.) and customer credit ratings (A, B, C, etc.). These features help the model distinguish different situations and patterns. By comprehensively extracting numerical and categorical features, a complete feature set is formed, providing a comprehensive and accurate data foundation for subsequent model predictions.
[0069] S400: Input the feature set into the preset repayment prediction model to obtain multiple repayment prediction features.
[0070] The repayment prediction model can specifically employ the Gradient Boosting Decision Tree (GBDT) model. As a powerful machine learning algorithm, GBDT offers numerous advantages, including the ability to accurately learn complex data calculation rules from historical data. Through multiple rounds of iterative training, each iteration generates a decision tree. These decision trees are combined according to certain rules to jointly analyze and predict the input feature set. During training, the model continuously adjusts its parameters to minimize prediction error, thereby improving prediction accuracy. After training, the GBDT model can output multiple repayment prediction features based on the feature set corresponding to the loan-related data. These features include, but are not limited to, the principal repayment amount; the interest repayment amount; the penalty interest repayment amount; and the new remaining principal amount after repayment.
[0071] S600: Input multiple repayment prediction features and feature sets into the preset journal entry prediction model to obtain the prediction account label set.
[0072] The journal entry prediction model employs a multi-label classification model. This model can simultaneously handle classification problems involving multiple labels. In this accounting test scenario, a single transaction may trigger records for multiple accounting subjects. Therefore, the multi-label classification model can intelligently predict all accounting subjects that a transaction should trigger. The multiple repayment prediction features obtained in step S400 and the feature set from step S200 are used as input data and fed into the preset multi-label classification journal entry prediction model. The multi-label classification model comprehensively analyzes the input data, predicting all accounting subjects that the transaction may trigger based on various feature information within the data, combined with patterns and rules learned by the model itself. This prediction is then output as a set of predicted accounting subject labels.
[0073] S800: Obtain the actual accounting entries after the repayment transaction occurs, compare the actual accounting entries with the predicted accounting item label set, and obtain the accounting test verification results.
[0074] After an actual repayment transaction occurs, the system records the actual accounting entries for that transaction. This includes the actual triggered accounting entries and the corresponding amount for each accounting entry. Specifically, after obtaining the actual accounting entries, they are compared with the predicted accounting entry label set obtained in step S600. This comparison process includes two aspects: first, comparing whether the actual triggered accounting entries match the predicted accounting entry labels; and second, comparing whether the actual amount corresponding to each accounting entry matches the predicted allocated amount. Through this end-to-end instantaneous comparison, the system can quickly and accurately determine whether the accounting system processing is correct. If the actual accounting entries and the predicted accounting entry label set are completely consistent, the accounting processing is correct, and the verification result is passed; if there are inconsistencies, it indicates that there may be a problem with the accounting system processing.
[0075] The aforementioned large-scale model-based accounting test and verification method retrieves loan document-related data from the accounting test database and obtains corresponding feature sets based on this data. These feature sets include both numerical and categorical features. The feature sets are then input into a pre-defined repayment prediction model to obtain multiple repayment prediction features. These features and the feature set are then input into a pre-defined journal entry prediction model to obtain a set of predicted accounting subject labels. The actual accounting entries after the repayment transaction occur are obtained, and these entries are compared with the predicted accounting subject label set to obtain the accounting test and verification results. Throughout this process, the pre-defined repayment prediction model and journal entry prediction model are used sequentially to process the feature sets, resulting in multiple repayment prediction features and sets of predicted accounting subject labels. Leveraging the model's powerful analytical and predictive capabilities, efficient acquisition of intermediate accounting test data is achieved. Finally, by comparing the predicted accounting subject labels with the actual journal entries, efficient and accurate accounting test and verification are achieved.
[0076] In one embodiment, such as Figure 3 As shown, before inputting the feature set into the preset repayment prediction model to obtain multiple repayment prediction features, the following steps are also included:
[0077] S320: Obtain the accounting system business rules.
[0078] Accounting system business rules are a series of guidelines and norms followed in the course of accounting-related business operations. These rules cover all aspects of lending business, such as repayment methods (calculation rules for different repayment methods such as equal principal and interest payments, equal principal payments, and interest-only payments followed by principal payments), interest rate calculation rules (calculation methods for fixed interest rates and floating interest rates), and prepayment rules (calculation of prepayment fees, handling of remaining principal, etc.). These business rules are obtained by interacting with the accounting system, from its configuration files, rule tables stored in the database, or system interfaces.
[0079] S340: Construct derived features based on the accounting system's business rules, numerical features, and categorical features.
[0080] After obtaining the business rules of the accounting system, derived features are constructed by combining the numerical and categorical features obtained in step S200. For example, based on the relationship between the remaining principal and the loan amount in the business rules, the derived feature "remaining principal as a percentage of the loan amount" is constructed. Specifically, the remaining principal and loan amount are obtained from the main table, and then calculated using the formula: Remaining principal as a percentage of the loan amount = Remaining principal / Loan amount. As another example, for the equal principal and interest repayment method, the derived feature "theoretical monthly repayment amount" is constructed based on its corresponding interest rate calculation rules and repayment calculation rules. The loan amount and current period are obtained from the main table, and the annual interest rate is obtained from the interest rate table. Then, the theoretical monthly repayment amount (equal principal and interest) is calculated using the formula: Loan amount × Annual interest rate / 12 × (1 + Annual interest rate / 12)^Number of periods / [(1 + Annual interest rate / 12)^Number of periods - 1]. By constructing these derived features, a business baseline and repayment progress reference can be provided.
[0081] S360: Add the derived features to the feature set.
[0082] After constructing the derived features, these derived features are merged with the original feature set from step S200. The merging can be done through simple data concatenation, adding the data columns of the derived features to the data table of the original feature set to form a new feature set containing more feature information. In subsequent processing, this new feature set is input into a preset bad debt prediction model.
[0083] In one embodiment, the feature set is input into a preset repayment prediction model to obtain multiple repayment prediction features, including:
[0084] Step 1: Standardize or normalize the numerical features in the feature set and perform one-hot encoding on the categorical features in the feature set to obtain the processed feature set.
[0085] In accounting testing scenarios, the feature set includes various numerical features, such as loan amount, loan term, interest rate, and derived features constructed in the steps, such as "theoretical monthly repayment amount" and "the proportion of remaining principal to loan amount." Because these numerical features differ significantly in their units and value ranges, directly inputting them into the model may affect the model's training performance and convergence speed. Therefore, they are standardized or normalized.
[0086] Categorical features in the feature set, such as repayment methods (equal principal and interest payments, equal principal payments, interest-only payments, etc.) and credit ratings (A, B, C, etc.), cannot be directly processed by the model and need to be converted into numerical form. This embodiment uses one-hot encoding to process categorical features. Alternatively, LightGBM native categorical feature support can also be used for processing.
[0087] Step 2: Input the processed feature set into the preset repayment prediction model to obtain multiple repayment prediction features.
[0088] The processed feature set obtained in step 1 is input into a preset repayment prediction model. In this embodiment, the repayment prediction model can be a gradient boosting decision tree model such as LightGBM. This model has efficient training speed and good prediction performance, and can handle large-scale data and complex feature relationships.
[0089] Internally, the basic numerical features, encoded categorical features, and derived features in the processed feature set are collaboratively processed through a "tree structure splitting + feature interaction learning" approach. For numerical features (including derived features), the model does not require manual normalization / standardization (even if processed in step 1, the model can still automatically optimize), but directly judges the splitting value based on the feature's "information gain" (such as Gini index, mean squared error). For example, "theoretical monthly repayment amount" and "remaining principal percentage" are treated as continuous values, and different thresholds (such as "theoretical monthly repayment amount > 5000 yuan" and "remaining principal percentage < 30%) are tried at each node of the tree to split the feature, selecting the threshold with the highest discriminative power for the predicted value (such as the principal repayment). In this way, the model can gradually build a complex tree structure to effectively partition and predict the data. For encoded categorical features, for one-hot encoded "repayment methods" (e.g., equal principal and interest = 100, equal principal = 010), the model treats each encoded dimension as an independent feature and evaluates the contribution of each dimension to the prediction result separately. For label encoding (e.g., credit rating A = 1, B = 2), if LightGBM native categorical feature support is used, the model will process them according to their numerical values and analyze the impact of different credit ratings on repayment. During the tree structure splitting process, different features will interact and learn from each other, jointly influencing the model's prediction results. Through the construction of multi-layered tree structures and feature interactions, the model can learn complex patterns and rules in the data, ultimately outputting multiple repayment prediction features, such as the amount of principal repaid; the amount of interest repaid; the amount of penalty interest repaid; and the new remaining principal amount after repayment.
[0090] In one embodiment, before inputting the feature set into a preset repayment prediction model to obtain multiple repayment prediction features, the method further includes:
[0091] Step 1: Obtain training sample data for repayment prediction. The training sample data includes the loan feature vector before the repayment transaction and multiple repayment labels after the repayment transaction.
[0092] The accounting system has accumulated a large number of repayment records that have been verified as correct by business operations. These records provide a rich data source for model training. The process of obtaining each training sample includes: Loan feature vector (input feature X): Relevant information before the repayment transaction is generated from the accounting system's database to construct the loan feature vector. This information covers various aspects of the loan, specifically numerical features, categorical features, and derived features corresponding to historical repayment records. Repayment label (training label Y): After the repayment transaction occurs, the system records and verifies the repayment-related information as training labels. These labels include principal repayment, interest repayment, penalty interest repayment, and the new remaining principal.
[0093] Step 2: Weight the loss values of each repayment tag to construct the total loss function.
[0094] To enable the model to comprehensively consider the prediction accuracy of multiple repayment tags, the loss values of each repayment tag are weighted to construct a total loss function. For each repayment tag, an appropriate loss function is selected to measure the difference between the model's predicted value and the actual value. In this embodiment, each sub-loss function is preferably Mean Squared Error (MSE) or Mean Absolute Error (MAE).
[0095] The weighted sum of the individual loss functions for each output task (i.e., the various repayment tags mentioned above) is used to construct the total loss function. The expression for the total loss function is: TotalLoss = α × Loss_principal + β × Loss_interest + γ × Loss_penalty + δ × Loss_new_principal. Here, α, β, γ, and δ are the weighting coefficients for the loss of each task, and these coefficients can be adjusted according to the importance of the business. For example, if principal repayment is the most important indicator in the business, the value of α can be increased accordingly, making the model focus more on the prediction accuracy of principal repayment during training. By adjusting the weighting coefficients, the importance of different repayment tags in model training can be flexibly balanced, allowing the model to better meet actual business needs.
[0096] Step 3: Train an initial LightGBM model based on the loan feature vector, multiple repayment labels, and the total loss function. During the model training process, the gradient descent algorithm is used to iteratively optimize the model parameters to obtain the preset repayment prediction model.
[0097] First, an initial LightGBM model is initialized. The loan feature vector obtained in step 1 is used as the input feature X, and the total loss function constructed in step 2 is used as the optimization objective to train the initial LightGBM model. During training, the gradient descent algorithm is used to iteratively optimize the parameters of the decision tree set. The gradient descent algorithm determines the direction and step size of parameter updates by calculating the gradient of the loss function with respect to the model parameters, gradually reducing the loss function value and thus improving the model's predictive performance. Specifically, for each iteration, the algorithm calculates the gradient of the loss function based on the current model parameters, and then updates the model parameters in the opposite direction of the gradient, adjusting the model in the direction of decreasing loss function, ultimately obtaining the preset repayment prediction model.
[0098] In one embodiment, an initial LightGBM model is trained based on the loan feature vector, multiple repayment tags, and a total loss function. During model training, a gradient descent algorithm is used to iteratively optimize the model parameters, resulting in a pre-defined repayment prediction model, including:
[0099] Step 1: Configure the hyperparameters of the initial LightGBM model based on multiple repayment labels and the total loss function. The initial LightGBM model is built based on the LightGBM multi-output regressor.
[0100] When constructing the initial LightGBM model, considering the need to simultaneously predict multiple repayment labels (principal repayment, interest repayment, penalty interest repayment, and new remaining principal), a LightGBM multi-output regressor was used. This model can handle multi-output regression problems, sharing feature processing logic while learning independent tree structures and parameters for each output task. Hyperparameter configuration was adjusted based on the characteristics of multiple repayment labels and the total loss function. For example, the learning rate was set, controlling the correction magnitude of each tree to the overall prediction result. A smaller learning rate makes the model training process more stable but requires more iterations; a larger learning rate may speed up training but may cause the model to have difficulty converging. Additionally, the maximum tree depth was set to limit the complexity of each decision tree and prevent overfitting. Furthermore, hyperparameters such as the minimum number of samples per leaf node were reasonably set according to the size of the training data and the feature dimension to ensure the model can fully learn the patterns in the data while avoiding overfitting.
[0101] Step 2: Iteratively train the initial LightGBM model based on the loan feature vector and multiple repayment tags; in each round of training iteration, calculate the negative gradient of the loss function between the current model's predicted repayment tags and the actual repayment tags, and construct a new decision tree to fit the negative gradient of the loss function; weight and superimpose the prediction results corresponding to the new decision tree into the current ensemble tree model; during the construction of each tree in the model, traverse all features in the loan feature vector, select the best split point based on the principle of minimizing squared error, and automatically learn the nonlinear relationship between the loan feature vector and multiple repayment tags.
[0102] The specific steps in each round of iterative training are as follows:
[0103] 1) Calculate the negative gradient of the loss function: First, calculate the negative gradient of the loss function between the repayment label predicted by the current model and the actual repayment label. The loss function used is the total loss function, i.e., TotalLoss = α*Loss_principal + β*Loss_interest + γ*Loss_penalty + δ*Loss_new_principal, where α, β, γ, and δ are the weight coefficients of the loss for each task. By calculating the negative gradient of this loss function with respect to the predicted value, the direction and magnitude of the model's prediction error are determined, providing a basis for subsequently constructing a new decision tree.
[0104] 2) Construct a new decision tree to fit the negative gradient of the loss function: Use the calculated negative gradient of the loss function as the target value to construct a new decision tree.
[0105] 3) Update the ensemble tree model: The prediction results corresponding to the new decision trees are weighted and superimposed into the current ensemble tree model. In this way, each new tree learns the error between the previous round of predictions and the actual values, and corrects the overall prediction result through weighting, similar to a process of "continuous error correction". For example, the first tree uses "remaining principal" and "theoretical monthly repayment amount" to initially predict the principal repayment; the second tree finds that the sample prediction error of "repayment method = equal principal" is large, so it focuses on correcting the error using the features of "repayment method" and "number of repayments"; finally, the prediction results of multiple trees are added together to obtain a more accurate output.
[0106] 4) Feature Splitting and Automatic Learning of Nonlinear Relationships: During the construction of each tree in the model, all features in the loan feature vector are traversed, and the optimal split point is selected based on the principle of minimizing squared error. Starting from the root node, the model calculates the "reduction in prediction error of child nodes after splitting" for each feature (e.g., the decrease in mean squared error in regression tasks), and selects the feature with the largest error reduction and the threshold for splitting. For example, when predicting "the amount of principal repaid," the model may first split by "the proportion of remaining principal to loan amount" (e.g., samples with a proportion >50% are assigned to the left subtree, and those ≤50% to the right subtree), because this feature can significantly distinguish between "early repayment period (lower principal repayment)" and "late repayment period (higher principal repayment)." Then, it further splits in the subtree using "theoretical monthly repayment amount" (e.g., samples with a theoretical value >3000 yuan have higher principal repayment), gradually refining the prediction rules. In this way, the model automatically learns the nonlinear relationship between the loan feature vector and multiple repayment labels. At the same time, features that are selected more often during splitting and have a greater error reduction are more important. Derivative features (such as theoretical monthly repayment amount) are usually of high importance because they are directly related to business rules, and become the "key nodes" for model splitting.
[0107] Step 3: When the preset iteration stop condition is reached, the final generated ensemble tree model is used as the preset repayment prediction model.
[0108] The preset iteration stopping condition can be one of several things, such as reaching a preset number of iterations, or the model's performance on the validation set (e.g., loss function value, prediction accuracy) stabilizing and no longer showing significant improvement. When the preset iteration stopping condition is met, the model training process stops, and the final generated ensemble tree model is used as the preset repayment prediction model.
[0109] In one embodiment, the above-described accounting test verification method based on a large model further includes:
[0110] Step 1: Obtain the accounting system business rules.
[0111] Business rules can be determined based on historical data from the accounting system, industry standards, and accounting principles, and they support user-defined settings and modifications. Specifically, business rules cover various conditions and constraints in the accounting process. For example, in loan repayment, they involve the calculation relationships and value range rules between core indicators such as new remaining principal, repaid principal, repaid interest, and repaid penalty interest. For instance, it may stipulate that "new remaining principal" must equal "remaining principal before repayment" minus "projected repaid principal," and that interest and penalty interest amounts must be within a reasonable range calculated based on the contractual interest rate.
[0112] Step 2: Extract the loan receipt information prior to the repayment transaction from the relevant loan receipt data.
[0113] Extract loan information prior to repayment from pre-stored loan-related data. This information includes various data related to repayment, such as the remaining principal before repayment, the contractual interest rate, the number of installments already repaid, and the repayment method. The extraction process can be performed using database queries to accurately retrieve the corresponding data records from the test database based on the loan's unique identifier (such as the loan number). For example, the SQL statement "SELECT Remaining Principal Before Repayment, Contractual Interest Rate, Number of Installments Already Repaid, Repayment Method FROM Loan Table WHERE Loan Number = [Target Loan Number]" can be used to retrieve the loan information prior to repayment for the target loan.
[0114] Step 3: Based on the accounting system's business rules and the loan receipt information before the repayment transaction, verify whether the repayment prediction characteristics meet the constraints corresponding to the accounting system's business rules.
[0115] Based on the accounting system business rules obtained in step 1 and the pre-repayment transaction loan information extracted in step 2, a comprehensive verification of the repayment prediction characteristics is performed. Specific verifications may include verifying the relationship between the new remaining principal and the principal to be repaid; and verifying the reasonableness of interest and penalty interest amounts.
[0116] Step 4: If the verification passes, proceed to the step of inputting the repayment prediction features and feature set into the preset journal entry prediction model to obtain the prediction account label set.
[0117] When the verification result in step 3 is passed, it indicates that the repayment prediction features meet the requirements of the accounting system's business rules. At this point, the subsequent journal entry prediction process can begin. The verified repayment prediction features, along with a pre-prepared feature set (which includes other necessary features related to accounting processing, such as basic information about the loan agreement and customer information), are input into a preset journal entry prediction model. This model, trained on a large amount of data, can accurately predict the corresponding set of accounting subject labels based on the input feature information. For example, based on information such as principal repayment and interest repayment in the repayment prediction features, combined with other relevant features in the feature set, the corresponding accounting subject for the repayment transaction is predicted, such as "Bank Deposit," "Accounts Receivable - Principal," or "Accounts Receivable - Interest." This approach achieves a seamless process from repayment prediction to journal entry prediction, providing strong support for the automated testing and verification of the accounting system.
[0118] In one embodiment, before inputting the repayment prediction features and feature set into a preset journal entry prediction model to obtain the predicted accounting item label set, the method further includes:
[0119] Step 1: Obtain historical repayment transaction data.
[0120] A large amount of historical repayment transaction data is obtained from the accounting system database. This data covers various types of repayment transactions, including normal repayments, overdue repayments, and early repayments, as well as repayment transactions under different product types (such as housing loans and consumer loans) and different customer types (such as individual customers and corporate customers).
[0121] Step 2: Extract the accounting entries corresponding to each repayment transaction in the historical repayment transaction data and determine the sample accounting subject labels.
[0122] For each historical repayment transaction data obtained in step 1, the corresponding accounting entry information is extracted from the accounting entry records of the accounting system. Based on business rules and accounting standards, the accounting subject label associated with each repayment transaction is determined. For example, a normal equal principal and interest repayment transaction might have corresponding accounting entries involving "Customer Deposit Account (Debit)," "Loan Principal Account (Credit)," and "Interest Income Account (Credit)." Therefore, the sample accounting subject label for this repayment transaction data is determined to be one of these three accounts.
[0123] Step 3: Convert the sample accounting item labels into multi-label 0-1 vectors to form the sample label matrix Y; the multi-label 0-1 vectors are used to characterize whether each accounting item appears in each repayment transaction data.
[0124] To facilitate model processing and learning, the sample accounting item labels determined in step 2 are converted into multi-label 0-1 vectors. Specifically, each accounting item in each transaction is labeled according to a predefined label index. If an accounting item appears in the transaction, the value at the corresponding position is 1; otherwise, the value is 0. For example, according to the label index defined in the technical disclosure (0-bank deposit debit, 1-loan principal credit, 2-interest income credit, 3-penalty interest income credit, 4-default fee income credit), for a normal repayment transaction (no penalty interest or default fee), its features are P=800, I=200, Pe=0, NP_new=9200, then the corresponding label vector is [1, 1, 1, 0, 0]. Arrange the tag vectors corresponding to all historical repayment transaction data in the order of samples to form a sample tag matrix Y. The shape of this matrix is the number of samples × the total number of tags (e.g., if the number of samples is N and the total number of tags is 5, then the shape of the matrix is N×5), which corresponds one-to-one with the sample feature matrix X constructed later.
[0125] Step 4: Extract the sample feature matrix X corresponding to the sample label matrix Y from the historical repayment transaction data. The sample feature matrix X is constructed based on the combination of the sample repayment prediction features and sample features corresponding to the historical repayment transaction data.
[0126] Extract sample features corresponding to the sample label matrix Y from the historical repayment transaction data obtained in Step 1. These sample features include core input features and auxiliary features. The core input features are the dynamic expected results output by the repayment prediction model, namely the predicted principal repayment (P), interest repayment (I), penalty interest repayment (Pe), and new remaining principal (NP_new). The auxiliary features are loan characteristics obtained from the database, such as the remaining principal before repayment (NP_before), repayment type (normal / overdue / early), overdue days, contract interest rate, and total repayment amount (P+I+Pe). These auxiliary features are used to determine the business scenario. Integrate the core input features and auxiliary features, arrange them in the same sample order as the sample label matrix Y, and construct the sample feature matrix X. For example, for each historical repayment transaction data, combine its core input features and auxiliary features into a feature vector in a certain order, and then arrange the feature vectors of all samples into a matrix to obtain the sample feature matrix X. The shape of this matrix is also the number of samples × feature dimension (the feature dimension is determined according to the number of extracted features).
[0127] Step 5: Based on the sample feature matrix X and the corresponding sample label matrix Y, train the initial multi-label classification model using weighted multi-label binary cross-entropy as the loss function to obtain the preset entry prediction model.
[0128] An initial multi-label classification model is constructed using a specialized multi-label learning algorithm. This model is responsible for learning the mapping relationship from business scenarios (represented by the sample feature matrix X) to accounting subject combinations (represented by the sample label matrix Y). During model training, weighted multi-label binary cross-entropy is used as the loss function to strengthen the learning of strong correlations between core features and specific subjects.
[0129] Specifically, for each sample i and label t (subject), the loss function is defined as:
[0130]
[0131] in: The weight of label t (core subjects have higher weights, such as penalty interest income loan w=2, ordinary subjects w=1); For real labels (1 = present, 0 = absent). The model predicts probabilities; N is the number of samples, and T is the total number of labels (e.g., 5 subjects).
[0132] By continuously adjusting the model's parameters, the loss function value is minimized, meaning the probability distribution predicted by the model is made as close as possible to the true label distribution. During training, iterative training is performed using the sample feature matrix X and the sample label matrix Y. After multiple iterations, training stops when the performance of the initial multi-label classification model reaches the preset performance requirements (such as accuracy and recall), resulting in the preset entry prediction model.
[0133] Furthermore, in practical applications, the interactive process of the accounting test and verification method based on a large model in this application is as follows: Figure 4 As shown.
[0134] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0135] Based on the same inventive concept, this application also provides a large-model-based accounting testing and verification device for implementing the aforementioned large-model-based accounting testing and verification method. The solution provided by this device is similar to the implementation described in the above method. Therefore, the specific limitations in one or more embodiments of the large-model-based accounting testing and verification device provided below can be found in the limitations of the large-model-based accounting testing and verification method described above, and will not be repeated here.
[0136] In one embodiment, such as Figure 5 As shown, an accounting test and verification device based on a large model is provided, including:
[0137] The feature acquisition module 200 is used to obtain relevant data of the loan receipts from the accounting test database and obtain the corresponding feature set based on the relevant data of the loan receipts. The feature set includes numerical features and categorical features.
[0138] The dual-model processing module 400 is used to input the feature set into a preset repayment prediction model to obtain multiple repayment prediction features; and to input the multiple repayment prediction features and feature set into a preset journal entry prediction model to obtain a set of predicted accounting item labels.
[0139] The verification module 600 is used to obtain the actual accounting entries after the repayment transaction occurs, compare the actual accounting entries with the predicted accounting subject label set, and obtain the accounting test verification results.
[0140] In one embodiment, the dual-model processing module 400 is further configured to acquire accounting system business rules; construct derived features based on accounting system business rules, numerical features, and categorical features; and add the derived features to the feature set.
[0141] In one embodiment, the dual-model processing module 400 is further configured to standardize or normalize the numerical features in the feature set and perform one-hot encoding on the categorical features in the feature set to obtain the processed feature set; the processed feature set is then input into a preset repayment prediction model to obtain multiple repayment prediction features.
[0142] In one embodiment, the dual-model processing module 400 is further configured to acquire training sample data for repayment prediction, including a loan feature vector before the repayment transaction and multiple repayment tags after the repayment transaction; weight the loss values of each repayment tag to construct a total loss function; train an initial LightGBM model based on the loan feature vector, multiple repayment tags, and the total loss function, and iteratively optimize the model parameters using a gradient descent algorithm during model training to obtain a preset repayment prediction model.
[0143] In one embodiment, the dual-model processing module 400 is further configured to configure the hyperparameters of the initial LightGBM model based on multiple repayment tags and the total loss function. The initial LightGBM model is constructed based on the LightGBM multi-output regressor. The initial LightGBM model is iteratively trained based on the loan feature vector and multiple repayment tags. In each round of iterative training, the negative gradient of the loss function between the current model's predicted repayment tags and the actual repayment tags is calculated, and a new decision tree is constructed to fit the negative gradient of the loss function. The prediction results corresponding to the new decision tree are weighted and superimposed into the current ensemble tree model. During the construction of each tree in the model, all features in the loan feature vector are traversed, and the best split point is selected based on the principle of minimizing the squared error. The nonlinear relationship between the loan feature vector and multiple repayment tags is automatically learned. When the preset iteration stopping condition is reached, the finally generated ensemble tree model is used as the preset repayment prediction model.
[0144] In one embodiment, the dual-model processing module 400 is also used to obtain the accounting system business rules; extract the pre-repayment transaction loan information from the loan information; verify whether the repayment prediction features meet the constraints corresponding to the accounting system business rules based on the accounting system business rules and the pre-repayment transaction loan information; if the verification is successful, proceed to the step of inputting the repayment prediction features and feature set into the preset entry prediction model to obtain the predicted accounting subject label set.
[0145] In one embodiment, the dual-model processing module 400 is further configured to acquire historical repayment transaction data; extract the accounting entries corresponding to each repayment transaction in the historical repayment transaction data to determine sample accounting subject labels; convert the sample accounting subject labels into multi-label 0-1 vectors to form a sample label matrix Y; the multi-label 0-1 vectors are used to characterize whether each accounting subject appears in each repayment transaction; extract the sample feature matrix X corresponding to the sample label matrix Y from the historical repayment transaction data, the sample feature matrix X is constructed based on the sample repayment prediction features and sample features corresponding to the historical repayment transaction data; and train an initial multi-label classification model based on the sample feature matrix X and the corresponding sample label matrix Y, using weighted multi-label binary cross-entropy as the loss function, to obtain a preset entry prediction model.
[0146] Each module in the aforementioned large-model-based accounting testing and verification device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the computer device's memory as software, so that the processor can call and execute the corresponding operations of each module.
[0147] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores preset data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements a large-scale accounting test and verification method.
[0148] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0149] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described accounting test and verification method based on a large model.
[0150] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described accounting test and verification method based on a large model.
[0151] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described accounting test and verification method based on a large model.
[0152] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0153] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0154] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for accounting testing and verification based on a large model, characterized in that, The method includes: Obtain relevant data on loan receipts from the accounting test database, and obtain a corresponding feature set based on the relevant data on loan receipts. The feature set includes numerical features and categorical features. The feature set is input into a preset repayment prediction model to obtain multiple repayment prediction features; The multiple repayment prediction features and the feature set are input into a preset journal entry prediction model to obtain a set of predicted accounting item labels; Obtain the actual accounting entries after the repayment transaction occurs, compare the actual accounting entries with the predicted accounting subject label set, and obtain the accounting test verification results.
2. The method according to claim 1, characterized in that, Before inputting the feature set into a preset repayment prediction model to obtain multiple repayment prediction features, the method further includes: Obtain the accounting system's business rules; Based on the accounting system business rules, the numerical features, and the categorical features, derived features are constructed. The derived features are added to the feature set.
3. The method according to claim 1, characterized in that, The step of inputting the feature set into a preset repayment prediction model to obtain multiple repayment prediction features includes: The numerical features in the feature set are standardized or normalized, and the categorical features in the feature set are one-hot encoded to obtain the processed feature set. The processed feature set is input into a preset repayment prediction model to obtain multiple repayment prediction features.
4. The method according to claim 1, characterized in that, Before inputting the feature set into a preset repayment prediction model to obtain multiple repayment prediction features, the method further includes: Obtain training sample data for repayment prediction, the training sample data including the loan feature vector before the repayment transaction and multiple repayment labels after the repayment transaction. The total loss function is constructed by weighting the loss values of each repayment tag. An initial LightGBM model is trained based on the loan feature vector, the multiple repayment tags, and the total loss function. During the model training process, the gradient descent algorithm is used to iteratively optimize the model parameters to obtain a preset repayment prediction model.
5. The method according to claim 4, characterized in that, The process of training an initial LightGBM model based on the loan feature vector, the multiple repayment tags, and the total loss function, and iteratively optimizing the model parameters using a gradient descent algorithm during model training to obtain a preset repayment prediction model includes: Based on the multiple repayment tags and the total loss function, the hyperparameters of the initial LightGBM model are configured, and the initial LightGBM model is constructed based on the LightGBM multi-output regressor; The initial LightGBM model is iteratively trained based on the loan feature vector and the multiple repayment tags. In each iteration, the negative gradient of the loss function between the current model's predicted repayment tags and the actual repayment tags is calculated, and a new decision tree is constructed to fit the negative gradient of the loss function. The prediction results corresponding to the new decision tree are weighted and superimposed into the current ensemble tree model. During the construction of each tree in the model, all features in the loan feature vector are traversed, and the optimal split point is selected based on the principle of minimizing the squared error. The nonlinear relationship between the loan feature vector and the multiple repayment tags is automatically learned. When the preset iteration stop condition is reached, the final generated ensemble tree model will be used as the preset repayment prediction model.
6. The method according to claim 1, characterized in that, Also includes: Obtain the accounting system's business rules; Extract the loan receipt information prior to the repayment transaction from the relevant data of the loan receipt; Based on the accounting system business rules and the loan receipt information before the repayment transaction, verify whether the repayment prediction feature meets the constraints corresponding to the accounting system business rules; If the verification passes, the process proceeds to the step of inputting the repayment prediction features and the feature set into the preset journal entry prediction model to obtain the predicted accounting subject label set.
7. The method according to claim 1, characterized in that, Before inputting the repayment prediction features and the feature set into the preset journal entry prediction model to obtain the predicted accounting item label set, the method further includes: Obtain historical repayment transaction data; Extract the accounting entries corresponding to each repayment transaction from the historical repayment transaction data to determine the sample accounting subject labels; The sample accounting item labels are converted into multi-label 0-1 vectors to form a sample label matrix Y; the multi-label 0-1 vectors are used to characterize whether each accounting item appears in each repayment transaction data; Extract the sample feature matrix X corresponding to the sample label matrix Y from the historical repayment transaction data. The sample feature matrix X is constructed based on the combination of sample repayment prediction features and sample features corresponding to the historical repayment transaction data. Based on the sample feature matrix X and the corresponding sample label matrix Y, an initial multi-label classification model is trained using weighted multi-label binary cross-entropy as the loss function to obtain a preset entry prediction model.
8. An accounting testing and verification device based on a large model, characterized in that, The device includes: The feature acquisition module is used to acquire loan receipt-related data from the accounting test database and acquire a corresponding feature set based on the loan receipt-related data. The feature set includes numerical features and categorical features. The dual-model processing module is used to input the feature set into a preset repayment prediction model to obtain multiple repayment prediction features; and to input the multiple repayment prediction features and the feature set into a preset journal entry prediction model to obtain a set of predicted accounting subject labels. The verification module is used to obtain the actual accounting entries after the repayment transaction occurs, compare the actual accounting entries with the predicted accounting subject label set, and obtain the accounting test verification results.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.