Contract pre-approval method, equipment and medium
The contract is preprocessed and pre-approved through the DQN network model, which solves the problems of lax and low efficiency caused by relying on manual experience in the existing technology, and achieves more efficient and accurate contract review.
Patent Information
- Application Number
- CN202510375917.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-11
AI Technical Summary
The existing contract management system relies on manual experience during auditing, which leads to lax audits and low efficiency, and is prone to errors.
The DQN network model is used to preprocess the contract data, generate pre-approval opinions, and combine manual approval to reduce manual intervention.
It improves the accuracy and efficiency of contract audits, reduces the time of manual intervention, and enhances the scientific nature of audits.
Smart Images

Figure CN120297905A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of contract management, and specifically relates to a contract pre-approval method, device, and medium. Background Art
[0002] In modern enterprises, there are a large number of contracts that need to be managed. A contract management system, that is, contract full life cycle management software, covers all key links of contract management, ensuring that the entire process of a contract from drafting to archiving is effectively managed and controlled. Through automated processes, manual intervention is reduced, and the efficiency and accuracy of contract management are improved. It can be integrated with other management systems of the enterprise (such as ERP, CRM, etc.) to achieve data sharing and process collaboration.
[0003] In the prior art, although the contract management system can reduce manual intervention, when reviewing a contract, staff still need to spend a lot of time determining the specific content of the contract. At the same time, the review process overly relies on the experience and on-the-spot judgment of the staff. If the staff have insufficient experience or make mistakes in judgment, it is easy to cause adverse effects such as lax review. Summary of the Invention
[0004] To solve the above problems, this application proposes a contract pre-approval method, device, and medium. The method includes:
[0005] Receiving a contract to be approved provided by a contract applicant; performing data preprocessing on the contract to be approved to obtain structured data; using the structured data as the input of a pre-trained DQN network model to obtain a pre-approval opinion on the contract to be approved; and outputting the pre-approval opinion for reference in manual approval.
[0006] In one example, the performing data preprocessing on the contract to be approved to obtain structured data specifically includes: extracting key field information of the contract to be approved and converting the key field information into first structured data, where the key field information includes at least one of contract type, signing date, performing unit, counterparty unit, and contract amount; extracting contract text information of the contract to be approved and converting the contract text information into second structured data; obtaining contract applicant information of the contract to be approved and historical approval records of the contract applicant within a preset time period; and converting the contract applicant information and the historical approval records into third structured data.
[0007] In one example, before using the structured data as the input of the pre-trained DQN network model, the method further includes: determining an untrained initial DQN network model; initializing the model parameters of the initial DQN network model, where the model parameters include main network parameters, target network parameters, Q-table parameters, memory bank parameters, and hyperparameters; the Q-table is used to store states and actions, the memory bank is used to store experience data, and the hyperparameters include at least one of a learning rate, a discount factor, and an exploration rate; after inputting sample data into the initial DQN network model, based on a preset action selection strategy, storing the interaction experiences of the sample data in different environments in the memory bank; the interaction experiences include the current state, the selected action, the resulting reward, and the new state; randomly selecting interaction experiences from the memory bank, and training the initial DQN network model based on a preset reward function and a loss function until the model converges to obtain the DQN network model.
[0008] In one example, the reward value of the preset reward function is related to the total approval time for the approval user to enter the approval interface and the approval results of each approval node; the preset action selection strategy is to select a random action with a first probability and select the action with the highest current Q value with a second probability; the sum of the first probability and the second probability is 1.
[0009] In one example, the loss function is the mean square error between the target Q value output by the target network and the actual Q value output by the main network; the loss function is used to update the parameters of the main network to minimize the loss function.
[0010] In one example, outputting the pre-approval opinion specifically includes: the pre-approval opinion includes at least one of approval, rejection, supplementary information, and return for modification; if the pre-approval opinion is return for modification, after the contract applicant modifies the contract to be approved, pre-approving the contract to be approved again; if the approval opinion is approval, rejection, or supplementary information, displaying the approval opinion on the pre-approval decision reference interface.
[0011] In one example, after outputting the pre-approval opinion for manual approval reference, the method further includes: obtaining the manual approval result and feedback information corresponding to the contract to be approved; based on the manual approval result and the feedback information, performing backpropagation training on the DQN network model.
[0012] In one example, using the structured data as the input of the pre-trained DQN network model specifically includes: determining historical approved contracts in the historical approval database based on the contract applicant and application time of the contract to be approved; respectively determining the similarity of the input data between the contract to be approved and the historical approved contracts; determining similar historical approved contracts whose similarity of the input data with the contract to be approved is higher than a preset threshold; and jointly inputting the approval data corresponding to the similar historical approved contracts and the structured data into the DQN network model.
[0013] The present application also provides a contract pre-approval device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to: receive a contract to be approved provided by a contract applicant; perform data preprocessing on the contract to be approved to obtain structured data; use the structured data as the input of the pre-trained DQN network model to obtain a pre-approval opinion for the contract to be approved; and output the pre-approval opinion for reference in manual approval.
[0014] The present application also provides a non-volatile computer storage medium storing computer-executable instructions, which are set to: receive a contract to be approved provided by a contract applicant; perform data preprocessing on the contract to be approved to obtain structured data; use the structured data as the input of the pre-trained DQN network model to obtain a pre-approval opinion for the contract to be approved; and output the pre-approval opinion for reference in manual approval.
[0015] The method proposed by the present application can bring the following beneficial effects: By providing a contract pre-approval mechanism, it can help staff quickly grasp the main content of the contract and give approval suggestions, reducing the approval time and effort of the staff. At the same time, by using the DQN algorithm model, the approval opinion of the contract can be closer to the actual situation, thereby increasing the accuracy of pre-approval. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0017] Figure 1 is a schematic flowchart of a contract pre-approval method in an embodiment of the present application;
[0018] Figure 2This is a schematic structural diagram of a contract pre - approval device in an embodiment of the present application. Detailed implementation manners
[0019] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0020] Before contract approval, a preliminary review of the applied contract can quickly identify potential problems or non - compliant areas. In this way, unnecessary delays can be avoided in the formal approval stage, thereby improving the efficiency of the entire approval process. In the pre - approval stage, a preliminary assessment can be made of the applicant's credit status, qualification conditions, the legality of contract terms, etc. This helps to identify potential risk points before formal approval, so as to take corresponding risk management measures and reduce the losses that may be brought about by approval mistakes. Through pre - approval, a preliminary screening of the applicant or the contract can be carried out to ensure that only the applications or contracts that meet the requirements enter the formal approval process. This helps to optimize the allocation of approval resources, enabling the limited approval resources to serve the applicants or contracts that truly meet the requirements more efficiently. Pre - approval can shorten the time for contract applicants to wait for approval results and improve their satisfaction. At the same time, the feedback and guidance provided during the pre - approval process also help applicants or contract submitters better understand the approval requirements and standards, thereby improving the quality of their applications.
[0021] The following will, with reference to the drawings, detail the technical solutions provided by each embodiment of the present application.
[0022] Figure 1 This is a schematic flowchart of a contract pre - approval method provided for one or more embodiments of this specification. This method can be applied to different business fields, such as the Internet finance business field, e - commerce business field, instant messaging business field, game business field, official business field, etc. This process can be executed by computing devices in the corresponding fields (such as the risk control server corresponding to payment services or intelligent mobile terminals, etc.), and some input parameters or intermediate results in the process allow manual intervention and adjustment to help improve accuracy.
[0023] The implementation of the analysis method involved in the embodiments of the present application can be a terminal device or a server, and the present application does not make special restrictions on this. For the convenience of understanding and description, the following embodiments will be described in detail taking the server as an example.
[0024] It should be noted that the server can be a single device or a system composed of multiple devices, that is, a distributed server. This application does not make specific limitations on this.
[0025] As Figure 1 shown, an embodiment of the present application provides a contract pre-approval method, including:
[0026] S101: Receive a contract to be approved provided by a contract applicant.
[0027] After the contract applicant submits the contract for approval, the contract enters the contract pre-approval stage. At this time, the server can receive the contract to be approved. It should be noted that the file format of the contract to be approved here can be a text type, a table type, or a picture type. For different types of contracts to be approved, different processing is required. For example, the text content of the contract to be approved in picture type needs to be extracted to obtain the clause information corresponding to the contract to be approved.
[0028] S102: Perform data preprocessing on the contract to be approved to obtain structured data.
[0029] In this stage, the server obtains information related to the contract as the input of the pre-approval model. Before inputting the data into the pre-approval model, it is necessary to perform data preprocessing on the relevant information to convert the contract to be approved into a data type that the model can recognize.
[0030] In one embodiment, when performing data preprocessing, first select the key field information of the contract, such as contract type, signing date, performing unit, counterparty unit, contract amount, etc., and convert it into structured data as the input of the Deep Q Network algorithm. In addition, the contract text information and contract clause content can use methods such as NLP natural language processing to extract the text information and use it as part of the input of the intelligent agent. Since there is a certain inertia when the contract applicant makes the contract form, in order to improve the accuracy of the pre-approval system, in the intelligent pre-approval, the information of the approval requester is added as a supplement, including the information of the applicant, such as the user id of the applicant. The historical approval records of the applicant. Since the contract form-making level of the contract applicant will change due to the progress habit, in order to ensure accuracy, only the historical approval records of the approver in the most recent month are taken.
[0031] S103: Use the structured data as the input of the pre-trained DQN network model to obtain the pre-approval opinion of the contract to be approved.
[0032] After taking the structured data as the input of the DQN network model, the DQN network model can output the approval opinion on the contract to be approved and enter the subsequent manual approval process stage. The DQN network model here is a pre-trained model. During its training process, to ensure accuracy, it will be repeatedly trained until the stopping condition is reached, such as reaching the maximum number of training steps or learning convergence.
[0033] Now, the DQN network model will be described: The Deep Q Network algorithm is a deep reinforcement learning algorithm. It includes three key elements: the action Action of the agent Agent, the state State of the environment Environment, and the environmental reward Reward. The reinforcement learning algorithm is that the agent Agent makes an action Action in the environment Environment according to the state State of the environment, which acts on the environment and changes its own state. The environment gives the agent a reward according to the quality of the agent's action. In reinforcement learning, the goal of the agent is to maximize the total sum of rewards obtained, that is, the return Return. Its specific value can be simply determined by the following calculation formula:
[0034]
[0035] Among them, R t is the return, r t+1 is the reward given by the environment for the action taken in the (t + 1)-th time period, n is the total number of time periods, and γ is the discount factor, indicating the preference for immediate rewards.
[0036] The Q-learning algorithm is the most basic reinforcement learning algorithm, which is derived from the Bellman equation. The Bellman equation is:
[0037] Q(s,a) = R(s,a) + γ∑ s‘∈s P(s′|s,a)max a‘∈A Q(s′,a′)
[0038] Among them, R(s,a) is the immediate reward obtained after taking the action a in the state s, P(s ′ |s,a) is the probability of transferring from the state s to the state s ′ by taking the action a, γ is the discount factor, which is used to balance the importance of the current reward and future rewards, and max a‘∈A Q(s′,a′) represents the Q value that can be obtained by taking the optimal action in the state s′. The goal of the Q-learning algorithm is to find a policy that selects the optimal action in each state, that is, to maximize the Q function. However, in practical applications, we often do not know the state transition probability P and the reward function R tThe exact form, so it is necessary to learn the Q-function by interacting with the environment. First, we need to initialize a Q-table, where each element Q(s,a) is set to an initial value (usually 0). Then, the agent iteratively updates the Q-table by interacting with the environment. In each iteration, the agent starts from a certain initial state, selects an action according to the current Q-table and a certain policy (such as the ε-greedy policy), observes the new state and the obtained reward after executing the action, and then updates the Q-table according to the approximate form of the Bellman equation. The approximate form of the Bellman equation is usually expressed in Q-learning as:
[0039] Q(s,a)←Q(s,a)+α[R(s,a)+γmax a‘∈A Q(s′,a′)-Q(s,a)]
[0040] where α is the learning rate, which is used to control the speed at which new information overrides old information. This update process can be understood as follows: After the agent takes action a in state s, it obtains an immediate reward R(s,a) and a new state s ′ . Then, it predicts the Q-value that can be obtained by taking the optimal action in state s ′ according to the Q-table, adds it to the immediate reward, and obtains an estimated cumulative discounted reward. Finally, it uses this estimated value to update the original Q(s,a) value. The agent iteratively updates the Q-table by continuously repeating the above process until the Q-table converges or reaches a preset number of iterations. The converged Q-table contains the estimated values of the cumulative discounted rewards that can be obtained by selecting the optimal action in each state. To solve the infeasibility of the Q-table in high-dimensional spaces, the method of value function approximation is introduced. Value function approximation uses a function to estimate the Q-value instead of storing the Q-values of each state-action pair. These functions can be linear, polynomial, or more complex machine learning models, such as neural networks.
[0041] Neural network design: Design a main network whose input is the state s and whose output is the Q-value corresponding to each action a (i.e., Q(s,a)). The parameters of the neural network are updated through the backpropagation algorithm to minimize the loss function. Definition of the loss function: The loss function is usually defined as the mean squared error between the predicted Q-value and the actual Q-value. The actual Q-value can be calculated by adding the current reward r to the maximum discounted Q-value in the future (for non-terminal states), and its formula is:
[0042] L=[r+γmax a‘∈A Q(s′,a′;θ′)-Q(s,a)] 2
[0043] where θ are the parameters of the current Q-network and θ′ are the parameters of the target Q-network. To break the correlation between data and improve the sample utilization efficiency, the DQN algorithm introduces an experience replay mechanism. The experiences obtained by the agent interacting with the environment (i.e., the state transition quadruple s t 、a t 、r t 、s t+1 ,s t+1 i.e., s ′ ) are stored in a fixed-size memory bank. During training, a batch of experiences is randomly sampled from the memory bank to update the Q-network.
[0044] Target Network: To stabilize the training process, two neural networks with the same structure are used: the current Q-network and the target Q-network. The current Q-network is used to calculate the predicted Q-values, while the target Q-network is used to calculate the maximum future Q-value in the actual Q-values. The parameters of the target network are periodically copied from the current Q-network to maintain its stability.
[0045] In one embodiment, for the DQN network model provided in the present application, its training process includes: determining an untrained initial DQN network model, initializing the model parameters of the initial DQN network model, where the model parameters include main network parameters, target network parameters, Q-table parameters, memory bank parameters, and hyperparameters; the Q-table is used to store states and actions, the memory bank is used to store experience data, and the hyperparameters include at least one of a learning rate, a discount factor, and an exploration rate;
[0046] At the same time, for each evaluation process, the state s needs to be initialized. For each time step t: select an action a according to the current Q-network and a preset policy, where the preset policy can be an ε-greedy policy. Then execute the action a, observe the reward r and the next state s ′ . Store the experience s t 、a t 、r t 、s t+1 ,s t+1 in the memory bank. Randomly sample a batch of experiences from the memory bank to update the parameters of the current Q-network. Every once in a while, copy the parameters of the current Q-network to the target Q-network. Train the initial DQN network model based on a preset reward function and a loss function until the model converges to obtain the DQN network model. This method helps to break the correlation between data and improve the stability of training. Here, the preset action selection policy is specifically described. When selecting an action, a random action is selected with a first probability and the action with the highest current Q-value is selected with a second probability; where the sum of the first probability and the second probability is one.
[0047] Calculate the target Q-value using the target network, which can be expressed as
[0048] Q target (s,a) = R + γmax a‘∈A Q(s′,a′)
[0049] Where s’ is the next state, R is the reward, and γ is the discount factor. When calculating the current Q value and loss, the main network can be used to calculate the current Q value Q(s,a). The loss function is calculated, usually the mean squared error (MSE) between the predicted Q value and the target Q value. Backpropagation and weight update: Use the gradient of the loss function to update the parameters of the main network to minimize the loss function. This is usually achieved through the gradient descent algorithm. Regularly update the target network: Regularly copy the parameters of the main network to the target network to maintain the stability of the target network. Repeat training until the stopping condition is reached (such as reaching the maximum number of training steps or learning convergence).
[0050] Specifically, when training, a reasonable reward function needs to be set for the Agent, which can enable the Agent to make fast and accurate actions. The reward function is used to evaluate the quality of the pre-approval method decision. In intelligent pre-approval, the reward is designed based on factors such as the accuracy, efficiency, and user satisfaction of the approval. If the pre-approval decision is consistent with the subsequent manual decision and the approval process is fast, a higher reward is given; if the approval decision is inconsistent with the manual approval decision, a penalty is given. The reward function for approval efficiency is measured according to the time taken for approval. The shorter the time, the more help the pre-review system provides to the user. User satisfaction depends on the user's satisfaction with the system and scores the pre-approval system. Specifically, in order for the pre-approval system to give the approval user reference information, a pre-approval decision reference interface is inserted before opening the contract approval interface, and the decision and suggestions of the pre-approval system are displayed as suggestions for manual approval. In terms of the decision accuracy reward, a proportional reward function is designed according to the importance of the approval node. A higher reward is given if the approval decision is consistent with the important node, a lower reward is given if the approval decision is consistent with the less important node, and a fixed reward is given if the decision is consistent with all nodes in the entire process. In terms of decision efficiency, the total time for the approval user to enter the approval interface is counted as a reference. The shorter the time, the higher the reward. Since the approval time is affected by many factors, the overall reward is less than the rewards for accuracy and satisfaction. In terms of satisfaction reward, there is a scoring session after each approval node to determine the user satisfaction reward. Finally, the reward given to the intelligent agent is the sum of the three parts of the reward.
[0051] The DQN network model structure provided by this application is usually a deep feedforward neural network for estimating the Q-value function. Its structure mainly includes the following parts: Input Layer: Receives the feature vector of the environmental state as input. The state features are represented as a vector, and the number of neurons in the input layer is equal to the dimension of the state feature vector. Hidden Layers: DQN can include one or more hidden layers for learning the abstract representation of state features. These hidden layers usually adopt fully connected layers. Activation functions (such as ReLU) are applied to the outputs of the hidden layers to introduce non-linearity. Output Layer: The number of neurons in the output layer is equal to the dimension of the action space, and each neuron corresponds to a possible action. The value of the output layer represents the Q-value of selecting each action in the given state. The network structure of DQN may vary due to the complexity of the problem, and the number of hidden layers and neurons can be adjusted to adapt to different tasks.
[0052] S104: Output the pre-approval opinion for reference in manual approval.
[0053] In the subsequent manual review stage, the pre-approval opinion can be seen as a reference, but the final decision on the approval opinion of the contract to be approved is still made manually.
[0054] In one embodiment, it is possible that the modified contract is too similar to or exactly the same as the original contract to be approved. At this time, to reduce the computational amount, the historical approval contracts in the historical approval database can be determined based on the contract applicant and application time of the contract to be approved. Then, the similarity of the input data between the contract to be approved and the corresponding historical approval contracts is determined respectively. If the similarity is too high, it means that the contract to be approved may have been approved before, or it is a contract that is being reviewed again after modification. At this time, the similar historical approval contracts with the input data similarity higher than the preset threshold to the contract to be approved can be selected, and when input into the DQN network model, the approval data corresponding to the similar historical approval contracts and the structured data are input together, so that the model can more easily judge whether to pass the review by comparing the differences between the two contracts.
[0055] As Figure 2 shown, an embodiment of this application also provides a contract pre-approval device, including:
[0056] At least one processor; and, a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor can:
[0057] Receive the contract to be approved provided by the contract applicant; perform data preprocessing on the contract to be approved to obtain structured data; use the structured data as the input of the pre-trained DQN network model to obtain the pre-approval opinion of the contract to be approved; output the pre-approval opinion for reference in manual approval.
[0058] The embodiment of the present application also provides a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are set as follows:
[0059] Receive the contract to be approved provided by the contract applicant; perform data preprocessing on the contract to be approved to obtain structured data; use the structured data as the input of the pre-trained DQN network model to obtain the pre-approval opinion of the contract to be approved; output the pre-approval opinion for reference in manual approval.
[0060] Each embodiment in the present application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.
[0061] The device and medium provided by the embodiment of the present application correspond one by one to the method. Therefore, the device and medium also have beneficial technical effects similar to those of the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the device and medium will not be elaborated here.
[0062] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0063] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0064] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0065] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0066] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0067] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.
[0068] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0069] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0070] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A contract pre-approval method, characterized in that, It includes: Receiving a contract to be approved provided by a contract applicant; Performing data preprocessing on the contract to be approved to obtain structured data; Using the structured data as the input of a pre-trained DQN network model to obtain a pre-approval opinion on the contract to be approved; Outputting the pre-approval opinion for reference in manual approval.
2. The method according to claim 1, wherein The performing data preprocessing on the contract to be approved to obtain structured data specifically includes: Extracting key field information of the contract to be approved and converting the key field information into first structured data, where the key field information includes at least one of contract type, signing date, performing unit, counterparty unit, and contract amount; Extracting contract text information of the contract to be approved and converting the contract text information into second structured data; Obtaining contract applicant information of the contract to be approved and historical approval records of the contract applicant within a preset time period; converting the contract applicant information and the historical approval records into third structured data.
3. The method according to claim 1, characterized in that, Before using the structured data as the input of the pre-trained DQN network model, the method further includes: Determining an initial DQN network model that has not been trained; Initializing model parameters of the initial DQN network model, where the model parameters include main network parameters, target network parameters, Q-table parameters, memory bank parameters, and hyperparameters; the Q-table is used to store states and actions, the memory bank is used to store experience data, and the hyperparameters include at least one of learning rate, discount factor, and exploration rate; After inputting sample data into the initial DQN network model, storing interaction experiences of the sample data in different environments in the memory bank based on a preset action selection strategy; the interaction experiences include the current state, the selected action, the result reward, and the new state; Randomly selecting interaction experiences from the memory bank and training the initial DQN network model until the model converges based on a preset reward function and loss function to obtain the DQN network model.
4. The method according to claim 3, characterized in that, The reward value of the preset reward function is related to the total approval time for the approval user to enter the approval interface and the approval results of each approval node; The preset action selection strategy is to select a random action with a first probability and select the action with the highest current Q value with a second probability; the sum of the first probability and the second probability is 1.
5. The method according to claim 3, wherein The loss function is the mean square error between the target Q value output by the target network and the actual Q value output by the main network; The loss function is used to update the parameters of the main network to minimize the loss function.
6. The method according to claim 1, wherein The outputting the pre-approval opinion specifically includes: The pre-approval opinion includes at least one of approval passed, approval rejected, supplementary information, and returned for modification; If the pre-approval opinion is returned for modification, after the contract applicant modifies the contract to be approved, the contract to be approved is pre-approved again; If the approval opinion is approval passed, approval rejected, or supplementary information, the approval opinion is displayed on the pre-approval decision reference interface.
7. The method according to claim 1, wherein Output the pre-approval opinion for reference in manual approval. After that, the method further includes: Obtain the manual approval result and feedback information corresponding to the contract to be approved; Based on the manual approval result and the feedback information, perform backpropagation training on the DQN network model.
8. The method according to claim 1, characterized in that, The step of using the structured data as the input of the pre-trained DQN network model specifically includes: Based on the contract applicant and application time of the contract to be approved, determine the historical approved contracts in the historical approval database; Respectively determine the similarity of the input data between the contract to be approved and the historical approved contracts; Determine the similar historical approved contracts whose similarity of input data with the contract to be approved is higher than a preset threshold; Input the approval data corresponding to the similar historical approved contracts and the structured data into the DQN network model together.
9. A contract pre-approval device, characterized in that, It includes: At least one processor; And a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute: Receive a contract to be approved provided by a contract applicant; Perform data preprocessing on the contract to be approved to obtain structured data; Use the structured data as the input of the pre-trained DQN network model to obtain the pre-approval opinion of the contract to be approved; Output the pre-approval opinion for reference in manual approval.
10. A non-volatile computer storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are set to: Receive a contract to be approved provided by a contract applicant; Perform data preprocessing on the contract to be approved to obtain structured data; Use the structured data as the input of the pre-trained DQN network model to obtain the pre-approval opinion of the contract to be approved; Output the pre-approval opinion for reference in manual approval.
Citation Information
Cited By
Workflow processing method, system and equipment based on routing algorithm and medium
CN120746254A