Structured business record generation and feature calculation system based on domain adaptive model
By using a domain-adaptive model and dynamic templates based on the Transformer architecture, the problems of inconsistent recording standards, missing information, and data distortion in financial business scenarios are solved, achieving an efficient adaptive business decision-making closed loop and improving business processing efficiency and data quality.
Patent Information
- Application Number
- CN202511392789.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-27
- Publication Date
- 2025-12-23
AI Technical Summary
Existing technologies in financial business scenarios suffer from problems such as inconsistent recording standards, missing information, data distortion, and inability to conduct quantitative analysis. Furthermore, they lack dynamic adaptability and cannot build an end-to-end adaptive business decision-making closed loop.
By adopting a domain-adaptive model based on the Transformer architecture, combined with dynamic templates and data security mechanisms, the system achieves automated generation of structured business records and high-quality feature calculation, forming an adaptive closed loop of record generation, feature calculation, model evaluation, and strategy formulation.
It achieved record format consistency of ≥98%, data reuse rate increased to 95%, data tampering rate approached 0, feature quality significantly improved, business processing efficiency increased by 80%, first-time performance rate increased by 35%-45%, and overdue amount recovery rate increased by 25%-30%.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
[0001] This invention relates to the interdisciplinary field of financial technology, natural language processing (NLP), and data mining, specifically to a structured business record generation and feature calculation system based on a domain adaptive model (IPC classification: G06Q40 / 02, financial data processing; G06N3 / 04, neural network). This system addresses the pain points of manual record-keeping in business scenarios, enabling standardized generation of structured business records and high-quality feature calculation, supporting the construction of business evaluation models, and ultimately improving the business processing efficiency of financial institutions and other entities. Background Technology
[0002] In financial and other business scenarios, business records (such as performance follow-up records and objection handling records) are the core basis for tracking the status of business objects, evaluating business effectiveness, and formulating follow-up strategies. Currently, the industry mainly relies on manual recording methods, which has the following significant problems: (1) Inconsistent recording standards: Different business processing personnel have different definitions of "key information" (such as some recording "performance commitment time", while others only recording "whether there is a commitment"), resulting in chaotic recording formats and a data reuse rate of less than 50% when collaborating across teams; (2) Frequent human error: There is a lot of key information in business calls (such as performance evaluation and communication response), and the omission rate of manual real-time recording is 15%-20%, and subsequent tracing cannot accurately restore the communication scene; (3) Data authenticity is affected: Business is often outsourced to multiple service providers, and competition leads some personnel to maliciously tamper with records (such as false records of "business object promises to fulfill the contract"), with a data distortion rate of about 8%-12%; (4) Inability to quantify analysis: Manual records are mostly unstructured text, and manual feature extraction is time-consuming and the missing value rate is ≥15%, making it difficult to support an efficient business evaluation model.
[0003] Existing technologies have significant shortcomings and lack a closed-loop solution: for example, related patents only use a general BERT model + static rules to generate records, lacking dynamic adaptation capabilities between "asset-business stages," templates cannot be updated according to business scenarios, and feature calculation and model application are not involved; or they only realize business dialogue generation, without feature calculation and model application, failing to build a business closed loop of "record-feature-model-strategy"; and although they involve financial speech to structured text conversion, feature selection relies solely on TF-IDF, lacking IV / PSI stability constraints, failing to output high-quality model features, and lacking data security mechanisms. None of the above solutions address the core pain points of manual recording, nor do they achieve an end-to-end adaptive business decision-making closed loop. This solution is the first to build such a closed-loop system, filling a technological gap in the industry. Summary of the Invention
[0004] Technical problems to be solved This invention aims to solve the problems of "disorganized standards, missing information, distorted data, and inability to quantify" in existing business record systems. At the same time, it meets the dynamic adaptation requirements of different asset types and different business stages. By outputting high-quality feature vectors with IV≥0.038 and PSI≤0.05, it supports the business evaluation model to achieve AUC≥0.75. For the first time, it constructs an end-to-end adaptive business decision-making closed loop of "record generation - feature calculation - model evaluation - strategy formulation", and ultimately achieves risk stratification of business objects and improves business processing efficiency (target: increase the average daily business volume per person by ≥50% and the first-time fulfillment rate by ≥30%).
[0005] Technical solution This invention provides a structured business record generation and feature calculation system based on a domain adaptive model, the core of which is as follows: (1) Constructing an adaptive model for the business domain: Based on the Transformer architecture, pre-training and fine-tuning are performed using business scenario corpus (including key information annotations such as performance evaluation and communication response) to ensure that the key information recognition accuracy is ≥95%; (2) Design a dynamic template for “asset-business stage”: establish a template library covering 6 types of assets and 3 business stages, support hot updates (synchronous response ≤ 5 minutes, including version conflict detection and rollback), align template fields with model feature requirements, and ensure that records are connected with subsequent feature calculations; (3) Achieve automated generation of structured business records: Combining the domain adaptive model information recognition capability and dynamic templates, output standardized records, and achieve a data tampering rate close to 0 (≤0.1%) through encrypted storage + verification code anti-tampering + real-time monitoring. (4) Construct a targeted feature calculation process: Generate high-quality feature vectors through feature cleaning (outlier removal rate ≤ 5%), transformation (normalization error ≤ 3%), and screening (IV ≥ 0.038 + PSI ≤ 0.05); (5) Improve data security: Adopt de-identification + national cryptographic standard encryption + RBAC permissions + tamper monitoring (based on the real-time comparison mechanism of dynamic verification code, combined with operation log analysis, alarm accuracy rate ≥95%), which complies with the Personal Information Protection Law and ensures the data security of business objects; (6) Forming an end-to-end adaptive closed loop: Feature vectors support the business evaluation model and output a risk score of 300-900 points. The risk score is directly fed back to the dynamic adaptation module to optimize the subsequent business processing strategy template and guide differentiated business strategies (high risk: high frequency follow-up + legal notification; low risk: low frequency reminder) to achieve adaptive decision-making.
[0006] Beneficial effects Compared with the prior art, the beneficial effects of the present invention are as follows: (1) Unified recording standard: Dynamic template + domain adaptive model automatically generated, recording format consistency ≥98%, data reuse rate increased to over 95%; (2) Reduce information omissions: The model automatically identifies key information such as performance evaluation and communication response, reducing the omission rate to below 2%, which is 80% lower than manual recording; (3) Ensure data authenticity: By using encrypted storage + verification code anti-tampering + real-time monitoring (based on dynamic verification code comparison), the data tampering rate is close to 0 (≤0.1%), while data security is ensured by encryption + desensitization; (4) Improve feature quality: The core feature IV value is 60%+ higher than the industry benchmark, PSI is 0.0027-0.025, the feature missing value rate is ≤3%, and the supporting model AUC reaches 0.78 (20%+ higher than the existing technology). (5) Achieve precise stratification: Based on a risk score of 300-900, the business objects of high risk (300-500 points, default rate ≥90%), medium risk (500-700 points, default rate 80%-90%), and low risk (700-900 points, default rate ≤70%) are clearly distinguished, and the stratification accuracy rate is ≥80%; (6) Improve business efficiency: After the application by financial institutions, the average daily business volume per person increased by 80%+, the first-time performance rate increased by 35%-45%, the recovery rate of overdue amount increased by 25%-30%, and the business processing cycle was shortened by 20%-25%; (7) Strong dynamic adaptability: The template library supports the adaptation period of new asset types ≤ 24 hours, without the need to reconstruct the system, and its adaptability is better than the existing static template solution; (8) Pioneering Adaptive Closed Loop: For the first time, an end-to-end adaptive business decision-making closed loop of "record-feature-model-strategy-template optimization" is constructed to solve the problem of fragmentation of existing technologies. Detailed Implementation
[0007] The following diagram is illustrated in conjunction with the attached figures (Figure 1: System Architecture Diagram; Figure 2: Domain Adaptive Model Training Flowchart; etc.). Figure 3 Figure 4 shows a comparison chart of feature IV / PSI; Figure 5 shows a comparison chart of business object risk score distribution and default rate trend; Figure 6 shows a comparison chart of business efficiency improvement of financial institutions. The specific implementation of the present invention is described in detail below.
[0008] Specific implementation of each module
[0009] Integration logic: It integrates with mainstream business telephone systems through a standardized API interface, and receives voice files (in WAV / MP3 format) within 10 seconds after the call ends. Transcription processing: Call the business scenario speech recognition service, load the business-specific vocabulary library (including terms such as performance evaluation and communication response), and achieve a transcription accuracy of ≥95%; remove invalid interjections such as "um" and "ah" through the rule engine (residual rate ≤1%) and correct recognition errors (error correction accuracy ≥98%). Data output: The standardized transcribed text is associated with business metadata (asset type, business stage, business handler ID, business object ID) and output to the domain adaptive model module, with a data transmission success rate of 100%.
[0010] Domain Adaptive Model Module Referring to Figure 2, the model training process is as follows: Step 1: Data preparation: Collect speech-to-text transcripts of 6 asset types and 3 business stages (covering millions of call durations), business rule documents, and annotated structured business records (annotation accuracy ≥ 99%). Step 2: Data preprocessing: Clean irrelevant content from the transcribed text (cleaning rate ≥ 99%), standardize business terms (consistency ≥ 98%), and construct training samples in an "input-output" format (80% training set, 10% validation set, and 10% test set). Step 3: Pre-training: Based on the Transformer base model, use unlabeled business domain corpus for domain adaptation pre-training to optimize the model's ability to understand the language of the target business scenario; Step 4: Fine-tuning training: Use labeled samples for supervised learning fine-tuning, iteratively optimizing model parameters until the key information recognition accuracy on the validation set is ≥95% and the structured output consistency is ≥98%; Step 5: Deployment and Iteration: Deploy the model as an API service to support high-concurrency calls (response latency ≤ 1 second); use newly added annotation records for incremental fine-tuning every month to ensure that the model performance degradation rate is ≤ 1% / month.
[0011] Dynamic Adaptation Module Template library construction: Define differentiated templates for different asset-business phase scenarios. The templates include basic information fields, business information fields (including performance evaluation and communication response related indicators), and plan fields. Field attributes include name, data type, required fields, and value range. Template hot update: Administrators can add / modify templates through the web interface. The system automatically generates update commands and synchronizes them to the domain adaptive model module via the interface. The synchronization response time is ≤5 minutes, and no system restart is required. Version conflicts are automatically detected during the synchronization process, and historical versions are backed up, with a rollback success rate of ≥99%. Strategy optimization reception: Receives risk score data output by the business assessment model and automatically adjusts the template field weights for the corresponding asset-business stage (e.g., adding a "legal notice" field for high-risk business objects) to achieve adaptive template optimization; Template matching: After receiving business metadata, template matching is completed within 100ms, and structured generation instructions are output to the domain adaptive model module.
[0012] Structured business record generation module Information integration: Receive key information from the domain adaptive model output (such as "performance behavior assessment: high, communication response time: 80 seconds, overdue days: 25 days") and align it with dynamic template fields (matching accuracy ≥ 99%). Standardized format: Generate structured business records in the format of "Record ID + Basic Information + Business Information + Plan Field". Example: Record ID Asset types Business Phase Name of the business target (anonymized) Performance evaluation Communication response time (seconds) Overdue days 123e4567-e89b credit card Early open* high 80 25 Anti-tampering measures: When generating structured business records, a dynamic verification code is generated synchronously (bound to the record content in real time), and the data is encrypted through a data security module during storage, making tampering detectable; Data storage: The encrypted structured business records are written to a relational database with a data write success rate of ≥99.99% and a storage latency of ≤1 second.
[0013] Historical Business Records Summary Module Summary configuration: Administrators configure summary dimensions (such as "Business Object ID + Month" or "Asset Type + Business Stage + Risk Score Range") through the web interface. The configuration takes effect in ≤5 minutes. Aggregation calculation: Data aggregation is performed via a scheduled task (executed daily at midnight). A distributed computing framework is used to process the data. The aggregation time for large-scale, multi-scenario samples is ≤10 minutes, and the statistical error of the aggregated data is ≤5%. Report generation: The summarized results are stored in a distributed data storage architecture, supporting the viewing of visual reports (such as pie charts of performance behavior assessment distribution, communication response time trend charts, and risk score stratification percentage charts), with a report loading time of ≤3 seconds.
[0014] Feature Calculation Module Referring to Figure 3, the feature extraction process is as follows: Data input: Reads encrypted structured business records (decrypted by the data security module) and summary reports, with a data read throughput of ≥5000 records / second; Feature cleaning: The 3σ principle is used to handle continuous feature outliers (removal rate ≤ 5%), and the mode is used to fill discrete feature missing values (filling accuracy ≥ 95%). Feature transformation: Continuous features are normalized (error ≤ 3%), discrete features are encoded, and derived features are calculated based on business logic (error ≤ 1%). Feature selection: Features were selected based on IV value (≥0.038) and PSI value (≤0.05), and 9 core features were retained. Among them, the IV value of key features was significantly higher than the selection threshold (≥0.038). Feature output: Features are integrated into a vector format and output to the business assessment model platform via API interface. Before output, the data is verified by the data security module to ensure data integrity, and the output latency is ≤500ms. At the same time, a risk score of 300-900 points is generated (scoring error ≤5 points).
[0015] Data security module De-identification processing: Sensitive information such as ID card numbers and mobile phone numbers of business users are de-identified (the first 6 digits and the last 4 digits of the ID card number are retained, and the first 3 digits and the last 4 digits of the mobile phone number are retained). The de-identification algorithm complies with the requirements of the Personal Information Protection Law, and the de-identified data cannot be reversed. Encrypted storage: Encryption algorithms conforming to national cryptographic standards (such as SM4 algorithm) are used to encrypt voice data, structured business records, and feature vectors. The keys are updated regularly through a key management system (KMS) (cycle ≤ 90 days) to prevent key leakage. Access control: Configure 3 types of role permissions (business processing personnel can only view the anonymized structured business records and corresponding risk scores under their responsibility, administrators can configure templates and view the full summary data, and model personnel can only obtain feature vectors). The permission verification is carried out through a dual mechanism of identity authentication and operation authorization, with a pass rate of 100%. Data access logs are retained for ≥6 months. Tamper monitoring: Based on a real-time comparison mechanism of dynamic check codes (automatically verifying the consistency between check codes and content when reading records), combined with operation log analysis (recording the account, time, IP address, and operation content of all data access and modification operations), it detects modification behavior of structured business records and feature data in real time, triggering alarms with an accuracy rate of ≥95% and a tampering behavior interception rate of ≥99%, ensuring that the data tampering rate approaches 0 (≤0.1%).
[0016] System test results
[0017] Test subjects: A large-scale, multi-scenario business sample of a financial institution (covering 2 types of assets and 3 business stages); Test results (see attached Figure 4): Test metrics System effect Industry standard effect Increase Core Feature IV Value Level 60%+ improvement over the benchmark benchmark value 60%+↑ Characteristic PSI value range 0.0027-0.025 0.015-0.07 45%-65%↓ Model AUC value 0.78 ≤0.70 11%↑ Model KS value 0.42 ≤0.35 20%↑ Default rate in high-risk areas ≥95% ≤90% 5%↑ Default rate in low-risk areas ≤65% ≥75% 13%↓ Data tampering rate ≤0.1% ≥1% 90%↓ Application effects in financial institutions Two different types of financial institutions were selected for a three-month pilot application, with a test sample of 100,000 overdue business cases. The application results (see Appendix 5) are as follows: 1. A consumer finance institution (Assets: Consumer finance / Auto finance; Business stage: Early to mid-stage)
[0018] Business efficiency indicators Before application After application Increase Average daily processing volume per person benchmark value Baseline value × 1.8 80%↑ First-time compliance rate 22.5% 35.8% 59%↑ Overdue amount recovery rate (30 days) 38.2% 58.5% 53%↑ Business processing cycle 28 days 21 days 25%↓ Template optimization response time - ≤1 hour -
[0019] Business efficiency indicators Before application After application Increase Average daily processing volume per person benchmark value Baseline value × 1.6 60%↑ First-time compliance rate 18.3% 28.5% 56%↑ Overdue amount recovery rate (30 days) 32.1% 48.6% 51%↑ Business processing cycle 35 days 28 days 20%↓ Template optimization response time - ≤1.5 hours - Attached Figure Description Figure 1: System architecture diagram, showing the connection relationship and data flow of 7 modules (business voice acquisition, domain adaptive model, dynamic adaptation, structured business record generation, historical business record summary, feature calculation, and data security), and marking the closed-loop link of "risk score feedback to dynamic adaptation module"; Figure 2: Flowchart of domain adaptive model training, showing the entire process of "data preparation → preprocessing → pretraining → fine-tuning → evaluation → deployment", with the core performance target marked (accuracy of key information identification ≥95%). Figure 3: Feature IV / PSI comparison chart. The left chart is a bar chart of core feature IV values (key features are significantly higher than the threshold), and the right chart is a line chart of feature PSI values (all ≤0.025). Figure 4: Distribution of risk scores and trend of default rate for business targets. The left figure shows the sample distribution of the scoring interval, and the right figure shows the decreasing trend of default rate as the score increases. Figure 5: Comparison of business efficiency improvements in financial institutions. The left figure shows the comparison of the multiple of increase in the average daily business volume per person, and the right figure shows the comparison of the recovery rate of overdue amounts.
Claims
1. A structured business record generation and feature calculation system based on a domain adaptive model, characterized in that, include: The business voice acquisition module is used to collect voice data from business calls and convert the voice data into text data. The domain adaptive model module is based on the Transformer architecture and is pre-trained and fine-tuned using business domain corpus (including speech-to-text of different asset types and different business stages, business rule documents, and annotated business records). It has the ability to identify key information and output structured data in the target business scenario. The dynamic adaptation module has a built-in dual-dimensional configuration template library of "asset type - business stage" (covering asset types such as credit cards, auto finance, consumer finance, agricultural loans, medical aesthetics loans, and education loans, as well as early, mid, and late business stages). It supports a hot template update mechanism (automatically synchronized to the model module after adding / modifying a template). It is used to call the matching record template according to the asset type and business stage of the business to be processed, and output structured generation instructions to the domain adaptive model module. The structured business record generation module receives key information (including business object identity information, business overdue information, core communication content, performance commitment, objection type, and next processing plan) output by the domain adaptive model module, and generates standardized structured business records by combining the matching template of the dynamic adaptation module. The historical business record summary module supports aggregation and statistics of the structured business records by dimensions such as business object ID, business period (day / week / month), and asset type, and generates a historical business behavior summary report. The feature calculation module transforms the information in the structured business records and summary reports into quantitative features, coded features, and derived features, and outputs a feature vector for constructing a business evaluation model. The feature vector includes core features such as performance behavior evaluation indicators and communication response characteristics, wherein the core features have an IV ≥ 0.038, PSI ≤ 0.05, overall feature PSI ≤ 0.05, and an average correlation ≥ 0.6 with the business target variables (performance willingness / performance rate). The data security module desensitizes sensitive information of business objects, uses encryption algorithms that comply with national cryptographic standards to store data, restricts the scope of data access based on role-based access control (RBAC), monitors malicious tampering behavior in real time, and triggers alarms with an accuracy rate of ≥95%, which complies with the data security requirements of the Personal Information Protection Law.
2. The system according to claim 1, characterized in that, The training process of the domain adaptive model module includes: Step 1: Data preprocessing, including noise cleaning (removing irrelevant speech-to-text content), terminology standardization (unifying business expressions such as "overdue days" and "performance commitment"), and annotation enhancement (semi-supervised learning to supplement key information annotations for unannotated data). Step 2: Pre-training. Based on the Transformer base model, domain-adaptive pre-training is performed using unlabeled business domain corpus to optimize the model's ability to understand the language of the target business scenario. Step 3: Fine-tune the training. Use annotated structured business records as training samples (input: speech-to-text + asset type + business stage; output: standardized structured business records). Optimize the model parameters through supervised learning. After iterative optimization, the model's key information recognition accuracy is ≥95%, and the consistency of structured output is ≥98%.
3. The system according to claim 1, characterized in that, The template hot update mechanism of the dynamic adaptation module is as follows: When a new asset type (such as supply chain finance loan) is added or the business stage definition is adjusted, the administrator can add / modify template fields (including field name, data type, whether it is required, and value range) through the visual configuration interface. The template update command is automatically synchronized to the domain adaptive model module through the interface, with a synchronization response time of ≤5 minutes and no need to restart the system. The hot update mechanism includes a version conflict detection function, which automatically backs up historical versions when the template is updated, with a rollback success rate of ≥99%.
4. The system according to claim 1, characterized in that, The structured business record generation module outputs records containing the following core fields: Basic information fields: Record ID, Business Processing Personnel ID, Business Object ID (anonymized), Asset Type, Business Stage, Call Time, Call Duration; Business information fields: Overdue Days, Amount Due This Period, Amount Paid, Remaining Amount, Performance Commitment (Commitment Time, Commitment Amount, Performance Status), Objection Type (Difficulty in Cash Flow / Objection to Statement / No Willingness to Perform, etc.), Communication Attitude (Active Cooperation / Neutral / Resistant), Performance Behavior Evaluation Indicators (High / Medium / Low), Communication Response Time (seconds); Planning fields: Next Contact Time, Next Business Processing Strategy (Reminder / Negotiation / Legal Notification, etc.).
5. The system according to claim 1, characterized in that, The aggregation and statistical functions of the historical business record summary module include: Business object dimension summary: cumulative overdue duration, cumulative number of performance commitments and performance rate, historical objection type distribution, cumulative number of calls and total duration, and historical rating distribution of performance behavior evaluation indicators for a specific business object; Business dimension summary: average overdue days, overall performance commitment performance rate, percentage of top 3 objection types, average call duration, and average communication response time for a specific asset type / business stage. The statistical error of the summary data is ≤5%.
6. The system according to claim 1, characterized in that, The features output by the feature calculation module include continuous features, discrete coded features, and derived features, and these features are directly connected to the business evaluation model to form an end-to-end decision-making closed loop. The continuous features include: overdue days (numerical value), amount due this period (numerical value), call duration (seconds), promised amount to be fulfilled (numerical value), historical fulfillment rate (percentage numerical value), and communication response time (seconds). These continuous features undergo outlier processing and normalization transformation, with a processing error ≤3%. The discrete coding features include: asset type (credit card = 1, auto finance = 2, consumer finance = 3, farmer loan = 4, medical aesthetics loan = 5, education loan = 6), business stage (early stage = 1, mid-stage = 2, late stage = 3), communication attitude (positive cooperation = 1, neutral = 2, resistant = 3), objection type (difficulty in cash flow = 1, statement objection = 2, no willingness to fulfill = 3), performance evaluation indicators (high = 3, medium = 2, low = 1), cooperation tendency (yes = 1, no = 0), and perfunctory performance (yes = 1, no = 0). Derivative features: The ratio of overdue days to amount due, the ratio of promised performance amount to amount due, the number of business transactions in the past 30 days, communication quality score, and problem-solving ability score, calculated based on business logic. The correlation between the derived features and the performance rate of the business object is ≥0.
5.
7. The system according to claim 1, characterized in that, The data security module specifically includes: a de-identification processing unit: de-identifying the ID card number (retaining the first 6 digits + last 4 digits) and mobile phone number (retaining the first 3 digits + last 4 digits) of the business object, ensuring that the de-identified data cannot be reversed; an encrypted storage unit: using an encryption algorithm conforming to national cryptographic standards (such as the SM4 algorithm) to store and encrypt voice data and structured business records, with the key updated periodically (update cycle ≤ 90 days); an access control unit: configuring 3 types of role permissions (business processing personnel can only view the structured business records they are responsible for, administrators configure templates and view summary data, and model personnel can only obtain feature vectors), with a 100% pass rate for access control verification and data access log retention for ≥ 6 months; and a tampering monitoring unit: based on a real-time comparison mechanism using dynamic verification codes, combined with operation log analysis, detecting modification behavior of structured business records and feature data in real time, with an alarm accuracy rate ≥ 95% and a tampering interception rate ≥ 99%.
8. The system according to claim 1, characterized in that, The feature vectors output by the feature calculation module can be directly used to train the business evaluation model. The model performance indicators reach AUC≥0.75 and KS≥0.4, and it supports generating business object risk scores of 300-900 points based on the feature vectors. The default rate shows a monotonically decreasing trend within the score range (default rate ≥90% in the high-risk range and ≤70% in the low-risk range).
9. The system according to claim 1, characterized in that, The business voice acquisition module supports integration with mainstream business telephone systems, receives voice files through a standardized API interface, removes invalid interjections and corrects recognition errors during the voice-to-text process, and achieves a transcription accuracy of ≥95%.