Method for carrying out intelligent abstracting on unhealthy asset association cases
Through multimodal data processing system and blockchain storage technology, the problems of low efficiency, poor accuracy, strong discontinuity and insufficient privacy protection in the intelligent summary and risk management of non-performing asset-related cases are solved, and efficient, secure and real-time data processing and risk assessment are achieved.
Patent Information
- Application Number
- CN202510724467.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology has problems such as low efficiency, poor accuracy, strong discontinuity, and insufficient privacy protection in the intelligent summary and risk management of non-performing asset-related cases, which is difficult to meet the needs of rapid processing and real-time updates.
The multimodal data processing system is adopted, including data standardization, encryption processing, graph construction, digest generation, feedback optimization and blockchain storage modules. Through BERT semantic similarity matching, U-Net image semantic extraction, Shamir secret sharding, Paillier homomorphic encryption, dynamic relationship map and blockchain storage technology, data is automated processing and real-time update.
It improves the processing efficiency and coverage of non-performing asset-related cases, ensures data privacy and security, promptly detects guarantee chain risks, provides accurate personalized summary and risk assessment, and reduces processing costs and communication overhead.
Smart Images

Figure CN120234411A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and specifically to a method for intelligent summarization of non-performing asset related cases. Background Art
[0002] With the development of the non-performing asset management industry, especially the increasing demand for the processing of non-performing asset related cases in the financial and legal fields, accurately and timely summarizing and analyzing these cases has become an important topic for improving asset disposal efficiency and reducing risks. Non-performing asset related cases involve complex multi-modal data, including legal documents, financial statements, and contract scans, and the key information therein (such as debtors, debt amounts, collateral, etc.) directly affects the case processing results and risk assessments. However, the existing technologies still face many challenges in the intelligent summarization and risk management of non-performing asset related cases.
[0003] Currently, the processing and analysis of non-performing asset related cases mainly rely on the following technical means:
[0004] Manual review: The traditional method extracts key case information and generates summary reports by manually reading legal documents and financial statements. This method relies on the experience and subjective judgment of professionals and is widely used in the processing of small cases.
[0005] Text keyword extraction technology: Structured information (such as case numbers, debt amounts) is extracted from legal documents and financial statements through simple keyword matching or regular expressions, and preliminary reports are generated, which is suitable for standardized data processing.
[0006] Image recognition technology: Optical Character Recognition (OCR) technology is used to process contract scan images, extract text content, and combined with manual proofreading to assist in case information entry.
[0007] Database query and statistical analysis: Case data is stored in a relational database, and SQL queries and basic statistical methods are used to analyze debtor credit or collateral value to support risk assessment.
[0008] Although these technologies have played a certain role in the processing of non-performing asset related cases, there are still the following significant deficiencies:
[0009] Problem 1: Manual review is time-consuming, inefficient, and limited by personnel experience and subjectivity, making it difficult to cover large-scale case data, resulting in the omission or misidentification of key information and making it difficult to meet the requirements of rapid processing.
[0010] Problem 2: The text keyword extraction technology relies on fixed rules, making it difficult to adapt to the semantic complexity in the legal and financial fields, with low accuracy, especially performing poorly when dealing with unstructured data (such as contract terms). In addition, OCR image recognition is limited by scanning quality and font diversity, with high maintenance costs and being vulnerable to environmental impacts.
[0011] Problem 3: Existing technologies mostly rely on periodic manual reviews or batch data processing, lacking the ability to update case status in real-time and dynamically. This discontinuity makes it difficult to detect potential risks (such as changes in the credit of guarantors) in a timely manner, missing the best opportunity to optimize the disposal plan.
[0012] Problem 4: Although database queries and statistical analysis can provide basic data support, they lack advanced data analysis means (such as machine learning or graph analysis), relying only on simple thresholds or statistical indicators for risk assessment, prone to false alarms or missed reports, and unable to accurately reflect the complex relationships between cases (such as guarantee chain risks), limiting the level of intelligent management.
[0013] Therefore, a method for intelligent summarization of non-performing asset related cases is needed to solve the above problems. Summary of the Invention
[0014] Technical Problems to be Solved
[0015] Aiming at the deficiencies of the existing technology, the present invention provides a method for intelligent summarization of non-performing asset related cases, solving the problems in the above background technology.
[0016] Technical Solution
[0017] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for intelligent summarization of non-performing asset related cases, the method is executed by a multi-modal data processing system, the multi-modal data processing system includes a data standardization module, an encryption processing module, a graph construction module, a summary generation module, a feedback optimization module, a verification module, and a blockchain storage module, and the method includes the following steps:
[0018] Sp1. The data standardization module obtains multi-modal data of non-performing asset related cases, including legal document texts, financial forms, and contract scanned image files, converts the text and table data into structured data through domain adaptation template matching technology, extracts the text in the image through image semantic extraction technology, generates structured data including case numbers, debtors, debt amounts, collateral, guarantors, and case statuses, and transmits it to the encryption processing module in the form of extensible markup language;
[0019] Sp2. The encryption processing module executes an encrypted sharding algorithm on the structured data, divides it into encrypted fragments and stores them locally, calculates statistical information through a multi-party secure aggregation protocol, and transmits the encrypted data to the graph construction module;
[0020] Sp3. The graph construction module constructs a dynamic relationship graph based on the structured data and statistical information, with debtors, guarantors, and collateral as nodes and guarantee and mortgage relationships as edges. It generates a risk assessment result through a path reasoning algorithm and transmits it to the summary generation module in key-value pair format;
[0021] Sp4. The summary generation module generates an initial summary through domain template filling and semantic priority extraction techniques, including case number, debtor, debt amount, collateral, guarantor, case status, and risk points, and transmits it to the feedback optimization module in text format;
[0022] Sp5. The feedback optimization module collects feedback through the user interface, optimizes the initial summary based on the content priority adjustment algorithm, generates a personalized summary, and transmits it to the verification module;
[0023] Sp6. The verification module verifies the personalized summary through formal rules and fact consistency verification techniques, generates a verified summary, and transmits it to the blockchain storage module in encrypted text format;
[0024] Sp7. The blockchain storage module stores the verified summary and metadata in the private chain, triggers generation and update through smart contracts, and shares the summary through the distributed ledger.
[0025] Preferably, in Sp1, the domain adaptation template matching technology matches the case number, debtor, and guarantor in the legal document text data through predefined templates in the legal and financial fields. The domain adaptation template is customized based on the semantic features of non-performing asset cases, and the template selection is optimized through a semantic similarity matching algorithm. The semantic similarity matching algorithm calculates the matching degree between the text data and the template based on the legal and financial term libraries of non-performing asset cases. The image semantic extraction technology extracts key text fields from the contract scanned image data through image segmentation and semantic parsing technologies. The structured data is transmitted to the encryption processing module in extensible markup language format through the internal data bus.
[0026] Preferably, in Sp2, the encrypted sharding algorithm divides the structured data into multiple encrypted fragments and stores them in local nodes through polynomial interpolation technology. The multi-party secure aggregation protocol calculates the aggregation values of the debt amount and risk level through weighted statistical technology. The encrypted fragments and aggregation values are encrypted and transmitted to the graph construction module through the transport layer security protocol, which ensures that the aggregation calculation does not disclose the original data.
[0027] Preferably, in the Sp3, the path inference algorithm identifies the guarantee chain risk in the dynamic relationship graph through the risk propagation analysis technology. The risk propagation analysis technology dynamically adjusts the weights based on the credit scores of the guarantors and the collateral. The dynamic relationship graph updates the status of nodes and edges in real time through the event listening technology. The risk assessment result is transmitted to the summary generation module in the form of key-value pairs through the internal data bus.
[0028] Preferably, in the Sp4, the semantic priority extraction technology extracts key phrases from structured data through syntactic dependency analysis and keyword priority ranking. The keyword priority ranking is based on the risk weights of non-performing asset cases. The domain template filling technology generates an initial summary through a predefined summary template, and the template is customized according to the case type and user role. The initial summary is transmitted to the feedback optimization module in text format through the internal data bus.
[0029] Preferably, in the Sp5, the content priority adjustment algorithm dynamically adjusts the field priorities of the summary template through the scores and annotations in the user feedback data. The user feedback data is stored in the Extensible Markup Language (XML) format. The optimized personalized summary is transmitted to the user through the user interface and to the verification module through the internal data bus.
[0030] Preferably, in the Sp6, the formal rule checks the logical relationship between the debt amount and the collateral value in the personalized summary through logical reasoning technology. The fact consistency verification technology verifies the field consistency between the personalized summary and the structured data through field matching technology. The verified summary is transmitted to the blockchain storage module in encrypted text format through the internal data bus.
[0031] Preferably, in the Sp7, the smart contract controls the summary generation and update through predefined event trigger conditions, and the event trigger conditions include case status changes and user requests. The distributed ledger technology synchronizes the verified summary data through a consensus algorithm. The metadata includes the case number, generation time, and content hash generated through the Secure Hash Algorithm (SHA).
[0032] Preferably, before the multi-modal data enters the domain adaptation template matching technology in the Sp1, the data cleaning module performs format normalization operations. The format normalization operations remove the irrelevant formats in the legal document text data and standardize the numerical values in the financial form data. The cleaned data is transmitted to the subsequent operations of Sp1 in XML format through the internal data bus.
[0033] Preferably, in Sp3, the dynamic relationship graph enhances the weights of edges in the graph database through the relationship priority assignment technology. The relationship priority assignment is based on the credit scores of the guarantor and the collateral and the statistical analysis of historical case data. The risk assessment results are transmitted to the summary generation module in the form of key-value pairs through the internal data bus. In Sp1, the domain adaptation template matching technology optimizes template selection through the semantic similarity matching algorithm. The semantic similarity matching algorithm calculates the matching degree between the text data and the template based on the legal and financial term libraries of non-performing asset cases. The structured data including case number, debtor, debt amount, collateral, guarantor, and case status is transmitted to the encryption processing module in the form of Extensible Markup Language through the internal data bus. In Sp2, the multi-party secure aggregation protocol optimizes the computing efficiency through the hierarchical aggregation technology. The hierarchical aggregation technology calculates the statistical information layer by layer to reduce the communication overhead. The encrypted fragments and aggregation values are transmitted to the graph construction module after being encrypted by the Transport Layer Security protocol. In Sp4, the domain template filling technology selects the summary template according to the user role through the dynamic template selection algorithm. The dynamic template selection algorithm is based on the user's historical feedback data and case types. The initial summary is transmitted to the feedback optimization module in text format through the internal data bus.
[0034] The control flow design of the whole method is as follows:
[0035] Initialization and input verification:
[0036] When the system starts, check whether the multimodal data (legal document text, financial form, contract scanned image) is complete. If the data is missing or in the wrong format, trigger the exception handling module, record the log and prompt the user to complete the data. Otherwise, enter Sp1.
[0037] Execute the main process sequentially:
[0038] Step Sp1 (data normalization module): Receive the multimodal data, perform data cleaning, domain adaptation template matching, and image semantic extraction to generate structured data (in XML format). After completion, check the integrity of the structured data (whether fields such as case number, debtor, etc. are complete). If complete, transmit it to Sp2. Otherwise, return to the exception handling.
[0039] Step Sp2 (encryption processing module): Receive the structured data, perform secret sharding and multi-party secure aggregation to generate encrypted fragments and statistical information. After encryption, verify whether the aggregation values are consistent. If consistent, transmit it to Sp3. Otherwise, retry the aggregation process up to 3 times.
[0040] Step Sp3 (Graph Construction Module): Receive encrypted fragments and statistical information, construct a dynamic relationship graph, and generate a risk assessment result. The graph update is triggered by Kafka event listening. If the event processing fails, record the error and continue to the next step.
[0041] Step Sp4 (Abstract Generation Module): Receive the risk assessment result and generate an initial abstract. After generation, check whether the abstract contains all key fields (case number, risk points, etc.). If any are missing, trigger template re-population.
[0042] Step Sp5 (Feedback Optimization Module): Receive the initial abstract, collect user feedback, and optimize to generate a personalized abstract. After optimization, if the user satisfaction score is less than 3 (out of 5), return to Sp4 for regeneration.
[0043] Step Sp6 (Verification Module): Receive the personalized abstract and perform logical checks and consistency verification. If the verification fails (logical relationship error or similarity less than 90%), return to Sp5 for re-optimization; otherwise, proceed to Sp7.
[0044] Step Sp7 (Blockchain Storage Module): Receive the verified abstract, store it in the private chain, and synchronize the distributed ledger. If the storage fails (e.g., network interruption), retry up to 5 times and then record the log.
[0045] Exception Handling and Rollback Mechanism:
[0046] When an exception occurs in each step (such as data loss, encryption failure, verification failure), call the exception handling module, record the error details (time, module, error type) in the log file, pause the current process, and notify the user.
[0047] If three consecutive exceptions are not resolved, the system automatically rolls back to the previous step, restores to the most recent successful state, and executes again.
[0048] Termination Condition:
[0049] The process terminates after Sp7 completes blockchain storage and returns a storage success confirmation. If the user requests an update (such as a change in case status), it is triggered by a smart contract and the loop execution starts again from the Sp3 graph construction module.
[0050] Beneficial Effects
[0051] The present invention provides a method for intelligent summarization of non-performing asset related cases. It has the following beneficial effects:
[0052] 1. Through the automated processing and real-time standardization of multi-modal data, the present invention can extract the key information of non-performing asset-related cases in a timely manner, effectively avoid the inefficiency and omissions of manual review, and improve the case processing efficiency and coverage. By using BERT semantic similarity matching and U-Net image semantic extraction technologies, the accuracy rates reach 95% and 90% respectively, achieving efficient and comprehensive data extraction compared to the time-consuming and information omission problems of traditional manual review.
[0053] 2. Through Shamir secret sharing and Paillier homomorphic encryption technologies, the present invention ensures the privacy and security of non-performing asset-related case data. At the same time, through MapReduce hierarchical aggregation to optimize the computing efficiency, it reduces the communication overhead of data transmission. Compared with the high maintenance cost and poor environmental adaptability of existing OCR and keyword extraction technologies, experimental data shows that the error rate of the present invention is only 0.5% - 2%, and it is transmitted through the TLS 1.3 protocol, reducing the processing time and providing a low-cost and highly robust data protection solution.
[0054] 3. Through Kafka event listening and PageRank path inference technologies, the present invention realizes the real-time update and risk assessment of the dynamic relationship graph, can timely detect the risks of the guarantee chain (such as changes in the credit of guarantors), and make up for the discontinuity deficiency of existing periodic reviews. In the experiment, the event listening throughput reaches 1000 messages per second, and the risk assessment accuracy rate reaches 90%. Compared with the static analysis of traditional database queries, it improves the timeliness and dynamicity of risk monitoring.
[0055] 4. Through HanLP semantic priority extraction, Q-learning feedback optimization and Drools verification technologies, combined with dynamic template selection and complex relationship analysis, the present invention provides accurate personalized summaries and risk assessments, overcoming the false alarm and missed alarm problems caused by the dependence on simple thresholds in existing technologies. Compared with the single-index analysis of traditional statistical methods, the present invention accurately identifies the risks of the guarantee chain through machine learning and graph technologies, improving the intelligent management level. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 is the specific flowchart of the present invention;
[0057] Figure 2 is the operation framework diagram of the present invention;
[0058] Figure 3 is the simulation diagram of the intelligent summary processing time of non-performing asset cases of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Specific Embodiment 1:
[0061] As Figures 1 to 3 shown, a method for intelligent summarization of non-performing asset related cases is executed by a multi-modal data processing system. The multi-modal data processing system includes a data standardization module, an encryption processing module, a graph construction module, a summary generation module, a feedback optimization module, a verification module, and a blockchain storage module. The method includes the following steps:
[0062] Sp1. The data standardization module obtains the multi-modal data of non-performing asset related cases, including legal document texts, financial forms, and contract scanned image. The text and table data are converted into structured data through domain adaptation template matching technology, and the text in the image is extracted through image semantic extraction technology to generate structured data including case number, debtor, debt amount, collateral, guarantor, and case status, and is transmitted to the encryption processing module in the format of extensible markup language; the multi-modal data is limited to legal document texts, financial forms, and contract scanned image, and the structured data is transmitted to the encryption processing module through RESTful API;
[0063] Sp2. The encryption processing module executes an encryption sharding algorithm on the structured data, divides it into encrypted fragments and stores them in the local node, calculates statistical information through a multi-party secure aggregation protocol, and is transmitted to the graph construction module through the TLS1.3 protocol after encryption;
[0064] Sp3. The graph construction module constructs a dynamic relationship graph based on the structured data and statistical information, with the debtor, guarantor, and collateral as nodes and the guarantee and mortgage relationships as edges, and generates a risk assessment result through a path reasoning algorithm, and is transmitted to the summary generation module in the format of key-value pairs;
[0065] Sp4. The summary generation module generates an initial summary through domain template filling and semantic priority extraction technology, including case number, debtor, debt amount, collateral, guarantor, case status, and risk points, and is transmitted to the feedback optimization module in text format;
[0066] Sp5. The feedback optimization module collects feedback through the user interface, optimizes the initial summary based on the content priority adjustment algorithm, generates a personalized summary, and transmits it to the verification module;
[0067] Sp6. The verification module verifies the personalized summary through formal rules and fact consistency verification technology, generates the verified summary, and transmits it to the blockchain storage module in encrypted text format;
[0068] Sp7. The blockchain storage module stores the verified summary and metadata in the private chain, triggers generation and update through smart contracts, and shares the summary through the distributed ledger.
[0069] In Sp1, the domain adaptation template matching technology matches case numbers, debtors, and guarantors in the legal document text data through predefined templates in the legal and financial domains. The domain adaptation template is customized based on the semantic features of non-performing asset cases, and the template selection is optimized through the semantic similarity matching algorithm. The semantic similarity matching algorithm calculates the matching degree between the text data and the template based on the legal and financial term libraries of non-performing asset cases. The image semantic extraction technology extracts key text fields from the contract scanned image data through image segmentation and semantic parsing technology. The structured data can be transmitted to the encryption processing module in the Extensible Markup Language format through the internal data bus.
[0070] In Sp2, the encryption sharding algorithm divides the structured data into multiple encrypted fragments through polynomial interpolation technology and stores them in local nodes. The multi-party secure aggregation protocol calculates the aggregation values of the debt amount and risk level through weighted statistical technology. The encrypted fragments and aggregation values are encrypted through the Transport Layer Security protocol and transmitted to the graph construction module. The Transport Layer Security protocol ensures that the aggregation calculation does not leak the original data.
[0071] In Sp3, the path inference algorithm identifies the guarantee chain risk in the dynamic relationship graph through risk propagation analysis technology. The risk propagation analysis dynamically adjusts the weight of the edge based on the credit scores of the guarantor and the collateral. The dynamic relationship graph updates the status of nodes and edges in real time through event listening technology. The risk assessment results are transmitted to the summary generation module in key-value pair format through the internal data bus.
[0072] In Sp4, the semantic priority extraction technology extracts key phrases from the structured data through syntactic dependency analysis and keyword priority ranking. The keyword priority ranking is based on the risk weights of non-performing asset cases. The domain template filling technology generates an initial summary through a predefined summary template. The template is customized according to the case type and user role. The initial summary is transmitted to the feedback optimization module in text format through the internal data bus.
[0073] In Sp5, the content priority adjustment algorithm dynamically adjusts the field priorities of the summary template through the scores and annotations in the user feedback data. The user feedback data is stored in the Extensible Markup Language format. The optimized personalized summary is transmitted to the user through the user interface and to the verification module through the internal data bus.
[0074] In Sp6, the formalization rules check the logical relationship between the debt amount and the collateral value in the personalized summary through logical reasoning techniques. The fact consistency verification technique verifies the field consistency between the personalized summary and the structured data through field matching techniques. The verified summary is transmitted to the blockchain storage module in encrypted text format through the internal data bus.
[0075] In Sp7, the smart contract controls the generation and update of the summary through predefined event trigger conditions, including case status changes and user requests. The distributed ledger technology synchronizes the verified summary data through a consensus algorithm. The metadata includes the case number, generation time, and content hash generated by the secure hash algorithm.
[0076] Before the multi-modal data enters the domain adaptation template matching technology in Sp1, the data cleaning module performs format normalization operations. The format normalization operations remove irrelevant formats from the legal document text data and standardize the numerical values in the financial form data. The cleaned data can be transmitted to the subsequent operations of Sp1 in Extensible Markup Language (XML) format through the internal data bus.
[0077] In Sp3, the dynamic relationship graph enhances the weights of the edges in the graph database through relationship priority assignment techniques. The relationship priority assignment is based on the credit scores of the guarantor and the collateral and the statistical analysis of historical case data. The risk assessment results are transmitted to the summary generation module in key-value pair format through the internal data bus. In Sp1, the domain adaptation template matching technology optimizes the template selection through a semantic similarity matching algorithm. The semantic similarity matching algorithm calculates the matching degree between the text data and the template based on the legal and financial term libraries of non-performing asset cases. The structured data including the case number, debtor, debt amount, collateral, guarantor, and case status is transmitted to the encryption processing module in XML format through the internal data bus. In Sp2, the multi-party secure aggregation protocol optimizes the calculation efficiency through hierarchical aggregation techniques. The hierarchical aggregation techniques calculate the statistical information layer by layer to reduce communication overhead. The encrypted fragments and aggregation values are transmitted to the graph construction module after being encrypted through the Transport Layer Security (TLS) protocol. In Sp4, the domain template filling technology selects the summary template according to the user role through a dynamic template selection algorithm. The dynamic template selection algorithm is based on the user's historical feedback data and case types. The initial summary is transmitted to the feedback optimization module in text format through the internal data bus. Specific Embodiment 2:
[0079] As Figures 1 to 3 shown, this method is executed by a multi-modal data processing system, which includes a data standardization module, an encryption processing module, a graph construction module, a summary generation module, a feedback optimization module, a verification module, and a blockchain storage module. The following details the implementation steps of this method, covering the entire process of data processing, encryption, graph construction, summary generation, feedback optimization, verification, and storage.
[0080] Step Sp1: Data Standardization:
[0081] Data Acquisition and Cleaning: The data standardization module obtains multimodal data from the non-performing asset management database through the file upload interface, limited to legal document texts (including court judgments, arbitration awards), financial forms (including debtor balance sheets), and contract scan images (including mortgage contract scans). To ensure data quality, the data cleaning module first performs format normalization operations on the multimodal data: uses regular expressions to remove HTML tags, redundant spaces, and irrelevant formats (including headers and footers) in legal documents, and standardizes the currency units in financial forms to RMB (such as converting US dollar amounts to RMB at the real-time exchange rate). The cleaned data is verified through XML Schema to ensure format consistency.
[0082] Data Structuring: The cleaned legal documents and financial forms are transformed into structured data through domain adaptation template matching technology. The templates are customized based on the semantic features of non-performing asset cases and contain predefined fields such as case number, debtor, debt amount, collateral, guarantor, and case status. The matching process uses the BERT model to calculate the cosine similarity between the text and the template based on a pre-trained legal and financial term library (constructed from historical legal document and financial statement corpora), and selects the template with the highest matching degree to generate structured data. For contract scan images, the U-Net model is used for image segmentation. After dividing the text area, the BiLSTM-CRF model is used to extract key fields (such as debtor name, collateral description). Finally, structured data in XML format containing case number, debtor, debt amount, collateral, guarantor, and case status is generated and transmitted to the encryption processing module through the RESTful API.
[0083] Step Sp2: Encryption Processing:
[0084] The encryption processing module receives the structured data in XML format and divides the data into n = 5 encrypted fragments through the Shamir secret sharing algorithm. k = 3 fragments are required for reconstruction, and the fragments are stored in the secure database of the local node. To calculate statistical information (such as the total debt amount, risk level distribution), a multi-party secure aggregation protocol with Paillier homomorphic encryption is adopted. Participants (such as banks, courts, asset management companies) calculate the weighted average of the debt amount and the risk level through weighted statistical techniques, and the weights are based on the debtor's credit score and case type. The aggregation calculation does not disclose the original data, and the results and encrypted fragments are transmitted to the graph construction module after being encrypted through the TLS1.3 protocol.
[0085] Step Sp3: Graph Construction:
[0086] The graph construction module constructs a dynamic relationship graph using the Neo4j graph database based on structured data and statistical information. The nodes are debtors, guarantors, and collateral (including real estate and vehicles), and the guarantee and mortgage relationships are the edges. The path inference algorithm uses a weighted PageRank model to identify guarantee chain risks through risk propagation analysis. The edge weights are dynamically adjusted based on the credit scores of the guarantors and collateral (calculated by combining historical default rates and financial data). The graph listens for changes in case status (such as lawsuit settlement and repayment completion) through the Kafka message queue and updates the status of nodes and edges in real time. The risk assessment results (such as high-risk guarantee chains) are transmitted to the summary generation module in JSON format through the RESTful API.
[0087] Step Sp4: Summary Generation:
[0088] The summary generation module generates an initial summary through domain template filling and semantic priority extraction techniques. Semantic priority extraction uses HanLP for syntactic dependency analysis to extract key phrases from structured data (such as "debtor default" and "insufficient collateral value", and these key phrases can be set and selected according to the actual situation). The keyword priority ranking is based on the TF-IDF algorithm combined with the risk weight of non-performing asset cases (the weight is determined by the debt amount and credit score). Domain template filling selects predefined summary templates according to the case type (bankruptcy liquidation, debt restructuring) and user role (asset manager, judge) through the decision tree algorithm to generate an initial summary containing the case number, debtor, debt amount, collateral, guarantor, case status, and risk points. The summary is transmitted to the feedback optimization module in text format through the RESTful API.
[0089] Step Sp5: Feedback Optimization:
[0090] The feedback optimization module collects user feedback through a web interface (an interactive page developed based on React), including ratings (such as 1-5 points) and text annotations (such as "highlight collateral risk"). The content priority adjustment algorithm uses Q-learning reinforcement learning to dynamically optimize the field priorities of the summary template based on user feedback (such as improving the sorting of high-risk fields). The user feedback data is stored in the local MongoDB database in XML format. The optimized personalized summary is displayed to the user through the web interface and transmitted to the verification module through the RESTful API.
[0091] Step Sp6: Verification:
[0092] The verification module verifies the personalized summary through formal rules and fact consistency verification techniques. The formal rules adopt the Drools rule engine to check the logical relationship between the debt amount and the value of the collateral (such as whether the value of the collateral is sufficient to cover the debt). The fact consistency verification uses the Levenshtein distance algorithm to compare the consistency of the summary fields with the fields of the structured data (such as whether the debtor's name matches). The verified summary is encrypted by AES-256 and transmitted to the blockchain storage module based on the RESTful API.
[0093] Step Sp7: Blockchain storage:
[0094] The blockchain storage module stores the encrypted summary and metadata in the Hyperledger Fabric private chain. The metadata includes the case number, generation time, and SHA-256 content hash to ensure data integrity. The smart contract is implemented in Solidity, and the triggering conditions include the case status changing to litigation closed or the user submitting an update request through the API. The distributed ledger adopts the PBFT consensus algorithm to synchronize the verified summary among multiple parties (such as banks, courts), and the metadata is stored in the on-chain JSON structure to achieve secure sharing.
[0095] In Sp1, the data cleaning module removes HTML tags and redundant spaces in legal documents through regular expressions, standardizes the numerical units of financial tables to RMB, and the cleaned data is verified by XML Schema and transmitted to the domain adaptation template matching step based on the RESTful API.
[0096] In Sp2, multi-party secure aggregation calculates statistical information hierarchically through the MapReduce framework to reduce communication overhead.
[0097] In Sp3, the relationship priority assignment uses a random forest model to adjust the weights by combining the credit scores of the guarantor and the collateral and the historical case default rate.
[0098] In Sp4, the dynamic template selection uses a collaborative filtering algorithm to optimize the template selection based on the user's historical feedback and case types.
[0099] The following content is a supplement to the above content:
[0100] Step Sp1: Data standardization:
[0101] Data acquisition interface: The file upload interface uses the HTTP protocol and supports batch upload. The data is obtained from the MySQL relational database, which stores the metadata of non-performing asset cases, financial records, and contract image paths. The table structure contains fields such as case numbers and debtor information.
[0102] Terminology Database Construction and Update: The legal and financial terminology database is constructed by extracting keywords from 100,000 historical legal documents and 50,000 financial statements. The terminology database is updated quarterly by analyzing newly added documents and statements, adding new terms such as new words in financial regulations.
[0103] Image Segmentation and Semantic Parsing Accuracy: The U-Net model is trained on a dataset of 10,000 annotated contract images, and the accuracy of segmenting text regions reaches 95%. The BiLSTM-CRF model is trained on a legal entity annotation dataset, supports the extraction of handwritten and printed contract texts, and the recognition accuracy reaches 90%. For blurred or low-quality images, the extraction effect is improved through preprocessing such as enhancing contrast.
[0104] Data Cleaning Error Handling: For missing fields such as the debtor's name, "unknown debtor" is used for filling. For format errors such as illegal dates, error logs are recorded and the fields are skipped, and a cleaning report is generated and stored in a local text file.
[0105] Step Sp2: Encryption Processing:
[0106] Security Database Type: Encrypted segments are stored in an SQLite database on the local node. The database file is protected by AES-256 encryption and can only be accessed through authentication.
[0107] Multi-Party Secure Aggregation Participant Interaction: Participants (such as banks, courts, asset management companies) interact through the gRPC protocol in a private network. The network adopts a star structure, and the central node is the asset management company's server, which is responsible for coordinating aggregation calculations.
[0108] Weighted Statistical Weight Calculation: The weights are determined based on the debtor's credit score and case type. The credit score ranges from 0 to 100. Case types such as bankruptcy liquidation have higher weights, while debt restructuring has lower weights. The weights are calculated by combining predefined rules.
[0109] Step Sp3: Graph Construction:
[0110] Credit Score Calculation: The credit score is generated based on the debtor's historical default times, asset-liability ratio, and income stability by analyzing 50,000 historical case records and is updated monthly.
[0111] Kafka Message Queue Configuration: Kafka is deployed as a 3-node distributed cluster. Messages are in JSON format, containing case numbers and status change information such as "closed", and 1000 messages are processed per second.
[0112] Risk Assessment Result Output: The JSON structure contains case numbers, risk levels, and node-edge information of high-risk guarantee chains. The fields are clearly defined as string and object types.
[0113] Step Sp4: Abstract Generation:
[0114] TF-IDF Weight Calculation: The risk weight is determined based on the proportion of the debt amount to the total debt and the credit score. TF-IDF calculates the importance of keywords from 100,000 case texts.
[0115] Decision Tree Algorithm Implementation: The decision tree is trained using 1,000 labeled abstract samples, and the classification rules are based on case types (such as bankruptcy liquidation, debt restructuring) and user roles (such as asset managers, judges).
[0116] Template Library Size and Update: The template library contains 50 abstract templates, which are updated every six months according to user feedback and new case types. Newly added templates are reviewed and confirmed by experts.
[0117] Step Sp5: Feedback Optimization:
[0118] Q-learning Training Details: Q-learning analyzes approximately 100 abstract field combinations, adjusts the field priorities, and the reward is based on user ratings from 1 to 5. Optimization is completed after 1,000 training sessions.
[0119] User Feedback Processing: Conflicting feedback is processed by taking the average. For example, if one user rates 4 and another rates 2, the average is 3. Feedback from high-authority users such as judges is preferred.
[0120] MongoDB Deployment: MongoDB is deployed as a single-node database with a storage capacity of 100GB, supporting high-concurrency read and write, and automatic daily backup.
[0121] Step Sp6: Verification:
[0122] Drools Rule Definition: The logical rules include checking whether the value of the collateral is at least 1.2 times the debt amount. Values lower than this are marked as high-risk. The rule set contains 10 similar rules formulated by experts.
[0123] Levenshtein Distance Threshold: Field matching requires a similarity of 90%. Fields with similarity lower than this are marked as inconsistent, triggering manual review.
[0124] Step Sp7: Blockchain Storage:
[0125] Hyperledger Fabric Configuration: The private blockchain contains 5 nodes, which are run by banks, courts, asset management companies, etc. Permissions are managed through the channel mechanism, and only authorized nodes can access the abstract data.
[0126] Solidity Contract Logic: The contract updates the abstract after verifying data integrity when the case status changes to "litigation closed" or the user submits an update request through the API.
[0127] PBFT Performance: The PBFT algorithm supports 100 transactions per second with an average latency of 200 milliseconds, suitable for the synchronization requirements of 10 participants.
[0128] Additional Details:
[0129] Sp1 Data Cleaning: PDF documents extract text through the PDFBox tool, and multilingual contracts are cleaned after being converted to Chinese through the Google Translate API, with an accuracy rate of 85%.
[0130] Sp2 Hierarchical Aggregation: The MapReduce hierarchical calculation is divided into 3 layers, including data collection, node aggregation, and final summary. Each layer is partitioned by case number, reducing the communication overhead by 50%.
[0131] Sp3 Random Forest Model: The model inputs include credit scores, default frequencies, and the proportion of collateral value. The training data consists of 50,000 historical cases, with an accuracy rate of 90%.
[0132] Sp4 Collaborative Filtering Algorithm: The user-based collaborative filtering is adopted. Based on the historical feedback of 1,000 users, templates matching the user roles and case types are recommended, with an accuracy rate of 80%. Specific Embodiment Three:
[0134] As Figures 1 to 3 shown below, the key algorithms mentioned in the above embodiments are analyzed in detail, including their core mathematical formulas and explanations:
[0135] 1. BERT is applied to the domain adaptation template matching in Sp1. By calculating the cosine similarity between legal documents and financial forms and the pre-trained legal and financial term library, the template selection is optimized.
[0136] The mathematical formula is as follows:
[0137] Word Embedding Generation: BERT maps the input text (sequence of words) to a context-related vector representation through multiple layers of Transformer encoders:
[0138]
[0139] where, is the embedding vector of the -th word, is the vector dimension (usually 768).
[0140] Cosine Similarity:
[0141]
[0142] Among them, is the embedding of the input text, is the embedding of the template.
[0143] BERT captures the semantic context of the text through bidirectional Transformer and generates high-quality word embeddings. Cosine similarity measures the semantic similarity between the input text and the template, and selects the template with the highest matching degree.
[0144] Problems to be solved:
[0145] Problem: Legal documents and financial forms contain complex legal and financial terms, and it is difficult to accurately extract fields such as case numbers and debtors through general template matching.
[0146] Solution: BERT utilizes a pre-trained legal and financial term library (100,000 documents and 50,000 statements) to convert the text into semantic vectors, and cosine similarity ensures that the template most conforming to the semantic features of non-performing asset cases is matched.
[0147] Specific means:
[0148] Training: Pre-train BERT on a corpus of 100,000 legal documents and 50,000 financial statements to generate exclusive word embeddings in the legal and financial fields.
[0149] Application: Input a legal document (such as "Debtor Zhang defaulted"), generate an embedding vector through BERT, calculate the cosine similarity with the embedding vectors of the template library (including predefined fields such as "debtor" and "case number"), select the template with the highest similarity (such as ≥0.9), and generate XML structured data.
[0150] Effect: The matching accuracy reaches 95%, ensuring the accurate extraction of fields such as case numbers and debtors.
[0151] 2. U-Net is applied to image segmentation in Sp1 to divide the text area from the contract scanned image.
[0152] The mathematical formula is as follows:
[0153] Convolution operation:
[0154]
[0155] Among them, is the input image, is the convolution kernel weight, is the bias, is the feature map.
[0156] Loss function (cross entropy):
[0157]
[0158] Among them, is the true label (text / non-text) of the pixel, is the predicted probability.
[0159] Explanation: U-Net extracts image features through an encoder-decoder structure. The encoding path downsamples to extract global features, and the decoding path upsamples to restore details, generating a segmentation mask for the text area. The cross-entropy loss optimizes the model to make the segmentation result closer to the true text area.
[0160] Problem to be solved: Contract scanned images (such as mortgage contracts) contain complex backgrounds, noise, or handwritten text, making it difficult to accurately separate the text area.
[0161] Solution: U-Net accurately segments the text area through multi-layer convolution and upsampling, supporting subsequent semantic parsing.
[0162] Specific means:
[0163] Training: U-Net is trained on a dataset of 10,000 annotated contract images. The annotations include text areas and background areas. After training, the segmentation accuracy reaches 95%. Application: Input a scanned mortgage contract, U-Net generates a text area mask and outputs a binary image (text area is 1, background is 0) for BiLSTM-CRF to extract text.
[0164] Effect: Accurately separate the text areas of handwritten and printed contracts, and enhance the effect through contrast enhancement preprocessing when dealing with blurred images.
[0165] 3. Shamir secret sharing is applied to the secret sharding in Sp2, splitting structured data into encrypted fragments and storing them in local nodes.
[0166] The mathematical formula is as follows:
[0167] Polynomial construction:
[0168]
[0169] Among them, is the original data, is the random coefficient, .
[0170] Fragment generation:
[0171]
[0172] Among them, , generating 5 fragments, and 3 fragments are required for reconstruction.
[0173] Explanation: Shamir secret sharing divides data into multiple fragments by constructing a polynomial. Any fewer than k fragments cannot recover the data. During reconstruction, the original data is recovered through Lagrange interpolation.
[0174] Problem to be solved: Structured data (such as debt amounts) may leak privacy when shared among multiple parties.
[0175] Solution: Shamir secret sharing divides the data into 5 encrypted fragments and stores them in local nodes, ensuring that the leakage of a single node does not affect data security.
[0176] Specific means:
[0177] Application: Input XML structured data (such as a debt amount of 1 million), construct a second-order polynomial (k = 3), generate 5 fragments and store them in an SQLite database. 3 fragments are required to reconstruct through Lagrange interpolation.
[0178] Effect: Even if the data of 2 nodes is leaked, the original data cannot be recovered, ensuring privacy security.
[0179] 4. Decision Tree (C4.5 Algorithm): Applied to the field template filling in Sp4, select the summary template according to the case type and user role.
[0180] The mathematical formula is as follows:
[0181] Information gain:
[0182]
[0183] Among them, , is the sample set, is the class proportion, is the attribute value subset of.
[0184] Explanation: The decision tree selects the best classification attribute (such as case type, user role) through information gain and constructs classification rules.
[0185] Problem to be solved: Different case types (such as bankruptcy liquidation) and user roles (such as judge) require different summary templates, which are difficult to select automatically.
[0186] Solution: The decision tree classifies according to the case type and user role and selects the most suitable template.
[0187] Specific means:
[0188] Training: Train a C4.5 decision tree based on 1000 labeled abstract samples (labeled case types and user roles). The attributes include case types and user roles.
[0189] Application: Input the case type (such as bankruptcy liquidation) and user role (such as judge), and the decision tree outputs a matching template (such as a template containing the "litigation status" field).
[0190] Effect: The accuracy of template selection reaches 90%, meeting personalized needs.
[0191] 5. Q-learning: Applied to the adjustment of content priorities in Sp5, optimizing the priorities of abstract fields based on user feedback.
[0192] The mathematical formula is as follows:
[0193] Q-value update:
[0194]
[0195] Among them, is the current state (field combination), is the action (adjusting the priority), is the reward (user rating), is the learning rate, is the discount factor.
[0196] Explanation: Q-learning iteratively updates the Q-value to learn the optimal field priority adjustment strategy.
[0197] Problem solved: Users' preferences for abstract fields (such as highlighting risk points) vary from person to person, and static templates are difficult to meet personalized needs.
[0198] Solution: Q-learning dynamically optimizes field priorities based on user ratings.
[0199] Specific means:
[0200] Training: The state is 100 field combinations, the actions include increasing and decreasing field priorities, the reward is based on user ratings (1-5 points), and train for 1000 iterations.
[0201] Application: Input the user rating (such as the rating of "highlighting the risk of collateral" is 4), Q-learning adjusts the field priority (such as increasing the "risk point" field) to generate a personalized abstract.
[0202] Effect: The satisfaction of the optimized abstract increases by 20%, and the accuracy of user feedback processing reaches 80%. Specific embodiment four:
[0204] Such as Figures 1 to 3As shown below, the following is a detailed description of the hardware components and hardware specifications of each module in Embodiment 1:
[0205] Module 1: Data Standardization Module (Sp1):
[0206] Hardware components: High-performance server (equipped with multi-core CPU and GPU), enterprise-level storage device (NAS), network switch.
[0207] Hardware description: The data standardization module is responsible for obtaining multi-modal data (legal document texts such as court judgments, financial forms such as debtor balance sheets, contract scan images such as mortgage contracts), performing cleaning, domain adaptation template matching (BERT), image segmentation (U-Net), semantic parsing (BiLSTM-CRF), and generating structured data in XML format. The high-performance server uses a dual-way Intel Xeon Gold 6348 CPU (28 cores, 2.6 GHz main frequency, 56 threads), paired with 2 NVIDIA A100 GPUs (40 GB video memory, 6912 CUDA cores), supporting the BERT model (processing a corpus of 100,000 documents and 50,000 reports, training takes about 48 hours) and U-Net image segmentation (processing 10,000 annotated images, segmenting a 5 MB JPEG contract takes 0.2 seconds, accuracy 95%). The BiLSTM-CRF model runs on the GPU, extracting key fields (such as debtor name, collateral description) takes 0.1 seconds / field, accuracy 90%, supporting handwritten and printed text. The enterprise-level NAS (NetApp FAS8300, 100 TB capacity, RAID6 configuration, read / write speed 2 GB / s, IOPS 100,000) stores the MySQL database (containing metadata such as case numbers, debtor information, about 50 GB) and the term library (100,000 documents, 50,000 reports, about 20 GB), supporting HTTP protocol batch upload (gigabit network, processing 100 MB of data per second). Data cleaning uses regular expressions to remove HTML tags and redundant spaces, the PDFBox tool to extract text from PDF documents, and the Google Translate API to process multi-language contracts (accuracy 85%), and the cleaned reports are stored in the NAS (1 MB of text per day). The network switch (Cisco Catalyst 9300, 48-port gigabit Ethernet, throughput 480 Gbps) connects the server and the NAS, supporting RESTful API to transfer XML data to the encryption module (bandwidth 1 Gbps, latency < 5 ms). The hardware processes blurred images (through contrast enhancement preprocessing) and non-standard formats (such as PDF) to ensure data quality and meet high-concurrency requirements.
[0208] Module 2: Encryption Processing Module (Sp2):
[0209] Hardware Composition: Encryption Server (with Hardware Security Module HSM), Local Storage Node (SSD Array), Network Security Device (Firewall).
[0210] Hardware Description: The encryption processing module performs Shamir secret sharing (dividing XML data into 5 fragments, requiring 3 for reconstruction), Paillier homomorphic encryption (securely aggregating debt amounts and risk levels for multiple parties), and MapReduce hierarchical computing to protect data privacy and transmit statistical results. The encryption server uses an AMD EPYC 7543 CPU (32 cores, 2.8 GHz main frequency, 64 MB cache), paired with a Thales Luna HSM (supporting AES-256 and Paillier encryption, processing 100 encryption tasks per second), generating Shamir secret shards (taking 0.05 seconds for 1 MB of data) and storing them on local nodes. The local storage node uses a Samsung enterprise-level SSD array (4 TB capacity, NVMe interface, read / write speed of 7 GB / s, IOPS of 1.5 million), running an SQLite database (AES-256 disk encryption, 100 GB capacity), storing encrypted fragments, and backing up 1 GB of data daily. MapReduce hierarchical computing (3 layers: data collection, node aggregation, final summary) runs on the server, taking 1 second to process 1000 case statistics (debt amount, risk level), reducing communication overhead by 50%. Participants (such as banks and courts) interact through the gRPC protocol in a private star network, with the central node being the asset management company's server. The network security device (Fortinet FortiGate 1000F firewall, throughput of 2 Gbps, latency < 10 ms) supports the TLS 1.3 protocol, protecting the transmission of encrypted fragments and aggregation results (such as the average debt amount of 750,000), filtering unauthorized access, and ensuring data non-disclosure. The hardware supports RESTful API transmission to the graph construction module (bandwidth of 1 Gbps), meeting the high security and efficiency requirements of multi-party collaboration scenarios.
[0211] Module 3: Graph Construction Module (Sp3):
[0212] Hardware Composition: Graph Database Server, Kafka Cluster Server, Storage Array, Network Switch.
[0213] Hardware Description: The graph construction module constructs a dynamic relationship graph (with debtors, guarantors, and collateral as nodes, and guarantee and mortgage relationships as edges) using the Neo4j graph database based on structured data and statistical information. It generates risk assessment results in JSON format through random forest (edge weight prediction), weighted PageRank (risk propagation analysis), and Kafka message queue (real-time update). The graph database server uses an Intel Xeon Platinum 8380 CPU (40 cores, 2.3 GHz main frequency, 80 MB cache), 128 GB DDR4 memory, runs Neo4j Enterprise Edition, stores 1 million nodes and 5 million edges, processes 1000 queries per second, and takes 0.5 seconds per case to construct the graph. The random forest model (100 decision trees, 50,000 historical case training data, 90% accuracy) predicts the edge weight (e.g., the guarantor's credit score is 80, and the weight is 0.85), taking 0.1 seconds per case. PageRank calculates high-risk guarantee chains (PR value > 0.5), and it takes 1 second to process a 1000-node graph. The Kafka cluster server (3 nodes, each node with an AMD EPYC 7313 CPU, 16 cores, 3.0 GHz, 64 GB memory) runs distributed Kafka, stores JSON messages (such as the case status "closed"), processes 1000 messages per second, with a latency < 50 ms, ensuring real-time graph updates. The storage array (HP E Nimble HF40, 50 TB capacity, read / write speed 5 GB / s, IOPS 500,000) stores historical case data (about 30 GB) and graph snapshots, with a daily update of 10 GB. The network switch (Arista 7050X, 48-port 10 Gigabit Ethernet, throughput 1.2 Tbps) connects the servers and the storage array, transmitting JSON risk assessment results (RESTful API, bandwidth 10 Gbps, latency < 5 ms). The hardware supports dynamic graph updates and high-concurrency risk assessment, meeting the needs of guarantee chain risk quantification.
[0214] Module 4: Summary Generation Module (Sp4):
[0215] Hardware Composition: Application server, storage device, network switch.
[0216] Hardware Description: The abstract generation module uses HanLP (syntactic dependency parsing), TF-IDF (keyword ranking), decision tree (C4.5 template selection), and collaborative filtering (optimizing templates based on user feedback) to generate an initial abstract in text format, including fields such as case number and debtor. The application server uses a dual-way Intel Xeon Silver 4314 CPU (16 cores, 2.4 GHz main frequency, 32 MB cache), 64 GB of DDR4 memory, runs HanLP, takes 2 seconds to process 1000 case texts, and has an accuracy of 85% for extracting key phrases (such as "debtor default"). The TF-IDF algorithm is based on a corpus of 100,000 case texts (about 20 GB), combines risk weights (debt amount, credit score), and takes 0.1 seconds per case to rank the top 10 phrases. The decision tree (trained with 1000 labeled samples) selects the optimal template from 50 templates, and collaborative filtering (1000 user feedbacks, accuracy 80%) optimizes the selection, taking 0.05 seconds to process 1 case. The storage device (Dell EMC PowerStore 5000T, 20 TB capacity, read / write speed 4 GB / s, IOPS 300,000) stores the corpus, template library (about 1 GB), and temporary abstract data, reads and writes 5 GB per day, and updates the template library every six months. The network switch (Juniper EX4600, 40-port 10 Gigabit Ethernet, throughput 960 Gbps) supports RESTful API to transmit text abstracts (bandwidth 10 Gbps, latency <5 ms), such as "Case 2025-NPL-001: Debtor Zhang, debt 1 million". The hardware supports the rapid generation of customized abstracts to meet the needs of case types such as bankruptcy liquidation and debt restructuring and user roles such as judges and asset managers.
[0217] Module 5: Feedback Optimization Module (Sp5):
[0218] Hardware Composition: Web server, MongoDB database server, network switch.
[0219] Hardware Description: The feedback optimization module collects user feedback (rating from 1 to 5, with annotations such as "highlight the collateral risk") through a React-developed web interface, uses the Q-learning algorithm to optimize the priority of summary fields, and generates personalized summaries. The web server uses an AMD EPYC 7443 CPU (24 cores, 3.0 GHz main frequency, 128 MB cache), 32 GB of DDR4 memory, runs React pages, supports 1000 concurrent user accesses, has a response time < 200 ms, and each feedback is about 1 KB (XML format). The Q-learning analyzes 100 field combinations, adjusts the priority based on user ratings, trains for 1000 iterations, takes 0.1 seconds to process one case, and the satisfaction rate is increased by 20%. Contradictory feedback is processed by taking the average (e.g., rating 4 and 2 results in 3), and high-privilege users (such as judges) are prioritized. The MongoDB database server (single node, Intel Xeon E5-2699 CPU, 18 cores, 2.2 GHz, 128 GB memory, 100 GB storage capacity, read / write speed 2 GB / s) stores feedback data, supports 5000 read / writes per second, and backs up 1 GB per day. The network switch (HPE Aruba 6300F, 48-port Gigabit Ethernet, throughput 176 Gbps) supports RESTful API to transmit personalized summaries (bandwidth 1 Gbps, latency < 5 ms), such as "Case 2025-NPL-001: Debtor Zhang, collateral value 500,000". The hardware ensures efficient feedback collection and summary optimization to meet personalized needs.
[0220] Module 6: Verification Module (Sp6):
[0221] Hardware Composition: Verification server, storage device, network switch.
[0222] Hardware Description: The verification module uses the Drools rule engine (for logical consistency checking) and the Levenshtein distance algorithm (for field consistency verification) to verify personalized summaries and generate AES-256 encrypted summaries. The verification server uses an Intel Xeon Gold 5318Y CPU (24 cores, 2.1 GHz main frequency, 36 MB cache), 64 GB of DDR4 memory, runs the Drools engine, checks 10 rules (such as the collateral value ≥ 1.2 times the debt amount), and it takes 1 second to verify 1000 summaries. The Levenshtein distance algorithm compares summary fields with structured data (such as "Zhang San" and "Zhang San San"), with a similarity threshold of 0.9, it takes 0.01 seconds to process 1 field, with an accuracy of 95%, and manual review is triggered when the threshold is exceeded. The storage device (WD Ultrastar DCSN840, 10 TB capacity, NVMe interface, read / write speed 6 GB / s, IOPS 1 million) stores structured data (about 5 GB) and verification logs, with 2 GB of read / write per day. The network switch (Cisco Nexus 9300, 36-port 10 Gigabit Ethernet, throughput 720 Gbps) supports RESTful API transmission of encrypted summaries (bandwidth 10 Gbps, latency < 5 ms). The hardware supports fast verification to ensure summary logic and data consistency and protect privacy.
[0223] Module 7: Blockchain Storage Module (Sp7):
[0224] Hardware Composition: Blockchain node server (Hyperledger Fabric), Hardware Security Module (HSM), network security device.
[0225] Hardware Description: The blockchain storage module stores the AES-256 encryption digest and metadata (case number, generation time, SHA-256 hash) on the Hyperledger Fabric private chain, and uses Solidity smart contracts and the PBFT consensus algorithm to achieve secure sharing. The blockchain node server (5 nodes, each node with an AMD EPYC 7643 CPU, 48 cores, 2.3 GHz main frequency, 256 MB cache, 128 GB memory) runs Fabric, stores 1 million digests (about 10 GB), supports 100 transactions per second, and has a latency of 200 ms. The HSM (SafeNet Luna SA, encryption speed of 500 times per second) generates SHA-256 hashes and AES-256 encryption to ensure data integrity. The Solidity contract is triggered by case status changes (such as "lawsuit closed") or user API requests, verifies the hash and then stores the JSON data, taking 0.1 seconds to process 1 digest. The PBFT algorithm synchronizes data among the 5 nodes, and the channel mechanism restricts access rights, allowing only authorized nodes (such as banks, courts) to read. The network security device (Palo Alto PA-5200 firewall, throughput of 5 Gbps, latency < 10 ms) supports the TLS 1.3 protocol to protect data transmission. The hardware ensures that the digest cannot be tampered with, meeting the needs of multi-party sharing.
[0226] Additional Details of Hardware Support:
[0227] Sp1 Data Cleaning: PDF document processing uses the PDFBox tool, running on a high-performance server (Xeon Gold 6348). Multilingual contracts are converted through the Google Translate API (cloud call, latency < 100 ms) and stored in the NAS (100 TB, 50 GB read / write per day).
[0228] Sp2 Hierarchical Aggregation: MapReduce runs on an encrypted server (AMD EPYC 7543), and the 3-layer computing structure is stored in an SSD array (4 TB, read / write speed of 7 GB / s), supporting 10 GB of data processing per day.
[0229] Sp3 Random Forest: The model runs on a graph database server (Xeon Platinum 8380), and 50,000 historical cases (30 GB) are stored in a storage array (50 TB).
[0230] Sp4 Collaborative Filtering: The algorithm runs on an application server (Xeon Silver 4314), and user feedback is stored in a storage device (20 TB, 5 GB read / write per day). Specific Embodiment Five:
[0232] As Figures 1 to 3 shown, the following are specific use cases:
[0233] Use Case 1: Handling Non-Performing Loan Cases of Banks:
[0234] Scenario Description: A bank needs to handle a non-performing loan case (case number 2025-NPL-001), involving debtor Zhang, with a debt amount of 1 million yuan, guarantor Li, and the collateral being a property (worth 500,000 yuan). The case status is in litigation. The bank hopes to generate a summary for asset managers and judges to use, while protecting data privacy and sharing it with the court.
[0235] Usage Process:
[0236] The data standardization module obtains case data from a MySQL database (stored in NetApp FAS8300 NAS with a capacity of 100 TB) through the HTTP protocol: court judgments (PDF format, 2 MB), balance sheets (Excel format, 1 MB), scanned copies of mortgage contracts (JPEG format, 5 MB). A high-performance server (Intel Xeon Gold 6348 CPU, 2 NVIDIA A100 GPUs) runs the data cleaning module, which uses regular expressions to remove HTML tags and redundant spaces in the judgments, the PDFBox tool to extract PDF text, the Google Translate API to convert English contract terms to Chinese (with an accuracy of 85%), and standardizes the currency unit of the balance sheet to RMB. The cleaned data is verified through XML Schema and stored in NAS. The BERT model (trained on 100,000 documents and 50,000 reports) calculates the cosine similarity between the judgments and tables and the legal and financial term library (with an accuracy of 95%), and matches templates to generate XML structured data. The U-Net model (trained on 10,000 images, with an accuracy of 95%) segments the text area of the contract scanned copy, and BiLSTM-CRF (with an accuracy of 90%) extracts fields such as "Debtor: Zhang", integrates them into the XML data, and transmits them to the encryption processing module through the RESTful API (Cisco Catalyst 9300 switch, 1 Gbps bandwidth).
[0237] The encryption processing module runs the Shamir secret sharing algorithm on an encryption server (AMDEPYC 7543 CPU, Thales Luna HSM), divides the XML data into 5 encrypted fragments (n = 5, k = 3, taking 0.05 seconds), and stores them in a Samsung SSD array (4TB, AES-256 encryption). The bank and the court perform Paillier homomorphic encryption aggregation through the gRPC protocol (Fortinet FortiGate 1000F firewall, TLS 1.3), and use MapReduce hierarchical calculation (3 layers, reducing the communication overhead by 50%) to calculate the average debt amount (1 million) and the risk level (high). The results are transmitted to the graph construction module through TLS 1.3.
[0238] The graph construction module runs Neo4j on a graph database server (Intel Xeon Platinum 8380 CPU, 128GB of memory) to construct a graph (nodes: Zhang, Li, real estate; edges: guarantee, mortgage). The random forest model (50,000 historical cases, accuracy rate of 90%) predicts the edge weights (Li's credit score is 80, weight 0.85), and the PageRank algorithm (taking 1 second) identifies high-risk guarantee chains (Li, PR value 0.6). The Kafka cluster (3 nodes, AMDEPYC 7313 CPU) listens for status changes (such as "in litigation"), processes 1000 JSON messages per second (delay < 50ms). The risk assessment results are transmitted to the summary generation module through the RESTful API (Arista 7050X switch, 10Gbps).
[0239] The summary generation module runs HanLP on an application server (Intel Xeon Silver 4314 CPU) to extract key phrases ("debtor default", accuracy rate of 85%), and combines TF-IDF (100,000 text corpora) with the debt amount to sort the phrases. The decision tree (1000 samples) selects a template for the judge (including the litigation status), and collaborative filtering (1000 user feedbacks, accuracy rate of 80%) optimizes the selection to generate a summary: "Case 2025-NPL-001: Debtor Zhang, debt 1 million, guarantor Li, mortgaged property real estate, status in litigation, risk point: low credit of the guarantor", which is transmitted to the feedback optimization module through the RESTful API (Juniper EX4600 switch).
[0240] The feedback optimization module collects judge feedback (rating 4 points, marked "highlight collateral risk") through the React interface on the Web server (AMD EPYC 7443 CPU). Q-learning (1000 iterations, taking 0.1 seconds) adjusts the field priorities to generate a personalized summary: "Case 2025-NPL-001: Debtor Zhang, debt of 1 million, collateral value of 500,000, high risk", which is stored in MongoDB (100GB, Intel Xeon E5-2699 CPU) and transmitted to the verification module through the RESTful API (HPE Aruba 6300F switch).
[0241] The verification module runs the Drools rule engine on the verification server (Intel Xeon Gold 5318Y CPU) to check if the collateral value (500,000) < 1.2 times the debt amount (1.2 million), and marks it as high risk. The Levenshtein distance algorithm verifies the field consistency ("Zhang" matches with an accuracy of 95%), and the encrypted summary (AES-256, Thales Luna HSM) is transmitted to the blockchain storage module (Cisco Nexus 9300 switch).
[0242] The blockchain storage module stores the encrypted summary and metadata on a 5-node Hyperledger Fabric server (AMD EPYC 7643 CPU). The Solidity contract is triggered in the "in litigation" state, and the PBFT algorithm synchronizes the data (100 transactions per second, latency of 200ms), which can be accessed by the bank and the court through the channel mechanism.
[0243] Output result: The bank and the judge obtain the summary: "Case 2025-NPL-001: Debtor Zhang, debt of 1 million, collateral value of 500,000, high risk", which is stored on the blockchain and shared securely.
[0244] Problem solved: Handling heterogeneous data (PDF, Excel, JPEG), protecting privacy, generating judge-customized summaries, updating risks in real-time, and sharing data securely.
[0245] Use case two: Handling enterprise bankruptcy liquidation cases:
[0246] Scenario description: An asset management company handles an enterprise bankruptcy liquidation case (case number 2025-NPL-002), where debtor company A has a debt of 5 million yuan, guarantor company B, and the collateral is a factory building (worth 3 million yuan), and the case status is bankruptcy acceptance. A summary needs to be generated and shared with the court and the bank.
[0247] Usage process:
[0248] The data normalization module retrieves data from a MySQL database (NetApp FAS8300 NAS): bankruptcy acceptance documents (PDF, 3MB), financial statements (Excel, 2MB), and factory building mortgage contracts (JPEG, 6MB). A high-performance server (Xeon Gold 6348, 2 A100 GPUs) runs regular expressions to clean the text, uses PDFBox to extract PDF content, and standardizes the currency unit to RMB. A BERT model matches templates (accuracy 95%), a U-Net divides contract images (accuracy 95%), a BiLSTM-CRF extracts fields, and transmits through a RESTful API (Cisco Catalyst 9300).
[0249] The encryption processing module (AMD EPYC 7543, Thales Luna HSM) uses the Shamir algorithm to divide the data into 5 segments (time-consuming 0.05 seconds) and stores them in a Samsung SSD (4TB). Paillier homomorphic encryption aggregates the debt amount (5 million) and the risk level (high), performs hierarchical calculation using MapReduce (time-consuming 1 second), and transmits through TLS 1.3 (Fortinet firewall).
[0250] The graph construction module (Xeon Platinum 8380, Neo4j) constructs a graph (nodes: Company A, Company B, factory building). Random forest predicts edge weights (Company B credit score 70, weight 0.7), and PageRank identifies high-risk chains (time-consuming 1 second). A Kafka cluster (3 nodes, AMD EPYC 7313) listens for the "bankruptcy acceptance" status and outputs a JSON result ({"case_id":"2025-NPL-002","risk_level":"high"}), and transmits through a RESTful API (Arista 7050X).
[0251] The abstract generation module (Xeon Silver 4314) uses HanLP to extract phrases ("enterprise bankruptcy"), sorts them using TF-IDF, selects a template for the asset manager using a decision tree, and optimizes through collaborative filtering (accuracy 80%) to generate an abstract: "Case 2025-NPL-002: Debtor Company A, debt 5 million, guarantor Company B, mortgaged factory building, status bankruptcy acceptance, risk point: low credit of the guarantor", and transmits through a RESTful API (Juniper EX4600).
[0252] The feedback optimization module (AMDEPYC7443, React interface) collects feedback (rating 3, labeled "Increase liquidation progress"), adjusts priorities with Q-learning, generates a summary: "Case 2025-NPL-002: Debtor company A, debt of 5 million, collateral value of 3 million, liquidation progress 20%, high risk", stores it in MongoDB (100GB), and transmits it to the verification module (HPE Aruba6300F).
[0253] The verification module (Xeon Gold 5318Y) uses Drools to check if the collateral value (3 million) < 1.2 times the debt (6 million), Levenshtein to verify field consistency (accuracy 95%), and AES-256 to encrypt and transmit the summary (Cisco Nexus9300).
[0254] The blockchain storage module (5-node Fabric, AMDEPYC7643) stores the summary and metadata. The Solidity contract is triggered by "bankruptcy acceptance", and PBFT synchronizes the data for access by banks and courts.
[0255] Output result: Summary: "Case 2025-NPL-002: Debtor company A, debt of 5 million, collateral value of 3 million, liquidation progress 20%, high risk", securely shared.
[0256] Problem solved: Handling complex enterprise bankruptcy data, generating customized summaries for asset managers, protecting privacy, and sharing in real time.
[0257] Use case three: Handling debt restructuring cases:
[0258] Scenario description: The court is handling a debt restructuring case (case number 2025-NPL-003). The debtor is Wang, with a debt amount of 2 million yuan. The guarantor is Zhao, and the collateral is a vehicle (worth 1 million yuan). The status is restructuring negotiation. A summary needs to be generated for use by the judge and the bank.
[0259] Usage process:
[0260] The data standardization module retrieves the restructuring agreement (PDF, 1MB), financial statements (Excel, 1MB), and vehicle mortgage contract (JPEG, 4MB) from MySQL (NetApp FAS8300). A high-performance server (Xeon Gold 6348, A100 GPU) cleans the data, PDFBox extracts the text, and transmits it via RESTful API (Cisco Catalyst9300).
[0261] The encryption processing module (AMDEPYC7543, Thales Luna HSM) splits data using Shamir (5 fragments), aggregates debt amounts using Paillier (2 million), performs MapReduce calculations (takes 1 second), and transmits via TLS1.3 (Fortinet firewall).
[0262] The graph construction module (Xeon Platinum 8380, Neo4j) constructs a graph, predicts edge weights using random forest (Zhao's credit score is 85, weight 0.85), and identifies risks using PageRank (takes 1 second). Kafka (3 nodes, AMDEPYC7313) listens to "restructuring negotiation" and outputs JSON (Arista 7050X).
[0263] The abstract generation module (Xeon Silver 4314) extracts phrases ("debt restructuring") using HanLP and TF-IDF, and generates an abstract using decision trees and collaborative filtering: "Case 2025-NPL-003: Debtor Wang, debt 2 million, guarantor Zhao, collateral vehicle, status restructuring negotiation, risk point: low value of collateral", and transmits via RESTful API (Juniper EX4600).
[0264] The feedback optimization module (AMDEPYC7443) collects feedback (rating 5, marked "retain restructuring details"), generates an abstract using Q-learning: "Case 2025-NPL-003: Debtor Wang, debt 2 million, value of collateral 1 million, restructuring progress 50%, medium risk", stores it in MongoDB, and transmits (HPE Aruba 6300F).
[0265] The verification module (Xeon Gold 5318Y) verifies using Drools and Levenshtein (accuracy 95%), and encrypts and transmits the abstract using AES-256 (Cisco Nexus 9300).
[0266] The blockchain storage module (5-node Fabric) stores the abstract, the Solidity contract is triggered by "restructuring negotiation", and PBFT synchronizes for access by the court and the bank.
[0267] Output result: Abstract: "Case 2025-NPL-003: Debtor Wang, debt 2 million, value of collateral 1 million, restructuring progress 50%, medium risk", securely shared.
[0268] Problems solved: Support dynamic updates of debt restructuring, generate customized abstracts for judges, and protect privacy. Specific embodiment six:
[0270] AsFigures 1 to 3 As shown below are the specific experimental data:
[0271] Technical module Comparison index Results of this method Results of traditional method Improvement rate Data structuring Time consumption for multi-modal data processing 2.1 minutes / 100 copies 5.8 minutes / 100 copies +176% Spectrum risk assessment Accuracy rate of guarantee chain risk identification 91.4% 78.2% +16.9% Abstract generation Omission rate of key information 3.2% 12.5% -74.4% Blockchain storage Multi-node synchronization delay 200ms 850ms -76.5%
[0272] The above table is for the comparative experiment of key technical performances, and the sources of the experimental data are as follows:
[0273] Data structuring - Multimodal data processing time consumption:
[0274] Results of this method (2.1 minutes / 100 copies): In the actual test environment, build a multimodal data processing system adapted to this method, including relevant components such as a data standardization module. Select 100 multimodal data samples of non-performing asset related cases covering legal document texts, financial tables, and contract scanned image. Start timing from receiving data at the data acquisition interface. After the format normalization operation of the data cleaning module (such as removing HTML tags from legal documents, standardizing the numerical units of financial tables, etc.), and then convert the data into structured data through domain adaptation template matching technology (for legal documents and financial tables) and image semantic extraction technology (for contract scanned images). Finally, record the total time consumption of the whole process, and take the average value of multiple tests to obtain that it takes 2.1 minutes to process 100 copies of data.
[0275] Results of the traditional method (5.8 minutes / 100 copies): Build a traditional non-performing asset data processing system for the same 100 multimodal data samples. For legal document texts, use a simple TF-IDF algorithm for word segmentation and keyword extraction. For financial tables, rely on manual data entry. For contract images, only use basic OCR to recognize the text and then manually proofread and supplement. Record the operation time of each link respectively, and summarize the total time consumption of processing 100 copies of data. The average value of multiple tests is 5.8 minutes.
[0276] Graph risk assessment - Accuracy rate of guarantee chain risk identification:
[0277] Results of this method (91.4%): Use the Neo4j graph database to build a dynamic relationship graph with debtors, guarantors, and mortgaged properties as nodes and guarantee and mortgage relationships as edges. Identify guarantee chain risks through a path reasoning algorithm (adopting a weighted PageRank model combined with risk propagation analysis technology). Select 1000 cases with clear actual guarantee chain risk situations (manually verified and marked by professionals) from historical non-performing asset case data as the test set, conduct risk assessment with this method, compare the assessment results with the actual risk situations, count the number of correctly identified cases, and calculate the proportion of the number of correctly identified cases to the total number of cases as 91.4%.
[0278] Results of the traditional method (78.2%): A static relationship graph was constructed in the traditional way (manually sorting out the association relationships), and a fixed-weight model was used for risk assessment (only based on a few factors such as the debt amount). For the same 1,000 test set cases, the guarantee chain risk was assessed using the traditional method, and the number of correctly identified cases was counted by comparing with the actual risk situation. After calculation, the proportion of the total number of cases was 78.2%.
[0279] Abstract generation - Omission rate of key information:
[0280] Results of this method (3.2%): The abstract generation module generates an initial abstract through domain template filling and semantic priority extraction techniques. 500 case data were selected from actual non-performing asset cases. After generating the abstract, professional non-performing asset domain experts (covering different roles such as asset managers and judges) were invited to manually verify each abstract and mark the abstracts that did not cover the key information (such as case numbers, debtors, debt amounts, collateral, guarantors, case status, and risk points, etc., pre-defined key information). The number of abstracts with missing key information was counted, and the proportion of the total number of abstracts was calculated. The average value was obtained through multiple sampling tests, and the omission rate of key information was 3.2%.
[0281] Results of the traditional method (12.5%): The traditional abstract generation method (fixed template filling, lacking semantic priority extraction) was used to generate abstracts for the same 500 case data. The same expert team conducted manual verification, marked the omission situations according to the same key information criteria, counted the number of abstracts with missing key information, and calculated the proportion of the total number of abstracts. The average value was obtained through multiple tests, and the omission rate of key information was 12.5%.
[0282] Blockchain storage - Multi-node synchronization delay:
[0283] Results of this method (200ms): A private chain based on Hyperledger Fabric was built, and 5 nodes were set up (simulating participants such as banks, courts, and asset management companies respectively). Encrypted abstracts and metadata were transmitted between the nodes, and the monitoring tool deployed on the nodes was used to record the time interval from the initiation of data storage operations on one node to the completion of data synchronization on all other nodes. Multiple synchronization operation tests with different data volumes (simulating the data scale in actual business) were carried out, and the synchronization delay time of each time was recorded. The average value was calculated to obtain the multi-node synchronization delay of 200ms.
[0284] Results of the traditional method (850 ms): Construct a traditional centralized database storage architecture, and simulate multiple clients (similar to nodes in a blockchain) interacting with the central database. Conduct data synchronization operation tests, and record the time interval from when a data update request is initiated by one client to when other clients obtain the updated data. Conduct multiple tests on the synchronization situation under different data volumes, and statistically calculate the average synchronization delay time, obtaining a result of 850 ms.
[0285] It should be noted that Figure 3 In [reference], the changing trend of the processing time of the method based on "A Method for Intelligent Summarization of Non-Performing Asset Related Cases" with the number of cases was simulated through simulation, aiming to verify the efficiency and stability of the method.
[0286] Where the X-axis is (the number of cases), ranging from 0 to 100 cases;
[0287] Meaning: The X-axis represents the number of non-performing asset related cases input, gradually increasing from 1 to 100, simulating the processing scenarios from small to medium-sized datasets, and reflecting the adaptability of the system to different data scales.
[0288] The Y-axis (processing time, seconds), range: approximately 2.50 to 2.535 seconds;
[0289] Meaning: The Y-axis represents the total processing time for each case from data standardization (Sp1) to blockchain storage (Sp7). The numerical range is relatively narrow (2.50 to 2.535 seconds), indicating that the processing time fluctuates slightly with the increase in the number of cases, reflecting the low time complexity of the method.
[0290] Data points and trends:
[0291] Data points: Each case corresponds to a processing time value (blue dots), showing slight fluctuations with the increase in the number of cases.
[0292] Trend: The overall processing time fluctuates between 2.50 and 2.53 seconds, with an average value of approximately 2.52 seconds (marked in the figure), indicating that the system maintains stable performance within 100 cases. The fluctuations may be caused by random factors such as the simulated debt amount, collateral value, or status, which conforms to the dynamic characteristics of multi-modal data processing.
[0293] Annotation: The "Average processing time: 2.52 seconds" marked in the figure provides a quantitative reference for evaluating efficiency.
[0294] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a reference structure" does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0295] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for intelligent summarization of non-performing asset related cases, characterized in that: The method is executed by a multi-modal data processing system, which includes a data standardization module, an encryption processing module, a graph construction module, a summary generation module, a feedback optimization module, a verification module, and a blockchain storage module. The method includes the following steps: Sp1. The data standardization module obtains multi-modal data of non-performing asset-related cases, including legal document texts, financial forms, and contract scan image files. It converts the text and form data into structured data through domain adaptation template matching technology, extracts the text in the images through image semantic extraction technology, and generates structured data including case numbers, debtors, debt amounts, collateral, guarantors, and case statuses, which is transmitted to the encryption processing module in Extensible Markup Language (XML) format; Sp2. The encryption processing module executes an encryption sharding algorithm on the structured data, divides it into encrypted fragments and stores them locally, calculates statistical information through a multi-party secure aggregation protocol, and transmits the encrypted data to the graph construction module; Sp3. The graph construction module constructs a dynamic relationship graph based on the structured data and statistical information, with debtors, guarantors, and collateral as nodes and guarantee and mortgage relationships as edges. It generates a risk assessment result through a path reasoning algorithm and transmits it to the summary generation module in key-value pair format; Sp4. The summary generation module generates an initial summary through domain template filling and semantic priority extraction technology, including case numbers, debtors, debt amounts, collateral, guarantors, case statuses, and risk points, which is transmitted to the feedback optimization module in text format; Sp5. The feedback optimization module collects feedback through a user interface, optimizes the initial summary based on a content priority adjustment algorithm, generates a personalized summary, and transmits it to the verification module; Sp6. The verification module verifies the personalized summary through formal rules and fact consistency verification technology, generates a verified summary, and transmits it to the blockchain storage module in encrypted text format; Sp7. The blockchain storage module stores the verified summary and metadata in a private chain, triggers generation and updates through smart contracts, and shares the summary through a distributed ledger.
2. The method for intelligently abstracting bad asset related cases according to claim 1, characterized in that: In Sp1, the domain adaptation template matching technology matches case numbers, debtors, and guarantors in legal document text data through predefined templates in the legal and financial fields. The domain adaptation template is customized based on the semantic characteristics of non-performing asset cases, and the template selection is optimized through a semantic similarity matching algorithm. The semantic similarity matching algorithm calculates the matching degree between text data and the template based on the legal and financial term libraries of non-performing asset cases. The image semantic extraction technology extracts key text fields from contract scan image data through image segmentation and semantic parsing technology. The structured data is transmitted to the encryption processing module in XML format through an internal data bus.
3. A method for intelligent summarization of non-performing asset related cases according to claim 1, characterized in that: In the Sp2, the encryption sharding algorithm divides the structured data into multiple encrypted fragments through polynomial interpolation technology and stores them in local nodes. The multi-party secure aggregation protocol calculates the aggregated values of the debt amount and risk level through weighted statistical technology. The encrypted fragments and aggregated values are transmitted to the graph construction module after being encrypted by the Transport Layer Security (TLS) protocol, and the TLS protocol ensures that the aggregation calculation does not disclose the original data.
4. A method for intelligently abstracting bad asset related cases according to claim 1, characterized in that: In the Sp3, the path inference algorithm identifies the guarantee chain risk in the dynamic relationship graph through risk propagation analysis technology. The risk propagation analysis technology dynamically adjusts the weights based on the credit scores of the guarantors and collateral. The dynamic relationship graph updates the status of nodes and edges in real time through event listening technology. The risk assessment results are transmitted to the summary generation module in key-value pair format through the internal data bus.
5. A method for intelligent summarization of non-performing asset related cases according to claim 1, characterized in that: In the Sp4, the semantic priority extraction technology extracts key phrases from structured data through syntactic dependency analysis and keyword priority ranking. The keyword priority ranking is based on the risk weights of non-performing asset cases. The domain template filling technology generates an initial summary through a predefined summary template, and the template is customized according to the case type and user role. The initial summary is transmitted to the feedback optimization module in text format through the internal data bus.
6. The method for intelligent summarization of non-performing asset related cases according to claim 1, characterized in that: In the Sp5, the content priority adjustment algorithm dynamically adjusts the field priorities of the summary template through the scores and annotations in the user feedback data. The user feedback data is stored in Extensible Markup Language (XML) format. The optimized personalized summary is transmitted to the user through the user interface and to the verification module through the internal data bus.
7. A method for intelligently abstracting bad asset related cases according to claim 1, characterized in that: In the Sp6, the formal rules check the logical relationship between the debt amount and the collateral value in the personalized summary through logical reasoning technology. The fact consistency verification technology verifies the field consistency between the personalized summary and the structured data through field matching technology. The verified summary is transmitted to the blockchain storage module in encrypted text format through the internal data bus.
8. A method for intelligent summarization of non-performing asset related cases according to claim 1, characterized in that: In the Sp7, the smart contract controls the summary generation and update through predefined event trigger conditions, and the event trigger conditions include case status changes and user requests. The distributed ledger technology synchronizes the verified summary data through a consensus algorithm. The metadata includes the case number, generation time, and content hash generated through the Secure Hash Algorithm (SHA).
9. A method for intelligently abstracting bad asset related cases according to claim 1, characterized in that: Before the multi-modal data enters the domain adaptation template matching technology in the Sp1, the data cleaning module performs format normalization operations. The format normalization operations remove the irrelevant formats from the legal document text data and standardize the numerical values in the financial table data. The cleaned data is transmitted to the subsequent operations of Sp1 in XML format through the internal data bus.
10. A method for intelligent summarization of non-performing asset related cases according to claim 1, characterized in that: In Sp3, the dynamic relationship graph enhances the weights of edges in the graph database through the relationship priority assignment technique. The relationship priority assignment is based on the statistical analysis of the credit scores of the guarantors and mortgaged properties and historical case data. The risk assessment results are transmitted to the summary generation module in the form of key-value pairs through the internal data bus. In Sp1, the domain adaptation template matching technique optimizes template selection through the semantic similarity matching algorithm. The semantic similarity matching algorithm calculates the matching degree between the text data and the template based on the legal and financial term libraries of non-performing asset cases. The structured data including the case number, debtor, debt amount, mortgaged property, guarantor, and case status is transmitted to the encryption processing module in the Extensible Markup Language (XML) format through the internal data bus. In Sp2, the multi-party secure aggregation protocol optimizes the computing efficiency through the hierarchical aggregation technique. The hierarchical aggregation technique calculates the statistical information in layers to reduce the communication overhead. The encrypted fragments and aggregation values are transmitted to the graph construction module after being encrypted by the Transport Layer Security (TLS) protocol. In Sp4, the domain template filling technique selects the summary template according to the user role through the dynamic template selection algorithm. The dynamic template selection algorithm is based on the user's historical feedback data and case types. The initial summary is transmitted to the feedback optimization module in the text format through the internal data bus.
Citation Information
Patent Citations
Cross-block chain interaction method and system, computer equipment and storage medium
CN110020956A
Case knowledge graph construction platform and method based on big data technology
CN116662559A
Intelligent case generation method based on judicial knowledge domain graph
CN119443249A
Archive management system based on cloud archive library
CN120011617A
Cited By
Smart energy management system based on multi-source heterogeneous data dynamic fusion
CN120781262A
Question answering method and device for beacon information, question answering equipment and storage medium
CN121166855A