Enterprise feature analysis method and device based on multi-document mapping, equipment and medium
Patent Information
- Application Number
- CN202610801983.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-09-11
AI Technical Summary
[0003]本发明提供一种基于多文档映射的企业特征分析方法、装置、计算机设备及介质,以解决目前市场上已有的企业特征分析方法准确度低的问题
[0008]In the above-mentioned enterprise feature analysis method, apparatus, computer equipment, and storage medium based on multi-document mapping, the original transaction document data of the target enterprise in multiple transaction scenarios can be obtained through a preset multi-source document mapping system. A preset standardized template is then used to perform unified encoding format conversion, field name normalization, and document category labeling on the original transaction document data to generate standardized document data. Based on preset entity recognition rules and numerical extraction algorithms, the numerical data, descriptive information, and key entities in the standardized document data are extracted in a structured manner to form structured information records. Finally, the numerical data, descriptive information, and key entities within the same transaction in the structured information records are aligned to obtain... After aligning the transaction information records, physical logic verification is performed on the original transaction document data to generate physical logic verification results. Temporal logic verification is then performed on the aligned transaction information records to generate temporal logic verification results. Based on the physical logic verification results and the temporal logic verification results, the transaction chain integrity index of the target enterprise is calculated. The target enterprise's historical repayment records, tax indicator data, and order completion rate are obtained. A pre-built multi-dimensional credit weighted calculation model is used to analyze the target enterprise's credit characteristic values based on the transaction chain integrity index, the enterprise's historical repayment records, tax indicator data, and the order completion rate, thereby obtaining credit characteristic value analysis results and improving the accuracy of enterprise credit characteristic analysis.
Smart Images

Figure CN122736774A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis technology, and in particular to a method, apparatus, device, and medium for enterprise feature analysis based on multi-document mapping. Background Technology
[0002] In the fintech sector, traditional risk control models for supply chain finance and credit risk control targeting micro and small enterprises (MSEs) rely heavily on standardized financial statements and fixed asset collateral, making it difficult for asset-light enterprises with fragmented operating data and inadequate financial systems to obtain financing. While existing technologies, such as OCR, are beginning to extract key fields from single documents like invoices and contracts, significant shortcomings remain: First, unstructured data is fragmented, lacking the ability to perform cross-document semantic association and cross-validation of multiple source documents like invoices, logistics orders, and contracts, resulting in the waste of a large amount of soft information reflecting the true operating situation. Second, transaction authenticity verification methods are limited, making it difficult to perform closed-loop verification of the transaction chain integrity from both physical logic (e.g., quantity and weight matching) and temporal logic (e.g., the interval between contract, delivery, and invoicing dates), making it difficult to identify fraudulent transactions. Third, risk assessment is time-consuming, failing to achieve real-time processing and dynamic credit granting, thus failing to meet the immediate financing needs of MSEs. Summary of the Invention
[0003] This invention provides a method, apparatus, computer equipment, and medium for enterprise feature analysis based on multi-document mapping, in order to solve the problem of low accuracy of existing enterprise feature analysis methods on the market.
[0004] Firstly, a method for enterprise feature analysis based on multi-document mapping is provided, including: The system acquires the target company's original transaction document data across multiple transaction scenarios based on a pre-defined multi-source document mapping system. The original transaction document data is converted into a unified encoding format, field names are normalized, and document category tags are labeled by calling a standardized template with a preset format, thereby generating standardized document data; Based on preset entity recognition rules and numerical extraction algorithms, the numerical data, descriptive information and key entities in the standardized document data are extracted in a structured manner to form a structured information record. Align the numerical data, descriptive information, and key entities within the same transaction in the structured information record to obtain the aligned transaction information record; Perform physical logic verification on the original transaction document data to generate physical logic verification results, and perform time-series logic verification on the aligned transaction information records to generate time-series logic verification results. Calculate the transaction chain integrity index of the target enterprise based on the physical logic verification results and the temporal logic verification results; The system obtains the target company's historical repayment records, tax indicator data, and order completion rate. Using a pre-built multi-dimensional credit weighted calculation model, it analyzes the target company's credit characteristic value based on the transaction chain integrity index, the company's historical repayment records, the tax indicator data, and the order completion rate, and obtains the credit characteristic value analysis results.
[0005] Secondly, a multi-document mapping-based enterprise feature analysis device is provided, including: The data acquisition module is used to acquire the original transaction document data of the target enterprise in multiple transaction scenarios based on a preset multi-source document mapping system; The data extraction module is used to call a standardized template with a preset format to perform unified encoding format conversion, field name normalization and document category labeling on the original transaction document data to generate standardized document data. Based on preset entity recognition rules and numerical extraction algorithms, the module performs structured extraction of numerical data, descriptive information and key entities in the standardized document data to form structured information records. The data alignment module is used to align numerical data, descriptive information, and key entities within the same transaction in the structured information record to obtain aligned transaction information records. The index calculation module is used to perform physical logic verification on the original transaction document data and generate physical logic verification results, perform temporal logic verification on the aligned transaction information records and generate temporal logic verification results, and calculate the transaction chain integrity index of the target enterprise based on the physical logic verification results and the temporal logic verification results. The feature analysis module is used to obtain the target company's historical repayment records, tax indicator data, and order completion rate. It uses a pre-built multi-dimensional credit weighted calculation model to analyze the target company's credit feature values based on the transaction chain integrity index, the company's historical repayment records, the tax indicator data, and the order completion rate, and obtains the credit feature value analysis results.
[0006] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described enterprise feature analysis method based on multi-document mapping.
[0007] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the enterprise feature analysis method based on multi-document mapping described above.
[0008] In the above-mentioned enterprise feature analysis method, apparatus, computer equipment, and storage medium based on multi-document mapping, the original transaction document data of the target enterprise in multiple transaction scenarios can be obtained through a preset multi-source document mapping system. A preset standardized template is then used to perform unified encoding format conversion, field name normalization, and document category labeling on the original transaction document data to generate standardized document data. Based on preset entity recognition rules and numerical extraction algorithms, the numerical data, descriptive information, and key entities in the standardized document data are extracted in a structured manner to form structured information records. Finally, the numerical data, descriptive information, and key entities within the same transaction in the structured information records are aligned to obtain... After aligning the transaction information records, physical logic verification is performed on the original transaction document data to generate physical logic verification results. Temporal logic verification is then performed on the aligned transaction information records to generate temporal logic verification results. Based on the physical logic verification results and the temporal logic verification results, the transaction chain integrity index of the target enterprise is calculated. The target enterprise's historical repayment records, tax indicator data, and order completion rate are obtained. A pre-built multi-dimensional credit weighted calculation model is used to analyze the target enterprise's credit characteristic values based on the transaction chain integrity index, the enterprise's historical repayment records, tax indicator data, and the order completion rate, thereby obtaining credit characteristic value analysis results and improving the accuracy of enterprise credit characteristic analysis. Attached Figure Description
[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a schematic diagram of an application environment for an enterprise feature analysis method based on multi-document mapping in one embodiment of the present invention; Figure 2 This is a flowchart illustrating an enterprise feature analysis method based on multi-document mapping in one embodiment of the present invention; Figure 3 This is a schematic diagram of a multi-document mapping-based enterprise feature analysis device according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0011] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0012] The enterprise feature analysis method based on multi-document mapping provided in this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can use the client to determine a forgetting reward value based on the target speech and preset timbre forgetting prompts using a preset first reward function. Based on the target speech and preset timbre retention prompts, it uses the first reward function to determine a retention reward value. A diverse speech sample set is generated based on the forgetting and retention reward values. The reward value for each speech sample in the diverse speech sample set is calculated using a preset second reward function, resulting in a reward value set. The parameters of a pre-acquired text-to-speech model are updated based on the reward value set to obtain an updated text-to-speech model. Finally, the server obtains the data generated based on the updated text-to-speech model. The synthesized speech is compared with the preset target prompt speech in terms of timbre similarity and accuracy. If the timbre similarity is less than a preset similarity threshold or the accuracy is less than a preset accuracy threshold, the forgetting reward value and the retention reward value are adjusted according to the timbre similarity and the accuracy. Then, the process returns to the step of generating a diverse set of speech samples based on the forgetting reward value and the retention reward value. If the timbre similarity is greater than or equal to the similarity threshold and the accuracy is greater than or equal to the accuracy threshold, the updated text-to-speech model is confirmed as the final model, thus achieving timbre forgetting during the text-to-speech process. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0013] Please see Figure 2 As shown, Figure 2 A flowchart illustrating the enterprise feature analysis method based on multi-document mapping provided in this embodiment of the invention includes the following steps: S1. Obtain the original transaction document data of the target enterprise in multiple transaction scenarios based on the preset multi-source document mapping system.
[0014] In the fintech field, particularly in supply chain finance, SME lending, and intelligent risk control scenarios, the raw transaction document data refers to original business vouchers collected directly from a company's daily operations without any standardized cleaning or structuring processing. These vouchers typically exist in unstructured or semi-structured form. The sources of this data are extremely diverse, including electronic invoices (PDF / OFD format) exported from a company's ERP system, purchase and sales contracts (Word / PDF scans), logistics waybills (images or Excel files), warehouse receipts, customs declarations, bank statements, etc. Its core characteristics lie in format heterogeneity (significant differences in field naming and layout structure for the same type of document across different companies or systems), unstructured content (a large amount of information embedded in free text, tables, stamps, or handwritten areas), and multimodality (including text, images, tables, and other formats). This type of data truthfully records the transaction trajectory of enterprises throughout the entire chain of procurement, production, sales, and logistics. It contains key soft information reflecting the enterprise's ability to fulfill its obligations (such as counterparties, product specifications, time nodes, physical quantities, etc.). However, due to the lack of unified standards, traditional risk control systems cannot use it directly. It needs to be deeply mined through technologies such as OCR recognition, NLP parsing, and entity alignment in order to transform it into quantifiable, verifiable, and assessable credit assets.
[0015] In this embodiment of the invention, the original transaction document data includes the target company's invoice documents, contract documents, logistics documents, warehousing documents, and order documents.
[0016] In this embodiment of the invention, the acquisition of original transaction document data of a target enterprise in multiple transaction scenarios based on a preset multi-source document mapping system is achieved by pre-constructing a multi-source document mapping configuration library. This configuration library defines the connection protocols, authentication methods, and data mapping rules with different external data sources (such as enterprise ERP systems, supply chain platforms, electronic invoice platforms, logistics company interfaces, bank transaction records, etc.). When a credit assessment of the target enterprise is required, the system automatically calls the API interfaces of each data source or accesses file transfer services through the multi-source document mapping system according to preset trigger conditions (such as the enterprise initiating a financing application) or scheduled tasks. The system then uniformly collects original transaction document data (including invoices, contracts, logistics orders, warehouse orders, customs declarations, etc.) scattered in different systems and formats according to the configured data mapping rules.
[0017] In this embodiment of the invention, the original transaction document data of the target enterprise in multiple transaction scenarios is obtained by using a preset multi-source document mapping system, thereby realizing the automated collection of heterogeneous document data of the enterprise and solving the problems of single data source and low acquisition efficiency.
[0018] S2. Call the preset standardized template to perform unified encoding format conversion, field name normalization and document category labeling on the original transaction document data to generate standardized document data.
[0019] In this embodiment of the invention, the step of calling a pre-formatted standardized template to perform unified encoding format conversion, field name normalization, and document category tagging on the original transaction document data to generate standardized document data includes: Identify text data and non-text data in the original transaction document data, use a preset parser to identify the text in the non-text data, and extract and parse the text data; The text data and the parsed text data are encoded into a preset standard character format to form unified encoded document data; The field names of the unified coded document data are normalized to generate normalized document data; The normalized document data is labeled with document category tags to generate standardized document data.
[0020] In detail, the process involves identifying text and non-text data in the original transaction document data, using a preset parser to identify the text in the non-text data, and extracting and parsing the text data. The non-text data can refer to image data and PDF documents. OCR technology can be used to recognize text in images, or embedded text can be extracted through a PDF parsing library, thereby converting the non-text data into processable text content and extracting and parsing the text data.
[0021] In detail, the process of encoding the text data and the parsed text data into a preset standard character format to form unified encoded document data uses UTF-8 encoding to eliminate differences in character sets caused by different systems or sources, avoid garbled characters for Chinese characters and special characters, and thus form unified encoded document data with uniform format and consistent character encoding, providing a standardized data foundation for subsequent field parsing and normalization.
[0022] In detail, the process of normalizing field names in the unified coded document data to generate normalized document data involves using a predefined field mapping template to uniformly map fields with the same meaning but different names (e.g., "invoice number", "ticket number", "invoice_no") in different document types (such as invoices, contracts, and logistics orders) to standard field names, while retaining the original field names as extended information. This eliminates the heterogeneity of field names and results in normalized document data with consistent field names.
[0023] In detail, the process of labeling the normalized document data with document category tags to generate standardized document data is based on preset document classification rules, combined with the document's structural features, key fields, keywords and other information, to assign corresponding category tags (such as "invoice", "contract", "logistics order") to each document, so that the corresponding parsing and verification logic can be called according to different document categories in the future.
[0024] In this embodiment of the invention, the original transaction document data is converted into a unified encoding format, field names are normalized, and document category tags are marked by calling a standardized template with a preset format, thereby generating standardized document data. This improves the accuracy and efficiency of subsequent information extraction and reduces the parsing error rate caused by format differences.
[0025] S3. Based on preset entity recognition rules and numerical extraction algorithms, the numerical data, descriptive information and key entities in the standardized document data are extracted in a structured manner to form a structured information record.
[0026] In this embodiment of the invention, the step of extracting structured data, descriptive information, and key entities from the standardized document data based on preset entity recognition rules and numerical extraction algorithms to form structured information records includes: A parser template is selected from a preset parser configuration library based on the category labels of the standardized document data; Based on the parser template, the key fields of the standardized document data are located to obtain the document data after field location; The parser template is used to extract numerical data and descriptive information from the document data after the field is located; Entity recognition is performed on the document data after the field is located to obtain the entity recognition result; The document data after the field is located, the numerical data, the descriptive information, and the entity recognition result are converted into JSON objects according to the preset data format template to form a structured information record.
[0027] In detail, the step of selecting a parser template from a pre-defined parser configuration library based on the category tags of the standardized document data involves reading the document category tags (such as "invoice," "contract," "logistics order," etc.) carried by each standardized document, using these tags as index keys, and matching them in the pre-built parser configuration library to select the parser template corresponding to that document type. Each parser template encapsulates a list of key fields to be extracted for that type of document, field location rules, numerical extraction rules, and entity recognition rules, providing targeted parsing guidance for subsequent steps.
[0028] In detail, said obtaining document data after field positioning by positioning the key fields of the standardized document data according to the parser template refers to traversing the standardized document and positioning the specific position of each key field in the document according to the field positioning rules defined in the parser template (such as keyword matching, XPath path, regular expression, layout coordinate range, etc.). For structured table documents, the system identifies the column where the field is located through header matching; for semi-structured or unstructured documents, the text block corresponding to the field is determined through adjacent keywords or layout analysis.
[0029] In detail, said extracting numeric data and descriptive information from the document data after field positioning respectively by using the parser template refers to, by traversing all fields marked as numeric type in the document data after field positioning (such as "amount", "quantity", "tax rate", "unit price", etc.), for the text content of each field, cleaning and removing non-numeric interference items such as currency symbols, thousands separators and unit characters through regular expressions; then performing digital conversion on numbers represented in Chinese characters (such as "twelve thousand yuan in full"); for a text block containing multiple numerical values, accurate interception can be performed in combination with field context; finally converting the extraction result into a standard numerical type; and for fields marked as descriptive type in the parser template (such as "cargo name and specification", "contract terms", "remark information", etc.), extraction is performed by using a rule-based method. The extraction rules comprehensively consider the field position, document structure boundaries (such as table row range, paragraph separators) and keyword anchoring, and intercept the corresponding free text content from the document. For descriptive information spanning multiple lines or multiple cells, the system automatically judges the start and end positions of the text through a boundary recognition algorithm, performs basic cleaning on the extraction result (such as removing redundant line breaks, merging line-broken text), retains the integrity of the original semantics, and marks whether there is truncation or an abnormality.
[0030] In detail, said performing entity recognition on the document data after field positioning to obtain an entity recognition result refers to extracting comprehensive key entities from the document data after field positioning based on preset entity recognition rules, wherein the entity types include but are not limited to order number, contract number, enterprise name, taxpayer identification number, date, address, etc., and recognition is performed by adopting a combination strategy: for entities with standard formats (such as order number, date), regular expressions are used for pattern matching; for open text entities such as enterprise names, extraction is performed by combining dictionary matching and a lightweight named entity recognition model; for all recognized entities, the system performs unified standardization processing (such as unifying upper and lower case, removing redundant spaces, standardizing abbreviated forms), and summarizes information such as entity type, entity value and occurrence position into an entity recognition result.
[0031] In detail, the step of converting the document data after field positioning, the numerical data, the descriptive information, and the entity recognition results into a JSON object according to a preset data format template to form a structured information record involves uniformly integrating all the extraction results generated in the above steps and organizing them into a standard JSON structure according to the preset data format template. This JSON object contains multiple layers: a document metadata layer records basic information such as document identifiers, category tags, and source files; a field extraction layer stores the name, extracted value, value type, and confidence level of each key field in key-value pairs; an entity list layer centrally stores all identified key entities and their attributes; and field positioning information is retained as optional traceability data.
[0032] In this embodiment of the invention, numerical data, descriptive information and key entities in the standardized document data are extracted in a structured manner based on preset entity recognition rules and numerical extraction algorithms to form structured information records, providing directly calculable input for subsequent transaction chain integrity verification.
[0033] S4. Align the numerical data, descriptive information and key entities within the same transaction in the structured information record to obtain the aligned transaction information record.
[0034] In this embodiment of the invention, aligning numerical data, descriptive information, and key entities within the same transaction in the structured information record to obtain an aligned transaction information record includes: Based on the key entities in the structured information records, the structured information records are clustered by transaction to form a structured information record after transaction clustering; The structured information records after transaction clustering are subjected to entity standardization and conflict resolution to obtain entity-standardized transaction data. The standardized transaction data of the entities is subjected to cross-document mapping and numerical alignment processing to generate aligned transaction data. Construct a complete view of the transaction chain based on the aligned transaction data; The aligned transaction information record is generated based on the transaction chain integrity view.
[0035] In detail, the process of clustering the structured information records based on key entities in the records to form clustered structured information records involves traversing all structured information records, extracting key entities (such as order number, contract number, buyer name, seller name, transaction date, etc.) from each record, and grouping documents sharing the same or highly similar key entities into the same transaction cluster according to preset transaction clustering rules. The clustering rules employ a strategy combining exact matching and fuzzy matching: for unique identifiers such as order number and contract number, exact matching is directly performed; when a unique identifier is missing, combined entity matching (such as "buyer name + transaction date + amount") is used, with similarity calculation and threshold setting for determination. After clustering, the system assigns a unique transaction identifier to each transaction cluster and populates it into all its associated structured information records.
[0036] In detail, the process of performing entity standardization and conflict resolution on the structured information records after transaction clustering to obtain entity-standardized transaction data involves uniformly standardizing key entities (such as company name, order number, date, tax ID, etc.) in all document records within the same transaction cluster. This normalizes different representations of the same entity in different documents into a standard format. For entities with similar names but different actual entities, ambiguity is eliminated by combining contextual information such as tax ID and address. Simultaneously, the system detects and handles entity conflicts: when different documents record the same entity inconsistently, conflict resolution is performed according to preset confidence rules (such as prioritizing invoice records, using a majority voting mechanism, or marking it as pending verification), and a conflict handling log is recorded to ensure the consistency and accuracy of entity information within the transaction.
[0037] In detail, the process of performing cross-document mapping and numerical alignment on the standardized transaction data to generate aligned transaction data involves identifying the field correspondences between different document types within the same transaction cluster and establishing a cross-document field mapping table. For example, the "total amount" in an invoice is mapped to the same semantic field as the "total contract price" in a contract, and the "weight of goods" in a logistics order is physically associated with the "quantity of goods" in an invoice. Based on this, the system performs alignment verification on the numerical data in each document: extracting the values of the same semantic field in different documents and verifying their logical consistency (e.g., whether the invoice amount equals the total contract price). For discrepancies, alignment is performed according to preset rules (e.g., based on the invoice, based on the contract, or marked as pending verification), and the difference value and alignment basis are recorded. For descriptive information, the system aggregates and associates texts describing the same transaction element in each document.
[0038] In detail, constructing a transaction chain integrity view based on the aligned transaction data involves integrating the alignment results of all documents within the same transaction cluster into a unified transaction chain integrity view. This view is organized according to a preset data model and includes multiple layers: a transaction basic information layer records transaction identifiers, transaction types, participant information, and a transaction timeline (contract signing date, delivery date, invoicing date); a document list layer lists all document types, document identifiers, and source information involved in the transaction; a key field alignment layer lists corresponding numerical and descriptive values from different documents based on semantic fields, and marks the alignment methods and differences; and an entity unified view layer stores standardized entity information and its mapping relationship in the original documents. Furthermore, the system performs an integrity check on the transaction chain view, identifies missing key document types (such as invoices without corresponding logistics orders), and clearly marks missing items in the view, providing an intuitive basis for subsequent transaction chain integrity verification.
[0039] In detail, generating aligned transaction information records based on the transaction chain integrity view involves transforming the transaction chain integrity view into standardized aligned transaction information records. Each record corresponds to a complete transaction and is stored in a structured JSON format. This record includes the following core modules: transaction metadata (transaction identifier, transaction status, standardized information of participants, transaction timeline), associated document index (a list of all document IDs involved in this transaction and their respective types), numerical alignment results (using semantic fields as keys to store aligned numerical values, original source document values, difference markers, and alignment confidence), entity alignment results (a collection of standardized entities, including entity type, standard value, and original document mapping relationship), description aggregation results (aligned text and source document references for each description field), and integrity markers (specific information identifying whether the transaction chain is complete, missing, or conflicting). The final output aligned transaction information records serve as a unified transaction-level data foundation, directly accessible to the subsequent transaction chain integrity verification module and performance credit index calculation module.
[0040] In this embodiment of the invention, by aligning the numerical data, descriptive information, and key entities within the same transaction in the structured information record, an aligned transaction information record is obtained, realizing the semantic association of multi-source data in the transaction dimension, and providing a complete transaction chain data foundation for subsequent physical logic verification and temporal logic verification.
[0041] S5. Perform physical logic verification on the original transaction document data to generate physical logic verification results, and perform time-series logic verification on the aligned transaction information records to generate time-series logic verification results.
[0042] In this embodiment of the invention, the step of performing physical logic verification on the original transaction document data and generating a physical logic verification result includes: Identify the document types related to physical logic verification in the original transaction document data, extract key physical quantities based on the document types, and obtain the physical verification input dataset; Perform physical mapping rule matching on the physical verification input dataset to generate physical verification data with rule annotations; Physical consistency calculation is performed on the physical verification data with rule annotations to obtain a consistency calculation result containing transaction ID, physical consistency score, and inconsistency details list; The physical logic verification result is generated based on the consistency calculation result.
[0043] In detail, the process of identifying document types related to physical logic verification in the original transaction document data, extracting key physical quantities based on the document types, and obtaining a physical verification input dataset involves filtering document types related to physical verification (mainly invoices and logistics orders) based on document category tags. For invoice documents, key physical quantities such as goods name, goods quantity, unit of measurement, and net / gross weight are extracted. For logistics order documents, physical information such as total weight, number of packages, and transportation method are extracted. The extraction process uses a preset physical field mapping template to obtain corresponding values from the structured information records of the documents. If the document is not yet structured, OCR and NLP modules are temporarily called for real-time extraction. Finally, the extracted physical data is associated and organized according to transaction identifiers and document IDs to form a structured physical verification input dataset.
[0044] In detail, the physical mapping rule matching of the physical verification input dataset to generate rule-annotated physical verification data is achieved by loading a pre-built physical mapping rule library. This rule library defines reasonable mapping relationships between invoice physical quantities and logistics physical quantities for different types of goods, packaging methods, and transportation scenarios (e.g., steel goods are weighed by ton, with an allowable error of ±5%; the number of pieces for ordinary goods must match the number of pieces exactly). The system matches the corresponding mapping rules based on the goods specification description in the invoice.
[0045] In detail, the physical consistency calculation of the rule-annotated physical verification data to obtain a consistency calculation result including transaction ID, physical consistency score, and inconsistency details list is performed for each transaction. Based on the matched physical mapping rules, the difference between the theoretical logistics physical quantity (e.g., invoice quantity multiplied by unit standard weight) and the actual logistics order physical quantity is calculated. For direct matching rules (e.g., number of pieces to number of pieces), the exact matching degree is calculated, with a perfect match score of 1. When a difference exists, points are deducted according to the difference rate using a preset piecewise function. For linear mapping rules (e.g., weight correspondence), the relative error is calculated. Full marks are awarded if the error is within the allowable range, and points are deducted proportionally up to zero if it exceeds the threshold. The system records the physical consistency score for each transaction and generates an inconsistency details list, detailing the difference type, difference value, and possible causes.
[0046] In detail, generating a physical logic verification result based on the consistency calculation result involves comparing the physical consistency score with a preset verification threshold. If the score is greater than or equal to the threshold, the physical logic verification result of the transaction is determined to be "passed"; otherwise, it is determined to be "failed" and a warning flag is triggered. For transactions that fail verification, the system summarizes all issues in the inconsistency details list and generates a structured problem description text (e.g., "The invoice quantity is 100 tons, but the logistics order weight is only 48 tons, the difference exceeds the allowable range"). The verification result status, physical consistency score, inconsistency details, verification time, rule version, and other information are associated with the transaction identifier to form a standardized physical logic verification result.
[0047] In this embodiment of the invention, the step of performing time-series logic verification on the aligned transaction information records and generating a time-series logic verification result includes: Extract the time-series related fields from the aligned transaction information records to form a time-series input dataset; The time series input dataset is matched with preset time series rules to obtain time series data with rule annotations; Perform time series consistency calculation on the time series data with rule annotations, and generate time series consistency calculation results; The timing consistency calculation results are used to generate timing logic verification results.
[0048] In detail, the step of extracting time-related fields from the aligned transaction information records to form a time-series input dataset involves traversing all aligned transaction information records and extracting key fields related to time nodes, mainly including contract signing date, shipping date (or the start date on the logistics bill), invoice issuance date, and receipt confirmation date. The extracted time fields are then standardized in format, uniformly converted to YYYY-MM-DD format, and the date validity is verified.
[0049] In detail, the process of matching the time-series input dataset with preset time-series rules to obtain time-series data with rule annotations is achieved by loading a pre-built time-series rule library. This rule library defines reasonable time interval ranges for different industries and business scenarios (e.g., shipping within 3-15 days after contract signing and invoicing within 5-10 days after shipment in the manufacturing industry). It also supports dynamic rule generation: for enterprises that cannot match specific industry rules, the system uses the statistical values of the time intervals of the enterprise's historical transactions (e.g., calculating the percentiles of the average shipping interval and invoicing interval over the past 12 months) as personalized rules. The system matches the corresponding time-series rule set according to the type of goods or enterprise profile of each transaction, and binds the matched rules to the transaction to generate time-series data with rule annotations.
[0050] In detail, the step of performing time-series consistency calculation on the rule-annotated time-series data and generating time-series consistency calculation results involves calculating key time intervals for each transaction (such as the number of days from contract signing to delivery, or the number of days from delivery to invoicing), comparing the calculated intervals with the reasonable ranges defined in the matching rules, and assigning a score of 1 to the corresponding sub-item if the interval is within the reasonable range, and deducting points according to a preset function if the interval exceeds the range (e.g., 0.8 for exceeding by less than 10%, and 0 for exceeding by more than 50%). If a time node is missing, the corresponding sub-item score is 0 or weighted according to the degree of impact of the missing node. The total time-series consistency score of the transaction is obtained by combining the scores of each sub-item using a weighted average or a preset weighted fusion. At the same time, the score of each sub-item and the details of inconsistencies are recorded (e.g., "delivery was only made 20 days after contract signing, exceeding the industry standard by 3-15 days"). The output is a time-series consistency calculation result containing the transaction ID, the total time-series consistency score, the sub-item score list, and the details of inconsistencies.
[0051] In detail, generating the timing logic verification result based on the timing consistency calculation result involves comparing the total timing consistency score with a preset verification threshold (e.g., 0.7). If the total score is greater than or equal to the threshold, the timing logic verification result of the transaction is determined to be "passed"; otherwise, it is determined to be "failed" and a warning flag is triggered. For transactions that fail verification, the system summarizes the inconsistency details to generate a structured problem description (e.g., "Shipping delay exceeds industry standards, which may affect the authenticity of the transaction"). At the same time, if there are missing key time nodes, the missing items are clearly marked in the result. Finally, the verification result status, total timing consistency score, inconsistency details, missing flags, and verification time are associated with the transaction identifier to form a standardized timing logic verification result.
[0052] In this embodiment of the invention, by performing physical logic verification on the original transaction document data to generate physical logic verification results, and performing temporal logic verification on the aligned transaction information records to generate temporal logic verification results, dual verification of the authenticity of the transaction is achieved, providing a scientific and quantifiable basis for determining the transaction chain integrity index.
[0053] S6. Calculate the transaction chain integrity index of the target enterprise based on the physical logic verification results and the timing logic verification results.
[0054] In this embodiment of the invention, calculating the transaction chain integrity index of the target enterprise based on the physical logic verification result and the temporal logic verification result includes: The physical logic verification results and the timing logic verification results are merged to obtain a verification dataset. Using a pre-defined calculation model, the physical consistency score and temporal consistency score corresponding to each transaction in the verification dataset are weighted and summed to generate a set of transaction chain integrity indices. The transaction chain integrity index set is dynamically weighted based on the time decay algorithm to obtain the weight set. The transaction chain integrity index is obtained by performing a weighted summation on the transaction chain integrity index set based on the weight set.
[0055] In detail, the step of using a preset calculation model to perform a weighted summation of the physical consistency score and the temporal consistency score corresponding to each transaction in the verification dataset to generate a transaction chain integrity index set involves calling a preset weighted summation model for each transaction in the verification dataset. This model configures the weight coefficients of physical verification and temporal verification according to the business scenario (e.g., each accounting for 0.5 by default), multiplies the physical consistency score by the weight coefficient, multiplies the temporal consistency score by the corresponding weight coefficient, and then adds the two products to obtain the transaction chain integrity index of the transaction.
[0056] In detail, the time decay algorithm is used to assign dynamic weights to the transaction chain integrity index set, resulting in a weight set. For each transaction in the transaction chain integrity index set, the corresponding transaction occurrence time (such as the invoice issuance date or contract signing date) is extracted, and a preset time decay function (such as an exponential decay function or a linear decay function) is used to calculate the dynamic weight for each transaction. The design principle of the time decay function is that the closer the transaction occurrence time is to the current assessment time, the higher the weight is assigned, reflecting that recent transactions are more representative of the company's current performance capability, and vice versa, to avoid the excessive influence of long-standing transactions on the current credit assessment. The system associates the calculated dynamic weights with the corresponding transactions to generate a weight set containing transaction identifiers, transaction chain integrity indices, and dynamic weights.
[0057] In this embodiment of the invention, the transaction chain integrity index of the target enterprise is calculated based on the physical logic verification results and the temporal logic verification results, transforming the originally scattered verification results into a unified evaluation index, enabling the system to intuitively measure the enterprise's performance credibility in the supply chain.
[0058] S7. Obtain the target company's historical repayment records, tax indicator data, and order completion rate. Utilize a pre-built multi-dimensional credit weighted calculation model to analyze the target company's credit characteristic value based on the transaction chain integrity index, the company's historical repayment records, the tax indicator data, and the order completion rate, and obtain the credit characteristic value analysis results.
[0059] In this embodiment of the invention, the method of using a pre-constructed multi-dimensional credit weighted calculation model to analyze the credit characteristic value of the target enterprise based on the transaction chain integrity index, the enterprise's historical repayment records, the tax indicator data, and the order completion rate, and obtaining the credit characteristic value analysis result, is achieved by pre-constructing a multi-dimensional credit weighted calculation model. This model defines the calculation formula for the credit characteristic value as follows: PCI=w1 TCC+w2 Prepay+w3 Ocomp Wherein, TCC is the transaction chain integrity index, Prepay is the quantitative score of the repayment performance of the enterprise's historical repayment records, Ocomp is the order completion rate, and w1, w2, and w3 are preset weight coefficients.
[0060] In this embodiment, by utilizing a pre-built multi-dimensional credit weighted calculation model to analyze the credit characteristic value of the target enterprise based on the transaction chain integrity index, the enterprise's historical repayment records, the tax indicator data, and the order completion rate, the credit characteristic value analysis results are obtained, thereby achieving a comprehensive and quantitative assessment of the enterprise's credit characteristics.
[0061] Furthermore, the tax indicator data can be used as an adjustment factor to be integrated into the multi-dimensional credit weighted calculation model. The transaction chain integrity index of the target enterprise, the enterprise's historical repayment records (converted into standardized repayment performance scores), tax indicator data (such as tax rating, tax amount stability, etc.) and order completion rate are normalized and then input into the model. The model is then weighted and fused according to the preset weight coefficients to obtain a comprehensive credit feature value.
[0062] In the fintech field, this solution is primarily applied to supply chain finance, SME lending, and intelligent risk control scenarios. It transforms unstructured transaction documents (such as invoices, contracts, and logistics documents) scattered throughout a company's operational chain into quantifiable, verifiable, and traceable credit assets, enabling accurate assessment of the "asset-light, operation-heavy" characteristics of SMEs. Specifically, this solution can be embedded into a bank's corporate lending system or supply chain finance platform. During the financing application process, it automatically collects multi-source transaction data from enterprises, using a dual verification mechanism of physical and temporal logic to identify fraudulent transactions and abnormal behavior. It outputs a real-time transaction chain integrity index and, combined with the enterprise's historical repayment records, tax indicators, and order completion rates, dynamically generates a performance credit index and differentiated credit limits and risk pricing through a multi-dimensional credit weighted model. This application allows banks to move away from reliance on fixed asset collateral and standard financial statements, transitioning from a "static approval" model to a "real-time, automated, and information-driven" risk control model. This significantly improves the accessibility and efficiency of financing for SMEs, while effectively reducing fraud risk and credit losses, providing an innovative solution for fintech to empower the real economy.
[0063] As can be seen, the above-mentioned scheme proposes a multi-document mapping-based enterprise feature analysis method, aiming to solve the problems of insufficient utilization of unstructured data, difficulty in verifying the authenticity of transactions, and poor timeliness of risk assessment in traditional micro and small enterprise credit risk control. The system first automatically collects original transaction document data from enterprises in various transaction scenarios such as procurement, sales, and logistics through a pre-set multi-source document mapping system, covering heterogeneous sources such as invoices, contracts, and logistics orders. Then, it calls standardized templates to uniformly encode and convert the data, normalize field names, and label document categories, forming standardized document data. Based on this, the system accurately extracts numerical data, descriptive information, and key entities from the standardized documents using pre-set entity recognition rules and numerical extraction algorithms, generating structured information records. Finally, through cross-document entity alignment and field mapping, it integrates information from the same transaction scattered across different documents into an aligned transaction information record. To verify the authenticity of transactions, the system performs physical logic checks on the original transaction document data (such as the matching of invoice quantity and logistics weight) and temporal logic checks on the aligned transaction information records (such as the reasonable interval between contract signing, delivery, and invoicing dates). Based on the results of these two checks, the system calculates the Transaction Chain Integrity Index (TCC), which is weighted and aggregated using a time decay mechanism to reflect the overall authenticity of the company's recent transactions. Finally, the system obtains the company's historical repayment records, tax indicator data, and order completion rate. Using a pre-built multi-dimensional credit weighted calculation model, it integrates the Transaction Chain Integrity Index with the above indicators to calculate the company's Performance Credit Index (PCI) in real time, outputting credit feature analysis results, including credit score, credit rating, and risk pricing recommendations. This invention moves traditional risk control from a static model relying on fixed assets and standard financial statements to a new stage of real-time assessment driven by unstructured soft information, effectively solving the financing difficulties of asset-light micro and small enterprises and significantly improving the ability to identify fraudulent transactions.
[0064] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0065] In one embodiment, a multi-document mapping-based enterprise feature analysis device is provided, which corresponds one-to-one with the multi-document mapping-based enterprise feature analysis method described in the above embodiments. For example... Figure 3 As shown, the enterprise feature analysis device based on multi-document mapping includes a data acquisition module 101, a data extraction module 102, a data alignment module 103, an index calculation module 104, and a feature analysis module 105. Detailed descriptions of each functional module are as follows: Data acquisition module 101 is used to acquire the original transaction document data of the target enterprise in multiple transaction scenarios based on a preset multi-source document mapping system; The data extraction module 102 is used to call a standardized template with a preset format to perform unified encoding format conversion, field name normalization and document category labeling on the original transaction document data to generate standardized document data. Based on preset entity recognition rules and numerical extraction algorithms, the module performs structured extraction of numerical data, descriptive information and key entities in the standardized document data to form structured information records. Data alignment module 103 is used to align numerical data, descriptive information and key entities within the same transaction in the structured information record to obtain aligned transaction information records; The index calculation module 104 is used to perform physical logic verification on the original transaction document data and generate physical logic verification results, perform temporal logic verification on the aligned transaction information records and generate temporal logic verification results, and calculate the transaction chain integrity index of the target enterprise based on the physical logic verification results and the temporal logic verification results. The feature analysis module 105 is used to obtain the target company's historical repayment records, tax indicator data, and order completion rate. It uses a pre-built multi-dimensional credit weighted calculation model to analyze the target company's credit feature values based on the transaction chain integrity index, the company's historical repayment records, the tax indicator data, and the order completion rate, and obtains the credit feature value analysis results.
[0066] In one embodiment, when the data extraction module 102 executes the process of calling a preset format standardized template to perform unified encoding format conversion, field name normalization, and document category tagging on the original transaction document data to generate standardized document data, it is specifically used for: Identify text data and non-text data in the original transaction document data, use a preset parser to identify the text in the non-text data, and extract and parse the text data; The text data and the parsed text data are encoded into a preset standard character format to form unified encoded document data; The field names of the unified coded document data are normalized to generate normalized document data; The normalized document data is labeled with document category tags to generate standardized document data.
[0067] In one embodiment, the data extraction module 102, when performing the structured extraction of numerical data, descriptive information, and key entities from the standardized document data based on preset entity recognition rules and numerical extraction algorithms to form structured information records, is specifically used for: A parser template is selected from a preset parser configuration library based on the category labels of the standardized document data; Based on the parser template, the key fields of the standardized document data are located to obtain the document data after field location; The parser template is used to extract numerical data and descriptive information from the document data after the field is located; Entity recognition is performed on the document data after the field is located to obtain the entity recognition result; The document data after the field is located, the numerical data, the descriptive information, and the entity recognition result are converted into JSON objects according to the preset data format template to form a structured information record.
[0068] In one embodiment, the data alignment module 103, when performing the alignment of numerical data, descriptive information, and key entities within the same transaction in the structured information record to obtain aligned transaction information records, is specifically used for: Based on the key entities in the structured information records, the structured information records are clustered by transaction to form a structured information record after transaction clustering; The structured information records after transaction clustering are subjected to entity standardization and conflict resolution to obtain entity-standardized transaction data. The standardized transaction data of the entities is subjected to cross-document mapping and numerical alignment processing to generate aligned transaction data. Construct a complete view of the transaction chain based on the aligned transaction data; The aligned transaction information record is generated based on the transaction chain integrity view.
[0069] In one embodiment, the index calculation module 104, when performing the physical logic verification on the original transaction document data and generating the physical logic verification result, is specifically used for: Identify the document types related to physical logic verification in the original transaction document data, extract key physical quantities based on the document types, and obtain the physical verification input dataset; Perform physical mapping rule matching on the physical verification input dataset to generate physical verification data with rule annotations; Physical consistency calculation is performed on the physical verification data with rule annotations to obtain a consistency calculation result containing transaction ID, physical consistency score, and inconsistency details list; The physical logic verification result is generated based on the consistency calculation result.
[0070] In one embodiment, the index calculation module 104, when performing the time-series logic verification on the aligned transaction information records and generating the time-series logic verification result, is specifically used for: Extract the time-series related fields from the aligned transaction information records to form a time-series input dataset; The time series input dataset is matched with preset time series rules to obtain time series data with rule annotations; Perform time series consistency calculation on the time series data with rule annotations, and generate time series consistency calculation results; The timing consistency calculation results are used to generate timing logic verification results.
[0071] In one embodiment, the index calculation module 104, when performing the calculation of the transaction chain integrity index of the target enterprise based on the physical logic verification result and the temporal logic verification result, is specifically used for: The physical logic verification results and the timing logic verification results are merged to obtain a verification dataset. Using a pre-defined calculation model, the physical consistency score and temporal consistency score corresponding to each transaction in the verification dataset are weighted and summed to generate a set of transaction chain integrity indices. The transaction chain integrity index set is dynamically weighted based on the time decay algorithm to obtain the weight set. The transaction chain integrity index is obtained by performing a weighted summation on the transaction chain integrity index set based on the weight set.
[0072] This invention provides a multi-document mapping-based enterprise feature analysis device, aiming to solve the problems of insufficient utilization of unstructured data, difficulty in verifying transaction authenticity, and poor timeliness of risk assessment in traditional micro and small enterprise credit risk control. The system first automatically collects original transaction document data from various transaction scenarios such as procurement, sales, and logistics through a pre-set multi-source document mapping system, covering heterogeneous sources such as invoices, contracts, and logistics orders. Then, it calls standardized templates to uniformly encode and convert the data, normalize field names, and label document categories, forming standardized document data. Based on this, the system accurately extracts numerical data, descriptive information, and key entities from the standardized documents using pre-set entity recognition rules and numerical extraction algorithms, generating structured information records. Finally, through cross-document entity alignment and field mapping, it integrates information from the same transaction scattered across different documents into an aligned transaction information record. To verify the authenticity of transactions, the system performs physical logic checks on the original transaction document data (such as the matching of invoice quantity and logistics weight) and temporal logic checks on the aligned transaction information records (such as the reasonable interval between contract signing, delivery, and invoicing dates). Based on the results of these two checks, the system calculates the Transaction Chain Integrity Index (TCC), which is weighted and aggregated using a time decay mechanism to reflect the overall authenticity of the company's recent transactions. Finally, the system obtains the company's historical repayment records, tax indicator data, and order completion rate. Using a pre-built multi-dimensional credit weighted calculation model, it integrates the Transaction Chain Integrity Index with the above indicators to calculate the company's Performance Credit Index (PCI) in real time, outputting credit feature analysis results, including credit score, credit rating, and risk pricing recommendations. This invention moves traditional risk control from a static model relying on fixed assets and standard financial statements to a new stage of real-time assessment driven by unstructured soft information, effectively solving the financing difficulties of asset-light micro and small enterprises and significantly improving the ability to identify fraudulent transactions.
[0073] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements server-side functions or steps of a multi-document mapping-based enterprise feature analysis method.
[0074] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a multi-document mapping-based enterprise feature analysis method.
[0075] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: The system acquires the target company's original transaction document data across multiple transaction scenarios based on a pre-defined multi-source document mapping system. The original transaction document data is converted into a unified encoding format, field names are normalized, and document category tags are labeled by calling a standardized template with a preset format, thereby generating standardized document data; Based on preset entity recognition rules and numerical extraction algorithms, the numerical data, descriptive information and key entities in the standardized document data are extracted in a structured manner to form a structured information record. Align the numerical data, descriptive information, and key entities within the same transaction in the structured information record to obtain the aligned transaction information record; Perform physical logic verification on the original transaction document data to generate physical logic verification results, and perform time-series logic verification on the aligned transaction information records to generate time-series logic verification results. Calculate the transaction chain integrity index of the target enterprise based on the physical logic verification results and the temporal logic verification results; The system obtains the target company's historical repayment records, tax indicator data, and order completion rate. Using a pre-built multi-dimensional credit weighted calculation model, it analyzes the target company's credit characteristic value based on the transaction chain integrity index, the company's historical repayment records, the tax indicator data, and the order completion rate, and obtains the credit characteristic value analysis results.
[0076] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: The system acquires the target company's original transaction document data across multiple transaction scenarios based on a pre-defined multi-source document mapping system. The original transaction document data is converted into a unified encoding format, field names are normalized, and document category tags are labeled by calling a standardized template with a preset format, thereby generating standardized document data; Based on preset entity recognition rules and numerical extraction algorithms, the numerical data, descriptive information and key entities in the standardized document data are extracted in a structured manner to form a structured information record. Align the numerical data, descriptive information, and key entities within the same transaction in the structured information record to obtain the aligned transaction information record; Perform physical logic verification on the original transaction document data to generate physical logic verification results, and perform time-series logic verification on the aligned transaction information records to generate time-series logic verification results. Calculate the transaction chain integrity index of the target enterprise based on the physical logic verification results and the temporal logic verification results; The system obtains the target company's historical repayment records, tax indicator data, and order completion rate. Using a pre-built multi-dimensional credit weighted calculation model, it analyzes the target company's credit characteristic value based on the transaction chain integrity index, the company's historical repayment records, the tax indicator data, and the order completion rate, and obtains the credit characteristic value analysis results.
[0077] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0078] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0079] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0080] Finally, it should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application is authorized (with knowledge and consent) by the relevant parties or fully authorized by all parties, and the executing entity can obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals. The above-described embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for enterprise feature analysis based on multi-document mapping, characterized in that, include: The system acquires the target company's original transaction document data across multiple transaction scenarios based on a pre-defined multi-source document mapping system. The original transaction document data is converted into a unified encoding format, field names are normalized, and document category tags are labeled by calling a standardized template with a preset format, thereby generating standardized document data; Based on preset entity recognition rules and numerical extraction algorithms, the numerical data, descriptive information and key entities in the standardized document data are extracted in a structured manner to form a structured information record. Align the numerical data, descriptive information, and key entities within the same transaction in the structured information record to obtain the aligned transaction information record; Perform physical logic verification on the original transaction document data to generate physical logic verification results, and perform time-series logic verification on the aligned transaction information records to generate time-series logic verification results. Calculate the transaction chain integrity index of the target enterprise based on the physical logic verification results and the temporal logic verification results; The system obtains the target company's historical repayment records, tax indicator data, and order completion rate. Using a pre-built multi-dimensional credit weighted calculation model, it analyzes the target company's credit characteristic value based on the transaction chain integrity index, the company's historical repayment records, the tax indicator data, and the order completion rate, and obtains the credit characteristic value analysis results.
2. The enterprise feature analysis method based on multi-document mapping as described in claim 1, characterized in that, The process of calling a pre-defined standardized template to perform unified encoding format conversion, field name normalization, and document category labeling on the original transaction document data to generate standardized document data includes: Identify text data and non-text data in the original transaction document data, use a preset parser to identify the text in the non-text data, and extract and parse the text data; The text data and the parsed text data are encoded into a preset standard character format to form unified encoded document data; The field names of the unified coded document data are normalized to generate normalized document data; The normalized document data is labeled with document category tags to generate standardized document data.
3. The enterprise feature analysis method based on multi-document mapping as described in claim 1, characterized in that, The process involves extracting numerical data, descriptive information, and key entities from the standardized document data based on preset entity recognition rules and numerical extraction algorithms to form structured information records, including: A parser template is selected from a preset parser configuration library based on the category labels of the standardized document data; Based on the parser template, the key fields of the standardized document data are located to obtain the document data after field location; The parser template is used to extract numerical data and descriptive information from the document data after the field is located; Entity recognition is performed on the document data after the field is located to obtain the entity recognition result; The document data after the field is located, the numerical data, the descriptive information, and the entity recognition result are converted into JSON objects according to the preset data format template to form a structured information record.
4. The enterprise feature analysis method based on multi-document mapping as described in claim 1, characterized in that, The process of aligning numerical data, descriptive information, and key entities within the same transaction in the structured information record to obtain aligned transaction information records includes: Based on the key entities in the structured information records, the structured information records are clustered by transaction to form a structured information record after transaction clustering; The structured information records after transaction clustering are subjected to entity standardization and conflict resolution to obtain entity-standardized transaction data. The standardized transaction data of the entities is subjected to cross-document mapping and numerical alignment processing to generate aligned transaction data. Construct a complete view of the transaction chain based on the aligned transaction data; The aligned transaction information record is generated based on the transaction chain integrity view.
5. The enterprise feature analysis method based on multi-document mapping as described in claim 1, characterized in that, The step of performing physical and logical verification on the original transaction document data and generating physical and logical verification results includes: Identify the document types related to physical logic verification in the original transaction document data, extract key physical quantities based on the document types, and obtain the physical verification input dataset; Perform physical mapping rule matching on the physical verification input dataset to generate physical verification data with rule annotations; Physical consistency calculation is performed on the physical verification data with rule annotations to obtain a consistency calculation result containing transaction ID, physical consistency score, and inconsistency details list; The physical logic verification result is generated based on the consistency calculation result.
6. The enterprise feature analysis method based on multi-document mapping as described in claim 1, characterized in that, The step of performing time-series logic verification on the aligned transaction information records and generating time-series logic verification results includes: Extract the time-series related fields from the aligned transaction information records to form a time-series input dataset; The time series input dataset is matched with preset time series rules to obtain time series data with rule annotations; Perform time series consistency calculation on the time series data with rule annotations, and generate time series consistency calculation results; The timing consistency calculation results are used to generate timing logic verification results.
7. The enterprise feature analysis method based on multi-document mapping as described in claim 1, characterized in that, The step of calculating the transaction chain integrity index of the target enterprise based on the physical logic verification results and the temporal logic verification results includes: The physical logic verification results and the timing logic verification results are merged to obtain a verification dataset. Using a pre-defined calculation model, the physical consistency score and temporal consistency score corresponding to each transaction in the verification dataset are weighted and summed to generate a set of transaction chain integrity indices. The transaction chain integrity index set is dynamically weighted based on the time decay algorithm to obtain the weight set. The transaction chain integrity index is obtained by performing a weighted summation on the transaction chain integrity index set based on the weight set.
8. An enterprise feature analysis device based on multi-document mapping, characterized in that, include: The data acquisition module is used to acquire the original transaction document data of the target enterprise in multiple transaction scenarios based on a preset multi-source document mapping system; The data extraction module is used to call a standardized template with a preset format to perform unified encoding format conversion, field name normalization and document category labeling on the original transaction document data to generate standardized document data. Based on preset entity recognition rules and numerical extraction algorithms, the module performs structured extraction of numerical data, descriptive information and key entities in the standardized document data to form structured information records. The data alignment module is used to align numerical data, descriptive information, and key entities within the same transaction in the structured information record to obtain aligned transaction information records. The index calculation module is used to perform physical logic verification on the original transaction document data and generate physical logic verification results, perform temporal logic verification on the aligned transaction information records and generate temporal logic verification results, and calculate the transaction chain integrity index of the target enterprise based on the physical logic verification results and the temporal logic verification results. The feature analysis module is used to obtain the target company's historical repayment records, tax indicator data, and order completion rate. It uses a pre-built multi-dimensional credit weighted calculation model to analyze the target company's credit feature values based on the transaction chain integrity index, the company's historical repayment records, the tax indicator data, and the order completion rate, and obtains the credit feature value analysis results.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the enterprise feature analysis method based on multiple document mapping as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the enterprise feature analysis method based on multiple document mapping as described in any one of claims 1 to 7.