Identity authentication and financial service access management method and system fusing multi-document information

By parsing the semantic units of multiple document information and establishing link relationships, the problem of inconsistent identity associations in the financial service access system has been solved, enabling accurate determination of subject attribution and structured access management.

CN122113070AActive Publication Date: 2026-05-29XIAMEN INFORMATION GRP BIG DATA OPERATION CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAMEN INFORMATION GRP BIG DATA OPERATION CO LTD
Filing Date
2026-04-28
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

The existing financial service access system has an inconsistency in identity association during multi-document authentication, which leads to legitimate companies being denied access due to inconsistent identity verification.

Method used

By acquiring basic data from multiple documents, analyzing the business meaning of fields, forming semantic units, and establishing identity identification links, account links, and entity registration links, the trajectory support quantity and attribution stability value of semantic units are calculated, and entity files are generated to determine the access status.

Benefits of technology

It enables unified processing of information from multiple documents and accurate determination of ownership, avoiding erroneous rejection of access due to a single document, and providing a clear data foundation and structured access management decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122113070A_ABST
    Figure CN122113070A_ABST
Patent Text Reader

Abstract

The application provides a method and system for identity authentication and financial service access management by fusing multi-document information, and relates to the technical field of data processing. The method comprises the following steps: analyzing the business meanings of fields in different documents, merging fields with the same business meanings into semantic units, reading the appearance order of each semantic unit in different documents to form constraint items, calculating the common support degree of semantic units by multiple documents to obtain a track support value, identifying and extracting semantic units with a track support value reaching a preset support threshold as line units, establishing identity link, account link and subject registration link between line units, calculating the stability degree of line units belonging to the same subject to obtain an attribution stability value, determining the access state of a target subject, and obtaining access management decision data. The application can solve the problem that a subject with multiple documents is misjudged as having inconsistent identities and cannot pass the financial service access.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and system for identity authentication and financial service access management that integrates multiple document information. Background Technology

[0002] Existing financial service access systems typically rely on single-document information for identity authentication. For example, when a user applies for a bank loan or opens an online payment account, the system primarily depends on their ID card number or bank card number as a unique identifier, verifying the user's identity by comparing it with information in the public security or UnionPay databases. Some improved systems have introduced a multi-document authentication approach, independently verifying information such as ID cards, driver's licenses, and business licenses, and comprehensively judging the user's identity based on the comparison results. However, such multi-source information verification methods generally use the document as the main thread, employing a step-by-step, independently invoked structure. Different document information is stored in separate databases, and data fusion relies on interface transmission.

[0003] However, the multi-document independent authentication method may have the drawback of inconsistent identity association. Due to the differences in data format, field definition, and verification logic between different documents, the system may not be able to automatically determine whether multiple documents belong to the same user entity. For example, in the scenario of corporate loan access, the legal representative's ID number, business license registration number, and corporate account opening information may have differences in field length or encoding method. When comparing them, the system may fail to match or make incorrect associations, resulting in legitimate companies being denied access due to inconsistent identity verification. Summary of the Invention

[0004] The purpose of this invention is to provide a method and system for identity authentication and financial service access management that integrates multiple document information, in order to solve the problems mentioned in the background art.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: Firstly, a method for identity authentication and financial service access management that integrates multiple document information, the method comprising: Acquire basic data for multiple certificates, parse the business meaning of each field in different certificates, group fields with the same business meaning into semantic units, record the certificate source of each semantic unit, and obtain semantic unit data; Based on the semantic unit data, the order in which each semantic unit appears in different documents is read to form constraint terms, and the constraint terms are compared with the actual values ​​of the semantic units to obtain constraint comparison data; Based on the constraint comparison data, the number of value segments, continuous coverage length and number of repetitions across documents of the constraint sequence are identified, the degree to which the semantic unit is supported by multiple documents is calculated, and the trajectory support quantity is obtained. By identifying and extracting semantic units whose trajectory support reaches a preset support threshold as clue units, and then aggregating the clue units according to the source of the documents, the main clue data is obtained. Based on the main clue data, establish identity identification links, account links, and main registration links between clue units, and check the link status of the endpoint fields of each link to obtain link status data; Based on the link status data, identify the number of loopback links, the number of broken links, and the number of repairs for jump links, calculate the stability of clue units belonging to the same subject, and obtain the belonging stability value. Based on the stable attribution value, the main clue data is solidified into main files, and the main attribution status and business attribute fields of the main files are identified to determine the access status of the target main and obtain access management decision data.

[0006] Furthermore, based on the semantic unit data, the order in which each semantic unit appears in different documents is read to form constraint terms. These constraint terms are then compared with the actual values ​​of the semantic units to obtain constraint comparison data, including: Based on the semantic unit data, the generation and update identifiers of each semantic unit in different documents are read to obtain the document sequence data; Based on the document sequence data, the order in which the same semantic unit appears in different documents is concatenated to determine the order of the semantic unit's cross-document relationship, thus obtaining the sequence structure data; By splitting sequential structure data into several constraint terms, the sequential relationship of each semantic unit in the document processing flow is determined, and constraint term data is obtained. Based on the constraint data, the actual values ​​of the semantic units of each constraint are compared with the constraint item by item, and the values ​​that satisfy the constraint and the values ​​that do not satisfy the constraint are recorded to obtain constraint comparison data.

[0007] Furthermore, based on the constraint comparison data, the number of value segments, continuous coverage length, and number of repetitions across documents in the constraint sequence are identified. The degree to which semantic units are supported by multiple documents is calculated to obtain the trajectory support quantity, including: Based on the constraint comparison data, the diffusion coverage of the effective value records of semantic units among multiple documents is calculated to obtain cross-document support items; the proportion of the coverage interval length of semantic units that satisfy the constraint items is calculated to obtain continuous coverage saturation items. Based on the constraint comparison data, calculate the normalized proportion of the continuous coverage length of each value segment to the total number of constraint terms to obtain the segment set term; calculate the degree of influence of the value segments that are valid within a single document on the formation of cross-document links to obtain the segment connectivity term. By fusing cross-document support items, continuous coverage saturation items, fragment concentration items, and fragment connectivity items, the degree to which a semantic unit is supported by multiple documents is calculated, thus obtaining the trajectory support quantity.

[0008] Furthermore, by identifying and extracting semantic units whose trajectory support reaches a preset support threshold as clue units, and aggregating the clue units according to the source of the documents, the main clue data is obtained, including: Based on the trajectory support amount, the preset support threshold is converted into a threshold range corresponding to different document sources, and the trajectory support amount of each semantic unit is judged to fall into the threshold range to obtain threshold judgment data. Based on the threshold determination data, semantic units whose trajectory support volume meets the threshold range are extracted as clue units, and clue identifiers and source tags are generated for each clue unit to obtain clue unit data. Based on the clue unit data, the clue units are divided into multiple source groups according to the source tags, and the correspondence between the clue identifier and the semantic unit identifier is recorded in each source group to obtain the source index data; Based on the source index data, the clue units in different source groups are aligned and grouped, and the clue units with alignment relationships are written into the same main set to obtain the main clue data.

[0009] Furthermore, based on the main clue data, identity identification links, account links, and main registration links are established between clue units, and the link status of the endpoint fields of each link is checked to obtain link status data, including: Based on the main clue data, extract the identity field, account field, and main registration field of each clue unit in the main set to obtain the link field data; Based on the link field data, the identity field, account field, and subject registration field are used as the identity endpoint, account endpoint, and registration endpoint, respectively, and the identity link, account link, and subject registration link are established to obtain the link construction data; Based on the link construction data, perform endpoint checks and record endpoint missing information, endpoint conflict information, and endpoint traceability information to obtain endpoint check data. Based on the endpoint inspection data, links containing endpoint missing information, endpoint conflict information, and endpoint traceability information are marked as broken links, skipped links, and loopback links, respectively, to obtain link status data.

[0010] Furthermore, based on the link status data, the number of loopback links, the number of broken links, and the number of repairs for skip links are identified. The stability of clue units belonging to the same entity is calculated, yielding a stability value, including: Based on link status data and link construction data, the support strength that can form a traceable closed loop within the subject set is calculated to obtain the closed loop support term; the proportion of the number of broken links to the total number of links is calculated to determine the degree of attribution uncertainty caused by missing endpoints, and the broken link suppression term is obtained. Based on the link status data, the degree of attribution fluctuation caused by repeated link alignment and replacement is calculated to obtain the repair disturbance term; the deviation of the proportion of identity link, account link and subject registration link in the total number of loopback links from the equilibrium proportion is calculated to obtain the link equilibrium term. By integrating the closed-loop support term, the link break suppression term, the disturbance repair term, and the link balancing term, the stability of the clue unit belonging to the same subject is calculated, and the belonging stability value is obtained.

[0011] Furthermore, based on the stable attribution value, the subject clue data is solidified into subject profiles, and the subject attribution status and business attribute fields of the subject profiles are identified to determine the access status of the target subject, thereby obtaining access management decision data, including: Based on the attribution stability value, the attribution stability value is mapped to a preset stability interval to determine the stability level of the main clue data and obtain the stability level data; Based on the stability level data, extract the main clues that meet the preset solidification conditions for stability level, and generate the main file based on the main clue data to obtain the main file data; Based on the main file data, determine the ownership status of the main file, and summarize the attribute fields related to financial services to obtain the main status data and business attribute data; Based on the entity status data and business attribute data, the access status of the target entity in financial services is determined, and access management decision data is obtained.

[0012] Secondly, an identity authentication and financial service access management system integrating multiple document information, the system comprising: The data module is used to acquire basic data of multiple certificates, parse the business meaning of each field in different certificates, group fields with the same business meaning into semantic units, record the certificate source of each semantic unit, and obtain semantic unit data. The constraint module is used to read the order of appearance of each semantic unit in different documents based on the semantic unit data to form constraint items, and compare the constraint items with the actual values ​​of the semantic units to obtain constraint comparison data; The support module is used to identify the number of value segments, continuous coverage length and number of times they are repeated across documents in the constraint sequence based on the constraint comparison data, calculate the degree to which the semantic unit is supported by multiple documents, and obtain the trajectory support quantity. The main module is used to identify and extract semantic units whose trajectory support reaches a preset support threshold as clue units, and to collect the clue units according to the source of the documents to obtain the main clue data. The link module is used to establish identity link, account link and subject registration link between clue units based on the main clue data, and check the link status of the endpoint fields of each link to obtain link status data. The stabilization module is used to identify the number of loopback links, the number of broken links, and the number of repairs for jump links based on link status data, calculate the stability of clue units belonging to the same subject, and obtain the belonging stability value. The access module is used to solidify the subject clue data into subject profiles based on the stable attribution value, identify the subject attribution status and business attribute fields of the subject profiles, determine the access status of the target subject, and obtain access management decision data.

[0013] The above-described solution of the present invention has at least the following beneficial effects: This invention forms constraint terms by reading the order in which semantic units appear in different documents, and compares these constraint terms with the actual values ​​of the semantic units. This ensures that the values ​​of the same semantic units in different documents are no longer isolated data, but are incorporated into a data structure with sequential constraints for unified processing. The implicit time and logical order information in the document generation, registration, or processing process is transformed into data elements that can be directly used in calculations, forming composite comparative data that includes both sequential and value relationships. This reflects the consistency or inconsistency of the generation logic of the same semantic unit in different document systems, providing clear structural evidence for subsequent analysis and avoiding the problem of being unable to distinguish the inherent logical relationships between data from different sources when judging solely based on static values.

[0014] This invention calculates the trajectory support quantity by measuring the degree to which a semantic unit is supported by multiple documents, transforming the correlation between multiple documents into quantifiable statistical results. This makes the existence state of a semantic unit in a multi-document environment no longer a simple presence or absence, but rather described by multi-dimensional statistical characteristics such as the distribution of value fragments, coverage continuity, and cross-document recurrence. It reflects the data quantification results of the stable existence state of semantic units across documents, enabling subsequent processing steps to distinguish and process different semantic units based on clear numerical results. This avoids mixing occasionally occurring or low-consistency data with data that persists in multiple documents, and clarifies the objective manifestation of the joint support relationship of multiple documents.

[0015] This invention transforms clue units into a set of data links with endpoint relationships by establishing identity identification links, account links, and subject registration links between clue units and checking the link status of each link endpoint field. This allows different types of subject association information to exist in a clear link form and be uniformly checked. By identifying endpoint missing, endpoint conflict, or endpoint traceability, potential data incompleteness or conflict during the subject association process is explicitly recorded as a structured status result. This reflects the intermediate result of the subject association reliability and provides a clear data input basis for subsequent stability calculations.

[0016] This invention obtains a belonging stability value by calculating the stability of clue units belonging to the same subject, and maps multiple link status statistical results into a single stability value. This gives the subject belonging relationship a clear quantitative expression at the data level. The belonging stability value comprehensively reflects the link integrity, link conflict situation and the occurrence of link adjustment behavior. This makes subject belonging no longer dependent on a single field or a single link for judgment, but uses multiple link statistical results as input to form a stability calculation output, providing data basis with numerical boundaries for subsequent subject file generation and admission judgment.

[0017] This invention determines the access status of a target entity by solidifying entity clue data into entity profiles and identifying the entity's ownership status and business attribute fields within these profiles. It establishes a complete data loop from original multi-document data to entity profiles and access decisions, ensuring that entity-related data is continuously stored and repeatedly accessed in structured archive form. Access management decision data is directly generated based on the ownership status and business attribute fields in the entity profiles, providing a clear data source and processing path for access judgment results. This avoids the problem of access results relying solely on real-time comparisons and lacking a reusable data foundation, thus achieving a direct mapping relationship between entity identification results and financial service access decision results. Attached Figure Description

[0018] Figure 1 This is a flowchart of an embodiment of the present invention providing a method for identity authentication and financial service access management that integrates multiple document information. Detailed Implementation

[0019] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0020] like Figure 1 As shown, embodiments of the present invention propose a method for identity authentication and financial service access management that integrates multiple document information. The method includes: Acquire basic data for multiple certificates, parse the business meaning of each field in different certificates, group fields with the same business meaning into semantic units, record the certificate source of each semantic unit, and obtain semantic unit data; Based on the semantic unit data, the order in which each semantic unit appears in different documents is read to form constraint terms, and the constraint terms are compared with the actual values ​​of the semantic units to obtain constraint comparison data; Based on the constraint comparison data, the number of value segments, continuous coverage length and number of repetitions across documents of the constraint sequence are identified, the degree to which the semantic unit is supported by multiple documents is calculated, and the trajectory support quantity is obtained. By identifying and extracting semantic units whose trajectory support reaches a preset support threshold as clue units, and then aggregating the clue units according to the source of the documents, the main clue data is obtained. Based on the main clue data, establish identity identification links, account links, and main registration links between clue units, and check the link status of the endpoint fields of each link to obtain link status data; Based on the link status data, identify the number of loopback links, the number of broken links, and the number of repairs for jump links, calculate the stability of clue units belonging to the same subject, and obtain the belonging stability value. Based on the stable attribution value, the main clue data is solidified into main files, and the main attribution status and business attribute fields of the main files are identified to determine the access status of the target main and obtain access management decision data.

[0021] In this embodiment of the invention, basic data of multiple certificates is acquired, and the business meaning of each field in different certificates is analyzed. Fields with the same business meaning are grouped into semantic units, and the certificate source of each semantic unit is recorded to obtain semantic unit data, providing a consistent data input basis for subsequent cross-certificate data processing. Based on the semantic unit data, the order of appearance of each semantic unit in different certificates is read to form constraint terms, and the constraint terms are compared with the actual values ​​of the semantic units to obtain constraint comparison data, so that subsequent processing can be based on constraint satisfaction rather than a single value. Based on the constraint comparison data, the number of value segments, continuous coverage length, and number of times they are repeated across certificates in the constraint sequence are identified, and the degree to which the semantic unit is supported by multiple certificates is calculated to obtain the trajectory support quantity, so that the consistency of the semantic unit in the multi-certificate environment has a clear data expression form. By identifying and extracting semantic units whose trajectory support quantity reaches a preset support threshold as clue units, the clue units are collected according to the certificate source to obtain main clue data, so that subsequent main identification is only based on a data set with clear multi-certificate support characteristics.

[0022] Based on the main clue data, identity identification links, account links, and main entity registration links are established between clue units. The link status of the endpoint fields of each link is checked to obtain link status data. The main entity-related clues are expressed in a structured form in the form of links, and the integrity and conflict of the links are recorded. Based on the link status data, the number of loopback links, the number of broken links, and the number of repairs for skipped links are identified. The stability of the clue units belonging to the same main entity is calculated to obtain the attribution stability value. This allows the main entity attribution relationship to be output in the form of quantitative data and has calculation results that can be directly used for subsequent judgments. Based on the attribution stability value, the main entity clue data is solidified into main entity files. The main entity attribution status and business attribute fields of the main entity files are identified to determine the access status of the target main entity and obtain access management decision data. This realizes a structured mapping from multi-document data to financial service access decisions, so that access decisions have a clear data generation path and reusable data storage results.

[0023] This process involves acquiring basic data from multiple certificates, parsing the business meaning of each field in different certificates, grouping fields with the same business meaning into semantic units, recording the certificate source of each semantic unit, and obtaining semantic unit data. Specifically, this includes: When a financial service access transaction is triggered, the system establishes a data collection session indexed by the target entity. Based on the transaction type, it determines the range of documents to be collected and initiates a data acquisition request to the corresponding document data source. This document data source includes at least two of the following: personal document data source, enterprise registration data source, and account opening data source. The system includes a session identifier, document type identifier, and the target entity's search elements in the acquisition request, causing each data source to return a set of document fields associated with that target entity. After receiving the returned results, the system performs integrity verification and structured encapsulation on the data returned by each data source. Each field record is uniformly encapsulated into a field entry containing document type, document source identifier, original field name, original field value, field location identifier, and field generation time identifier, forming multi-document basic data.

[0024] The system performs field preprocessing operations on each field entry in the basic data of multiple certificates to facilitate the parsing of business meaning. The field preprocessing operations include character set unification, null value standardization, separator normalization, and format regularization of numbered fields. For fields with check bits, mask bits, or segmented encoding, they are split into parsable fragments according to preset rules and the fragment sequence identifier is retained. For fields such as addresses, names, and certificate numbers that may contain multiple segments of information, the system segments the original values ​​of the fields according to a preset parsing template and generates candidate subfields, while recording the mapping relationship before and after the segmentation. For cases where the same field with the same name appears multiple times in the same certificate source, the system determines the field instance sequence based on the field position identifier and the generation time identifier and records the sequence number to ensure that the subsequent parsing process can distinguish different instances of the same field.

[0025] The system parses and determines the business meaning of each field in different documents. First, it reads the original field name, field location identifier, and document type information. It then calls a preset business meaning mapping table to perform initial matching of the fields. This mapping table uses document type as the entry point and maintains the correspondence between field name patterns and business meaning identifiers. When the original field name satisfies multiple mapping rules or the mapping table cannot uniquely determine the business meaning, the system further utilizes the structural features of the original field value to perform secondary discrimination. These structural features include character composition patterns, length ranges, checksum rule matching, and whether a preset regular expression pattern is satisfied. The system then determines the target business meaning identifier from multiple candidate business meaning identifiers. For cases where the same field uses different expressions in different documents, the system maps multiple expressions to the same business meaning identifier based on the synonym set in the business meaning mapping table. The system records the mapping path during the parsing process, enabling subsequent tracing of the parsing basis from the original field name to the business meaning identifier.

[0026] The system uses business meaning identifiers as the merge key to aggregate field entries from different document sources. Field entries with the same business meaning are grouped into the same semantic unit, and a semantic unit identifier is assigned to each semantic unit. During the merging process, the system retains the document source identifier and document type information for each field entry and writes this as the source set of the semantic unit into the semantic unit metadata to establish a correspondence between semantic units and document sources. When multiple field instances of the same semantic unit exist in the same document source, the system stores the multiple instances in an ordered manner according to the field instance sequence number and records the order of the instances. The system summarizes all semantic unit records to form semantic unit data and writes it to session-level cache or persistent storage for subsequent steps to read. Each semantic unit record in the semantic unit data carries a session identifier to associate with the subsequent constraint construction process, and maintains a traceable relationship from multiple document basic data to semantic unit data through the source set and field entry list. In case of subsequent comparison anomalies or link conflicts, it can trace back to the specific document source and field instance.

[0027] In a preferred embodiment of the present invention, based on semantic unit data, the order of appearance of each semantic unit in different documents is read to form constraint terms, and the constraint terms are compared with the actual values ​​of the semantic units to obtain constraint comparison data, including: Based on the semantic unit data, the generation and update identifiers of each semantic unit in different documents are read to obtain the document sequence data; Based on the document sequence data, the order in which the same semantic unit appears in different documents is concatenated to determine the order of the semantic unit's cross-document relationship, thus obtaining the sequence structure data; By splitting sequential structure data into several constraint terms, the sequential relationship of each semantic unit in the document processing flow is determined, and constraint term data is obtained. Based on the constraint data, the actual values ​​of the semantic units of each constraint are compared with the constraint item by item, and the values ​​that satisfy the constraint and the values ​​that do not satisfy the constraint are recorded to obtain constraint comparison data.

[0028] In this embodiment of the invention, based on semantic unit data, the generation identifier and update identifier of each semantic unit in different documents are read to obtain document sequence data, providing an objective data foundation in terms of time and process dimensions for cross-document relationship identification; based on the document sequence data, the order of appearance of the same semantic unit in different documents is concatenated to determine the order of appearance of semantic units across documents, obtaining sequence structure data, avoiding the problem that multiple document data exist only in parallel and cannot reflect their inherent sequential logic; by splitting the sequence structure data into several constraint terms, the sequential relationship of each semantic unit in the document processing flow is determined to obtain constraint term data, realizing the discretized expression of cross-document sequence relationship, providing standardized data objects for subsequent item-by-item comparison and verification; based on the constraint term data, the actual value of the semantic unit of each constraint term is compared with the constraint term item by item, recording the value records that satisfy the constraint term and the value records that do not satisfy the constraint term, obtaining constraint comparison data, clearly distinguishing the consistent and inconsistent states of the values ​​of different semantic units in the cross-document sequence relationship, providing direct data input for subsequent processing.

[0029] Specifically, based on the document sequence data, the order in which the same semantic unit appears in different documents is concatenated to determine the order of the semantic unit's cross-document relationship, resulting in sequence structure data, which includes: The system uses semantic units as processing objects, aggregating and reading multiple document records corresponding to the same semantic unit in the document sequence data, and obtaining the generation identifier and update identifier of the semantic unit in different documents. The system first standardizes the generation identifier and update identifier according to a unified rule, so that the sequence identifiers under different document systems can be compared in the same sorting dimension. Then, it sorts the order of appearance of the same semantic unit in each document according to the standardized generation identifier and update identifier. After sorting, the system concatenates the sorting results according to the document source, connecting the order of appearance of the same semantic unit in multiple documents into a continuous sequence path, and records the sequential relationship between adjacent documents and the corresponding document identifiers in the sequence path, forming a sequence structure data describing the cross-document appearance relationship of the semantic unit. This sequence structure data is used to reflect the order of appearance of the semantic unit in different document systems and its overall sequence form.

[0030] Specifically, by splitting sequential structure data into several constraint terms, the sequential relationship of each semantic unit in the document processing flow is determined, resulting in constraint term data, which includes: The system parses the sequential structure corresponding to each semantic unit, splitting the consecutively arranged document nodes in the sequential structure according to their adjacency relationship, and taking two adjacent document nodes as a basic sequential unit. During the splitting process, the system generates corresponding constraint terms for each basic sequential unit, and records the preceding document identifier, the following document identifier, and the sequence type information corresponding to the preceding-following relationship in the constraint terms, so that each constraint term can independently represent the sequential relationship of the semantic unit between two documents. Through the above splitting operation, the sequential structure that originally existed in the form of a whole path is transformed into a set of multiple independently processable constraint terms, forming constraint term data. This constraint term data is used to clarify the specific sequential relationship of the semantic unit in the document processing flow and serves as a direct basis for subsequent value comparison.

[0031] Specifically, based on the constraint data, the actual value of the semantic unit of each constraint is compared with the constraint item by item, and the value records that satisfy the constraint and the value records that do not satisfy the constraint are recorded to obtain constraint comparison data, which specifically includes: For each constraint in the constraint data, the system reads the actual values ​​of the semantic units involved in that constraint in the preceding and subsequent documents, and compares the actual values ​​with the order relationship described by the constraint item. During the comparison, the system determines whether the occurrence of the actual values ​​of the semantic units in the preceding and subsequent documents conforms to the order relationship defined by the constraint item. Value combinations that conform to the order relationship are recorded as value records that satisfy the constraint item, while value combinations that do not conform to the order relationship are recorded as value records that do not satisfy the constraint item. During the recording process, the system associates and labels each value record with the corresponding semantic unit identifier, document source information, and constraint item identifier, forming constraint comparison data containing the results of satisfaction and non-satisfaction. This constraint comparison data is used to fully reflect the compliance of the semantic units' values ​​under the multiple document order constraints and provides a direct data basis for subsequent statistical analysis based on the constraint satisfaction status.

[0032] In a preferred embodiment of the present invention, based on constraint comparison data, the number of value segments, continuous coverage length, and number of repetitions across documents in the constraint sequence are identified. The degree to which semantic units are supported by multiple documents is calculated to obtain the trajectory support quantity, including: Based on the constraint comparison data, the diffusion coverage of the effective value records of semantic units among multiple documents is calculated to obtain cross-document support items; the proportion of the coverage interval length of semantic units that satisfy the constraint items is calculated to obtain continuous coverage saturation items. Based on the constraint comparison data, calculate the normalized proportion of the continuous coverage length of each value segment to the total number of constraint terms to obtain the segment set term; calculate the degree of influence of the value segments that are valid within a single document on the formation of cross-document links to obtain the segment connectivity term. By fusing cross-document support items, continuous coverage saturation items, fragment concentration items, and fragment connectivity items, the degree to which a semantic unit is supported by multiple documents is calculated, thus obtaining the trajectory support quantity.

[0033] In this embodiment of the invention, based on constraint comparison data, the diffusion coverage of valid value records of semantic units across multiple documents is calculated to obtain cross-document support items, clearly distinguishing between value cases that are valid only in a single document and value cases that are valid simultaneously in multiple document sources; the proportion of the coverage interval length of semantic units satisfying constraint items is calculated to obtain continuous coverage saturation items, avoiding the problem of failing to reflect the overall sequential consistency based solely on single-point validity or non-continuous validity; based on constraint comparison data, the normalized proportion of the continuous coverage length of each value segment to the total number of constraint items is calculated to obtain segment concentration items, avoiding data with different distribution patterns in the same location. The theoretical level is treated equally; the influence of the value fragments that are valid within a single document on the formation of cross-document links is calculated to obtain the fragment connectivity term, which describes the connectable relationship between single document fragments and other document fragments, so that cross-document association analysis no longer simply excludes data fragments that are valid within a single document; the cross-document support term, continuous coverage saturation term, fragment concentration term and fragment connectivity term are integrated to calculate the degree to which semantic units are supported by multiple documents, and the trajectory support quantity is obtained, so that the existence state of semantic units in the multi-document environment is compressed into a data quantification result with a clear calculation source and composition structure, providing a numerical basis that can be directly referenced for the subsequent process.

[0034] In a preferred embodiment of the present invention, semantic units whose trajectory support reaches a preset support threshold are identified and extracted as clue units. These clue units are then aggregated according to their document origin to obtain main clue data, including: Based on the trajectory support amount, the preset support threshold is converted into a threshold range corresponding to different document sources, and the trajectory support amount of each semantic unit is judged to fall into the threshold range to obtain threshold judgment data. Based on the threshold determination data, semantic units whose trajectory support volume meets the threshold range are extracted as clue units, and clue identifiers and source tags are generated for each clue unit to obtain clue unit data. Based on the clue unit data, the clue units are divided into multiple source groups according to the source tags, and the correspondence between the clue identifier and the semantic unit identifier is recorded in each source group to obtain the source index data; Based on the source index data, the clue units in different source groups are aligned and grouped, and the clue units with alignment relationships are written into the same main set to obtain the main clue data.

[0035] In this embodiment of the invention, based on the trajectory support quantity, a preset support threshold is converted into a threshold range corresponding to different document sources. The trajectory support quantity of each semantic unit is then determined to fall within this threshold range to obtain threshold determination data. This avoids data filtering inconsistencies caused by distortion in direct comparison of trajectory support quantities due to differences in document sources. Based on the threshold determination data, semantic units whose trajectory support quantities satisfy the threshold range are extracted as clue units. A clue identifier and source label are generated for each clue unit to obtain clue unit data. The calculated filtering results are solidified into structured objects that can directly participate in subsequent grouping and aggregation, eliminating the need for repeated operations in subsequent steps. The process of calculating the reference trajectory support quantity can complete the clue-level data operation. Based on the clue unit data, the clue units are divided into multiple source groups according to the source label, and the correspondence between the clue identifier and the semantic unit identifier is recorded in each source group to obtain the source index data. This provides clear data boundaries for subsequent cross-source alignment operations and avoids data ambiguity caused by disordered mixing of clues from different document sources in the same set. Based on the source index data, the clue units in different source groups are aligned and aggregated. Clue units with alignment relationships are written into the same subject set to obtain the subject clue data, realizing the structural transformation from clue-level evidence to subject-level evidence set.

[0036] Specifically, based on the trajectory support quantity, the preset support threshold is converted into a threshold range corresponding to different document sources, and the trajectory support quantity of each semantic unit is determined to fall within the threshold range to obtain threshold determination data, which specifically includes: The system first reads the trajectory support quantity corresponding to each semantic unit and simultaneously obtains the document source information associated with each semantic unit. Then, based on the preset support threshold configuration rules, the preset support threshold is processed into intervals according to the source dimension, so that the same preset support threshold is mapped into multiple threshold intervals corresponding to different document source types. Each threshold interval is used to limit the range of trajectory support quantity that a semantic unit from the corresponding document source must meet to be judged as valid supporting evidence. After completing the threshold interval mapping, the system takes the semantic unit as the processing object, and judges whether the trajectory support quantity of the semantic unit falls into the threshold interval applicable to its corresponding document source. The system records whether the semantic unit meets the threshold interval conditions and the corresponding interval judgment result, and generates threshold judgment data containing the semantic unit identifier, document source identifier, trajectory support quantity value and threshold judgment result. The threshold judgment data reflects the support validity status of different semantic units under their respective document source conditions in a structured form.

[0037] Specifically, based on the threshold determination data, semantic units whose trajectory support values ​​satisfy the threshold range are extracted as clue units, and a clue identifier and source label are generated for each clue unit to obtain clue unit data, which specifically includes: The system displays the threshold judgment results as semantic units that meet the corresponding threshold interval conditions, filtering them from the entire semantic unit set and identifying these semantic units as clue units. After the clue units are identified, the system assigns a unique clue identifier to each clue unit for independent reference and differentiation in subsequent data processing. At the same time, it adds a source tag to the clue unit based on the document source information corresponding to the clue unit to clearly identify the source attribution of the clue unit in the multi-document system. Subsequently, the system integrates the clue identifier, semantic unit identifier, source tag, and corresponding trajectory support information to generate clue unit data, transforming the semantic units that originally existed only as judgment results into data objects with unique identifiers, source attributes, and supporting evidence.

[0038] In a preferred embodiment of the present invention, based on the main clue data, an identity identification link, an account link, and a main registration link are established between clue units, and the link status of the endpoint fields of each link is checked to obtain link status data, including: Based on the main clue data, extract the identity field, account field, and main registration field of each clue unit in the main set to obtain the link field data; Based on the link field data, the identity field, account field, and subject registration field are used as the identity endpoint, account endpoint, and registration endpoint, respectively, and the identity link, account link, and subject registration link are established to obtain the link construction data; Based on the link construction data, perform endpoint checks and record endpoint missing information, endpoint conflict information, and endpoint traceability information to obtain endpoint check data. Based on the endpoint inspection data, links containing endpoint missing information, endpoint conflict information, and endpoint traceability information are marked as broken links, skipped links, and loopback links, respectively, to obtain link status data.

[0039] In this embodiment of the invention, based on the subject clue data, the identity field, account field, and subject registration field of each clue unit in the subject set are extracted to obtain link field data, avoiding repeated parsing of clue unit data or field misuse issues during the link construction stage; based on the link field data, the identity field, account field, and subject registration field are respectively used as identity endpoint, account endpoint, and registration endpoint, and identity link, account link, and subject registration link are established to obtain link construction data, so that the subject association relationship is explicitly represented in the form of a topological structure, providing operable data objects for subsequent status checks based on the link structure; based on the link construction data, endpoint checks are performed, recording endpoint missing information, endpoint conflict information, and endpoint traceability information to obtain endpoint check data, solidifying potential incompleteness and conflict of the link in the form of explicit data marking, so that the link quality is no longer implicit in the field values; based on the endpoint check data, the links containing endpoint missing information, endpoint conflict information, and endpoint traceability information are marked as broken link links, skipped link links, and loopback links respectively to obtain link status data, so that the quality differences between different links can be expressed through a limited number of status categories.

[0040] Specifically, based on the link field data, the identity identifier field, account field, and entity registration field are used as the identity identifier endpoint, account endpoint, and registration endpoint, respectively, and identity identifier links, account links, and entity registration links are established to obtain link construction data, which specifically includes: First, the system reads the field type identifiers from the link field data to distinguish between three field sets: identity identifier fields, account fields, and subject registration fields. Within each subject set, these three types of fields are processed independently. For identity identifier fields, the system uses the value of the identity identifier field as the identity identifier endpoint identifier, mapping identity identifier fields with the same value or traceable association within the subject set to the same identity identifier endpoint. An identity identifier link is then established between corresponding clue units, with each end of the link pointing to the identity identifier endpoint in a different clue unit. For account fields, the system uses the account identifier value as the account endpoint identifier, mapping fields within the subject set to the same account identifier. The account fields with values ​​are merged into the same account endpoint, and an account link is established between the corresponding clue units so that the account link reflects the relationship between different clue units at the account level. For the subject registration field, the system uses the field value that represents the subject registration relationship in the subject registration field as the registration endpoint identifier. Within the subject set, subject registration fields with the same registration relationship identifier are mapped to the same registration endpoint, and subject registration links are established between the corresponding clue units. Each link is assigned a link identifier, a link type identifier, and a link endpoint identifier, and is associated with the clue unit identifier that generated the link for storage, forming link construction data containing link type, endpoint information, and clue source information.

[0041] Specifically, endpoint checks are performed based on the link construction data, recording endpoint missing information, endpoint conflict information, and endpoint traceability information to obtain endpoint check data, which includes: The system performs endpoint checks on each constructed identity link, account link, and entity registration link. Endpoint checks use the endpoint identifiers recorded in the link construction data as the inspection object, and perform cross-validation with data from other links and clue units within the entity set. During endpoint missing checks, the system determines whether the starting or target endpoint of the link has a valid field value. If, during link construction, an endpoint is found to have a missing field value within the entity set or the field value is empty, the system records the corresponding endpoint missing information in the endpoint check results for that link. During endpoint conflict checks, the system compares multiple link endpoints of the same link type within the entity set. When a conflict is found between the same clue unit or the same entity set... When multiple endpoint values ​​exist that are inconsistent and cannot be traced, the system records the inconsistency as endpoint conflict information and associates it with the corresponding link. During the endpoint traceability check, the system checks whether the current link endpoint can form a continuous associated path through other links or clue units based on the link endpoint identifier recorded in the link construction data. When an endpoint can establish a continuous connection with other endpoints in the subject set through at least one other link, the system records the endpoint traceability information in the endpoint check result of that link. Through the above check operations, the system writes the endpoint integrity, consistency, and traceability check results of each link into the endpoint check data in a structured form, so that potential anomalies or connectivity characteristics of the link are explicitly recorded.

[0042] Based on the endpoint inspection data, links containing endpoint missing information, endpoint conflict information, and endpoint traceability information are marked as broken links, skipped links, and loopback links, respectively, to obtain link status data, which specifically includes: The system performs status determination processing on links according to preset link status marking rules. These rules use endpoint missing information, endpoint conflict information, and endpoint traceability information recorded in the endpoint inspection data as input conditions. When the endpoint inspection data of a link contains endpoint missing information, the system marks the link as a broken link and records its broken link status identifier in the link status data, thus identifying the link as a subject association path with incomplete endpoint information. When the endpoint inspection data of a link contains endpoint conflict information, the system marks the link as a skip link and records its skip link status identifier in the link status data, thus identifying the link as a subject association path with endpoint value conflicts or inconsistencies. When the endpoint inspection data of a link contains endpoint traceability information and the link can form a closed association path with other links in the subject set, the system marks the link as a loopback link and records its loopback status identifier in the link status data, thus identifying the link as a subject association path with complete endpoint relationships and continuous traceability characteristics. The system converts the connection quality and completeness of different links in the subject set into enumerable link status data.

[0043] In a preferred embodiment of the present invention, based on link status data, the number of loopback links, the number of broken links, and the number of repairs for skip links are identified. The stability of clue units belonging to the same entity is calculated to obtain a stability value, including: Based on link status data and link construction data, the support strength that can form a traceable closed loop within the subject set is calculated to obtain the closed loop support term; the proportion of the number of broken links to the total number of links is calculated to determine the degree of attribution uncertainty caused by missing endpoints, and the broken link suppression term is obtained. Based on the link status data, the degree of attribution fluctuation caused by repeated link alignment and replacement is calculated to obtain the repair disturbance term; the deviation of the proportion of identity link, account link and subject registration link in the total number of loopback links from the equilibrium proportion is calculated to obtain the link equilibrium term. By integrating the closed-loop support term, the link break suppression term, the disturbance repair term, and the link balancing term, the stability of the clue unit belonging to the same subject is calculated, and the belonging stability value is obtained.

[0044] In this embodiment of the invention, based on link status data and link construction data, the support strength for forming a traceable closed loop within the subject set is calculated to obtain a closed loop support item, reflecting the degree to which multiple document clues form a self-consistent relationship within the subject, providing a clear positive support data foundation for subsequent stability calculations; the proportion of broken links to the total number of links is calculated to determine the degree of attribution uncertainty caused by missing endpoints, resulting in a broken link suppression item, explicitly reflecting the data incompleteness caused by missing document fields, unregistered information, or failed cross-document mapping during the subject association process as a calculable result, so that the impact of such incomplete links on the stability of subject attribution is separately characterized at the data level; based on link status data, the degree of attribution fluctuation caused by repeated link alignment and replacement is calculated to obtain a repair disturbance item, ensuring that the subject clues are properly attributed. Data fluctuations caused by multiple alignments, replacements, or reconstructions during the data collection process are clearly recorded and quantified, providing objective data evidence for the consistency of the data collection process in stability assessment. The deviation of the proportion of identity identification links, account links, and subject registration links in the total number of loop links from the equilibrium proportion is calculated to obtain the link equilibrium term, avoiding the situation where the subject attribution results are overly concentrated in a certain type of document link without the support of other links. The closed-loop support term, link break suppression term, repair disturbance term, and link equilibrium term are integrated to calculate the stability of clue units belonging to the same subject, obtaining the attribution stability value. The stability quantification results comprehensively reflect the link integrity, link uncertainty, link volatility, and link structure distribution characteristics, providing input basis with clear data sources and calculation paths for subsequent decision-making.

[0045] In a preferred embodiment of the present invention, the subject clue data is solidified into subject profiles based on the attribution stability value, and the subject attribution status and business attribute fields of the subject profiles are identified to determine the access status of the target subject, thereby obtaining access management decision data, including: Based on the attribution stability value, the attribution stability value is mapped to a preset stability interval to determine the stability level of the main clue data and obtain the stability level data; Based on the stability level data, extract the main clues that meet the preset solidification conditions for stability level, and generate the main file based on the main clue data to obtain the main file data; Based on the main file data, determine the ownership status of the main file, and summarize the attribute fields related to financial services to obtain the main status data and business attribute data; Based on the entity status data and business attribute data, the access status of the target entity in financial services is determined, and access management decision data is obtained.

[0046] In this embodiment of the invention, based on the attribution stability value, the attribution stability value is mapped to a preset stability interval to determine the stability level of the subject clue data, thereby obtaining stability level data. This achieves a hierarchical expression of the stability state of the subject clues, reducing the impact of numerical fluctuations on subsequent subject solidification operations. Based on the stability level data, subject clues whose stability levels meet preset solidification conditions are extracted, and subject profiles are generated based on these subject clue data, resulting in subject profile data. This achieves solidified storage of subject identification results, preventing data with insufficient stability from entering the subject profile system. Based on the subject profile data, the subject attribution status of the subject profile is determined, and attribute fields related to financial services are summarized to obtain subject status data and business attribute data. This achieves the separation of subject identity attributes and business decision attributes, providing clear data input classification for subsequent processing and reducing data processing ambiguity caused by attribute mixing. Based on the subject status data and business attribute data, the access status of the target subject in financial services is determined, resulting in access management decision data. This achieves a mapping relationship from subject identification results to access decision results, preventing access status from being generated independently of subject data.

[0047] Specifically, based on the attribution stability value, the attribution stability value is mapped to a preset stability interval to determine the stability level of the main clue data, thus obtaining stability level data, which includes: After obtaining the attribution stability value calculated for each subject clue data, the system first reads the pre-set stability interval configuration data from the storage unit. The stability interval configuration data includes multiple stability intervals divided according to numerical ranges, and each stability interval corresponds to a unique stability level identifier. The system sequentially compares the attribution stable value with the boundaries of each stable interval to identify the target stable interval into which the attribution stable value falls, and extracts the stability level identifier corresponding to the target stable interval. The system establishes a correlation between the stability level identifier and the corresponding main clue data, and writes the correlation result into the stability level dataset to form a data record with the main clue data as the index and the stability level as the attribute. Through this mapping process, the continuous attribution stable value calculated based on the link state is converted into discrete level data with clear interval boundaries, forming a hierarchical representation of the stability state of the main clue at the data structure level, providing a unified data judgment basis for subsequent data filtering and processing based on level conditions.

[0048] Specifically, based on the stability level data, main clues that meet the preset solidification conditions are extracted, and main profiles are generated based on these main clue data, resulting in main profile data, which specifically includes: The system reads pre-configured solidification condition rules, which use stability level as a judgment parameter to limit the range of data allowed to enter the main file. The system matches the stability level corresponding to each main clue data with the solidification condition rules. When the stability level of the main clue data meets the solidification conditions, the main clue data is marked as a solidifiable object. For the main clue data marked as solidifiable objects, the system starts the main file generation process, writing the clue unit identifier, document source information, associated field values, and corresponding stability level contained in the main clue data into the main file data structure, and assigning a unique main file identifier to the main file. Through the above processing, the main clue data that meets the stability conditions is transformed into structured main file data.

[0049] Specifically, based on the entity file data, the entity ownership status is determined, and attribute fields related to financial services are summarized to obtain entity status data and business attribute data, including: The system parses and processes the main file data, reading the distribution of clue units, document source information, and status identifiers related to subject identification. Based on the parsing results, the system determines the subject's ownership status corresponding to the main file and writes the determination result into the subject status dataset to represent the subject's ownership status at the subject identification level. Simultaneously, the system filters field information related to financial service processing from the main file. This field information includes account-related fields, registration-related fields, and subject attribute fields associated with financial business rules. The system then organizes and structures these fields in a unified format to form business attribute data.

[0050] Specifically, based on entity status data and business attribute data, the access status of the target entity in financial services is determined, resulting in access management decision data, which includes: The system uses entity status data and business attribute data as joint input data, and calls preset financial service access judgment rules to process the access status of target entities. The access judgment rules are based on the matching relationship between the entity's ownership status and business attribute fields to determine whether the target entity meets the financial service access conditions and generate the corresponding access status result. The system outputs the access status result as access management decision data and associates and stores the access management decision data with the corresponding entity file identifier. Through the above processing, the financial service access status is directly driven by the entity file data, forming a complete mapping path from the entity identification result to the access decision result at the data processing level, ensuring that the access management decision data has a clear data source and a traceable data processing process.

[0051] Embodiments of the present invention also provide an identity authentication and financial service access management system that integrates multiple document information, the system comprising: The data module is used to acquire basic data of multiple certificates, parse the business meaning of each field in different certificates, group fields with the same business meaning into semantic units, record the certificate source of each semantic unit, and obtain semantic unit data. The constraint module is used to read the order of appearance of each semantic unit in different documents based on the semantic unit data to form constraint items, and compare the constraint items with the actual values ​​of the semantic units to obtain constraint comparison data; The support module is used to identify the number of value segments, continuous coverage length and number of times they are repeated across documents in the constraint sequence based on the constraint comparison data, calculate the degree to which the semantic unit is supported by multiple documents, and obtain the trajectory support quantity. The main module is used to identify and extract semantic units whose trajectory support reaches a preset support threshold as clue units, and to collect the clue units according to the source of the documents to obtain the main clue data. The link module is used to establish identity link, account link and subject registration link between clue units based on the main clue data, and check the link status of the endpoint fields of each link to obtain link status data. The stabilization module is used to identify the number of loopback links, the number of broken links, and the number of repairs for jump links based on link status data, calculate the stability of clue units belonging to the same subject, and obtain the belonging stability value. The access module is used to solidify the subject clue data into subject profiles based on the stable attribution value, identify the subject attribution status and business attribute fields of the subject profiles, determine the access status of the target subject, and obtain access management decision data.

[0052] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0053] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0054] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0055] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for identity authentication and financial service access management that integrates multiple document information, characterized in that: The method includes: Acquire basic data for multiple certificates, parse the business meaning of each field in different certificates, group fields with the same business meaning into semantic units, record the certificate source of each semantic unit, and obtain semantic unit data; Based on the semantic unit data, the order in which each semantic unit appears in different documents is read to form constraint terms, and the constraint terms are compared with the actual values ​​of the semantic units to obtain constraint comparison data; Based on the constraint comparison data, the number of value segments, continuous coverage length and number of repetitions across documents of the constraint sequence are identified, the degree to which the semantic unit is supported by multiple documents is calculated, and the trajectory support quantity is obtained. By identifying and extracting semantic units whose trajectory support reaches a preset support threshold as clue units, and then aggregating the clue units according to the source of the documents, the main clue data is obtained. Based on the main clue data, establish identity identification links, account links, and main registration links between clue units, and check the link status of the endpoint fields of each link to obtain link status data; Based on the link status data, identify the number of loopback links, the number of broken links, and the number of repairs for jump links, calculate the stability of clue units belonging to the same subject, and obtain the belonging stability value. Based on the stable attribution value, the main clue data is solidified into main files, and the main attribution status and business attribute fields of the main files are identified to determine the access status of the target main and obtain access management decision data.

2. The method for identity authentication and financial service access management integrating multiple document information as described in claim 1, characterized in that, Based on the semantic unit data, the order in which each semantic unit appears in different documents is read to form constraint terms. These constraint terms are then compared with the actual values ​​of the semantic units to obtain constraint comparison data, including: Based on the semantic unit data, the generation and update identifiers of each semantic unit in different documents are read to obtain the document sequence data; Based on the document sequence data, the order in which the same semantic unit appears in different documents is concatenated to determine the order of the semantic unit's cross-document relationship, thus obtaining the sequence structure data; By splitting sequential structure data into several constraint terms, the sequential relationship of each semantic unit in the document processing flow is determined, and constraint term data is obtained. Based on the constraint data, the actual values ​​of the semantic units of each constraint are compared with the constraint item by item, and the values ​​that satisfy the constraint and the values ​​that do not satisfy the constraint are recorded to obtain constraint comparison data.

3. The method for identity authentication and financial service access management integrating multiple document information as described in claim 2, characterized in that, Based on the constraint comparison data, the number of value segments, continuous coverage length, and number of repetitions across documents in the constraint sequence are identified. The degree to which semantic units are supported by multiple documents is calculated to obtain the trajectory support quantity, including: Based on the constraint comparison data, the diffusion coverage of the effective value records of semantic units among multiple documents is calculated to obtain cross-document support items; the proportion of the coverage interval length of semantic units that satisfy the constraint items is calculated to obtain continuous coverage saturation items. Based on the constraint comparison data, calculate the normalized proportion of the continuous coverage length of each value segment to the total number of constraint terms to obtain the segment set term; calculate the degree of influence of the value segments that are valid within a single document on the formation of cross-document links to obtain the segment connectivity term. By fusing cross-document support items, continuous coverage saturation items, fragment concentration items, and fragment connectivity items, the degree to which a semantic unit is supported by multiple documents is calculated, thus obtaining the trajectory support quantity.

4. The method for identity authentication and financial service access management integrating multiple document information as described in claim 3, characterized in that, By identifying and extracting semantic units whose trajectory support reaches a preset support threshold as clue units, and aggregating these clue units according to document source, the main clue data is obtained, including: Based on the trajectory support amount, the preset support threshold is converted into a threshold range corresponding to different document sources, and the trajectory support amount of each semantic unit is judged to fall into the threshold range to obtain threshold judgment data. Based on the threshold determination data, semantic units whose trajectory support volume meets the threshold range are extracted as clue units, and clue identifiers and source tags are generated for each clue unit to obtain clue unit data. Based on the clue unit data, the clue units are divided into multiple source groups according to the source tags, and the correspondence between the clue identifier and the semantic unit identifier is recorded in each source group to obtain the source index data; Based on the source index data, the clue units in different source groups are aligned and grouped, and the clue units with alignment relationships are written into the same main set to obtain the main clue data.

5. The method for identity authentication and financial service access management integrating multiple document information as described in claim 4, characterized in that, Based on the main clue data, establish identity identification links, account links, and main registration links between clue units, and check the link status of the endpoint fields of each link to obtain link status data, including: Based on the main clue data, extract the identity field, account field, and main registration field of each clue unit in the main set to obtain the link field data; Based on the link field data, the identity field, account field, and subject registration field are used as the identity endpoint, account endpoint, and registration endpoint, respectively, and the identity link, account link, and subject registration link are established to obtain the link construction data; Based on the link construction data, perform endpoint checks and record endpoint missing information, endpoint conflict information, and endpoint traceability information to obtain endpoint check data. Based on the endpoint inspection data, links containing endpoint missing information, endpoint conflict information, and endpoint traceability information are marked as broken links, skipped links, and loopback links, respectively, to obtain link status data.

6. The method for identity authentication and financial service access management integrating multiple document information as described in claim 5, characterized in that, Based on link status data, the number of loopback links, the number of broken links, and the number of repairs for skip links are identified. The stability of clue units belonging to the same entity is calculated, and the belonging stability value is obtained, including: Based on link status data and link construction data, the support strength that can form a traceable closed loop within the subject set is calculated to obtain the closed loop support term; the proportion of the number of broken links to the total number of links is calculated to determine the degree of attribution uncertainty caused by missing endpoints, and the broken link suppression term is obtained. Based on the link status data, the degree of attribution fluctuation caused by repeated link alignment and replacement is calculated to obtain the repair disturbance term; the deviation of the proportion of identity link, account link and subject registration link in the total number of loopback links from the equilibrium proportion is calculated to obtain the link equilibrium term. By integrating the closed-loop support term, the link break suppression term, the disturbance repair term, and the link balancing term, the stability of the clue unit belonging to the same subject is calculated, and the belonging stability value is obtained.

7. The method for identity authentication and financial service access management integrating multiple document information as described in claim 6, characterized in that, Based on the stable attribution value, the subject lead data is solidified into subject profiles, and the subject attribution status and business attribute fields of the subject profiles are identified to determine the access status of the target subject, thereby obtaining access management decision data, including: Based on the attribution stability value, the attribution stability value is mapped to a preset stability interval to determine the stability level of the main clue data and obtain the stability level data; Based on the stability level data, extract the main clues that meet the preset solidification conditions for stability level, and generate the main file based on the main clue data to obtain the main file data; Based on the main file data, determine the ownership status of the main file, and summarize the attribute fields related to financial services to obtain the main status data and business attribute data; Based on the entity status data and business attribute data, the access status of the target entity in financial services is determined, and access management decision data is obtained.

8. An identity authentication and financial service access management system integrating multiple document information, characterized in that: The system is used to perform the method as described in any one of claims 1 to 7, the system comprising: The data module is used to acquire basic data of multiple certificates, parse the business meaning of each field in different certificates, group fields with the same business meaning into semantic units, record the certificate source of each semantic unit, and obtain semantic unit data. The constraint module is used to read the order of appearance of each semantic unit in different documents based on the semantic unit data to form constraint items, and compare the constraint items with the actual values ​​of the semantic units to obtain constraint comparison data; The support module is used to identify the number of value segments, continuous coverage length and number of times they are repeated across documents in the constraint sequence based on the constraint comparison data, calculate the degree to which the semantic unit is supported by multiple documents, and obtain the trajectory support quantity. The main module is used to identify and extract semantic units whose trajectory support reaches a preset support threshold as clue units, and to collect the clue units according to the source of the documents to obtain the main clue data. The link module is used to establish identity link, account link and subject registration link between clue units based on the main clue data, and check the link status of the endpoint fields of each link to obtain link status data. The stabilization module is used to identify the number of loopback links, the number of broken links, and the number of repairs for jump links based on link status data, calculate the stability of clue units belonging to the same subject, and obtain the belonging stability value. The access module is used to solidify the subject clue data into subject profiles based on the stable attribution value, identify the subject attribution status and business attribute fields of the subject profiles, determine the access status of the target subject, and obtain access management decision data.

9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.